【问题标题】:Linux Bash: cURL - how to pass variables to the URLLinux Bash:cURL - 如何将变量传递给 URL
【发布时间】:2017-11-09 07:51:41
【问题描述】:

我想做 cURL GET 请求。应使用以下网址:

https://iant.toulouse.inra.fr/bacteria/annotation/cgi/rhime.cgi' -H 'Host: iant.toulouse.inra.fr' -H 'User-Agent: Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:56.0) Gecko/20100101 Firefox/56.0' -H 'Accept: text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8' -H 'Accept-Language: de,en-US;q=0.7,en;q=0.3' --compressed -H 'Referer: https://iant.toulouse.inra.fr/bacteria/annotation/cgi/rhime.cgi?__wb_cookie=&__wb_cookie_name=auth.rhime&__wb_cookie_path=/bacteria/annotation/cgi&__wb_session=WB84Qfsf&__wb_main_menu=Genome&__wb_function=$parent' -H 'Content-Type: application/x-www-form-urlencoded' -H 'Connection: keep-alive' -H 'Upgrade-Insecure-Requests: 1' -H 'Pragma: no-cache' -H 'Cache-Control: no-cache' --data '__wb_function=PortalExtractSeq&mode=run&species=rhime&fastafile=%2Fwww%2Fbacteria%2Fannotation%2F%2Fsite%2Fprj%2Frhime%2F%2Fdb%2F$ab.genomic&begin=$start&end=$end&strand=$strand

在 URL 的末尾,我有一些单词,我想将它们设计为变量,因此根据输入,URL 会有所不同,然后我请求另一个资源。

网址的结尾。 $ab, $start, $end 和 $strand 是变量,都是字符串。

...2Frhime%2F%2Fdb%2F$ab.genomic&begin=$start&end=$end&strand=$strand

我遇到了“urlencode”,我想将我的 URL 作为一个大字符串存储在一个变量中并将其传递给 URL 编码,但我不确定该怎么做。

我试过这个/我正在寻找这样的东西:

#!bin/bash
[...]
cURL="https://iant.toulouse.inra.fr/bacteria/annotation/cgi/rhime.cgi' -H 'Host: iant.toulouse.inra.fr' -H 'User-Agent: Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:56.0) Gecko/20100101 Firefox/56.0' -H 'Accept: text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8' -H 'Accept-Language: de,en-US;q=0.7,en;q=0.3' --compressed -H 'Referer: https://iant.toulouse.inra.fr/bacteria/annotation/cgi/rhime.cgi?__wb_cookie=&__wb_cookie_name=auth.rhime&__wb_cookie_path=/bacteria/annotation/cgi&__wb_session=WB84Qfsf&__wb_main_menu=Genome&__wb_function=$parent' -H 'Content-Type: application/x-www-form-urlencoded' -H 'Connection: keep-alive' -H 'Upgrade-Insecure-Requests: 1' -H 'Pragma: no-cache' -H 'Cache-Control: no-cache' --data '__wb_function=PortalExtractSeq&mode=run&species=rhime&fastafile=%2Fwww%2Fbacteria%2Fannotation%2F%2Fsite%2Fprj%2Frhime%2F%2Fdb%2F$ab.genomic&begin=$start&end=$end&strand=$strand"

# storing HTTP response code in variable response. Only if the
# reponse code is OK (200), we move on
  response=$(curl -X HEAD -I --header 'Accept:txt/html' "https://iant.toulouse.inra.fr/bacteria/annotation/cgi/rhime.cgi?__wb_cookie=&__wb_cookie_name=auth.rhime&__wb_cookie_path=/bacteria/annotation/cgi&__wb_session=WB8jqwTM&__wb_main_menu=Genome&__wb_function="$location""|head -n1|awk '{print $2}')

  echo "$response"

# getting information via curl request
  if [ $response = 200 ] ; then
    info=$(curl -G "$ (urlencode "$cURL")")
  fi

  echo $info

对于我的响应代码检查,直接传递 $location 的方法似乎有效,但是使用更多变量时,我得到一个错误(响应代码 100,而我通过代码检查得到 200)

我在理解 curl/urlencode 时是否存在一般性错误?我错过了什么?

提前感谢您的时间和精力:)

更新

#!/bin/sh
# handling command-line input
file=$1
ecf=$2


# iterating through file and pulling out
# information for the GET- and POST-request

while read -r line
  do
    parent=$(echo $line | awk '{print substr($1,2,3)}')
    start=$(echo $line | awk '{print substr($2,2,6)}')
    end=$(echo $line | awk '{print substr($3,2,6)}')
    strand=$(echo $line | awk '{print substr($4,2,1)}')
    locus=$(echo $line | awk '{print substr($6,2,8)}')

# depending on $parent, the right insertion for the URL is generated
    if [ $parent = "SMc" ] ; then
      location="Genome"
      ab="SMc"
    elif [ $parent = "SMa" ] ; then
      location="PrintPsyma"
      ab="pSymA"
    else [ $parent = "SMb" ]
      location="PrintPsymb"
      ab="pSymB"
    fi
# building variables for curl content request


  options=( --compressed)

  headers=(
    -H 'Host: iant.toulouse.inra.fr'
    -H 'User-Agent: Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:56.0) Gecko/20100101 Firefox/56.0'
    -H 'Accept: txt/html,application/xhtml+xml,application/xml;1=0.9,*/*;q=0.8'
    -H 'Accept-Language: de,en-US;q=0.7,en;q=0.3'
    -H 'Referer: https://iant.toulouse.inra.fr/bacteria/annotation/cgi/rhime.cgi?__wb_cookie=&__wb_cookie_name=auth.rhime&__wb_cookie_path=/bacteria/annotation/cgi&__wb_session=WB84Qfsf&__wb_main_menu=Genome&__wb_function=$parent'
    -H 'Content-Type: application/x-www-form-urlencoded'
    -H 'Connection: keep-alive'
    -H 'Upgrade-Insecure-Requests: 1'
    -H 'Pragma: no-cache'
    -H 'Cache-Control: no-cache'
  )

    url='https://iant.toulouse.inra.fr/bacteria/annotation/cgi/rhime.cgi'

    ab=$(urlencode "${ab}")
    start=$(urlencode "${start}")
    end=$(urlencode "${end}")
    strand=$(urlencode "${strand}")
    data="__wb_function=PortalExtractSeq&mode=run&species=rhime&fastafile=%2Fwww%2Fbacteria%2Fannotation%2F%2Fsite%2Fprj%2Frhime%2F%2Fdb%2F$ab.genomic&begin=$start&end=$end&strand=$strand"




# storing HTTP response code in variable response. Only if the
# reponse code is OK (200), we move on
    response=$(curl -X HEAD -I --header 'Accept:txt/html' "https://iant.toulouse.inra.fr/bacteria/annotation/cgi/rhime.cgi?__wb_cookie=&__wb_cookie_name=auth.rhime&__wb_cookie_path=/bacteria/annotation/cgi&__wb_session=WB8jqwTM&__wb_main_menu=Genome&__wb_function="$location""|head -n1|awk '{print $2}')

    echo "$response"

# getting information via curl request
    if [ $response = 200 ] ; then
        info=$(curl -G "${options[@]}" "${headers[@]}" --data "${data}" "${url}")
    fi

    echo $info

done < $file

【问题讨论】:

    标签: linux bash curl urlencode http-get


    【解决方案1】:

    您需要分离概念。您放入 cURL 变量中的字符串不是 URL,它是 URL + 标题集 + 参数 + 一个压缩选项。它们都是不同的东西。

    像这样分别定义它们:

    url='https://iant.toulouse.inra.fr/bacteria/annotation/cgi/rhime.cgi'
    headers=(
        -H 'Host: iant.toulouse.inra.fr'
        -H 'User-Agent: ...'
        -H 'Accept: ...'
        -H 'Accept-Language: ...'
        ... other headers from your example ...
    )
    options=(
        --compressed
    )
    data="__wb_function=PortalExtractSeq&mode=run&species=rhime&fastafile=%2Fwww%2Fbacteria%2Fannotation%2F%2Fsite%2Fprj%2Frhime%2F%2Fdb%2F$ab.genomic&begin=$start&end=$end&strand=$strand"
    

    然后以这种方式运行 curl:

    curl -G "${options[@]}" "${headers[@]}" --data "${data}" "${url}"
    

    这将扩展为正确的 curl 命令。

    关于 urlencode 部分:您需要分别对 $ab、$start、$end 和 $strand 进行编码。如果您将它们插入字符串然后编码,那么该字符串中的所有特殊字符(如 &amp;= 也将被编码,而那些已经编码的字符(如您的示例中的 %2F)将被编码两次(将变为%252F)。

    为了保持代码整洁,您可以预先对其进行编码:

    ab=$(urlencode "${ab}")
    start=$(urlencode "${start}")
    end=$(urlencode "${end}")
    strand=$(urlencode "${strand}")
    data="__wb_function=PortalExtractSeq&mode=run&species=rhime&fastafile=%2Fwww%2Fbacteria%2Fannotation%2F%2Fsite%2Fprj%2Frhime%2F%2Fdb%2F$ab.genomic&begin=$start&end=$end&strand=$strand"
    

    ...或者用繁琐的方式来做:

    data="__wb_function=PortalExtractSeq&mode=run&species=rhime&fastafile=%2Fwww%2Fbacteria%2Fannotation%2F%2Fsite%2Fprj%2Frhime%2F%2Fdb%2F$(urlencode "${ab}").genomic&begin=$(urlencode "${start}")&end=$(urlencode "${end}")&strand=$(urlencode "${strand}")"
    

    我希望这会有所帮助。

    【讨论】:

    • 谢谢!确实如此。我通过使用浏览器中的开发人员视图手动复制它获得了“URL”。我手动完成了请求,找到了正确的请求,并选择了“复制为 cURL 地址”。所以我想,我得到的东西会立即起作用^^"。
    • 不知何故,options=(--compressed) 上的括号给了我一个错误:“(”意外。这是从哪里来的?
    • @Shushiro 在= 运算符之前或之后是否有空格字符?
    • 检查了空格,没有。它可能与循环有关吗?我走得更远,我正在从文件输入中读取参数(使用 while 循环进行循环)。如果你有兴趣,我更新了代码。
    • 您使用的是sh 而不是bash。将脚本中的 she-bang 从#!/bin/sh 更改为#!/bin/bash,它应该可以工作。我用来定义数组的语法是特定于 bash 的。你的问题的主题和标签让我觉得我可以使用它。
    猜你喜欢
    • 2018-05-17
    • 1970-01-01
    • 1970-01-01
    • 2013-11-02
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-07-28
    • 2018-04-12
    相关资源
    最近更新 更多