【问题标题】:Javascript: REGEX to change all relative Urls to AbsoluteJavascript:正则表达式将所有相对 URL 更改为绝对
【发布时间】:2011-11-24 13:33:33
【问题描述】:

我目前正在创建一个 Node.js 网络爬虫/代理,但我无法解析在源代码的脚本部分中找到的相关 URL,我认为 REGEX 可以解决问题。 虽然不知道我将如何实现这一目标。

无论如何我可以解决这个问题吗?

我也愿意接受一种更简单的方法,因为我对其他代理如何解析网站感到很困惑。我认为大多数只是美化的网站抓取工具,它们可以读取网站的源并将所有链接/表单转发回代理。

【问题讨论】:

  • 我会使用真正的解析器,而不是正则表达式。节点有 html 解析器。

标签: javascript regex node.js proxy web-scraping


【解决方案1】:

根据上面 Rob W 的评论,我写了一个注入函数:

function injectBase(html, base) {
  // Remove any <base> elements inside <head>     
  html = html.replace(/(<[^>/]*head[^>]*>)[\s\S]*?(<[^>/]*base[^>]*>)[\s\S]*?(<[^>]*head[^>]*>)/img, "$1 $3");

  // Add <base> just before </head>  
  html = html.replace(/(<[^>/]*head[^>]*>[\s\S]*?)(<[^>]*head[^>]*>)/img, "$1 " + base + " $2");  
  return(html);
}

【讨论】:

    【解决方案2】:

    这是当前线程中的Rob W answer "Advanced HTML string replacement functions" 加上我重构的一些代码以使 JSLint 满意。

    我应该将它作为答案的评论发布,但我没有足够的声望点。

    /*jslint browser: true */
    /*jslint regexp: true */
    /*jslint unparam: true*/
    /*jshint strict: false */
    
    /**
     * convertRelToAbsUrl
     *
     * https://stackoverflow.com/a/7544757/1983903
     * 
     * @param  {String} url
     * @return {String} updated url
     */
    function convertRelToAbsUrl(url) {
        var baseUrl = null;
    
        if (/^(https?|file|ftps?|mailto|javascript|data:image\/[^;]{2,9};):/i.test(url)) {
            return url; // url is already absolute
        }
    
        baseUrl = location.href.match(/^(.+)\/?(?:#.+)?$/)[0] + '/';
    
        if (url.substring(0, 2) === '//') {
            return location.protocol + url;
        }
        if (url.charAt(0) === '/') {
            return location.protocol + '//' + location.host + url;
        }
        if (url.substring(0, 2) === './') {
            url = '.' + url;
        } else if (/^\s*$/.test(url)) {
            return ''; // empty = return nothing
        }
    
        url = baseUrl + '../' + url;
    
        while (/\/\.\.\//.test(url)) {
            url = url.replace(/[^\/]+\/+\.\.\//g, '');
        }
    
        url = url.replace(/\.$/, '').replace(/\/\./g, '').replace(/"/g, '%22')
                .replace(/'/g, '%27').replace(/</g, '%3C').replace(/>/g, '%3E');
    
        return url;
    }
    
    /**
     * convertAllRelativeToAbsoluteUrls
     *
     * https://stackoverflow.com/a/7544757/1983903
     * 
     * @param  {String} html
     * @return {String} updated html
     */
    function convertAllRelativeToAbsoluteUrls(html) {
        var me = this,
            att = '[^-a-z0-9:._]',
            entityEnd = '(?:;|(?!\\d))',
            ents = {
                ' ' : '(?:\\s|&nbsp;?|&#0*32' + entityEnd + '|&#x0*20' + entityEnd + ')',
                '(' : '(?:\\(|&#0*40' + entityEnd + '|&#x0*28' + entityEnd + ')',
                ')' : '(?:\\)|&#0*41' + entityEnd + '|&#x0*29' + entityEnd + ')',
                '.' : '(?:\\.|&#0*46' + entityEnd + '|&#x0*2e' + entityEnd + ')'
            },
            charMap = {},
            s = ents[' '] + '*', // short-hand for common use
            any = '(?:[^>\"\']*(?:\"[^\"]*\"|\'[^\']*\'))*?[^>]*',
            slashRE = null,
            dotRE = null;
    
        function ae(string) {
            var allCharsLowerCase = string.toLowerCase(),
                allCharsUpperCase = string.toUpperCase(),
                reRes = '',
                charLowerCase = null,
                charUpperCase = null,
                reSub = null,
                i = null;
    
            if (ents[string]) {
                return ents[string];
            }
    
            for (i = 0; i < string.length; i++) {
                charLowerCase = allCharsLowerCase.charAt(i);
                if (charMap[charLowerCase]) {
                    reRes += charMap[charLowerCase];
                    continue;
                }
                charUpperCase = allCharsUpperCase.charAt(i);
                reSub = [charLowerCase];
                reSub.push('&#0*' + charLowerCase.charCodeAt(0) + entityEnd);
                reSub.push('&#x0*' + charLowerCase.charCodeAt(0).toString(16) + entityEnd);
    
                if (charLowerCase !== charUpperCase) {
                    reSub.push('&#0*' + charUpperCase.charCodeAt(0) + entityEnd);
                    reSub.push('&#x0*' + charUpperCase.charCodeAt(0).toString(16) + entityEnd);
                }
                reSub = '(?:' + reSub.join('|') + ')';
                reRes += (charMap[charLowerCase] = reSub);
            }
            return (ents[string] = reRes);
        }
    
        function by(match, group1, group2, group3) {
            return group1 + me.convertRelToAbsUrl(group2) + group3;
        }
    
        slashRE = new RegExp(ae('/'), 'g');
        dotRE = new RegExp(ae('.'), 'g');
    
        function by2(match, group1, group2, group3) {
            group2 = group2.replace(slashRE, '/').replace(dotRE, '.');
            return group1 + me.convertRelToAbsUrl(group2) + group3;
        }
    
        function cr(selector, attribute, marker, delimiter, end) {
            var re1 = null,
                re2 = null,
                re3 = null;
    
            if (typeof selector === 'string') {
                selector = new RegExp(selector, 'gi');
            }
    
            attribute = att + attribute;
            marker = typeof marker === 'string' ? marker : '\\s*=\\s*';
            delimiter = typeof delimiter === 'string' ? delimiter : '';
            end = typeof end === 'string' ? '?)(' + end : ')(';
    
            re1 = new RegExp('(' + attribute + marker + '")([^"' + delimiter + ']+' + end + ')', 'gi');
            re2 = new RegExp('(' + attribute + marker + '\')([^\'' + delimiter + ']+' + end + ')', 'gi');
            re3 = new RegExp('(' + attribute + marker + ')([^"\'][^\\s>' + delimiter + ']*' + end + ')', 'gi');
    
            html = html.replace(selector, function (match) {
                return match.replace(re1, by).replace(re2, by).replace(re3, by);
            });
        }
    
        function cri(selector, attribute, front, flags, delimiter, end) {
            var re1 = null,
                re2 = null,
                at1 = null,
                at2 = null,
                at3 = null,
                handleAttr = null;
    
            if (typeof selector === 'string') {
                selector = new RegExp(selector, 'gi');
            }
    
            attribute = att + attribute;
            flags = typeof flags === 'string' ? flags : 'gi';
            re1 = new RegExp('(' + attribute + '\\s*=\\s*")([^"]*)', 'gi');
            re2 = new RegExp("(" + attribute + "\\s*=\\s*')([^']+)", 'gi');
            at1 = new RegExp('(' + front + ')([^"]+)(")', flags);
            at2 = new RegExp("(" + front + ")([^']+)(')", flags);
    
            if (typeof delimiter === 'string') {
                end = typeof end === 'string' ? end : '';
                at3 = new RegExp('(' + front + ')([^\"\'][^' + delimiter + ']*' + (end ? '?)(' + end + ')' : ')()'), flags);
                handleAttr = function (match, g1, g2) {
                    return g1 + g2.replace(at1, by2).replace(at2, by2).replace(at3, by2);
                };
            } else {
                handleAttr = function (match, g1, g2) {
                    return g1 + g2.replace(at1, by2).replace(at2, by2);
                };
            }
            html = html.replace(selector, function (match) {
                return match.replace(re1, handleAttr).replace(re2, handleAttr);
            });
        }
    
        cri('<meta' + any + att + 'http-equiv\\s*=\\s*(?:\"' + ae('refresh')
            + '\"' + any + '>|\'' + ae('refresh') + '\'' + any + '>|' + ae('refresh')
            + '(?:' + ae(' ') + any + '>|>))', 'content', ae('url') + s + ae('=') + s, 'i');
    
        cr('<' + any + att + 'href\\s*=' + any + '>', 'href'); /* Linked elements */
        cr('<' + any + att + 'src\\s*=' + any + '>', 'src'); /* Embedded elements */
    
        cr('<object' + any + att + 'data\\s*=' + any + '>', 'data'); /* <object data= > */
        cr('<applet' + any + att + 'codebase\\s*=' + any + '>', 'codebase'); /* <applet codebase= > */
    
        /* <param name=movie value= >*/
        cr('<param' + any + att + 'name\\s*=\\s*(?:\"' + ae('movie') + '\"' + any + '>|\''
            + ae('movie') + '\'' + any + '>|' + ae('movie') + '(?:' + ae(' ') + any + '>|>))', 'value');
    
        cr(/<style[^>]*>(?:[^"']*(?:"[^"]*"|'[^']*'))*?[^'"]*(?:<\/style|$)/gi,
            'url', '\\s*\\(\\s*', '', '\\s*\\)'); /* <style> */
        cri('<' + any + att + 'style\\s*=' + any + '>', 'style',
            ae('url') + s + ae('(') + s, 0, s + ae(')'), ae(')')); /*< style=" url(...) " > */
    
        return html;
    }

    【讨论】:

      【解决方案3】:

      将网址从相对网址转换为绝对网址的可靠方法是使用内置的url module

      例子:

      var url = require('url');
      url.resolve("http://www.example.org/foo/bar/", "../baz/qux.html");
      
      >> gives 'http://www.example.org/foo/baz/qux.html' 
      

      【讨论】:

      • “要求”从何而来?
      • @ajkochanowicz:问题是关于 Node.js 应用程序的。 require() 是 C #include &lt;...&gt; 的 Node.js 等价物。 (嗯,不完全是。)所以,我的答案在编写 JS 代码以在浏览器中运行时不能使用。
      【解决方案4】:

      高级 HTML 字符串替换功能

      请注意 OP,因为他请求了这样一个功能:将 base_url 更改为您代理的基本 URL 以达到预期的结果。

      下面将显示两个函数(使用指南包含在代码中)。确保不要跳过此答案的任何解释部分,以完全理解函数的行为。

      • rel_to_abs(urL) - 此函数返回绝对 URL。当传递一个具有普遍信任协议的绝对 URL 时,它会立即返回这个 URL。否则,从 base_url 和函数参数生成一个绝对 URL。正确解析了相对 URL(../././/)。
      • replace_all_rel_by_abs - 此函数将解析 所有 在 HTML 中具有重要意义的 URL,例如 CSS url()、链接和外部资源。有关已解析实例的完整列表,请参阅代码。请参阅 this answer 了解调整后的实现,以从外部源清理 HTML 字符串(以嵌入到文档中)。
      • 测试用例(答案底部):要测试功能的有效性,只需将小书签粘贴到位置栏即可。


      rel_to_abs - 解析相对 URL
      function rel_to_abs(url){
          /* Only accept commonly trusted protocols:
           * Only data-image URLs are accepted, Exotic flavours (escaped slash,
           * html-entitied characters) are not supported to keep the function fast */
        if(/^(https?|file|ftps?|mailto|javascript|data:image\/[^;]{2,9};):/i.test(url))
               return url; //Url is already absolute
      
          var base_url = location.href.match(/^(.+)\/?(?:#.+)?$/)[0]+"/";
          if(url.substring(0,2) == "//")
              return location.protocol + url;
          else if(url.charAt(0) == "/")
              return location.protocol + "//" + location.host + url;
          else if(url.substring(0,2) == "./")
              url = "." + url;
          else if(/^\s*$/.test(url))
              return ""; //Empty = Return nothing
          else url = "../" + url;
      
          url = base_url + url;
          var i=0
          while(/\/\.\.\//.test(url = url.replace(/[^\/]+\/+\.\.\//g,"")));
      
          /* Escape certain characters to prevent XSS */
          url = url.replace(/\.$/,"").replace(/\/\./g,"").replace(/"/g,"%22")
                  .replace(/'/g,"%27").replace(/</g,"%3C").replace(/>/g,"%3E");
          return url;
      }
      

      案例/例子:

      • http://foo.bar。已经是绝对 URL,因此立即返回。
      • /doo 相对于根:返回当前根 + 提供的相对 URL。
      • ./meh 相对于当前目录。
      • ../booh 相对于父目录。

      该函数将相对路径转换为../,并执行搜索和替换(http://domain/sub/anything-but-a-slash/../mehttp://domain/sub/me)。


      replace_all_rel_by_abs - 转换所有相关的 URL 出现
      脚本实例中的 URL(&lt;script&gt;,事件处理程序被替换,因为几乎不可能创建一个快速且安全的过滤器来解析 JavaScript。

      这个脚本里面有一些 cmets。正则表达式是动态创建的,因为单个 RE 可以有 3000 个字符的大小。 &lt;meta http-equiv=refresh content=.. &gt; 可以通过各种方式进行混淆,因此是 RE 的大小。

      function replace_all_rel_by_abs(html){
          /*HTML/XML Attribute may not be prefixed by these characters (common 
             attribute chars.  This list is not complete, but will be sufficient
             for this function (see http://www.w3.org/TR/REC-xml/#NT-NameChar). */
          var att = "[^-a-z0-9:._]";
      
          var entityEnd = "(?:;|(?!\\d))";
          var ents = {" ":"(?:\\s|&nbsp;?|&#0*32"+entityEnd+"|&#x0*20"+entityEnd+")",
                      "(":"(?:\\(|&#0*40"+entityEnd+"|&#x0*28"+entityEnd+")",
                      ")":"(?:\\)|&#0*41"+entityEnd+"|&#x0*29"+entityEnd+")",
                      ".":"(?:\\.|&#0*46"+entityEnd+"|&#x0*2e"+entityEnd+")"};
                      /* Placeholders to filter obfuscations */
          var charMap = {};
          var s = ents[" "]+"*"; //Short-hand for common use
          var any = "(?:[^>\"']*(?:\"[^\"]*\"|'[^']*'))*?[^>]*";
          /* ^ Important: Must be pre- and postfixed by < and >.
           *   This RE should match anything within a tag!  */
      
          /*
            @name ae
            @description  Converts a given string in a sequence of the original
                            input and the HTML entity
            @param String string  String to convert
            */
          function ae(string){
              var all_chars_lowercase = string.toLowerCase();
              if(ents[string]) return ents[string];
              var all_chars_uppercase = string.toUpperCase();
              var RE_res = "";
              for(var i=0; i<string.length; i++){
                  var char_lowercase = all_chars_lowercase.charAt(i);
                  if(charMap[char_lowercase]){
                      RE_res += charMap[char_lowercase];
                      continue;
                  }
                  var char_uppercase = all_chars_uppercase.charAt(i);
                  var RE_sub = [char_lowercase];
                  RE_sub.push("&#0*" + char_lowercase.charCodeAt(0) + entityEnd);
                  RE_sub.push("&#x0*" + char_lowercase.charCodeAt(0).toString(16) + entityEnd);
                  if(char_lowercase != char_uppercase){
                      /* Note: RE ignorecase flag has already been activated */
                      RE_sub.push("&#0*" + char_uppercase.charCodeAt(0) + entityEnd);   
                      RE_sub.push("&#x0*" + char_uppercase.charCodeAt(0).toString(16) + entityEnd);
                  }
                  RE_sub = "(?:" + RE_sub.join("|") + ")";
                  RE_res += (charMap[char_lowercase] = RE_sub);
              }
              return(ents[string] = RE_res);
          }
      
          /*
            @name by
            @description  2nd argument for replace().
            */
          function by(match, group1, group2, group3){
              /* Note that this function can also be used to remove links:
               * return group1 + "javascript://" + group3; */
              return group1 + rel_to_abs(group2) + group3;
          }
          /*
            @name by2
            @description  2nd argument for replace(). Parses relevant HTML entities
            */
          var slashRE = new RegExp(ae("/"), 'g');
          var dotRE = new RegExp(ae("."), 'g');
          function by2(match, group1, group2, group3){
              /*Note that this function can also be used to remove links:
               * return group1 + "javascript://" + group3; */
              group2 = group2.replace(slashRE, "/").replace(dotRE, ".");
              return group1 + rel_to_abs(group2) + group3;
          }
          /*
            @name cr
            @description            Selects a HTML element and performs a
                                      search-and-replace on attributes
            @param String selector  HTML substring to match
            @param String attribute RegExp-escaped; HTML element attribute to match
            @param String marker    Optional RegExp-escaped; marks the prefix
            @param String delimiter Optional RegExp escaped; non-quote delimiters
            @param String end       Optional RegExp-escaped; forces the match to end
                                    before an occurence of <end>
           */
          function cr(selector, attribute, marker, delimiter, end){
              if(typeof selector == "string") selector = new RegExp(selector, "gi");
              attribute = att + attribute;
              marker = typeof marker == "string" ? marker : "\\s*=\\s*";
              delimiter = typeof delimiter == "string" ? delimiter : "";
              end = typeof end == "string" ? "?)("+end : ")(";
              var re1 = new RegExp('('+attribute+marker+'")([^"'+delimiter+']+'+end+')', 'gi');
              var re2 = new RegExp("("+attribute+marker+"')([^'"+delimiter+"]+"+end+")", 'gi');
              var re3 = new RegExp('('+attribute+marker+')([^"\'][^\\s>'+delimiter+']*'+end+')', 'gi');
              html = html.replace(selector, function(match){
                  return match.replace(re1, by).replace(re2, by).replace(re3, by);
              });
          }
          /* 
            @name cri
            @description            Selects an attribute of a HTML element, and
                                      performs a search-and-replace on certain values
            @param String selector  HTML element to match
            @param String attribute RegExp-escaped; HTML element attribute to match
            @param String front     RegExp-escaped; attribute value, prefix to match
            @param String flags     Optional RegExp flags, default "gi"
            @param String delimiter Optional RegExp-escaped; non-quote delimiters
            @param String end       Optional RegExp-escaped; forces the match to end
                                      before an occurence of <end>
           */
          function cri(selector, attribute, front, flags, delimiter, end){
              if(typeof selector == "string") selector = new RegExp(selector, "gi");
              attribute = att + attribute;
              flags = typeof flags == "string" ? flags : "gi";
              var re1 = new RegExp('('+attribute+'\\s*=\\s*")([^"]*)', 'gi');
              var re2 = new RegExp("("+attribute+"\\s*=\\s*')([^']+)", 'gi');
              var at1 = new RegExp('('+front+')([^"]+)(")', flags);
              var at2 = new RegExp("("+front+")([^']+)(')", flags);
              if(typeof delimiter == "string"){
                  end = typeof end == "string" ? end : "";
                  var at3 = new RegExp("("+front+")([^\"'][^"+delimiter+"]*" + (end?"?)("+end+")":")()"), flags);
                  var handleAttr = function(match, g1, g2){return g1+g2.replace(at1, by2).replace(at2, by2).replace(at3, by2)};
              } else {
                  var handleAttr = function(match, g1, g2){return g1+g2.replace(at1, by2).replace(at2, by2)};
          }
              html = html.replace(selector, function(match){
                   return match.replace(re1, handleAttr).replace(re2, handleAttr);
              });
          }
      
          /* <meta http-equiv=refresh content="  ; url= " > */
          cri("<meta"+any+att+"http-equiv\\s*=\\s*(?:\""+ae("refresh")+"\""+any+">|'"+ae("refresh")+"'"+any+">|"+ae("refresh")+"(?:"+ae(" ")+any+">|>))", "content", ae("url")+s+ae("=")+s, "i");
      
          cr("<"+any+att+"href\\s*="+any+">", "href"); /* Linked elements */
          cr("<"+any+att+"src\\s*="+any+">", "src"); /* Embedded elements */
      
          cr("<object"+any+att+"data\\s*="+any+">", "data"); /* <object data= > */
          cr("<applet"+any+att+"codebase\\s*="+any+">", "codebase"); /* <applet codebase= > */
      
          /* <param name=movie value= >*/
          cr("<param"+any+att+"name\\s*=\\s*(?:\""+ae("movie")+"\""+any+">|'"+ae("movie")+"'"+any+">|"+ae("movie")+"(?:"+ae(" ")+any+">|>))", "value");
      
          cr(/<style[^>]*>(?:[^"']*(?:"[^"]*"|'[^']*'))*?[^'"]*(?:<\/style|$)/gi, "url", "\\s*\\(\\s*", "", "\\s*\\)"); /* <style> */
          cri("<"+any+att+"style\\s*="+any+">", "style", ae("url")+s+ae("(")+s, 0, s+ae(")"), ae(")")); /*< style=" url(...) " > */
          return html;
      }
      

      私有函数的简短总结:

      • rel_to_abs(url) - 将相对/未知 URL 转换为绝对 URL
      • replace_all_rel_by_abs(html) - 用绝对 URL 替换 HTML 字符串中所有相关的 URL。
        1. ae - Any Entity - 返回一个 RE 模式来处理 HTML 实体。
        2. by - 替换 by - 这个简短的函数请求实际的 url 替换 (rel_to_abs)。这个函数可能被调用数百次,甚至数千次。请注意不要向此函数添加慢速算法(自定义)。
        3. cr - Create Replace - 创建并执行搜索和替换。
          示例:href="..."(在任何 HTML 标记中)。李>
        4. cri - Create Replace Inline - 创建并执行搜索和替换。
          示例:url(..)在 HTML 标记内的所有 style 属性中。

      测试用例

      打开任意页面,在地址栏中粘贴以下书签:

      javascript:void(function(){var s=document.createElement("script");s.src="http://rob.lekensteyn.nl/rel_to_abs.js";document.body.appendChild(s)})();
      

      注入的代码包含上面定义的两个函数,以及如下所示的测试用例。 注意:测试用例不会修改页面的 HTML,而是在文本区域中显示解析结果(可选)。

      var t=(new Date).getTime();
        var result = replace_all_rel_by_abs(document.documentElement.innerHTML);
        if(confirm((new Date).getTime()-t+" milliseconds to execute\n\nPut results in new textarea?")){
          var txt = document.createElement("textarea");
          txt.style.cssText = "position:fixed;top:0;left:0;width:100%;height:99%"
          txt.ondblclick = function(){this.parentNode.removeChild(this)}
          txt.value = result;
          document.body.appendChild(txt);
      }
      

      另见:

      【讨论】:

      • 谢谢,你知道有什么方法可以匹配脚本中的所有相对网址吗?
      • 在你的代码中包含我的函数,并在你想从一个可能的相对 URL 中获取绝对 URL 时调用 rel_to_abs,例如:var some_url = ".././callback/xhr.php";rel_to_abs(some_url);
      • 是的,我明白了。但我的意思是,我如何能够扫描我正在代理的网站以找到这些网址?
      • @Alex 不,我不会使用它们。 4 年前,我很高兴编写复杂的正则表达式来解析 URL,但现在我建议使用专用的 URL 解析器或 DOM 解析器(现在所有现代浏览器都具有这些功能)。如果您使用的是 Node.js,请使用 url 模块。如果您使用浏览器,请使用URL constructordocument.createElement('a') 或任何其他 url 库。
      • @Alex 如果结果将显示为 HTML 页面,只需在 HTML 中添加 &lt;base&gt; 标记,然后您根本不需要解析页面(它会也为外部资源工作)。否则,您可以使用DOMParser API 创建一个文档,然后遍历文档树并替换您关心的所有内容(例如文本节点、HTML 属性、样式表...)。
      【解决方案5】:

      如果您使用正则表达式来查找所有非绝对 URL,那么您只需在它们前面加上当前 URL 就可以了。

      您需要修复的网址不是以/http(s)://(或其他协议标记,如果您关心它们)开头的网址

      例如,假设您正在抓取 http://www.example.com/。如果你遇到一个相对 URL,比如说foo/bar,你只需像这样在被抓取的 URL 前加上前缀:http://www.example.com/foo/bar

      对于一个从页面中抓取 URL 的正则表达式,如果你用谷歌搜索一下,可能有很多好的可用的,所以我不会在这里开始发明一个糟糕的 :)

      【讨论】:

      猜你喜欢
      • 2018-12-13
      • 1970-01-01
      • 2011-01-29
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多