【问题标题】:C# code to linkify urls in a string用于链接字符串中的 url 的 C# 代码
【发布时间】:2009-04-16 21:28:00
【问题描述】:

是否有任何好的 c# 代码(和正则表达式)可以解析字符串并“链接”字符串中可能存在的任何 url?

【问题讨论】:

  • 这似乎是基于规范正则表达式的解决方案的问题。也许有人可以编辑标题以帮助搜索者找到它?

标签: c# asp.net regex linkify


【解决方案1】:

这是一个非常简单的任务,您可以使用Regex 和一个现成的正则表达式来完成它:

类似:

var html = Regex.Replace(html, @"^(http|https|ftp)\://[a-zA-Z0-9\-\.]+" +
                         "\.[a-zA-Z]{2,3}(:[a-zA-Z0-9]*)?/?" +
                         "([a-zA-Z0-9\-\._\?\,\'/\\\+&%\$#\=~])*$",
                         "<a href=\"$1\">$1</a>");

您可能不仅对创建链接感兴趣,而且对缩短 URL 感兴趣。这是一篇关于这个主题的好文章:

另见

【讨论】:

  • 嗨。大反响。您帖子(和链接)中的大多数建议似乎都有效,但它们似乎都破坏了正在评估的文本中的任何现有链接。
  • VSmith,您可以从 regixlib.com 尝试不同的 reg 表达式,然后找到最适合您的。
  • 嗯,效果很好。从而证明我们在这里提出的所有观点;)
  • 好东西!但是,要使用 .Net 正则表达式 (System.Text.RegularExpressions.Regex),并使用位于文本行中间的 url,我需要像这样修改代码:Regex.Replace(answer, @"(( http|https|ftp)\://[a-zA-Z0-9\-\.]+\.[a-zA-Z]{2,3}(:[a-zA-Z0-9]* )?/?([a-zA-Z0-9\-\._\?\,\'/\\\+&%\$#\=~])*)", @"$1");
  • 我使用以下正则表达式来说明不完全限定的主机名:@"((http|https|ftp)\://[a-zA-Z0-9\-\.]+(\.[a-zA-Z]{2,3})?(:[a-zA-Z0-9]*)?/?([a-zA-Z0-9\-\._\?\,\'/\\\+&amp;amp;%\$#\=~])*)"
【解决方案2】:

好吧,在对此进行了大量研究之后,并多次尝试修复时间

  1. 人们在同一帖子中输入 http://www.sitename.com 和 www.sitename.com
  2. 修复括号,如 (http://www.sitename.com) 和 http://msdn.microsoft.com/en-us/library/aa752574(vs.85).aspx
  3. 长网址如:http://www.amazon.com/gp/product/b000ads62g/ref=s9_simz_gw_s3_p74_t1?pf_rd_m=atvpdkikx0der&pf_rd_s=center-2&pf_rd_r=04eezfszazqzs8xfm9yd&pf_rd_t=101&pf_rd_p=470938631&pf_rd_i=507846

我们现在正在使用这个 HtmlHelper 扩展......我想我会分享并获得任何 cmets:

    private static Regex regExHttpLinks = new Regex(@"(?<=\()\b(https?://|www\.)[-A-Za-z0-9+&@#/%?=~_()|!:,.;]*[-A-Za-z0-9+&@#/%=~_()|](?=\))|(?<=(?<wrap>[=~|_#]))\b(https?://|www\.)[-A-Za-z0-9+&@#/%?=~_()|!:,.;]*[-A-Za-z0-9+&@#/%=~_()|](?=\k<wrap>)|\b(https?://|www\.)[-A-Za-z0-9+&@#/%?=~_()|!:,.;]*[-A-Za-z0-9+&@#/%=~_()|]", RegexOptions.Compiled | RegexOptions.IgnoreCase);

    public static string Format(this HtmlHelper htmlHelper, string html)
    {
        if (string.IsNullOrEmpty(html))
        {
            return html;
        }

        html = htmlHelper.Encode(html);
        html = html.Replace(Environment.NewLine, "<br />");

        // replace periods on numeric values that appear to be valid domain names
        var periodReplacement = "[[[replace:period]]]";
        html = Regex.Replace(html, @"(?<=\d)\.(?=\d)", periodReplacement);

        // create links for matches
        var linkMatches = regExHttpLinks.Matches(html);
        for (int i = 0; i < linkMatches.Count; i++)
        {
            var temp = linkMatches[i].ToString();

            if (!temp.Contains("://"))
            {
                temp = "http://" + temp;
            }

            html = html.Replace(linkMatches[i].ToString(), String.Format("<a href=\"{0}\" title=\"{0}\">{1}</a>", temp.Replace(".", periodReplacement).ToLower(), linkMatches[i].ToString().Replace(".", periodReplacement)));
        }

        // Clear out period replacement
        html = html.Replace(periodReplacement, ".");

        return html;
    }

【讨论】:

    【解决方案3】:
    protected string Linkify( string SearchText ) {
        // this will find links like:
        // http://www.mysite.com
        // as well as any links with other characters directly in front of it like:
        // href="http://www.mysite.com"
        // you can then use your own logic to determine which links to linkify
        Regex regx = new Regex( @"\b(((\S+)?)(@|mailto\:|(news|(ht|f)tp(s?))\://)\S+)\b", RegexOptions.IgnoreCase );
        SearchText = SearchText.Replace( "&nbsp;", " " );
        MatchCollection matches = regx.Matches( SearchText );
    
        foreach ( Match match in matches ) {
            if ( match.Value.StartsWith( "http" ) ) { // if it starts with anything else then dont linkify -- may already be linked!
                SearchText = SearchText.Replace( match.Value, "<a href='" + match.Value + "'>" + match.Value + "</a>" );
            }
        }
    
        return SearchText;
    }
    

    【讨论】:

    • 为发布那个干杯:)
    • 我们最终使用了非常相似的东西,只是做了一个修改。我们最终确保更换只发生一次。这意味着我们最终会丢失一些链接(多次出现的链接),但在两种情况下消除了乱码链接的可能性:1)当有两个链接时,其中一个比另一个更详细。例如"google.comgoogle.com/reader" 2) 当 HTML 链接与纯文本链接混合时。例如"google.com google.com">Google</a>" if (input.IndexOf(match.Value) == input.LastIndexOf(match.Value)) { ... }
    【解决方案4】:

    这并不像您在blog post by Jeff Atwood 中看到的那么容易。检测 URL 的结束位置尤其困难。

    例如,是否是 URL 的尾括号部分:

    • http://en.wikipedia.org/wiki/PCTools(CentralPointSoftware)
    • 括号中的 URL (http://en.wikipedia.org) 更多文本

    在第一种情况下,括号是 URL 的一部分。在第二种情况下,它们不是!

    【讨论】:

    • 正如您从这个答案中的链接 URL 中看到的那样,并不是每个人都做对了 :)
    • 其实我不希望这两个 URL 被链接。但似乎不支持。
    • Jeff 的正则表达式似乎在我的浏览器中显示不好,我认为应该是:"(?\bhttp://[-A-Za-z0-9+&@#/%?=~ _()|!:,.;]*[-A-Za-z0-9+&@#/%=~_()|]"
    【解决方案5】:

    找到以下正则表达式 http://daringfireball.net/2010/07/improved_regex_for_matching_urls

    对我来说看起来很不错。 Jeff Atwood 解决方案不能处理很多情况。 josefresno 在我看来处理所有案件。但是当我试图理解它时(如果有任何支持请求),我的大脑就沸腾了。

    【讨论】:

      【解决方案6】:

      有课:

      public class TextLink
      {
          #region Properties
      
          public const string BeginPattern = "((http|https)://)?(www.)?";
      
          public const string MiddlePattern = @"([a-z0-9\-]*\.)+[a-z]+(:[0-9]+)?";
      
          public const string EndPattern = @"(/\S*)?";
      
          public static string Pattern { get { return BeginPattern + MiddlePattern + EndPattern; } }
      
          public static string ExactPattern { get { return string.Format("^{0}$", Pattern); } }
      
          public string OriginalInput { get; private set; }
      
          public bool Valid { get; private set; }
      
          private bool _isHttps;
      
          private string _readyLink;
      
          #endregion
      
          #region Constructor
      
          public TextLink(string input)
          {
              this.OriginalInput = input;
      
              var text = Regex.Replace(input, @"(^\s)|(\s$)", "", RegexOptions.IgnoreCase);
      
              Valid = Regex.IsMatch(text, ExactPattern);
      
              if (Valid)
              {
                  _isHttps = Regex.IsMatch(text, "^https:", RegexOptions.IgnoreCase);
                  // clear begin:
                  _readyLink = Regex.Replace(text, BeginPattern, "", RegexOptions.IgnoreCase);
                  // HTTPS
                  if (_isHttps)
                  {
                      _readyLink = "https://www." + _readyLink;
                  }
                  // Default
                  else
                  {
                      _readyLink = "http://www." + _readyLink;
                  }
              }
          }
      
          #endregion
      
          #region Methods
      
          public override string ToString()
          {
              return _readyLink;
          }
      
          #endregion
      }
      

      在这个方法中使用它:

      public static string ReplaceUrls(string input)
      {
          var result = Regex.Replace(input.ToSafeString(), TextLink.Pattern, match =>
          {
              var textLink = new TextLink(match.Value);
              return textLink.Valid ?
                  string.Format("<a href=\"{0}\" target=\"_blank\">{1}</a>", textLink, textLink.OriginalInput) :
                  textLink.OriginalInput;
          });
          return result;
      }
      

      测试用例:

      [TestMethod]
      public void RegexUtil_TextLink_Parsing()
      {
          Assert.IsTrue(new TextLink("smthing.com").Valid);
          Assert.IsTrue(new TextLink("www.smthing.com/").Valid);
          Assert.IsTrue(new TextLink("http://smthing.com").Valid);
          Assert.IsTrue(new TextLink("http://www.smthing.com").Valid);
          Assert.IsTrue(new TextLink("http://www.smthing.com/").Valid);
          Assert.IsTrue(new TextLink("http://www.smthing.com/publisher").Valid);
      
          // port
          Assert.IsTrue(new TextLink("http://www.smthing.com:80").Valid);
          Assert.IsTrue(new TextLink("http://www.smthing.com:80/").Valid);
          // https
          Assert.IsTrue(new TextLink("https://smthing.com").Valid);
      
          Assert.IsFalse(new TextLink("").Valid);
          Assert.IsFalse(new TextLink("smthing.com.").Valid);
          Assert.IsFalse(new TextLink("smthing.com-").Valid);
      }
      
      [TestMethod]
      public void RegexUtil_TextLink_ToString()
      {
          // default
          Assert.AreEqual("http://www.smthing.com", new TextLink("smthing.com").ToString());
          Assert.AreEqual("http://www.smthing.com", new TextLink("http://www.smthing.com").ToString());
          Assert.AreEqual("http://www.smthing.com/", new TextLink("smthing.com/").ToString());
      
          Assert.AreEqual("https://www.smthing.com", new TextLink("https://www.smthing.com").ToString());
      }
      

      【讨论】:

      • 这很好用,但是它匹配诸如 o.context 之类的东西,或者其他有句点的字符串。在字符串中的某处强制 .com/.org/.net 等会很好
      • 它还强制使用 www,但并非总是如此。
      【解决方案7】:

      这对我有用:

      str = Regex.Replace(str,
                      @"((http|ftp|https):\/\/[\w\-_]+(\.[\w\-_]+)+([\w\-\.,@?^=%&amp;:/~\+#]*[\w\-\@?^=%&amp;/~\+#])?)",
                      "<a target='_blank' href='$1'>$1</a>");
      

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 2020-08-04
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2019-07-08
        • 1970-01-01
        • 2010-10-05
        • 1970-01-01
        相关资源
        最近更新 更多