【问题标题】:converting text file to html file with python使用python将文本文件转换为html文件
【发布时间】:2014-07-12 16:34:04
【问题描述】:

我有一个文本文件,其中包含:

JavaScript              0
/AA                     0
OpenAction              1
AcroForm                0
JBIG2Decode             0
RichMedia               0
Launch                  0
Colors>2^24             0
uri                     0

我编写了这段代码来将文本文件转换为 html:

contents = open("C:\\Users\\Suleiman JK\\Desktop\\Static_hash\\test","r")
    with open("suleiman.html", "w") as e:
        for lines in contents.readlines():
            e.write(lines + "<br>\n")

但我在 html 文件中遇到的问题是,在每一行中,两列之间没有空格:

JavaScript 0
/AA 0
OpenAction 1
AcroForm 0
JBIG2Decode 0
RichMedia 0
Launch 0
Colors>2^24 0
uri 0 

我应该怎么做才能拥有相同的内容和文本文件中的两列

【问题讨论】:

  • 发布所需的输出
  • html的内容应该和上面两列的文本文件一样

标签: python html text


【解决方案1】:

只需更改您的代码以包含 &lt;pre&gt;&lt;/pre&gt; 标记,以确保您的文本保持格式与您在原始文本文件中的格式相同。

contents = open"C:\\Users\\Suleiman JK\\Desktop\\Static_hash\\test","r")
with open("suleiman.html", "w") as e:
    for lines in contents.readlines():
        e.write("<pre>" + lines + "</pre> <br>\n")

【讨论】:

  • 您的意思是编辑原始文本文件以减少每列之间的空格吗?
  • 原始文本文件每行之间没有空格,但在html中,行之间有双空格
  • 从 e.write() 调用中的第二个文字中删除您的
    标签。
【解决方案2】:

这是 HTML -- 使用 BeautifulSoup

from bs4 import BeautifulSoup

soup = BeautifulSoup()
body = soup.new_tag('body')
soup.insert(0, body)
table = soup.new_tag('table')
body.insert(0, table)

with open('path/to/input/file.txt') as infile:
    for line in infile:
        row = soup.new_tag('tr')
        col1, col2 = line.split()
        for coltext in (col2, col1): # important that you reverse order
            col = soup.new_tag('td')
            col.string = coltext
            row.insert(0, col)
        table.insert(len(table.contents), row)

with open('path/to/output/file.html', 'w') as outfile:
    outfile.write(soup.prettify())

【讨论】:

    【解决方案3】:

    这是因为 HTML 解析器会折叠所有空格。有两种方法可以做到(可能还有更多)。

    一种方法是将其标记为“预格式化文本”,方法是将其放入&lt;pre&gt;...&lt;/pre&gt; 标签中。

    另一个是一张桌子(这就是一张桌子的用途):

    <table>
      <tr><td>Javascript</td><td>0</td></tr>
      ...
    </table>
    

    手动输入相当繁琐,但很容易从您的脚本中生成。这样的事情应该可以工作:

    contents = open("C:\\Users\\Suleiman JK\\Desktop\\Static_hash\\test","r")
    with open("suleiman.html", "w") as e:
        e.write("<table>\n")   
        for lines in contents.readlines():
            e.write("<tr><td>%s</td><td>%s</td></tr>\n"%lines.split())
        e.write("</table>\n")
    

    【讨论】:

      【解决方案4】:

      您可以使用独立的模板库,例如 makojinja。以下是 jinja 的示例:

      from jinja2 import Template
      c = '''<!doctype html>
      <html>
      <head>
          <title>My Title</title>
      </head>
      <body>
      <table>
         <thead>
             <tr><th>Col 1</th><th>Col 2</th></tr>
         </thead>
         <tbody>
             {% for col1, col2 in lines %}
             <tr><td>{{ col 1}}</td><td>{{ col2 }}</td></tr>
             {% endfor %}
         </tbody>
      </table>
      </body>
      </html>'''
      
      t = Template(c)
      
      lines = []
      
      with open('yourfile.txt', 'r') as f:
          for line in f:
              lines.append(line.split())
      
      with open('results.html', 'w') as f:
          f.write(t.render(lines=lines))
      

      如果您无法安装jinja,那么这里有一个替代方案:

      header = '<!doctyle html><html><head><title>My Title</title></head><body>'
      body = '<table><thead><tr><th>Col 1</th><th>Col 2</th></tr>'
      footer = '</table></body></html>'
      
      with open('input.txt', 'r') as input, open('output.html', 'w') as output:
         output.writeln(header)
         output.writeln(body)
         for line in input:
             col1, col2 = line.rstrip().split()
             output.write('<tr><td>{}</td><td>{}</td></tr>\n'.format(col1, col2))
         output.write(footer)
      

      【讨论】:

        【解决方案5】:

        我已经添加了标题,在这里逐行循环并将每一行附加到

        和 标记上,它应该作为没有列的单个表工作。 col1 和 col2 不需要使用这些标签( 和 [为可读性留出空格])。

        日志:sn-p:

        MUTHU 页面

        2019/08/19 19:59:25 MUTHUKUMAR_TIME_DATE,行:118 INFO |记录器 创建对象:MUTHUKUMAR_APP_USER_SIGNUP_LOG 2019/08/19 19:59:25 MUTHUKUMAR_DB_USER_SIGN_UP,行:48 信息 | ***** 用户注册页面 开始 ***** 2019/08/19 19:59:25 MUTHUKUMAR_DB_USER_SIGN_UP,行:49 信息 |输入名字:[仅允许字母字符,最少 3 个 字符最多 20 个字符]

        html源页面:

        '''

        <?xml version="1.0" encoding="utf-8"?>
        <body>
         <table>
          <p>
           MUTHU PAGE
          </p>
          <tr>
           <td>
            2019/08/19 19:59:25 MUTHUKUMAR_TIME_DATE,line: 118     INFO | Logger object created for: MUTHUKUMAR_APP_USER_SIGNUP_LOG
           </td>
          </tr>
          <tr>
           <td>
            2019/08/19 19:59:25 MUTHUKUMAR_DB_USER_SIGN_UP,line: 48     INFO | ***** User SIGNUP page start *****
           </td>
          </tr>
          <tr>
           <td>
            2019/08/19 19:59:25 MUTHUKUMAR_DB_USER_SIGN_UP,line: 49     INFO | Enter first name: [Alphabet character only allowed, minimum 3 character to maximum 20 chracter]
        

        '''

        代码:

        from bs4 import BeautifulSoup
        
        soup = BeautifulSoup(features='xml')
        body = soup.new_tag('body')
        soup.insert(0, body)
        table = soup.new_tag('table')
        body.insert(0, table)
        
        with open('C:\\Users\xxxxx\\Documents\\Latest_24_may_2019\\New_27_jun_2019\\DB\\log\\input.txt') as infile:
            title_s = soup.new_tag('p')
            title_s.string = " MUTHU PAGE "
            table.insert(0, title_s)
            for line in infile:
                row = soup.new_tag('tr')
                col1 = list(line.split('\n'))
                col1 = [ each for each in col1 if each != '']
                for coltext in col1:
                    col = soup.new_tag('td')
                    col.string = coltext
                    row.insert(0, col)
                table.insert(len(table.contents), row)
        
        with open('C:\\Users\xxxx\\Documents\\Latest_24_may_2019\\New_27_jun_2019\\DB\\log\\output.html', 'w') as outfile:
            outfile.write(soup.prettify())
        

        【讨论】:

          猜你喜欢
          • 2018-02-07
          • 2019-03-14
          • 1970-01-01
          • 2022-01-17
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 2013-01-19
          • 1970-01-01
          相关资源
          最近更新 更多