【问题标题】:extract code from code directive from restructuredtext using docutils使用 docutils 从 restructuredtext 中的 code 指令中提取代码
【发布时间】:2015-02-02 00:34:38
【问题描述】:

我想从重构文本字符串中的代码指令中逐字提取源代码。

接下来是我第一次尝试这样做,但我想知道是否有更好的(即更健壮,或更通用,或更直接)的方法。

假设我在 python 中有以下第一个文本作为字符串:

s = '''

My title
========

Use this to square a number.

.. code:: python

   def square(x):
       return x**2

and here is some javascript too.

.. code:: javascript

    foo = function() {
        console.log('foo');
    }

'''

要获得这两个代码块,我可以这样做

from docutils.core import publish_doctree

doctree = publish_doctree(s)
source_code = [child.astext() for child in doctree.children 
if 'code' in child.attributes['classes']]

现在 source_code 是一个列表,其中仅包含来自两个代码块的逐字源代码。如果需要,我也可以使用 childattributes 属性来找出代码类型。

它可以完成工作,但有更好的方法吗?

【问题讨论】:

    标签: python restructuredtext docutils


    【解决方案1】:

    您的解决方案只会在文档的顶层找到代码块,如果在其他元素上使用“代码”类(不太可能,但可能),它可能会返回误报。我还会检查元素/节点的类型,在其 .tagname 属性中指定。

    节点上有一个“遍历”方法(文档/文档树只是一个特殊的节点),它可以对文档树进行完整的遍历。它将查看文档中的所有元素并仅返回与用户指定条件匹配的元素(返回布尔值的函数)。方法如下:

    def is_code_block(node):
        return (node.tagname == 'literal_block'
                and 'code' in node.attributes['classes'])
    
    code_blocks = doctree.traverse(condition=is_code_block)
    source_code = [block.astext() for block in code_blocks]
    

    【讨论】:

    • 我如何(我可以?)对代码块进行修改,然后将更改写回原始文件?
    【解决方案2】:

    这可以进一步简化为:

    source_code = [block.astext() for block in doctree.traverse(nodes.literal_block)
                   if 'code' in block.attributes['classes']]
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2012-04-07
      • 2012-04-04
      • 2012-12-23
      • 2023-04-05
      • 2012-09-14
      • 1970-01-01
      • 2012-06-01
      • 1970-01-01
      相关资源
      最近更新 更多