【问题标题】:Unlabeled instances in DOCCANO and SpaCY. Do they offer any value?DOCCANO 和 SpaCY 中未标记的实例。他们提供任何价值吗?
【发布时间】:2021-06-11 01:37:48
【问题描述】:

我正在使用 doccano 进行序列标记,并使用 spacy 进行进一步建模。我标记的一些句子不包含我感兴趣的任何标签,因此它们保持“未标记”,即。没有标签。

{"id": 79, "data": "This powerful charm would protect him until he became of age, or no longer called his aunt's house home.", "label": []}
{"id": 82, "data": "He began attending Hogwarts School of Witchcraft and Wizardry in 1991.", "label": []}
{"id": 85, "data": "He later became the youngest Quidditch Seeker in over a century and eventually the captain of the Gryffindor House Quidditch Team in his sixth year, winning two Quidditch Cups.", "label": []}

我想训练 SpaCy 识别所有变体中的角色名称。

现在问题:

  • 为了训练 SpaCy 模型而包含未标记的实例是否有任何价值?
  • 如果有,我应该将此数据声明为“不平衡数据集”并采取相应措施吗? (提升?重击?过采样?等等)
  • 在这种情况下,最佳做法是什么?

【问题讨论】:

  • 我投票结束这个问题,因为它与 help center 中定义的编程无关,而是关于 ML 理论和/或方法 - 请参阅 machine-learning @ 中的介绍和注意事项987654322@.
  • @Sergey Bushmanov 感谢您的编辑。这是我的第一个问题。您的编辑受到赞赏,对于像我这样的小妞来说,这很有教育意义。 Спасибо,Сергей! Я благодарен Вам за правку。

标签: machine-learning nlp spacy doccano


【解决方案1】:

是的,您需要包含一些没有标记的示例,以便模型可以了解要标记的内容。例如,如果在所有示例句子中标记了所有大写单词,则模型可能会学会始终标记大写单词。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2023-03-28
    • 1970-01-01
    • 2018-03-18
    • 2019-04-19
    • 1970-01-01
    相关资源
    最近更新 更多