【问题标题】:What is the difference between corpus and lexicon in NLTK (python) [closed]NLTK(python)中的语料库和词典有什么区别[关闭]
【发布时间】:2015-10-10 04:11:32
【问题描述】:

谁能告诉我NLTK中Corporacorpuslexicon之间的区别?

什么是电影数据集

什么是Wordnet

【问题讨论】:

  • 如果您可以发布单独的问题而不是将您的问题合并为一个问题,则最好。这样,它可以帮助人们回答您的问题,也可以帮助其他人至少寻找您的一个问题。谢谢!
  • 嘿 Rohit,谢谢您的评论...我添加了这个,因为它们都是相关的...在其他上下文中回答一个会帮助我相信...
  • 这不是 machine-learning 本身,而是更多的 NLTK 和 nlp。

标签: machine-learning nlp nltk corpus lexical


【解决方案1】:

Corpora 是语料库的复数

语料库基本上是指正文,在自然语言处理 (NLP) 的上下文中,它表示正文。

(来源:https://www.google.com.sg/search?q=corpora


Lexicon是一个词汇,一个单词列表,一个字典(来源:https://www.google.com.sg/search?q=lexicon

在 NLTK 中,任何词典都被视为语料库,因为单词列表也是文本主体。例如。可以在 NLTK 语料库 API 中找到停用词列表:

>>> from nltk.corpus import stopwords
>>> print stopwords.words('english')
[u'i', u'me', u'my', u'myself', u'we', u'our', u'ours', u'ourselves', u'you', u'your', u'yours', u'yourself', u'yourselves', u'he', u'him', u'his', u'himself', u'she', u'her', u'hers', u'herself', u'it', u'its', u'itself', u'they', u'them', u'their', u'theirs', u'themselves', u'what', u'which', u'who', u'whom', u'this', u'that', u'these', u'those', u'am', u'is', u'are', u'was', u'were', u'be', u'been', u'being', u'have', u'has', u'had', u'having', u'do', u'does', u'did', u'doing', u'a', u'an', u'the', u'and', u'but', u'if', u'or', u'because', u'as', u'until', u'while', u'of', u'at', u'by', u'for', u'with', u'about', u'against', u'between', u'into', u'through', u'during', u'before', u'after', u'above', u'below', u'to', u'from', u'up', u'down', u'in', u'out', u'on', u'off', u'over', u'under', u'again', u'further', u'then', u'once', u'here', u'there', u'when', u'where', u'why', u'how', u'all', u'any', u'both', u'each', u'few', u'more', u'most', u'other', u'some', u'such', u'no', u'nor', u'not', u'only', u'own', u'same', u'so', u'than', u'too', u'very', u's', u't', u'can', u'will', u'just', u'don', u'should', u'now']

NLTK 中的电影评论数据集(规范称为Movie Reviews Corpus)是一个带有情感极性分类的 2k 电影评论的文本数据集(来源:http://www.nltk.org/book/ch02.html

它通常用于介绍 NLP 和情感分析的教程,请参阅 http://www.nltk.org/book/ch06.htmlnltk NaiveBayesClassifier training for sentiment analysis


WordNet英语词汇数据库(它就像一个具有词对词关系的词典/字典)(来源:https://wordnet.princeton.edu/)。

在 NLTK 中,它结合了开放多语言 WordNet (http://compling.hss.ntu.edu.sg/omw/),允许您查询其他语言的单词。

由于它也是一个单词列表(在这种情况下还包括许多其他内容,关系、引理、POS 等),它也可以在 NLTK 中使用 nltk.corpus 调用。

在 NLTK 中使用 wordnet 的规范习语如下:

>>> from nltk.corpus import wordnet as wn
>>> wn.synsets('dog')
[Synset('dog.n.01'), Synset('frump.n.01'), Synset('dog.n.03'), Synset('cad.n.01'), Synset('frank.n.02'), Synset('pawl.n.01'), Synset('andiron.n.01'), Synset('chase.v.01')]

理解/学习 NLP 术语和基础知识的最简单方法是阅读 NLTK 书中的这些教程:http://www.nltk.org/book/

【讨论】:

    猜你喜欢
    • 2020-08-17
    • 1970-01-01
    • 1970-01-01
    • 2020-08-05
    • 1970-01-01
    • 2021-03-24
    • 2019-07-01
    • 2018-04-22
    • 1970-01-01
    相关资源
    最近更新 更多