【问题标题】:stemDocment in tm package not working on past tense wordtm 包中的 stemDocment 不适用于过去时词
【发布时间】:2016-03-26 01:33:30
【问题描述】:

我有一个文件“check_text.txt”,其中包含“said say says make made”。我想对其进行词干提取以获得“say say say make make”。我尝试在tm 包中使用stemDocument,如下所示,但只得到“说说制造”。有没有办法对过去时词进行词干提取?在现实世界的自然语言处理中是否有必要这样做?谢谢!

filename = 'check_text.txt'
con <- file(filename, "rb")
text_data <- readLines(con,skipNul = TRUE)
close(con)
text_VS <- VectorSource(text_data)
text_corpus <- VCorpus(text_VS)
text_corpus <- tm_map(text_corpus, stemDocument, language = "english")
as.data.frame(text_corpus)$text

编辑:我也在SnowballC包中尝试了wordStem

> library(SnowballC)
> wordStem(c("said", "say", "says", "make", "made"))
[1] "said" "sai"  "sai"  "make" "made"

【问题讨论】:

    标签: r nlp tm stemming snowball


    【解决方案1】:

    如果一个包中有一个不规则英语动词的数据集,这个任务就很容易了。我只是不知道任何包含此类数据的包,所以我选择通过抓取创建自己的数据库。我不确定这个网站是否涵盖所有不规则的单词。如有必要,您想搜索更好的网站以创建自己的数据库。一旦你有了你的数据库,你就可以开始你的任务了。

    首先,我使用stemDocument() 并使用-s 清理当前表单。然后,我在words(即past)中收集了过去形式,过去形式的不定式(即inf1),在temp中确定了过去形式的顺序。我进一步确定了temp中过去表格的位置。我终于用不定式形式替换了 sat 形式。我对过去分词重复了同样的过程。

    library(tm)
    library(rvest)
    library(dplyr)
    library(splitstackshape)
    
    
    ### Create a database
    x <- read_html("http://www.englishpage.com/irregularverbs/irregularverbs.html")
    
    x %>%
    html_table(header = TRUE) %>%
    bind_rows %>%
    rename(Past = `Simple Past`, PP = `Past Participle`) %>%
    filter(!Infinitive %in% LETTERS) %>%
    cSplit(splitCols = c("Past", "PP"),
           sep = " / ", direction = "long") %>%
    filter(complete.cases(.)) %>%
    mutate_each(funs(gsub(pattern = "\\s\\(.*\\)$|\\s\\[\\?\\]",
                          replacement = "",
                          x = .))) -> mydic
    
    ### Work on the task
    
    words <- c("said", "drawn", "say", "says", "make", "made", "done")
    
    ### says to say
    temp <- stemDocument(words)
    
    ### past forms become present form
    ### Collect past forms
    past <- mydic$Past[which(mydic$Past %in% temp)]
    
    ### Collect infinitive forms of past forms
    inf1 <- mydic$Infinitive[which(mydic$Past %in% temp)]
    
    ### Identify the order of past forms in temp
    ind <- match(temp, past)
    ind <- ind[is.na(ind) == FALSE]
    
    ### Where are the past forms in temp?
    position <- which(temp %in% past)
    
    temp[position] <- inf1[ind]
    
    ### Check
    temp
    #[1] "say"   "drawn" "say"   "say"   "make"  "make"  "done" 
    
    
    ### PP forms to infinitive forms (same as past forms)
    
    pp <- mydic$PP[which(mydic$PP %in% temp)]
    inf2 <- mydic$Infinitive[which(mydic$PP %in% temp)]
    ind <- match(temp, pp)
    ind <- ind[is.na(ind) == FALSE]
    position <- which(temp %in% pp)
    temp[position] <- inf2[ind]
    
    ### Check
    temp
    #[1] "say"  "draw" "say"  "say"  "make" "make" "do" 
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2016-12-19
      • 2014-07-22
      • 1970-01-01
      • 1970-01-01
      • 2021-09-27
      • 1970-01-01
      相关资源
      最近更新 更多