【问题标题】:Counting words in a single document from corpus in R and putting it in dataframe从R中的语料库中计算单个文档中的单词并将其放入数据框中
【发布时间】:2013-06-25 10:23:03
【问题描述】:

我有文本文档,在每个文档中都有包含电视剧剧透的文本。每个文件都是一个不同的系列。我想比较每个系列中最常用的单词,我想我可以使用 ggplot 绘制它们,并且在一个轴上有“至少出现 x 次的系列 1 术语”和“至少出现 x 次的系列 2 术语” ' 另外一个。我希望我需要的是一个包含 3 列“术语”、“系列 x”、“系列 Y”的数据框。系列 x 和 y 具有该单词出现的次数。

我尝试了多种方法来做到这一点,但都失败了。我得到的最接近的是我可以读取语料库并创建一个数据框,其中所有术语在一列中,如下所示:

library("tm")

corpus <-Corpus(DirSource("series"))
corpus.p <-tm_map(corpus, removeWords, stopwords("english"))  #removes stopwords
corpus.p <-tm_map(corpus.p, stripWhitespace)  #removes stopwords
corpus.p <-tm_map(corpus.p, tolower)  
corpus.p <-tm_map(corpus.p, removeNumbers)
corpus.p <-tm_map(corpus.p, removePunctuation)
dtm <-DocumentTermMatrix(corpus.p)
docTermMatrix <- inspect(dtm)
termCountFrame <- data.frame(Term = colnames(docTermMatrix))

然后我知道我可以添加一列将这些单词相加:

termCountFrame$seriesX <- colSums(docTermMatrix)

但是当我只想要一个时,这会增加两个文档中的出现次数。

所以我的问题是:

1) 是否可以在单个文档上使用 colSums,如果没有,是否有另一种方法可以将 doctermmatrix 转换为每个文档的术语计数的数据框

2) 有谁知道我可以如何限制这一点,以便我获得每个文档中最常用的术语

【问题讨论】:

    标签: r dataframe text-mining corpus


    【解决方案1】:

    如果您的数据在文档术语矩阵中,您可以使用tm::findFreqTerms 来获取文档中最常用的术语。这是一个可重现的示例:

    require(tm)
    data(crude)
    dtm <- DocumentTermMatrix(crude)
    dtm
    A document-term matrix (20 documents, 1266 terms)
    
    Non-/sparse entries: 2255/23065
    Sparsity           : 91%
    Maximal term length: 17 
    Weighting          : term frequency (tf)
    
    # find most frequent terms in all 20 docs
    findFreqTerms(dtm, 2, 100)
    
    # find the doc names
    dtm$dimnames$Docs
     [1] "127" "144" "191" "194" "211" "236" "237" "242" "246" "248" "273" "349" "352" "353" "368" "489" "502"
    [18] "543" "704" "708"
    
    # do freq words on one doc
    findFreqTerms(dtm[dtm$dimnames$Docs == "127"], 2, 100)
     [1] "crude"     "cut"       "diamond"   "dlrs"      "for"       "its"       "oil"       "price"    
     [9] "prices"    "reduction" "said."     "that"      "the"       "today"     "weak"
    

    以下是查找 dtm 中每个文档的最常用词的方法,一次一个文档:

    # find freq words for each doc, one by one
    list_freqs <- lapply(dtm$dimnames$Docs, 
                  function(i) findFreqTerms(dtm[dtm$dimnames$Docs == i], 2, 100))
    
    
    list_freqs
    [[1]]
     [1] "crude"     "cut"       "diamond"   "dlrs"      "for"       "its"       "oil"       "price"    
     [9] "prices"    "reduction" "said."     "that"      "the"       "today"     "weak"     
    
    [[2]]
     [2] "\"opec"       "\"the"        "15.8"         "ability"      "above"        "address"      "agreement"   
     [8] "analysts"     "and"          "before"       "bpd"          "but"          "buyers"       "current"     
    [15] "demand"       "emergency"    "energy"       "for"          "has"          "have"         "higher"      
    [22] "hold"         "industry"     "its"          "keep"         "market"       "may"          "meet"        
    [29] "meeting"      "mizrahi"      "mln"          "must"         "next"         "not"          "now"         
    [36] "oil"          "opec"         "organization" "prices"       "problem"      "production"   "said"        
    [43] "said."        "set"          "that"         "the"          "their"        "they"         "this"        
    [50] "through"      "will"        
    
    [[3]]
    [3] "canada"   "canadian" "crude"    "for"      "oil"      "price"    "texaco"   "the"     
    
    [[4]]
    [4] "bbl."    "crude"   "dlrs"    "for"     "price"   "reduced" "texas"   "the"     "west"   
    
    [[5]]
     [5] "and"        "discounted" "estimates"  "for"        "mln"        "net"        "pct"        "present"   
     [9] "reserves"   "revenues"   "said"       "study"      "that"       "the"        "trust"      "value"     
    
    [[6]]
     [6] "ability"       "above"         "ali"           "and"           "are"           "barrel."      
     [7] "because"       "below"         "bpd"           "bpd."          "but"           "daily"        
    [13] "difficulties"  "dlrs"          "dollars"       "expected"      "for"           "had"          
    [19] "has"           "international" "its"           "kuwait"        "last"          "local"        
    [25] "march"         "markets"       "meeting"       "minister"      "mln"           "month"        
    [31] "official"      "oil"           "opec"          "opec\"s"       "prices"        "producing"    
    [37] "pumping"       "qatar,"        "quota"         "referring"     "said"          "said."        
    [43] "sheikh"        "such"          "than"          "that"          "the"           "their"        
    [49] "they"          "this"          "was"           "were"          "which"         "will"         
    
    [[7]]
     [7] "\"this"        "and"           "appears"       "are"           "areas"         "bank"         
     [7] "bankers"       "been"          "but"           "crossroads"    "crucial"       "economic"     
    [13] "economy"       "embassy"       "fall"          "for"           "general"       "government"   
    [19] "growth"        "has"           "have"          "indonesia\"s"  "indonesia,"    "international"
    [25] "its"           "last"          "measures"      "nearing"       "new"           "oil"          
    [31] "over"          "rate"          "reduced"       "report"        "say"           "says"         
    [37] "says."         "sector"        "since"         "the"           "u.s."          "was"          
    [43] "which"         "with"          "world"        
    
    [[8]]
     [8] "after"      "and"        "deposits"   "had"        "oil"        "opec"       "pct"        "quotes"    
     [9] "riyal"      "said"       "the"        "were"       "yesterday."
    
    [[9]]
     [9] "1985/86"     "1986/87"     "1987/88"     "abdul-aziz"  "about"       "and"         "been"       
     [8] "billion"     "budget"      "deficit"     "expenditure" "fiscal"      "for"         "government" 
    [15] "had"         "its"         "last"        "limit"       "oil"         "projected"   "public"     
    [22] "qatar,"      "revenue"     "riyals"      "riyals."     "said"        "sheikh"      "shortfall"  
    [29] "that"        "the"         "was"         "would"       "year"        "year's"     
    
    [[10]]
     [10] "15.8"      "about"     "above"     "accord"    "agency"    "ali"       "among"     "and"      
     [9] "arabia"    "are"       "dlrs"      "for"       "free"      "its"       "kuwait"    "market"   
    [17] "market,"   "minister," "mln"       "nazer"     "oil"       "opec"      "prices"    "producing"
    [25] "quoted"    "recent"    "said"      "said."     "saudi"     "sheikh"    "spa"       "stick"    
    [33] "that"      "the"       "they"      "under"     "was"       "which"     "with"     
    
    [[11]]
     [11] "1.2"        "and"        "appeared"   "arabia's"   "average"    "barrel."    "because"    "below"     
     [9] "bpd"        "but"        "corp"       "crude"      "december"   "dlrs"       "export"     "exports"   
    [17] "february"   "fell"       "for"        "four"       "from"       "gulf"       "january"    "january,"  
    [25] "last"       "mln"        "month"      "month,"     "neutral"    "official"   "oil"        "opec"      
    [33] "output"     "prices"     "production" "refinery"   "said"       "said."      "saudi"      "sell"      
    [41] "sources"    "than"       "the"        "they"       "throughput" "week"       "yanbu"      "zone"      
    
    [[12]]
     [12] "and"       "arab"      "crude"     "emirates"  "gulf"      "ministers" "official"  "oil"      
     [9] "states"    "the"       "wam"      
    
    [[13]]
     [13] "accord" "agency" "and"    "arabia" "its"    "nazer"  "oil"    "opec"   "prices" "saudi"  "the"   
    [12] "under" 
    
    [[14]]
     [14] "crude"   "daily"   "for"     "its"     "oil"     "opec"    "pumping" "that"    "the"     "was"    
    
    [[15]]
     [15] "after"   "closed"  "new"     "nuclear" "oil"     "plant"   "port"    "power"   "said"    "ship"   
    [11] "the"     "was"     "when"   
    
    [[16]]
     [16] "about"       "and"         "development" "exploration" "for"         "from"        "help"       
     [8] "its"         "mln"         "oil"         "one"         "present"     "prices"      "research"   
    [15] "reserve"     "said"        "strategic"   "the"         "u.s."        "with"        "would"      
    
    [[17]]
     [17] "about"       "and"         "benefits"    "development" "exploration" "for"         "from"       
     [8] "group"       "help"        "its"         "mln"         "oil"         "one"         "policy"     
    [15] "present"     "prices"      "protect"     "research"    "reserve"     "said"        "strategic"  
    [22] "study"       "such"        "the"         "u.s."        "with"        "would"      
    
    [[18]]
     [18] "1.50"    "company" "crude"   "dlrs"    "for"     "its"     "lowered" "oil"     "posted"  "prices" 
    [11] "said"    "said."   "the"     "union"   "west"   
    
    [[19]]
     [19] "according"    "and"          "april"        "before"       "can"          "change"       "efp"         
     [8] "energy"       "entering"     "exchange"     "for"          "futures"      "has"          "hold"        
    [15] "increase"     "into"         "mckiernan"    "new"          "not"          "nymex"        "oil"         
    [22] "one"          "position"     "prices"       "rule"         "said"         "spokeswoman." "that"        
    [29] "the"          "traders"      "transaction"  "when"         "will"        
    
    [[20]]
     [20] "1986,"        "1987"         "billion"      "cubic"        "fiscales"     "january"      "mln"         
     [8] "pct"          "petroliferos" "yacimientos"  
    

    如果你想在数据框中输出这个输出,你可以这样做:

    # from here http://stackoverflow.com/a/7196565/1036500
    L <- list_freqs
    cfun <- function(L) {
      pad.na <- function(x,len) {
        c(x,rep(NA,len-length(x)))
      }
      maxlen <- max(sapply(L,length))
      do.call(data.frame,lapply(L,pad.na,len=maxlen))
    }
    # make dataframe of words (but probably you want words as rownames and cells with counts?)
    tab_freqa <- cfun(L)
    

    但是,如果您想绘制“doc 1 高频项与 doc 2 高频项”,那么我们需要一种不同的方法...

    # convert dtm to matrix
    mat <- as.matrix(dtm)
    
    # make data frame similar to "3 columns 'Terms', 
    # 'Series x', 'Series Y'. With series x and y 
    # having the number of times that word occurs"
    cb <- data.frame(doc1 = mat['127',], doc2 = mat['144',])
    
    # keep only words that are in at least one doc
    cb <- cb[rowSums(cb)  > 0, ]
    
    # plot
    require(ggplot2)
    ggplot(cb, aes(doc1, doc2)) +
      geom_text(label = rownames(cb), 
               position=position_jitter())
    

    或者也许效率更高一点,我们可以将所有文档制作成一个大数据框并从中绘制图表:

    # this is the typical method to turn a 
    # dtm into a df...
    df <- as.data.frame(as.matrix(dtm))
    # and transpose for plotting
    df <- data.frame(t(df))
    # plot
    require(ggplot2)
    ggplot(df, aes(X127, X144)) +
      geom_text(label = rownames(df), 
               position=position_jitter())
    

    删除停用词后效果会更好,但这是一个很好的概念证明。这就是你所追求的吗?

    【讨论】:

    • 这太棒了。从这个答案中学到的东西不多
    • 很高兴它有帮助! R 非常擅长处理文字,这是一种专为数字工作而设计的语言。
    • 如果我只想要每个文档中 the 最常用术语的数据框,我该怎么做?特别是如果我事先不知道该术语的频率范围......
    • 如果您可以ask a question 了解这一点,您可能会得到几种不同的方法(使用我在此处回答中的可重现数据来设置您的问题)...
    • 对不起,老话题了。我尝试使用findFreqTerms(dtm[dtm$dimnames$Docs == "127], 2, 100),但它总是返回Error in x$j : $ operator is invalid for atomic vectors。可以拨打dtm$dimnames$Docs,但我想[] 不是很准确。
    【解决方案2】:

    对于问题 1)我使用 t(docTermMatrix) 创建了我想要的数据框,然后使用 as.data.frame

    dtm.frame <- as.data.frame(t(docTermMatrix))
    

    【讨论】:

    • 唯一的问题是你将数字作为列名,这对很多事情都有问题,尤其是使用ggplot进行绘图
    • 本的答案是正确的。
    • 我希望您不介意,但我已将答案写在这里。我给了你充分的信任:paddytherabbit.com/…
    猜你喜欢
    • 2021-03-16
    • 2021-05-10
    • 2014-05-18
    • 2020-10-28
    • 1970-01-01
    • 1970-01-01
    • 2017-04-17
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多