【问题标题】:How do I find most frequent words by each observation in R?如何通过 R 中的每个观察找到最常见的单词?
【发布时间】:2022-05-08 04:27:51
【问题描述】:

我对 NLP 很陌生。请不要严格评判我。

我有一个关于客户反馈的非常大的数据框,我的目标是分析反馈。我在反馈中标记了单词,删除了停用词(SMART)。现在,我需要接收最常用词和最不常用词的表格。

代码如下:

library(tokenizers)
library(stopwords)
words_as_tokens <- 
     tokenize_words(dat$description, 
                    stopwords = stopwords(language = "en", source = "smart"))

数据框如下所示:有很多反馈(可变“描述”)和提供反馈的客户(每个客户不是唯一的,可以重复)。我想收到一个包含 3 列的表格:a)客户名称 b)单词 c)它的频率。这个“排名”应该是按降序排列的。

【问题讨论】:

    标签: r nlp text-mining


    【解决方案1】:

    试试这个

    library(tokenizers)
    library(stopwords)
    library(tidyverse)
    
    # count freq of words
    words_as_tokens <- setNames(lapply(sapply(dat$description, 
                                     tokenize_words, 
                                     stopwords = stopwords(language = "en", source = "smart")), 
                              function(x) as.data.frame(sort(table(x), TRUE), stringsAsFactors = F)), dat$name)
    
    # tidyverse's job
    df <- words_as_tokens %>%
      bind_rows(, .id = "name") %>%
      rename(word = x)
    
    # output
    df
    
    #    name          word Freq
    # 1  John    experience    2
    # 2  John          word    2
    # 3  John    absolutely    1
    # 4  John        action    1
    # 5  John        amazon    1
    # 6  John     amazon.ae    1
    # 7  John     answering    1
    # ....
    # 42 Alex         break    2
    # 43 Alex          nice    2
    # 44 Alex         times    2
    # 45 Alex             8    1
    # 46 Alex        accent    1
    # 47 Alex        africa    1
    # 48 Alex        agents    1
    # ....
    

    数据

    dat <- data.frame(name = c("John", "Alex"),
                      description = c("Unprecedented. The perfect word to describe Amazon. In every positive sense of that word! All because of one man - Jeff Bezos. What an entrepreneur! What a vision! This is from personal experience. Let me explain. I had given up all hope, after a horrible experience with Amazon.ae (formerly Souq.com) - due to a Herculean effort to get an order cancelled and the subsequent refund issued. I have never faced such a feedback-resistant team in my life! They were robotically answering my calls and sending me monotonous, unhelpful emails, followed by absolutely zero action!",
                                     "Not only does Amazon have great products but their Customer Service for the most part is wonderful. Although most times you are outsourced to a different country, I personally have found that when I call it's either South Africa or Philippines and they speak so well, understand me and my NY accent and are quite nice. Let’s face it. Most times you are calling CS with a problem or issue. These agents have to listen to 8 hours of complaints so they themselves need a break. No matter how annoyed I am I try to be on my best behavior and as nice as can be because they too need a break with how nasty we as a society can be."), stringsAsFactors = F)
    
    

    【讨论】:

    • 抱歉,“x”是什么?什么功能?
    • @k1rgas 这是您通常在函数中使用的 lambda 函数的参数。例如,如果您想知道 data.frame 的每一列有多少缺失值,您可以这样做:apply(df, 2, function(x) sum(is.na(x))。在这种情况下,x 是 data.frame df 的每一列。
    【解决方案2】:

    您可以尝试使用quanteda 以及如下:

    library(quanteda)
    library(quanteda.textstats)
    # define a corpus object to store your initial documents
    mycorpus = corpus(dat$description)
    # convert the corpus to a Document-Feature Matrix
    mydfm = dfm( mycorpus, 
                 tolower = TRUE, 
                 remove = stopwords(),  # this removes English stopwords
                 remove_punct = TRUE,   # this removes punctuation
                 remove_numbers = TRUE, # this removes digits
                 remove_symbol = TRUE,  # this removes symbols 
                 remove_url = TRUE )    # this removes urls
    
    # calculate word frequencies and return a data.frame
    word_frequencies = textstat_frequency( mydfm )
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2016-11-09
      • 2018-09-10
      • 1970-01-01
      • 2021-10-09
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2014-12-14
      相关资源
      最近更新 更多