【问题标题】:How can I find the most common sequences in my data using R?如何使用 R 在我的数据中找到最常见的序列?
【发布时间】:2021-04-22 17:31:03
【问题描述】:

我试图弄清楚如何使用rollapply 函数(来自Zoo 包)来查找数据集中最常见的字符串序列,但我还需要对某些变量进行分组(例如日期, 行等)

在我继续之前,值得注意的是,此查询基于我之前在此处发布的一个问题:How can I find most common sequences (of strings) in my data using Tableau?

那里提供的解决方案非常有效,但我现在想将它应用到一个不同的数据集,这会带来一些新的挑战!以下是我在这个新数据集中使用的数据示例:

structure(list(Title = c("Dragons' Den", "One Hot Summer", "Keeping Faith", 
"Cuckoo", "Match of the Day", "Sportscene", "Sportscene", "The Irish League Show", 
"Match of the Day", "EastEnders", "Dragons' Den", "Fake or Fortune?", 
"Asian Provocateur", "In The Flesh", "Two Pints of Lager and a Packet of Crisps", 
"Travels in Trumpland with Ed Balls", "Hidden", "Train Surfing Wars: A Matter of Life and Death", 
"Bollywood: The World's Biggest Film Industry", "One Hot Summer", 
"Asian Provocateur", "In The Flesh", "Two Pints of Lager and a Packet of Crisps", 
"Travels in Trumpland with Ed Balls", "EastEnders", "Match of the Day", 
"Dragons' Den", "The Next Step", "Doctor Who Series 11 Trailer", 
"Doctor Who", "Doctor Who", "Doctor Who", "Picnic at Hanging Rock", 
"Sylvia", "Keeping Faith", "Cardinal: Blackfly Season", "Picnic at Hanging Rock", 
"Age Before Beauty", "One Hot Summer", "Stewart Lee's Comedy Vehicle", 
"Asian Provocateur", "In The Flesh", "Two Pints of Lager and a Packet of Crisps", 
"Travels in Trumpland with Ed Balls", "EastEnders", "Age Before Beauty", 
"Holby City", "Who Do You Think You Are?", "Louis Theroux: Dark States", 
"Louis Theroux: Dark States", "Louis Theroux", "Louis Theroux's Weird Weekends", 
"Picnic at Hanging Rock", "Sylvia", "Keeping Faith", "Cardinal: Blackfly Season"
), Programme_Genre = c("Entertainment", "Documentary", "Drama", 
"New SeriesComedy", "Sport", "Sport", "Sport", "Sport", "Sport", 
"Drama", "Entertainment", "Documentary", "Comedy", "Drama", "Comedy", 
"Documentary", "Crime Drama", "Documentary", "Documentary", "Documentary", 
"Comedy", "Drama", "Comedy", "Documentary", "Drama", "Sport", 
"Entertainment", "CBBC", "Sci-Fi", "Sci-Fi", "Sci-Fi", "Sci-Fi", 
"Drama", "Film", "Drama", "Crime Drama", "On Now", "Drama", "Documentary", 
"Comedy", "Comedy", "Drama", "Comedy", "Documentary", "Drama", 
"Drama", "Drama", "History", "Documentary", "Documentary", "Documentary", 
"Archive", "Drama", "Film", "Drama", "Crime Drama"), Programme_Category = c("Featured", 
"Featured", "Featured", "Featured", "This Weekend's Football", 
"This Weekend's Football", "This Weekend's Football", "This Weekend's Football", 
"Most Popular", "Most Popular", "Most Popular", "Most Popular", 
"Box Sets", "Box Sets", "Box Sets", "Box Sets", "Featured", "Featured", 
"Featured", "Featured", "Box Sets", "Box Sets", "Box Sets", "Box Sets", 
"Most Popular", "Most Popular", "Most Popular", "Most Popular", 
"Doctor Who S1-S10", "Doctor Who S1-S10", "Doctor Who S1-S10", 
"Doctor Who S1-S10", "Drama", "Drama", "Drama", "Drama", "Featured", 
"Featured", "Featured", "Featured", "Box Sets", "Box Sets", "Box Sets", 
"Box Sets", "Most Popular", "Most Popular", "Most Popular", "Most Popular", 
"Louis Theroux", "Louis Theroux", "Louis Theroux", "Louis Theroux", 
"Drama", "Drama", "Drama", "Drama"), date = c("13/08/2018", "13/08/2018", 
"13/08/2018", "13/08/2018", "13/08/2018", "13/08/2018", "13/08/2018", 
"13/08/2018", "13/08/2018", "13/08/2018", "13/08/2018", "13/08/2018", 
"13/08/2018", "13/08/2018", "13/08/2018", "13/08/2018", "14/08/2018", 
"14/08/2018", "14/08/2018", "14/08/2018", "14/08/2018", "14/08/2018", 
"14/08/2018", "14/08/2018", "14/08/2018", "14/08/2018", "14/08/2018", 
"14/08/2018", "14/08/2018", "14/08/2018", "14/08/2018", "14/08/2018", 
"14/08/2018", "14/08/2018", "14/08/2018", "14/08/2018", "15/08/2018", 
"15/08/2018", "15/08/2018", "15/08/2018", "15/08/2018", "15/08/2018", 
"15/08/2018", "15/08/2018", "15/08/2018", "15/08/2018", "15/08/2018", 
"15/08/2018", "15/08/2018", "15/08/2018", "15/08/2018", "15/08/2018", 
"15/08/2018", "15/08/2018", "15/08/2018", "15/08/2018"), column = c("1", 
"2", "3", "4", "1", "2", "3", "4", "1", "2", "3", "4", "1", "2", 
"3", "4", "1", "2", "3", "4", "1", "2", "3", "4", "1", "2", "3", 
"4", "1", "2", "3", "4", "1", "2", "3", "4", "1", "2", "3", "4", 
"1", "2", "3", "4", "1", "2", "3", "4", "1", "2", "3", "4", "1", 
"2", "3", "4"), row = c("1", "1", "1", "1", "2", "2", "2", "2", 
"3", "3", "3", "3", "4", "4", "4", "4", "1", "1", "1", "1", "2", 
"2", "2", "2", "3", "3", "3", "3", "4", "4", "4", "4", "5", "5", 
"5", "5", "1", "1", "1", "1", "2", "2", "2", "2", "3", "3", "3", 
"3", "4", "4", "4", "4", "5", "5", "5", "5")), row.names = c(NA, 
-56L), class = "data.frame") 

抱歉,但我不太确定共享数据的最佳做法。希望以上工作。它应该看起来像这样:

   Title            Programme_Genre     Programme_Category  date         column row
1   Dragons Den     Entertainment       Featured            13/08/2018      1   1
2  One Hot Summer   Documentary         Featured            13/08/2018      2   1
3  Keeping Faith    Drama               Featured            13/08/2018      3   1
4  Cuckoo           New Series Comedy   Featured            13/08/2018      4   1
5  Match of the Day Sport               This Weekends...    13/08/2018      1   2
6  Sportscene       Sport               This Weekends...    13/08/2018      2   2

我想要做的是使用rollapply 函数,类似于我在上一个问题中的建议(请参阅上面的链接),但仅用于查找出现在同一日期和特定列范围内的序列.例如,我想知道最常见的流派序列(“Programme_Genre”)是什么,但我只希望 rollapply 函数在每个日期的每一行的 1-4 列中执行此操作。我敢肯定我没有很好地解释这一点(如果你没有猜到,我不是来自数据科学背景)所以如果有必要我很乐意详细说明。提前致谢!

【问题讨论】:

  • 您的示例数据很棒(它可能看起来很糟糕,但这是分享它的好方法)。你的预期输出是什么?当您想要做的事情看起来很复杂时,获取样本数据并手动确定输出应该是什么样子通常会有很大帮助。即使你不是对所有行都这样做,对 6 行左右也会有很大帮助。
  • 你好,Japes。将其应用于第 1-4 列是什么意思?您是否期望 (1) 'Title'、(2) Program_Genre、(3) 节目类别和 (4) 日期的每个 Program_Genre 的结果?也就是说,每个节目类型有 4 个输出?还是要将 Program Genre、Title、Program_Category 和 Date 的值分组,并从 4 的每个组合中得到一个输出?
  • @r2evans 很难说上面的预期输出是什么,因为样本实际上很小,但如果我正在查看最常见的 4 序列,那么“运动”的组合“ sport" "sport" "sport" 将出现一次(在上面提供的数据中 - 它出现在数据框的第 5 到第 8 行中,并且是同一行 "2" 的一部分)。这有帮助吗?
  • @NicolásVelásquez 对不起 - 我的两个变量被称为“列”和“行”,这可能会造成一些混乱!让我再次尝试解释一下:我想搜索 4 种流派(或 3 或 2,但现在假设为 4)的最常见序列。但是,我只想在单个“行”(“列”变量中的数字 1-4)中查找此序列。是不是更清楚了?
  • Japes,如果你很难说出输出应该是什么,那么你怎么知道任何努力是否正确?获得输出是一回事,获得 正确 输出完全是另一回事......获得不正确(但好看)的输出充其量是误导,否则会默默地破坏数据。

标签: r sequence zoo rollapply


【解决方案1】:

使用 tidyverse、zoo 和 lubridate,尝试:

library(tidyverse)
library(zoo)
library(lubridate)

df %>% 
  mutate(date = lubridate::dmy(date)) %>% # Optional. Properly parses date as Date class. Makes sorting easier.
  filter(column <= 4) %>% # Step 1. Exclude observations with `column` values above 4.
  group_split(row, date) %>% # Step 2. Splits the DF into smaller DFs representing row and date groups.
  # Step 3 (below). Loops the solution to the previous question, gets a DF, and assigns the date and row signals to each observation.
  map_df(.x = . ,
         .f = ~(rollapply(data = .x$Programme_Genre , 3, c) %>% 
                  as_tibble() %>% 
                  mutate(date = unique(.x$date), row = unique(.x$row)))) %>% 
  group_by_all() %>% 
  tally() %>% 
  arrange(date, row, n)

    # A tibble: 26 x 6
# Groups:   V1, V2, V3, date [26]
   V1            V2            V3               date       row       n
   <chr>         <chr>         <chr>            <date>     <chr> <int>
 1 Documentary   Drama         New SeriesComedy 2018-08-13 1         1
 2 Entertainment Documentary   Drama            2018-08-13 1         1
 3 Sport         Sport         Sport            2018-08-13 2         2
 4 Drama         Entertainment Documentary      2018-08-13 3         1
 5 Sport         Drama         Entertainment    2018-08-13 3         1
 6 Comedy        Drama         Comedy           2018-08-13 4         1
 7 Drama         Comedy        Documentary      2018-08-13 4         1
 8 Crime Drama   Documentary   Documentary      2018-08-14 1         1
 9 Documentary   Documentary   Documentary      2018-08-14 1         1
10 Comedy        Drama         Comedy           2018-08-14 2         1
# ... with 16 more rows

【讨论】:

  • 谢谢尼古拉斯。这看起来很有希望。直到明天我才有机会对此进行测试,但是一旦我尝试了就会报告!感谢您抽出宝贵时间回复此问题。
  • 好的,终于有机会试试这个,但我收到以下错误消息:“错误:无法在位置 1 排列类‘函数’的列”。有任何想法吗?我一直在寻找答案,但无法弄清楚!
【解决方案2】:

在这种情况下,我也建议您使用链接问题中建议的类似策略。

首先加载库

library(tidyverse)
library(runner)

n=3 的策略

n <- 3

data %>% 
  group_by(date) %>%
  mutate(l_seq = runner(x = Programme_Genre, 
                        k = n, 
                        function(x) ifelse(length(x) == n, list(x), list(rep(NA, n)))
  )
  ) %>%
  ungroup() %>%
  group_split(date) %>%
  map_df(., ~ map_df(.x$l_seq, ~setNames(.x, paste0('Col', seq_len(n)))) %>%
           mutate(date = .x$date) %>% 
           na.omit() %>%
           group_by_all() %>%
           summarise(m = n(), .groups = 'drop') %>%
           filter(m == max(m) & m > 1)
  )

# A tibble: 2 x 5
  Col1   Col2   Col3   date           m
  <chr>  <chr>  <chr>  <chr>      <int>
1 Sport  Sport  Sport  13/08/2018     3
2 Sci-Fi Sci-Fi Sci-Fi 14/08/2018     2

不用说m 是在该特定日期为您提供最大序列计数的列

如果n=4,上面的语法会给你以下结果

# A tibble: 1 x 6
  Col1  Col2  Col3  Col4  date           m
  <chr> <chr> <chr> <chr> <chr>      <int>
1 Sport Sport Sport Sport 13/08/2018     2

样本中长度5不存在长度大于1的序列

【讨论】:

  • 嗨@AnilGoyal。谢谢您的帮助。我已经尝试了代码,但出现以下错误:“错误:参数 1 必须有名称”。有什么想法吗?
  • 您是否创建了cols,如示例所示。还是我应该修改我的代码?
  • 嗨,是的,我根据您的代码创建了cols。如果您的慷慨,我不想利用,但如果您确实知道解决方案,那就太棒了:)
  • 感谢您的尝试 - 但仍然遇到同样的错误! :(
  • 老实说,我不记得我从哪里得到的。我想我设法想出了一个修复方法,但它并不完美。但是,很快我会回到这个特别的项目,所以当我知道更多时会更新你:)
猜你喜欢
  • 1970-01-01
  • 2021-10-09
  • 2021-09-22
  • 1970-01-01
  • 2021-02-14
  • 1970-01-01
  • 1970-01-01
  • 2015-04-22
  • 2012-11-12
相关资源
最近更新 更多