【发布时间】:2021-04-22 17:31:03
【问题描述】:
我试图弄清楚如何使用rollapply 函数(来自Zoo 包)来查找数据集中最常见的字符串序列,但我还需要对某些变量进行分组(例如日期, 行等)
在我继续之前,值得注意的是,此查询基于我之前在此处发布的一个问题:How can I find most common sequences (of strings) in my data using Tableau?
那里提供的解决方案非常有效,但我现在想将它应用到一个不同的数据集,这会带来一些新的挑战!以下是我在这个新数据集中使用的数据示例:
structure(list(Title = c("Dragons' Den", "One Hot Summer", "Keeping Faith",
"Cuckoo", "Match of the Day", "Sportscene", "Sportscene", "The Irish League Show",
"Match of the Day", "EastEnders", "Dragons' Den", "Fake or Fortune?",
"Asian Provocateur", "In The Flesh", "Two Pints of Lager and a Packet of Crisps",
"Travels in Trumpland with Ed Balls", "Hidden", "Train Surfing Wars: A Matter of Life and Death",
"Bollywood: The World's Biggest Film Industry", "One Hot Summer",
"Asian Provocateur", "In The Flesh", "Two Pints of Lager and a Packet of Crisps",
"Travels in Trumpland with Ed Balls", "EastEnders", "Match of the Day",
"Dragons' Den", "The Next Step", "Doctor Who Series 11 Trailer",
"Doctor Who", "Doctor Who", "Doctor Who", "Picnic at Hanging Rock",
"Sylvia", "Keeping Faith", "Cardinal: Blackfly Season", "Picnic at Hanging Rock",
"Age Before Beauty", "One Hot Summer", "Stewart Lee's Comedy Vehicle",
"Asian Provocateur", "In The Flesh", "Two Pints of Lager and a Packet of Crisps",
"Travels in Trumpland with Ed Balls", "EastEnders", "Age Before Beauty",
"Holby City", "Who Do You Think You Are?", "Louis Theroux: Dark States",
"Louis Theroux: Dark States", "Louis Theroux", "Louis Theroux's Weird Weekends",
"Picnic at Hanging Rock", "Sylvia", "Keeping Faith", "Cardinal: Blackfly Season"
), Programme_Genre = c("Entertainment", "Documentary", "Drama",
"New SeriesComedy", "Sport", "Sport", "Sport", "Sport", "Sport",
"Drama", "Entertainment", "Documentary", "Comedy", "Drama", "Comedy",
"Documentary", "Crime Drama", "Documentary", "Documentary", "Documentary",
"Comedy", "Drama", "Comedy", "Documentary", "Drama", "Sport",
"Entertainment", "CBBC", "Sci-Fi", "Sci-Fi", "Sci-Fi", "Sci-Fi",
"Drama", "Film", "Drama", "Crime Drama", "On Now", "Drama", "Documentary",
"Comedy", "Comedy", "Drama", "Comedy", "Documentary", "Drama",
"Drama", "Drama", "History", "Documentary", "Documentary", "Documentary",
"Archive", "Drama", "Film", "Drama", "Crime Drama"), Programme_Category = c("Featured",
"Featured", "Featured", "Featured", "This Weekend's Football",
"This Weekend's Football", "This Weekend's Football", "This Weekend's Football",
"Most Popular", "Most Popular", "Most Popular", "Most Popular",
"Box Sets", "Box Sets", "Box Sets", "Box Sets", "Featured", "Featured",
"Featured", "Featured", "Box Sets", "Box Sets", "Box Sets", "Box Sets",
"Most Popular", "Most Popular", "Most Popular", "Most Popular",
"Doctor Who S1-S10", "Doctor Who S1-S10", "Doctor Who S1-S10",
"Doctor Who S1-S10", "Drama", "Drama", "Drama", "Drama", "Featured",
"Featured", "Featured", "Featured", "Box Sets", "Box Sets", "Box Sets",
"Box Sets", "Most Popular", "Most Popular", "Most Popular", "Most Popular",
"Louis Theroux", "Louis Theroux", "Louis Theroux", "Louis Theroux",
"Drama", "Drama", "Drama", "Drama"), date = c("13/08/2018", "13/08/2018",
"13/08/2018", "13/08/2018", "13/08/2018", "13/08/2018", "13/08/2018",
"13/08/2018", "13/08/2018", "13/08/2018", "13/08/2018", "13/08/2018",
"13/08/2018", "13/08/2018", "13/08/2018", "13/08/2018", "14/08/2018",
"14/08/2018", "14/08/2018", "14/08/2018", "14/08/2018", "14/08/2018",
"14/08/2018", "14/08/2018", "14/08/2018", "14/08/2018", "14/08/2018",
"14/08/2018", "14/08/2018", "14/08/2018", "14/08/2018", "14/08/2018",
"14/08/2018", "14/08/2018", "14/08/2018", "14/08/2018", "15/08/2018",
"15/08/2018", "15/08/2018", "15/08/2018", "15/08/2018", "15/08/2018",
"15/08/2018", "15/08/2018", "15/08/2018", "15/08/2018", "15/08/2018",
"15/08/2018", "15/08/2018", "15/08/2018", "15/08/2018", "15/08/2018",
"15/08/2018", "15/08/2018", "15/08/2018", "15/08/2018"), column = c("1",
"2", "3", "4", "1", "2", "3", "4", "1", "2", "3", "4", "1", "2",
"3", "4", "1", "2", "3", "4", "1", "2", "3", "4", "1", "2", "3",
"4", "1", "2", "3", "4", "1", "2", "3", "4", "1", "2", "3", "4",
"1", "2", "3", "4", "1", "2", "3", "4", "1", "2", "3", "4", "1",
"2", "3", "4"), row = c("1", "1", "1", "1", "2", "2", "2", "2",
"3", "3", "3", "3", "4", "4", "4", "4", "1", "1", "1", "1", "2",
"2", "2", "2", "3", "3", "3", "3", "4", "4", "4", "4", "5", "5",
"5", "5", "1", "1", "1", "1", "2", "2", "2", "2", "3", "3", "3",
"3", "4", "4", "4", "4", "5", "5", "5", "5")), row.names = c(NA,
-56L), class = "data.frame")
抱歉,但我不太确定共享数据的最佳做法。希望以上工作。它应该看起来像这样:
Title Programme_Genre Programme_Category date column row
1 Dragons Den Entertainment Featured 13/08/2018 1 1
2 One Hot Summer Documentary Featured 13/08/2018 2 1
3 Keeping Faith Drama Featured 13/08/2018 3 1
4 Cuckoo New Series Comedy Featured 13/08/2018 4 1
5 Match of the Day Sport This Weekends... 13/08/2018 1 2
6 Sportscene Sport This Weekends... 13/08/2018 2 2
我想要做的是使用rollapply 函数,类似于我在上一个问题中的建议(请参阅上面的链接),但仅用于查找出现在同一日期和特定列范围内的序列.例如,我想知道最常见的流派序列(“Programme_Genre”)是什么,但我只希望 rollapply 函数在每个日期的每一行的 1-4 列中执行此操作。我敢肯定我没有很好地解释这一点(如果你没有猜到,我不是来自数据科学背景)所以如果有必要我很乐意详细说明。提前致谢!
【问题讨论】:
-
您的示例数据很棒(它可能看起来很糟糕,但这是分享它的好方法)。你的预期输出是什么?当您想要做的事情看起来很复杂时,获取样本数据并手动确定输出应该是什么样子通常会有很大帮助。即使你不是对所有行都这样做,对 6 行左右也会有很大帮助。
-
你好,Japes。将其应用于第 1-4 列是什么意思?您是否期望 (1) 'Title'、(2) Program_Genre、(3) 节目类别和 (4) 日期的每个 Program_Genre 的结果?也就是说,每个节目类型有 4 个输出?还是要将 Program Genre、Title、Program_Category 和 Date 的值分组,并从 4 的每个组合中得到一个输出?
-
@r2evans 很难说上面的预期输出是什么,因为样本实际上很小,但如果我正在查看最常见的 4 序列,那么“运动”的组合“ sport" "sport" "sport" 将出现一次(在上面提供的数据中 - 它出现在数据框的第 5 到第 8 行中,并且是同一行 "2" 的一部分)。这有帮助吗?
-
@NicolásVelásquez 对不起 - 我的两个变量被称为“列”和“行”,这可能会造成一些混乱!让我再次尝试解释一下:我想搜索 4 种流派(或 3 或 2,但现在假设为 4)的最常见序列。但是,我只想在单个“行”(“列”变量中的数字 1-4)中查找此序列。是不是更清楚了?
-
Japes,如果你很难说出输出应该是什么,那么你怎么知道任何努力是否正确?获得输出是一回事,获得 正确 输出完全是另一回事......获得不正确(但好看)的输出充其量是误导,否则会默默地破坏数据。