【问题标题】:R Programming using "dplyr" to select rows and return the index of the rows foundR编程使用“dplyr”选择行并返回找到的行的索引
【发布时间】:2015-03-16 07:00:55
【问题描述】:

设置/问题:

使用 dplyr - 我无法确定返回过滤行的行索引的最佳方式,而不是返回过滤行的内容。

问题:

我可以使用 dplyr::filter() 从数据框中提取行...问题是要提取过滤行的索引值并将其添加到满足搜索条件的索引条目列表中.

问题:

是否有一种简单的方法可以使用 dplyr 针对特定条件搜索数据框并返回找到的每一行的数字索引?下面的代码使用 r::which() 将索引行提取到列表中...

    requiredPackages <- c("dplyr")

    ipak <- function(pkg){
            new.pkg <- pkg[!(pkg %in% installed.packages()[, "Package"])]
            if (length(new.pkg))
                    install.packages(new.pkg, dependencies = TRUE)
            sapply(pkg, require, character.only = TRUE)
    }

    ipak(requiredPackages)

    if (!file.exists("./week3/data")) {
            dir.create("./week3/data")
    }

    # CSV Download
    if (!file.exists("./week3/data/americancommunitySurvey.csv")) {
            fileUrl <- "https://d396qusza40orc.cloudfront.net/getdata%2Fdata%2Fss06hid.csv?accessType=DOWNLOAD"
            download.file(fileUrl, destfile = "./week3/data/americancommunitySurvey.csv", method = "curl")
    }

    housingData <- tbl_df(read.csv("./week3/data/americancommunitySurvey.csv"
                                   , stringsAsFactors = TRUE))

 Now we have to extract the relevant data
#
# Create a logical vector that identifies the households on greater than 10
# acres who sold more than $10,000 worth of agriculture products. Assign that
# logical vector to the variable agricultureLogical. Apply the which() function
# like this to identify the rows of the data frame where the logical vector is
# TRUE. which(agricultureLogical) What are the first 3 values that result?
#
# ACR 1
# Lot size
# b .N/A (GQ/not a one-family house or mobile home)
# 1 .House on less than one acre
# 2 .House on one to less than ten acres
# 3 .House on ten or more acres                 ACR == 3
#
# AGS 1
# Sales of Agriculture Products
# b .N/A (less than 1 acre/GQ/vacant/
#                 .2 or more units in structure)
# 1 .None
# 2 .$ 1 - $ 999
# 3 .$ 1000 - $ 2499
# 4 .$ 2500 - $ 4999
# 5 .$ 5000 - $ 9999
# 6 .$10000+                                    AGS == 6
#
# Thus, we need to select only the results that have a ACR == 3 AND a AGS == 6
#
agricultureLogical <- which(housingData$ACR == 3 & housingData$AGS == 6)
agricultureLogical
# Now we can display the first three values of the resulting list
head(agricultureLogical[1:3])

上面的代码给了我想要的结果,但我想了解如何使用 dplyr 来做到这一点。它困扰着我......我可以使用 dplyr::filter() 如下提取行 - 我如何提取找到的每一行的索引????

agricultureLogical <- filter(housingData, ACR == 3 & housingData$AGS == 6)

R 设置

版本 _
平台 x86_64-apple-darwin13.4.0
拱 x86_64
操作系统 darwin13.4.0
系统 x86_64,darwin13.4.0
状态
专业 3
次要 1.2
2014 年
第 10 个月
第 31 天
svn 版本 66913
语言 R
version.string R 版本 3.1.2 (2014-10-31) 昵称南瓜头盔

dplyr 版本 0.3.0.2

设置 Mac OS X

型号名称:MacBook Pro 型号标识符:MacBookPro10,1 处理器名称:英特尔酷睿 i7 处理器速度:2.7 GHz 处理器数量:1 核心总数:4 L2 缓存(每核):256 KB 三级缓存:8 MB 内存:16 GB

【问题讨论】:

  • 我不确定是否有专门用于此的 dplyr 函数,但您可能可以使用 1:n() 的逻辑子集
  • Richard - 感谢您的帖子 - 您能提供更多信息吗?
  • 可能类似于do(mtcars, data.frame(x = which(.$cyl == 4)))mtcars 数据集为例并查找哪些行包含等于4 的柱面。您可以在调用后添加%&gt;% .$x 以获取向量而不是data.frame 如果您选择
  • 如果你有这么简单的矢量化解决方案,为什么要使用dplyr
  • David Arenburg - 我不清楚如何使用 dplyr 返回数据框中的行索引。我想知道该怎么做。我了解如何使用 R::which() 执行此操作 本质上,我正在使用 dplyr 针对一组标准搜索数据框,而不是返回我想要返回行索引的行数据。我看不出如何用 dplyr 做到这一点,因此这个问题。我知道如何根据代码使用 which 来做到这一点。

标签: r dplyr


【解决方案1】:

更简单的解决方案是使用 with 包裹 which

agricultureLogical &lt;- housingData %&gt;% with(which(ACR == 3 &amp; AGS == 6))

【讨论】:

    【解决方案2】:
    housingData %>%
      mutate(test = ACR == 3 & AGS == 6) %>%
      pull(test) %>%
      which
    

    【讨论】:

      【解决方案3】:

      由于不推荐使用 add_rownames(),您可以使用 rownames_to_column()。 Ista 的解决方案将采用以下格式:

      housingData %>%
      rownames_to_column() %>%
      filter(ACR == 3 & AGS == 6) %>%
      `[[`("rowname") %>%
      as.numeric() -> agricultureLogical
      

      【讨论】:

      • 我需要安装 tidyverse 包来获取 rownames_to_column() 函数。 install.packages("tidyverse"),然后是 library(tidyverse)。
      • 还有一个 rowid_to_column() 可能自动为数字,因此可能不需要 as.numeric。 (也在 tidyverse 中)
      【解决方案4】:

      建议的解决方案

      这是我正在尝试做的一个示例...这是一种解决方案,但我不喜欢它。感谢 Richard Scriven 提供指向 1:n()...

      手动向数据框添加索引列...

      我还没有弄清楚如何为匹配特定条件集的每一行返回单独的索引号...

      所以我使用 dplyr:mutate() 向示例数据框添加了一个索引列。然后我在数据框上使用 dplyr::filter() 来根据所需条件应用过滤器。这给我留下了我想玩的行列表...包括原始数据框的索引...我现在使用 dplyr::select() 仅提取索引列符合条件的每一行的原始数据框条目...

      h1 <- housingData
      # Add an index column to the dataframe h1...
      h1 <- mutate(h1, IDX = 1:n())
      # Filter the h1 dataframe using the criteria defined...
      h1 <- filter(h1, ACR == 3 & housingData$AGS == 6)
      # Extract the index 
      h1 <- select(h1, IDX)
      # Convert to an integer list...
      agricultureLogical <- as.integer(as.character(h1$IDX))
      head(agricultureLogical[1:3])
      

      以上对我来说是重复的努力,因为索引隐含在原始数据框中。因此,我的感觉是必须有一种方法来返回过滤器识别的项目的索引集......感谢答案:-)

      【讨论】:

        【解决方案5】:

        如果您使用 dplyr >= 0.4,您可以执行以下操作

        housingData %>%
          add_rownames() %>%
          filter(ACR == 3 & AGS == 6) %>%
          `[[`("rowname") %>%
          as.numeric() -> agricultureLogical
        

        虽然你为什么会认为这是一个改进

        agricultureLogical <- which(housingData$ACR == 3 & housingData$AGS == 6)
        

        逃离我。

        【讨论】:

        • Ista,我不认为它优越我只是想确定如何更优化地使用 dplyr 包。即如何返回索引偏移量而不是行数据。
        • dplyr 更可取,尤其是对于大型数据集,因为它经过优化以更快地运行
        • @Servet 通常声称一种方法比另一种方法更快,并得到了基准的支持。愿意提供一个吗?
        • @Servet 我看不出那篇博文与这个问题有什么关系。 dplyr 在某些方面可能更快,但这个问题是关于提取符合某些条件的行号。
        猜你喜欢
        • 1970-01-01
        • 2021-10-29
        • 2012-03-25
        • 1970-01-01
        • 2022-08-02
        • 1970-01-01
        • 2013-08-21
        • 2023-03-30
        • 1970-01-01
        相关资源
        最近更新 更多