【问题标题】:R column filtrationR柱过滤
【发布时间】:2021-03-08 15:25:17
【问题描述】:

我在工作中一直在处理一些数据,我正在尝试根据特定行过滤列,但到目前为止我一直没有成功。谁能帮帮我?让我解释一下我想要实现的目标。我有一个显示以下信息的数据集

person_id|custody_start|custody_end|contact_month|month_start |month_end |contact_date 13126321 |02/23/2020 |07/17/2020 |2 月 20 日 |03/01/2020 |02/28/2020|26/02/2020 13126321 |02/23/2020 |07/17/2020 |3 月 20 日 |03/01/2020 |03/31/2020|12/03/2020 13126321 |02/23/2020 |07/17/2020 |4 月 20 日 |04/01/2020 |04/30/2020|11/04/2020 13126321 |02/23/2020 |07/17/2020 |5 月 20 日 |05/01/2020 |05/31/2020|12/05/2020 13126321 |02/23/2020 |07/17/2020 |6 月 20 日 |06/01/2020 |06/30/2020|11/06/2020 13126321 |02/23/2020 |07/17/2020 |7 月 20 日 |07/01/2020 |07/31/2020|12/07/2020

数据多次显示相同的记录,但它是每月联系的记录和每月联系的日期。基本上我在这里想要实现的是过滤掉该人在整个日历月(从每个月的 1 日到 30 日或 31 日)没有被拘留的列,我们不是从任何日期开始看 30 天,但日历月。所以在 2 月 20 日和 7 月 20 日,这个人整个月都没有被拘留,因为你可以看到这个人在 2 月 23 日被拘留,并在 7 月 17 日离开拘留所。因此,在这种情况下,第一个月和最后一个月不计算在内。每个 person_id 都有多个这样的记录,所以我不能只删除每个孩子的第一列和最后一列。我只需要保存该人整个日历月的拘留记录

我的最终结果应该是这样的

person_id|custody_start|custody_end|contact_month|month_start |month_end |contact_date 26321 |02/23/2020 |07/17/2020 |3 月 20 日 |03/01/2020 |03/31/2020|12/03/2020 26321 |02/23/2020 |07/17/2020 |4 月 20 日 |04/01/2020 |04/30/2020|11/04/2020 26321 |02/23/2020 |07/17/2020 |5 月 20 日 |05/01/2020 |05/31/2020|12/05/2020 26321 |02/23/2020 |07/17/2020 |6 月 20 日 |06/01/2020 |06/30/2020|11/06/2020

如果能提供任何帮助,我将不胜感激。谢谢

【问题讨论】:

    标签: r dplyr


    【解决方案1】:

    这样的?

    dat = read.table(text='person_id|custody_start|custody_end|contact_month|month_start     |month_end |contact_date
        13126321 |02/23/2020   |07/17/2020 |February 20  |02/01/2020      |02/28/2020|26/02/2020    
        13126321 |02/23/2020   |07/17/2020 |March 20     |03/01/2020      |03/31/2020|12/03/2020    
        13126321 |02/23/2020   |07/17/2020 |April 20     |04/01/2020      |04/30/2020|11/04/2020  
        13126321 |02/23/2020   |07/17/2020 |May 20       |05/01/2020      |05/31/2020|12/05/2020 
        13126321 |02/23/2020   |07/17/2020 |June 20      |06/01/2020      |06/30/2020|11/06/2020  
        13126321 |02/23/2020   |07/17/2020 |July 20      |07/01/2020      |07/31/2020|12/07/2020',sep="|",header = TRUE)
    
    
    dat %>% 
      mutate_at(vars(contains("custody"),contains("month_")),
                function(x) as.character(x) %>% mdy(.)) %>% 
      mutate(contact_date = dmy(as.character(contact_date))) %>% 
      dplyr::filter(month_start >= custody_start & month_end <= custody_end)
    
      person_id custody_start custody_end contact_month month_start  month_end contact_date
    1  13126321    2020-02-23  2020-07-17 March 20       2020-03-01 2020-03-31   2020-03-12
    2  13126321    2020-02-23  2020-07-17 April 20       2020-04-01 2020-04-30   2020-04-11
    3  13126321    2020-02-23  2020-07-17 May 20         2020-05-01 2020-05-31   2020-05-12
    4  13126321    2020-02-23  2020-07-17 June 20        2020-06-01 2020-06-30   2020-06-11
    

    【讨论】:

    • 非常感谢您的回复。我试过了,虽然代码确实有效,但由于某种原因,它使所有contact_date数据都为“NA”
    • 现在可以试试吗?
    猜你喜欢
    • 1970-01-01
    • 2021-02-04
    • 2021-04-21
    • 2021-07-04
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-08-04
    • 2013-07-08
    相关资源
    最近更新 更多