【发布时间】:2014-12-14 01:17:53
【问题描述】:
用read.csv.ffdf 读入一个大型数据集后,其中一列是时间。例如2014-10-18 00:01:02,表示该列中的 100 万行。该列是一个因素。如何将其转换为ff 支持的POSIXct?只需使用as.POSIXct() 只需将值转换为NA
或者当我在开始读取数据集时,我可以指定该列为POSIXct吗?
我的目标是获取月份和日期(甚至小时)。因此,除了转换为 POSIXct 之外,我对解决方案持开放态度。
例如,我们有 9 x 2 表,
test <- read.csv.ffdf(file="test.csv", header=T, first.rows=-1)
两列分别是ID(数值类)和时间(因子类)
这是输入
structure(list(virtual = structure(list(VirtualVmode = c("integer",
"integer"), AsIs = c(FALSE, FALSE), VirtualIsMatrix = c(FALSE,
FALSE), PhysicalIsMatrix = c(FALSE, FALSE), PhysicalElementNo = 1:2,
PhysicalFirstCol = c(1L, 1L), PhysicalLastCol = c(1L, 1L)), .Names = c("VirtualVmode",
"AsIs", "VirtualIsMatrix", "PhysicalIsMatrix", "PhysicalElementNo",
"PhysicalFirstCol", "PhysicalLastCol"), row.names = c("ID", "time"
), class = "data.frame", Dim = c(9L, 2L), Dimorder = 1:2), physical = structure(list(
ID = structure(list(), physical = <pointer: 0x000000000821ab20>, virtual = structure(list(), Length = 9L, Symmetric = FALSE), class = c("ff_vector",
"ff")), time = structure(list(), physical = <pointer: 0x000000000821abb0>, virtual = structure(list(), Length = 9L, Symmetric = FALSE, Levels = c("10/17/2003 0:01",
"12/5/1999 0:02", "2/1/2000 0:01", "3/23/1998 0:01", "3/24/2013 0:00",
"5/29/2004 0:00", "5/9/1985 0:01", "6/14/2010 0:01", "6/25/2008 0:02"
), ramclass = "factor"), class = c("ff_vector", "ff"))), .Names = c("ID",
"time")), row.names = NULL), .Names = c("virtual", "physical",
"row.names"), class = "ffdf")
【问题讨论】:
-
请提供一小部分数据样本,输出为
dput(head(data)) -
对于因子转换,您需要先在列上执行
as.character。然后你可以将它传递给as.POSIXct。 -
好像应用as.character后,列还是因子类。我认为问题在于 ff 不支持字符......也许我弄错了......
-
K 忘记
dput因为指针我们不能使用它。我的错