一般来说,R 并不适用于就地文件编辑,而且我知道没有(当前可用的)工具在任何情况下都支持它。即使像sed 这样的unixy 工具也可以进行快速编辑,但在技术上仍然没有“就地”进行(即使它隐藏了它是如何做到的)。 (可能有一些这样做,但可能不是您想要的易于访问。)
有一个值得注意的例外,一种为就地编辑(嗯,交互)而设计的文件格式。它包括重要的就地添加、过滤、替换和删除运算符。在大多数情况下,它通常会这样做而无需增加文件大小。我是 SQLite。
例如,
library(DBI)
# library(RSQLite) # don't need to load it, just need to have it available
fname <- "./iris.sqlite3"
con <- dbConnect(RSQLite::SQLite(), fname)
file.info(fname)$size
# [1] 0
dbWriteTable(con, "iris", iris)
# [1] TRUE
file.info(fname)$size
# [1] 16384
dbGetQuery(con, "select * from iris where [Sepal.Length]=4.7 and [Sepal.Width]=3.2 and [Petal.Length]=1.6")
# Sepal.Length Sepal.Width Petal.Length Petal.Width Species
# 1 4.7 3.2 1.6 0.2 setosa
file.info(fname)$size
# [1] 16384
dbExecute(con, "update iris set [Species]='virginica' where [Sepal.Length]=4.7 and [Sepal.Width]=3.2 and [Petal.Length]=1.6")
# [1] 1
dbGetQuery(con, "select * from iris where [Sepal.Length]=4.7 and [Sepal.Width]=3.2 and [Petal.Length]=1.6")
# Sepal.Length Sepal.Width Petal.Length Petal.Width Species
# 1 4.7 3.2 1.6 0.2 virginica
dbDisconnect(con)
file.info(fname)$size
# [1] 16384
优点
缺点
- 文件大小会有“开销”。值得注意的是
iris 占用不到 7K 的内存(请参阅object.size(iris)),但文件大小从 16K 开始。对于更大的数据,差距ratio(文件大小与实际数据)会缩小。 (我对 ggplot2::diamonds 做了同样的事情;对象为 3456376 字节,文件大小为 3780608,不到 10% 大。)
- 当 SQLite 认为有必要时,文件大小会增加。这是基于 R 和此问题/答案范围之外的许多因素。
- 如果删除大量数据,文件大小不会立即减小以适应...请参阅change sqlite file size after "DELETE FROM table"(提示:
vacuum)
- 有许多工具可以轻松/立即从此文件格式导入数据,但 Excel 和 Access 尤其缺席。使用SQLite-ODBC 是可行的,但需要一点肘部润滑脂才能做到。 (我可以接受,但并非所有用户都可以,而且一些企业网络使这一步变得困难或明确禁止。)
SQLite 文件作为 CSV
如果要全部导入,可以在导入时将其视为文件:
con <- dbConnect(RSQLite::SQLite(), fname)
iris2 <- dbGetQuery(con, "select * from iris")
dbDisconnect(con)
相比
iris2 <- read.csv("iris.csv", stringsAsFactors = FALSE)
如果你想变得花哨:
import_sqlite <- function(fname, tablename = NA) {
if (length(tablename) > 1L) {
warning("the condition has length > 1 and only the first element will be used")
tablename <- tablename[[1L]]
}
con <- DBI::dbConnect(RSQLite::SQLite(), fname)
on.exit(DBI::dbDisconnect(con), add = TRUE)
available_tables <- DBI::dbListTables(con)
if (length(available_tables) == 0L) {
stop("no tables found")
} else if (is.na(tablename)) {
if (length(available_tables) == 1L) {
tablename <- available_tables
}
}
if (tablename %in% available_tables) {
tablename <- DBI::dbQuoteIdentifier(con, tablename)
qry <- sprintf("select * from %s", tablename)
out <- tryCatch(list(data = DBI::dbGetQuery(con, DBI::SQL(qry)),
err = NULL),
error = function(e) list(data = NULL, err = e))
if (! is.null(out$err)) {
stop("[sqlite error] ", out$err$message)
} else {
return(out$data)
}
} else {
stop(sprintf("table %s not found", DBI::dbQuoteIdentifier(con, tablename)))
}
}
head(import_sqlite("iris.sqlite3"))
# Sepal.Length Sepal.Width Petal.Length Petal.Width Species
# 1 5.1 3.5 1.4 0.2 setosa
# 2 4.9 3.0 1.4 0.2 setosa
# 3 4.7 3.2 1.3 0.2 setosa
# 4 4.6 3.1 1.5 0.2 setosa
# 5 5.0 3.6 1.4 0.2 setosa
# 6 5.4 3.9 1.7 0.4 setosa
(除了概念验证之外,我没有提供该功能,您可以像与 CSV 一样与单个文件进行交互。其中有一些保护措施,但实际上只是对为了这个问题。)