【发布时间】:2018-11-15 13:48:37
【问题描述】:
我必须平滑一个大的时间序列,我正在使用“raster”包中的movingFun 函数。我根据以前的帖子测试了几个选项(请参阅下面的选项)。前 2 个工作,但在使用真实数据时非常慢(所有澳大利亚的所有 MOD13Q1 时间序列)。所以我尝试了选项 3,但失败了。如果有人可以帮助指出该功能中的问题,我将不胜感激。我可以访问内存,我正在使用具有 700GB 内存的 RStudio 服务器,但我不确定完成这项工作的最佳方法是什么。提前致谢。
a) 使用movingFun 和叠加
library(raster)
r <- raster(ncol=10, nrow=10)
r[] <- runif(ncell(r))
s <- brick(r,r*r,r+2,r^5,r*3,r*5)
ptm <- proc.time()
v <- overlay(s, fun=function(x) movingFun(x, fun=mean, n=3, na.rm=TRUE, circular=TRUE)) #works
proc.time() - ptm
user system elapsed
0.140 0.016 0.982
b) 创建一个函数并使用 clusterR。我认为这会比 (a) 更快。
fun1 = function(x) {overlay(x, fun=function(x) movingFun(x, fun=mean, n=6, na.rm=TRUE, circular=TRUE))}
beginCluster(4)
ptm <- proc.time()
v = clusterR(s, fun1, progress = "text")
proc.time() - ptm
endCluster()
user system elapsed
0.124 0.012 4.069
c) 我找到了由 Robert J. Hijmans 编写的this document,我尝试(但失败了)编写了一个小插曲中描述的函数。我无法完全遵循该功能中的所有步骤,这就是失败的原因。
smooth.fun <- function(x, filename='', smooth_n ='',...) { #x could be a stack or list of rasters
out <- brick(x)
big <- ! canProcessInMemory(out, 3)
filename <- trim(filename)
if (big & filename == '') {
filename <- rasterTmpFile()
}
if (filename != '') {
out <- writeStart(out, filename, ...)
todisk <- TRUE
} else {
vv <- matrix(ncol=nrow(out), nrow=ncol(out))
todisk <- FALSE
}
bs <- blockSize(out)
pb <- pbCreate(bs$n)
if (todisk) {
for (i in 1:bs$n) {
v <- getValues(out, row=bs$row[i], nrows=bs$nrows[i] )
v <- movingFun(v, fun=mean, n=smooth_n, na.rm=TRUE, circular=TRUE)
out <- writeValues(out, v, bs$row[i])
pbStep(pb, i)
}
out <- writeStop(out)
} else {
for (i in 1:bs$n) {
v <- getValues(out, row=bs$row[i], nrows=bs$nrows[i] )
v <- movingFun(v, fun=mean, n=smooth_n, na.rm=TRUE, circular=TRUE)
cols <- bs$row[i]:(bs$row[i]+bs$nrows[i]-1)
vv[,cols] <- matrix(v, nrow=out@ncols)
pbStep(pb, i)
}
out <- setValues(out, as.vector(vv))
}
pbClose(pb)
return(out)
}
s <- smooth.fun(s, filename='test.tif', smooth_n = 6, format='GTiff', overwrite=TRUE)
Error in .local(.Object, ...) :
`/path-to-dir/test.tif' does not exist in the file system,
and is not recognised as a supported dataset name.
【问题讨论】:
-
您是否尝试过在
rasterOptions中增加maxmemory和chunksize? -
@Geo-sp 是的,我都增加了,虽然我说服务器有 700GB 我只使用了一小部分 (
maxmemory=10e10),我得到了canProcessInMemory=TRUE;还有其他人使用相同的设施。我将chunksize增加到 2e+08。我尝试了 ff 包,但无法创建如此大的矩阵。我认为要走的路是创建一个文件并将结果写在上面(使用ff)。但我不确定如何实现这一目标。选项 (c) 是否旨在实现这一目标? -
我不确定。我建议使用
earth engine进行此类处理。 -
感谢@Geo-sp,这绝对是另一种选择。昨天一位同事想出了一个使用很少内存的合理解决方案。我正在尝试实现它,一旦我得到它的工作,我会在这里发布。
标签: r time-series raster smoothing r-raster