【问题标题】:Performance problems when reading from a socket从套接字读取时的性能问题
【发布时间】:2020-04-14 16:48:29
【问题描述】:

我编写的用于与 BaseX XML 数据库通信的包(请参阅https://CRAN.R-project.org/package=RBaseX)似乎很稳定,上个月我没有看到任何错误。唯一的问题是性能。 执行此查询:

let $words := for $text in collection('IncidentRemarks/IncidentRemarks.csv')/csv/record/INC_RM
  return ft:tokenize($text)
return $words

大约需要 48 毫秒。从套接字读取生成的 350000 字节需要 > 100 秒。

我使用这个函数从套接字读取:

str_receive = function(input, output, bin = FALSE) {
  if (missing(input)) input   <- self$get_socket()
  if (missing(output)) output <- raw(0)
  while ((rd <- readBin(input, what = "raw", n =1)) > 0) {
    if (rd == 0xff) rd <- readBin(input, what = "raw", n =1)
    output <- c(output, rd)
  }
  # The 'Full'-method embeds a \0 in the output
  if (!bin) ret <- strip_CR_NUL(output) %>% rawToChar()
  else ret <- output
  return(ret)
  }

该软件包使用 R6。由于我还没有找到分析 R6 方法的好方法,所以我使用 browser() 进行调试。它表明while循环导致延迟。 (我猜想尤其是output &lt;- c(output, rd) 是主要问题)。
加快从套接字读取的最佳方法是什么?

这个包的最新源代码可以在https://github.com/BenEngbers/RBaseX找到


PS。请不要告诉我必须使用“C”或“CPP”。我一直成功地避免使用这些语言 ;-)

4 月 6 日,

我隔离了从套接字读取的代码:

socket_reader <- function(socket_in) {
  string_read <- raw(0)
  while ((rd <- readBin(socket_in, what = "raw", n =1)) > 0) {
    if (rd == 0xff) rd <- readBin(socket_in, what = "raw", n =1)
    string_read <- c(string_read, rd)
  }
  return(string_read)
}

并将该代码替换为:

socket_reader <- function(socket_in) {
  string_read <- raw(0)
  CONT <- TRUE
  Buf_Size <- 4096
  while (CONT) {
    read_buffer <- readBin(socket_in, what = "raw", n = Buf_Size)
    if (length(read_buffer) < Buf_Size) CONT <- FALSE
    string_read <- c(string_read, read_buffer)
  }
  string_read <- strip_FF(string_read)
  string_read <- string_read[-(length(string_read))] %>% as.raw()
  return(string_read)
}

这段代码应该快很多。 strip_FF() 函数从 string_read 中删除 (the) \0xFF 字节,因此两个版本应该给出相同的结果。
然而,在几个读取操作之间,我必须从连接中读取一个 (1) 状态字节。 \0x00 表示成功,\0x01 表示失败。

我的新版本无法读取该状态字节。

如何从连接中读取 1 个字节并移动连接中的位置?

【问题讨论】:

    标签: r sockets


    【解决方案1】:

    我的假设是函数“readLines”会比“readBin”更快,这就是我尝试应用该函数的原因。顺便说一句,根据这篇文章Reading a complete file with R,这个假设是不合理的。我现在了解到 readLines 只能用于读取字符串,不能用于读取二进制数据。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2015-07-17
      • 1970-01-01
      • 1970-01-01
      • 2012-10-04
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2012-04-10
      相关资源
      最近更新 更多