【问题标题】:How to import a json from a file on cloud storage to Bigquery如何将 json 从云存储上的文件导入 Bigquery
【发布时间】:2013-11-28 08:43:08
【问题描述】:

我正在尝试通过 api 将文件 (json.txt) 从云存储导入 Bigquery,但出现错误。当通过 web ui 完成此操作时,它可以正常工作并且没有错误(我什至设置了 maxBadRecords=0)。有人可以告诉我我在这里做错了什么吗?是代码错误,还是我需要更改 Bigquery 中的某些设置?

该文件是一个纯文本 utf-8 文件,内容如下:我保留了有关 bigquery 和 json 导入的文档。

{"person_id":225,"person_name":"John","object_id":1}
{"person_id":226,"person_name":"John","object_id":1}
{"person_id":227,"person_name":"John","object_id":null}
{"person_id":229,"person_name":"John","object_id":1}

在导入作业时会引发以下错误:“值无法转换为预期类型。”对于每一行。

    {
    "reason": "invalid",
    "location": "Line:15 / Field:1",
    "message": "Value cannot be converted to expected type."
   },
   {
    "reason": "invalid",
    "location": "Line:16 / Field:1",
    "message": "Value cannot be converted to expected type."
   },
   {
    "reason": "invalid",
    "location": "Line:17 / Field:1",
    "message": "Value cannot be converted to expected type."
   },
  {
    "reason": "invalid",
    "location": "Line:18 / Field:1",
    "message": "Value cannot be converted to expected type."
   },
   {
    "reason": "invalid",
    "message": "Too many errors encountered. Limit is: 10."
   }
  ]
 },
 "statistics": {
  "creationTime": "1384484132723",
  "startTime": "1384484142972",
  "endTime": "1384484182520",
  "load": {
   "inputFiles": "1",
   "inputFileBytes": "960",
   "outputRows": "0",
   "outputBytes": "0"
  }
 }
}

可以在此处访问该文件: http://www.sendspace.com/file/7q0o37

我的代码和架构如下:

def insert_and_import_table_in_dataset(tar_file, table, dataset=DATASET)
config= {
  'configuration'=> {
      'load'=> {
        'sourceUris'=> ["gs://test-bucket/#{tar_file}"],
        'schema'=> {
          'fields'=> [
            { 'name'=>'person_id', 'type'=>'INTEGER', 'mode'=> 'nullable'},
            { 'name'=>'person_name', 'type'=>'STRING', 'mode'=> 'nullable'},
            { 'name'=>'object_id',  'type'=>'INTEGER', 'mode'=> 'nullable'}
          ]
        },
        'destinationTable'=> {
          'projectId'=> @project_id.to_s,
          'datasetId'=> dataset,
          'tableId'=> table
        },
        'sourceFormat' => 'NEWLINE_DELIMITED_JSON',
        'createDisposition' => 'CREATE_IF_NEEDED',
        'maxBadRecords'=> 10,
      }
    },
  }

result = @client.execute(
  :api_method=> @bigquery.jobs.insert,
  :parameters=> {
     #'uploadType' => 'resumable',          
      :projectId=> @project_id.to_s,
      :datasetId=> dataset},
  :body_object=> config
)

# upload = result.resumable_upload
# @client.execute(upload) if upload.resumable?

puts result.response.body
json = JSON.parse(result.response.body)    
while true
  job_status = get_job_status(json['jobReference']['jobId'])
  if job_status['status']['state'] == 'DONE'
    puts "DONE"
    return true
  else
   puts job_status['status']['state']
   puts job_status 
   sleep 5
  end
end
end

有人可以告诉我我做错了什么吗?我在哪里修什么?

另外,在未来的某个时候,我希望使用压缩文件并从中导入——“tar.gz”是否可以,或者我只需要将其设为“.gz”吗?

提前感谢您的所有帮助。欣赏它。

【问题讨论】:

    标签: python ruby json import google-bigquery


    【解决方案1】:

    你被很多人(包括我)被击中的同样的事情被击中—— 您正在导入 json 文件但未指定导入格式,因此默认为 csv。

    如果您将 configuration.load.sourceFormat 设置为 NEWLINE_DELIMITED_JSON,您应该可以继续使用。

    我们有一个错误,使它更难执行,或者至少能够检测文件何时是错误类型,但我会提高优先级。

    【讨论】:

    • 谢谢乔丹。就是这样。它现在工作正常,导入到 bigquery 中。真的很感谢所有的帮助!祝你有个美好的一天。我将使用额外的配置详细信息更新问题,以便其他人可以使用它。
    猜你喜欢
    • 2018-05-25
    • 1970-01-01
    • 2021-04-23
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2014-01-07
    • 2012-10-27
    • 2016-06-07
    相关资源
    最近更新 更多