【问题标题】:How to use SQL Server Bulk Insert in Kedro Node?如何在 Kedro 节点中使用 SQL Server 批量插入?
【发布时间】:2021-07-13 13:16:14
【问题描述】:

我正在使用 Kedro 管理数据管道,在最后一步我有一个巨大的 csv 文件存储在 S3 存储桶中,我需要将其加载回 SQL Server。

我通常会使用bulk insert 来解决这个问题,但不太确定如何将其放入 kedro 模板中。这是catalog.yml 中配置的目标表和 S3 存储桶

flp_test:
  type: pandas.SQLTableDataSet
  credentials: dw_dev_credentials
  table_name: flp_tst
  load_args:
    schema: 'dwschema'
  save_args:
    schema: 'dwschema'
    if_exists: 'replace'

bulk_insert_input:
   type: pandas.CSVDataSet
   filepath: s3://your_bucket/data/02_intermediate/company/motorbikes.csv
   credentials: dev_s3


def insert_data(self, conn, csv_file_nm, db_table_nm):
    qry = "BULK INSERT " + db_table_nm + " FROM '" + csv_file_nm + "' WITH (FORMAT = 'CSV')"
    # Execute the query
    cursor = conn.cursor()
    success = cursor.execute(qry)
    conn.commit()
    cursor.close
  • 如何将csv_file_nm 指向我的bulk_insert_input S3 目录?
  • 是否有适当的方法可以间接访问dw_dev_credentials 进行插入?

【问题讨论】:

    标签: sql-server bulkinsert bulk-load kedro


    【解决方案1】:

    Kedro 的pandas.SQLTableDataSet.html 按原样使用pandas.to_sql 方法。要按原样使用它,您需要一个pandas.CSVDataSetnode,然后写入目标pandas.SQLDataTable 数据集以便将其写入SQL。如果你有 Spark,这将比 Pandas 更快。

    为了使用内置的BULK INSERT 查询,我认为您需要定义一个custom dataset

    【讨论】:

      猜你喜欢
      • 2011-06-17
      • 2010-09-24
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2021-06-05
      相关资源
      最近更新 更多