【问题标题】:Use Python to parse multiple json items into a SQLite database使用 Python 将多个 json 项解析成 SQLite 数据库
【发布时间】:2017-06-05 23:23:14
【问题描述】:

我正在寻找一种非显式的方式来将 JSON 列表(即 [] 与 {} 内部的项目)解析到 sqlite 数据库中。

具体来说,我卡在我想说的地方

INSERT into MYTABLE (col 1, col2, ...) this datarow in this jsondata

我觉得应该有一种方法来抽象事物,使得上面的行几乎就是它所需要的。我的数据是一个 JSON 列表,包含多个 JSON 字典。每个字典都有十几个键:值对。没有嵌套。

{"object_id":3          ,"name":"rsid"          ,"column_id":1
,"system_type_id":127   ,"user_type_id":127     ,"max_length":8
,"precision":19         ,"scale":0              ,"collation_name":null
,"is_nullable":false    ,"is_ansi_padded":false ,"is_rowguidcol":false
,"is_identity":false    ,"is_computed":false    ,"is_filestream":false
,"is_replicated":false  ,"is_non_sql_subscribed":false
,"is_merge_published":false                     ,"is_dts_replicated":false
,"is_xml_document":false                        ,"xml_collection_id":0
,"default_object_id":0  ,"rule_object_id":0     ,"is_sparse":false
,"is_column_set":false}

json 列表[{k1:v1a, k2:v2a}, {k1:v1b,k2:v2b},...] 中的每个项目都将具有完全相同的 KEY 名称。我将这些作为我的 sqlite 数据库中列的名称。不过,VALUES 会有所不同。因此,通过将每个 KEY/COLUMN 的 VALUES 放入该项目的该行来填充表。

k1  | k2  | k3  | ... | km
v1a | v2a | v3a | ... | vma
v1b | v2b | v3b | ... | vmb
...
v1n | v2n | v3n | ... | vmn

在 SQL 中,插入语句不必按照与数据库列完全相同的顺序编写。这是因为您在 INSERT 声明中指定了要插入的列(以及顺序)。这对于 JSON 来说似乎是完美的,其中 JSON 列表中的每个 row/item 都包含其列名(键)。因此,我想要一个语句说“给定这行 JSON,通过将 JSON 键名与 SQL 表列名协调将其所有数据插入 SQL 表”。这就是我所说的非显式。

import json

r3 = some data file you read and close
r4 = json.loads(r3)
# let's dump this into SQLite

import sqlite3

the_database = sqlite3.connect("sys_col_database.sqlite")
the_cursor = the_database.cursor()

row_keys = r4[0].keys()
# all of the key are below for reference. 25 total keys.
'''
'is_merge_published',   'rule_object_id',   'system_type_id',
'is_xml_document',      'user_type_id',     'is_ansi_padded',
'column_id',            'is_column_set',    'scale',
'is_dts_replicated',    'object_id',        'xml_collection_id',
'max_length',           'collation_name',   'default_object_id',
'is_rowguidcol',        'precision',        'is_computed',
'is_sparse',            'is_filestream',    'name',
'is_nullable',          'is_identity',      'is_replicated',
'is_non_sql_subscribed'
'''

sys_col_table_statement = """create table sysColumns (
    is_merge_published text,
    rule_object_id integer,
    system_type_id integer,
    is_xml_document text,
    user_type_id integer,
    is_ansi_padded text,
    column_id integer,
    is_column_set text,
    scale integer,
    is_dts_replicated text,
    object_id integer,
    xml_collection_id integer,
    max_length integer,
    collation_name text,
    default_object_id integer,
    is_rowguidcol text,
    precision integer,
    is_computed text,
    is_sparse text,
    is_filestream text,
    name text,
    is_nullable text,
    is_identity text,
    is_replicated text,
    is_non_sql_subscribed text
    )
"""

the_cursor.execute(sys_col_table_statement)

insert_statement = """insert into sysColumns values (
{0},{1},{2},{3},{4},
{5},{6},{7},{8},{9},
{10},{11},{12},{13},{14},
{15},{16},{17},{18},{19},
{20},{21},{22},{23},{24})""".format(*r4[0].keys())

这是我卡住的地方。 insert_statementexecuted 的字符串构造。现在我需要执行它,但要从 r4 中的每个 JSON 项中为其提供正确的数据。我不知道怎么写。

【问题讨论】:

    标签: python json sqlite insert


    【解决方案1】:

    这里有两个选项。我还一步加载了 JSON 数据。

    #python 3.4.3
    
    import json
    import sqlite3
    
    with open("data.json",'r') as f:
        r4 = json.load(f)
    
    the_database = sqlite3.connect("sys_col_database.sqlite")
    the_cursor = the_database.cursor()
    
    row_keys = list(r4[0].keys())
    # all of the key are below for reference. 25 total keys.
    '''
    'is_merge_published',   'rule_object_id',   'system_type_id',
    'is_xml_document',      'user_type_id',     'is_ansi_padded',
    'column_id',            'is_column_set',    'scale',
    'is_dts_replicated',    'object_id',        'xml_collection_id',
    'max_length',           'collation_name',   'default_object_id',
    'is_rowguidcol',        'precision',        'is_computed',
    'is_sparse',            'is_filestream',    'name',
    'is_nullable',          'is_identity',      'is_replicated',
    'is_non_sql_subscribed'
    '''
    
    sys_col_table_statement = """create table sysColumns (
    is_merge_published text,
    rule_object_id integer,
    system_type_id integer,
    is_xml_document text,
    user_type_id integer,
    is_ansi_padded text,
    column_id integer,
    is_column_set text,
    scale integer,
    is_dts_replicated text,
    object_id integer,
    xml_collection_id integer,
    max_length integer,
    collation_name text,
    default_object_id integer,
    is_rowguidcol text,
    precision integer,
    is_computed text,
    is_sparse text,
    is_filestream text,
    name text,
    is_nullable text,
    is_identity text,
    is_replicated text,
    is_non_sql_subscribed text
    )
    """
    # method 1
    # create a string with the keys prepopulated
    prepared_insert_statement = """insert into sysColumns (\
    {0}, {1}, {2}, {3}, {4}, \
    {5}, {6}, {7}, {8}, {9}, \
    {10}, {11}, {12}, {13}, {14}, \
    {15}, {16}, {17}, {18}, {19}, \
    {20}, {21}, {22}, {23}, {24}) values (\
    '{{{0}}}', '{{{1}}}', '{{{2}}}', '{{{3}}}', '{{{4}}}', \
    '{{{5}}}', '{{{6}}}', '{{{7}}}', '{{{8}}}', '{{{9}}}', \
    '{{{10}}}', '{{{11}}}', '{{{12}}}', '{{{13}}}', '{{{14}}}', \
    '{{{15}}}', '{{{16}}}', '{{{17}}}', '{{{18}}}', '{{{19}}}', \
    '{{{20}}}', '{{{21}}}', '{{{22}}}', '{{{23}}}', '{{{24}}}'\
    )""".format(*r4[0].keys())
    
    # unpack the dictionary to extract the values
    insert_statement = prepared_insert_statement.format(**r4[0])
    
    # method 2
    # get the keys and values in one step
    r4_0_items = list(r4[0].items())
    insert_statement2 = """insert into sysColumns (\
    {0[0]}, {1[0]}, {2[0]}, {3[0]}, {4[0]}, \
    {5[0]}, {6[0]}, {7[0]}, {8[0]}, {9[0]}, \
    {10[0]}, {11[0]}, {12[0]}, {13[0]}, {14[0]}, \
    {15[0]}, {16[0]}, {17[0]}, {18[0]}, {19[0]}, \
    {20[0]}, {21[0]}, {22[0]}, {23[0]}, {24[0]}) values (\
    '{0[1]}', '{1[1]}', '{2[1]}', '{3[1]}', '{4[1]}', \
    '{5[1]}', '{6[1]}', '{7[1]}', '{8[1]}', '{9[1]}', \
    '{10[1]}', '{11[1]}', '{12[1]}', '{13[1]}', '{14[1]}', \
    '{15[1]}', '{16[1]}', '{17[1]}', '{18[1]}', '{19[1]}', \
    '{20[1]}', '{21[1]}', '{22[1]}', '{23[1]}', '{24[1]}'\
    )""".format(*r4_0_items)
    
    
    the_cursor.execute(sys_col_table_statement)
    #the_cursor.execute(insert_statement)
    the_cursor.execute(insert_statement2)
    the_database.commit()
    

    供参考:
    unpack dictionary
    string formatting examples

    【讨论】:

    • 我今天将对此进行测试。谢谢。我认为方法 2 对我目前的水平更有意义。
    • 再次感谢。我意识到我的 INSERT 语句格式不正确。你的榜样帮助我克服了障碍。此外,注意到更深层次的东西——我的一些“价值观”应该在 INSERT 语句中不被引用。最后,我个人不喜欢“With”语句,因为它缩进了它创建的变量,所以不清楚该变量现在是全局的。你用它是对的,只是让我头晕目眩。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2017-03-01
    • 2011-07-08
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多