【问题标题】:KeyError: "None of ['index'] are in the columns"KeyError:“['index'] 均不在列中”
【发布时间】:2021-08-14 18:53:46
【问题描述】:

这是一个 json 文件:

{
    "id": "68af48116a252820a1e103727003d1087cb21a32",
    "article": [
        "by mark duell .",
        "published : .",
        "05:58 est , 10 september 2012 .",
        "| .",
        "updated : .",
        "07:38 est , 10 september 2012 .",
        "a pet owner starved her two dogs so badly that one was forced to eat part of his mother 's dead body in a desperate attempt to survive .",
        "the mother died a ` horrendous ' death and both were in a terrible state when found after two weeks of starvation earlier this year at the home of katrina plumridge , 31 , in grimsby , lincolnshire .",
        "the barely-alive dog was ` shockingly thin ' and the house had a ` nauseating and overpowering ' stench , grimsby magistrates court heard .",
        "warning : graphic content .",
        "horrendous : the male dog , scrappy -lrb- right -rrb- , was so badly emaciated that he ate the body of his mother ronnie -lrb- centre -rrb- to try to survive at the home of katrina plumridge in grimsby , lincolnshire .",
        "the suffering was so serious that the female staffordshire bull terrier , named ronnie , died of starvation , nigel burn , prosecuting , told the court last friday .",
        "suspended jail term : the dogs were in a terrible state when found after two weeks of starvation at the home of katrina plumridge , 31 -lrb- pictured -rrb- .",
        "the male dog , her son scrappy , was so badly emaciated that he ate her body to try to survive .",
    ],
    "abstract": [
        "neglect by katrina plumridge saw staffordshire bull terrier ronnie die .",
        "dog 's son scrappy was forced to eat her to survive at grimsby house .",
        "alarm raised by letting agent shocked by ` thinnest dog he 'd ever seen '",
    ]
}

我已经运行df = pd.read_json('100252.json'),但我得到了错误:ValueError: arrays must all be same length

然后我尝试了

with open('100252.json') as json_data: 
    data = json.load(json_data) 

pd.DataFrame.from_dict(data, orient='index').T.set_index('index')

但我收到了错误KeyError: "None of ['index'] are in the columns"

我该如何解决这个问题?我不知道我的错误在哪里。所以我需要你的帮助

编辑

来源:https://huggingface.co/docs/datasets/loading_datasets.html

从这个网站,我想做一些类似的事情

>>> from datasets import Dataset
>>> import pandas as pd
>>> df = pd.DataFrame({"a": [1, 2, 3]})
>>> dataset = Dataset.from_pandas(df)

我必须将 json 文件传输到数据帧中,然后使用数据集库从 pandas 获取数据集

【问题讨论】:

  • 你想达到什么目的?
  • @AlexanderVolkovsky 让我编辑我的代码来解释我想要什么。
  • @AlexanderVolkovsky 有没有更好的理解
  • 我不明白想要的输出。您是否正在尝试创建一个包含 ["id", "article", "abstract"] 列的数据框?如果是这样,您只需将数组替换为连接字符串
  • 请发布所需输出数据帧的 sn-p。文章和摘要似乎是按句子拆分的单个文档。你想在一行中加载每个句子,是否应该将所有句子连接到一个单元格中?目前还不清楚输出应该是什么样子。

标签: python pandas huggingface-datasets


【解决方案1】:

Dataset 输入必须是一个dict,并且具有相同大小的列表作为值。所以,

  1. 将句子连接成一个字符串并创建一个单元素列表。
from datasets import Dataset
with open('100252.json') as json_data: 
    data = json.load(json_data)

data['id'] = [data['id']]
data['article'] = ["\n".join(data['article'])]
data['abstract'] = ["\n".join(data['abstract'])]

Dataset.from_dict(data)

您的数据集将包含一行。

  1. 对齐列表。例如用空字符串填充
max_len = max([len(data[col]) for col in ['article', 'abstract'] ])

data['id'] = [data['id']] * max_len
data['article'] = data['article'] + [""] * (max_len - len(data['article'])) 
data['abstract'] = data['abstract'] + [""] * (max_len - len(data['abstract'])) 
Dataset.from_dict(data)

【讨论】:

  • 感谢您的回答!但是,我不必加入这句话。它必须保留一个句子列表。你要修改它吗
猜你喜欢
  • 2021-01-08
  • 1970-01-01
  • 2021-08-04
  • 2022-11-18
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2019-06-09
  • 2021-11-19
相关资源
最近更新 更多