【发布时间】:2017-01-04 03:18:11
【问题描述】:
我有一个没有电影信息的数据集;我想从 OMDBapi 以 json 格式向它添加电影信息
我在 python 3.5 中编写这段代码来为我做这件事:
import urllib.request
import csv
import json
import datetime
from collections import defaultdict
from urllib import response
i=0
columns = defaultdict(list)
with open('C:\dataset\dataset.dat') as f:
reader = csv.DictReader(f)
for row in reader:
for (k,v) in row.items():
columns[k].append(v)
with open('C:\dataset\dataset.dat','r',encoding='utf-8') as csvinput:
with open('C:\dataset\dataset_edited.dat', 'w',encoding='utf-8') as csvoutput:
writer = csv.writer(csvoutput)
for row in csv.reader(csvinput):
if row[0] == "user_id":
writer.writerow(row+["movie_in_json_format"])
else:
movieJson=urllib.request.urlopen("http://www.omdbapi.com/?i=tt"+str(columns['item_id'][i])+"&y=&plot=short&r=json").read()
movieJson=movieJson.decode('utf-8')
writer.writerow(row+[movieJson])
i=i+1
json 格式以这种格式写入文件:
"{""Title"":""CitizenDog"",""Year"":""2004"",""Rated"":""N/A"",""Released"":""09 Mar 2006"",""Runtime"":""100 min"",""Genre"":""Comedy, Fantasy, Romance"",""Director"":""Wisit Sasanatieng"",""Writer"":""Koynuch (novel), Wisit Sasanatieng"",""Actors"":""Mahasamut Boonyaruk, Saengthong Gate-Uthong, Sawatwong Palakawong Na Autthaya, Nattha Wattanapaiboon"",""Plot"":""Pod is a man without a dream. He's a country bumpkin who comes to work at a tinned sardine factory in Bangkok. One day, Pod chops off his finger and packs it in the can, prompting him to go..."",""Language"":""Thai, English, Mandarin"",""Country"":""Thailand"",""Awards"":""2 wins & 1 nomination."",""Poster"":""http://ia.media-imdb.com/images/M/MV5BY2VlNDQwZTctMjBlNy00ZjYyLWEwYzAtNjA1YTNjNjVlMjU1XkEyXkFqcGdeQXVyMTIxMDUyOTI@._V1_SX300.jpg"",""Metascore"":""N/A"",""imdbRating"":""7.5"",""imdbVotes"":""1,544"",""imdbID"":""tt0444778"",""Type"":""movie"",""Response"":""True""}"
应该是这样的:
{"Title":"Citizen Dog","Year":"2004","Rated":"N/A","Released":"09 Mar 2006","Runtime":"100 min","Genre":"Comedy, Fantasy, Romance","Director":"Wisit Sasanatieng","Writer":"Koynuch (novel), Wisit Sasanatieng","Actors":"Mahasamut Boonyaruk, Saengthong Gate-Uthong, Sawatwong Palakawong Na Autthaya, Nattha Wattanapaiboon","Plot":"Pod is a man without a dream. He's a country bumpkin who comes to work at a tinned sardine factory in Bangkok. One day, Pod chops off his finger and packs it in the can, prompting him to go...","Language":"Thai, English, Mandarin","Country":"Thailand","Awards":"2 wins & 1 nomination.","Poster":"http://ia.media-imdb.com/images/M/MV5BY2VlNDQwZTctMjBlNy00ZjYyLWEwYzAtNjA1YTNjNjVlMjU1XkEyXkFqcGdeQXVyMTIxMDUyOTI@._V1_SX300.jpg","Metascore":"N/A","imdbRating":"7.5","imdbVotes":"1,544","imdbID":"tt0444778","Type":"movie","Response":"True"}
我该怎么做才能以正确的格式将此 json 写入文件?
~请注意,由于这个错误,“encoding='utf-8'”被添加到文件 i/o 中:
'charmap' codec can't encode character '\xf3' in position 3152: character maps to <undefined>
【问题讨论】:
-
我猜您使用的特定 CSV 方言需要用两个引号转义引号。想一想,CSV 解析器如何读取生成的 CSV 文件?
-
@roeland 我不知道 :(
-
尝试使用 CSV 解析器模块再次读取该文件,您应该得到原始字符串。作为替代方案,您可以将数据文件完全编写为 JSON 文件,而不是将 JSON 包装在 CSV 中,这样以后解析起来会更简单。
标签: python json unicode double-quotes