【发布时间】:2018-03-20 05:47:57
【问题描述】:
我有一个数据框df 如下:
+---+--------+----+
| Id| Size| Amt|
+---+--------+----+
| a1| 1|55.0|
| a2| 2|48.0|
| a3| 3|28.0|
+---+--------+----+
这个数据框的架构是:
StructType([
StructField("Id", StringType(), True),
StructField("Size", IntegerType(), True),
StructField("Amt", FloatType(), True)
])
当我使用df.write.json("my_output_path") 时,json 文件看起来像:
{"Id":"a1", "Size":1, "Amt":55.0}
{"Id":"a2", "Size":2, "Amt":48.0}
{"Id":"a3", "Size":3, "Amt":28.0}
使用df,我想创建df1,使其具有一个新的数组列(Arr),其中包含现有列的键值对。
df1.write.json("my_new_output_path") 的输出文件应如下所示:
{"Id":"a1", "Size":1, "Amt":55.0, "Arr":[{"Id":"a1","Size":1,"Amt":55.0 }] }
{"Id":"a2", "Size":2, "Amt":48.0, "Arr":[{"Id":"a2","Size":2,"Amt":48.0 }] }
{"Id":"a3", "Size":3, "Amt":28.0, "Arr":[{"Id":"a3","Size":3,"Amt":28.0 }] }
我尝试了以下方法,但它给了我不同的输出:
df1 = df.select('Id', 'Size', 'Amt', array('Id','Size','Amt').alias("Arr"))
df1.write.json("my_new_output_path")
电流输出:
{"Id":"a1", "Size":1, "Amt":55.0, "Arr":["a1", 1 ,55.0] }
{"Id":"a2", "Size":2, "Amt":48.0, "Arr":["a2", 2 ,48.0] }
{"Id":"a3", "Size":3, "Amt":28.0, "Arr":["a3", 3 ,28.0] }
我怎样才能得到预期的输出?任何解决方案或指针将不胜感激。
【问题讨论】:
标签: python arrays apache-spark pyspark pyspark-sql