【发布时间】:2021-05-13 13:07:57
【问题描述】:
在 Dataframe 中,我有一列包含以下数据
('Rated 3.0', "RATED\n \nWent there for a quick bite with friends.\nThe ambience had more of corporate feel. I would say it was unique.\nTried nachos, pasta churros and lasagne.\n\nNachos were pathetic.( Seriously don't order)\nPasta was okayish.\nLasagne was good.\nNutella churros were the best.\nOverall an okayish experience!\nPeace ??"), ('Rated 4.0', "RATED\n First of all, a big thanks to the staff of this Cafe. Very polite and courteous.\n\nI was there 15mins before their closing time. Without any discomfort or hesitation, the staff welcomed me with a warm smile and said they're still open, though they were preparing to close the cafe for the day.\n\nQuickly ordered the Thai green curry, which is served with rice. They got it for me within 10mins, hot and freshly made.\n\nIt was tasty with the taste of coconut milk. Not very spicy, it was mild spicy.\n\nI saw they had yummy looking dessert menu, should go there to try them out!\n\nA good spacious place to hang out for coffee, pastas, pizza or Thai food.")
我需要从每条记录中取出Rated 3.0 部分。这是一个 StringType 列。如何删除多余的数据并提取?
【问题讨论】:
-
你试过什么?例如,您是否尝试过使用regexp_extract 之类的方法?
标签: scala dataframe apache-spark apache-spark-sql dataset