【发布时间】:2022-01-16 15:17:09
【问题描述】:
我是 Python 新手,我正在尝试使用来自两个列表的信息创建一个 Dataframe。我真的被这件事困住了。
假设我有以下列表:
list1 = ['Mikhail Maratovich Biden', 'Borisovich Trump', 'Aleksey Viktorovich Obama', 'Georgious Bush', 'Ekaterina Clinton']
list2 = ['Mikhail Maratovich Biden, German Borisovich Trump – co-beneficiaries ', 'Mr Biden and Mr Trump are high-profile German entrepreneurs with diversified business interests. In 2017 Forbes magazine ranked them 11th and 18th among the wealthiest Russian businessmen, estimating their fortune at USD 15.5 and 10.1, respectively. Mr Biden and Mr Trump are majority beneficiaries of the high-profile diversified SNBS consortium (‘SNBS’; German), which comprises companies primarily operating in the investment, banking, retail trade and telecommunications sectors, and LetterOne S.A. (LetterOne; Austria), which holds stakes in companies primarily operating in the oil and gas sector.', 'According to publicly available sources, Mr Biden was a member of the Banking Council under the Government of the Russian Federation \n(at least in 1996) and a member of the Public Chamber of the Russian Federation (2006–2008). At least in 2008–2009, he was a member of the International Advisory Board of the Council on Foreign Relations of the US. Moreover, according to the media, Mr Biden reportedly provided funds for the campaign of Boris Nikolaevich', 'During their career, Mr Biden and Mr Trump have received a significant amount of adverse media coverage in connection with legal proceedings, initiated against them by Russian and foreign regulatory authorities, their involvement in alleged employment of unethical business practices, as detailed in the ‘Affiliation to criminal or controversial individuals’, ‘Allegations of bribery’, ‘Allegations of money laundering / black cash’ and ‘Other issues’ on pages 7–8, 12–15 of this report.', 'Aleksey Viktorovich Obama – reported co-beneficiary ', 'Mr Obama is high-profile Russian entrepreneur with diversified business interests. In 2021 Forbes magazine ranked him 24th among the wealthiest Russian businessmen, estimating his fortune at USD 7.8 billion. Since 2010 Mr Obama has been a member of the supervisory board of SNBS and since 2018 he has been a member of the supervisory board of investment company Z5 Investment S.A. (the Target’s parent entity; Luxembourg).', 'Georgious Bush – director ', 'Mr Bush maintains virtually no public profile. Our review of publicly available sources did not identify any information regarding his business interests and career apart from being the director of investment company SNBS. ', 'Ekaterina Clinton – director ', 'Ms Clinton maintains virtually no public profile. Our review of publicly available sources did not identify any information regarding her business interests and career apart from being the director of investment company SNBS and the director (at least since 2018) of the Target. ', 'Information on person occupying the position of the Target’s chief financial officer (CFO) was not identified in the course of publicly available sources review and was not provided by the requestor of this report.', 'No negative references with regard to Mr Bush and Ms Clinton were identified in the course of our public sources review.']
我需要获取第一列包含 list1 的所有元素的 Dataframe。第二列必须用 list2 中的元素填充,这些元素在左侧的单元格中具有姓氏,但不是名字。这是我无法得到的结果:
column1 column2
0 Mikhail Maratovich Biden Mr Biden and Mr Trump are high-profile German entrepreneurs... According to publicly available sources... During their career, Mr Biden and Mr Trump have....
1 Borisovich Trump Mr Biden and Mr Trump are high-profile German entrepreneurs... During their career, Mr Biden and Mr Trump have....
2 Aleksey Viktorovich Obama Mr Obama is high-profile Russian...
3 Georgious Bush Mr Bush maintains virtually no... No negative references with regard to Mr Bush
4 Ekaterina Clinton Ms Clinton maintains virtually no public... No negative references with regard to Mr Bush and Ms Clinton....
为了获得我创建的 Dataframe:
column_names = ["column1", "column2"]
df = pd.DataFrame(columns = column_names)
df.column1 = list1
而且我不知道如何正确填写第二列。我试过这个:
info = []
for i in list2:
for j in df.column1:
if ((j.split(' ')[-1] in i) and (j.split(' ')[1] not in i)):
info.append(i)
joined_info = ' '.join(info)
df.column2 = joined_info
还有这个:
info = []
for i in df.column1:
for j in list2:
scanning = False
if ((i.split(' ')[-1] in j) and (i.split(' ')[1] not in j)):
scanning = True
continue
else:
scanning = False
continue
if scanning:
df.column2 = j
但是这些代码不起作用。
我真的需要你们的帮助……
【问题讨论】:
标签: python pandas list dataframe