我有兩個資料框:
import pandas as pd
from numpy import nan
df1 = pd.DataFrame({'key':[1,2,3,4],
'only_at_df1':['a','b','c','d'],
'col2':['e','f','g','h'],})
df2 = pd.DataFrame({'key':[1,9],
'only_at_df2':[nan,'x'],
'col2':['e','z'],})
如何獲得這個:
df3 = pd.DataFrame({'key':[1,2,3,4,9],
'only_at_df1':['a','b','c','d',nan],
'only_at_df2':[nan,nan,nan,nan,'x'],
'col2':['e','f','g','h','z'],})
任何幫助表示贊賞。
uj5u.com熱心網友回復:
最好的可能是combine_first在臨時將“key”設定為索引后使用:
df1.set_index('key').combine_first(df2.set_index('key')).reset_index()
輸出:
key col2 only_at_df1 only_at_df2
0 1 e a NaN
1 2 f b NaN
2 3 g c NaN
3 4 h d NaN
4 9 z NaN x
uj5u.com熱心網友回復:
這似乎是 with 的直接merge使用how="outer":
df1.merge(df2, how="outer")
輸出:
key only_at_df1 col2 only_at_df2
0 1 a e NaN
1 2 b f NaN
2 3 c g NaN
3 4 d h NaN
4 9 NaN z x
轉載請註明出處,本文鏈接:https://www.uj5u.com/gongcheng/455773.html
