成為熊貓中的以下DataFrame。
| 國家 | 嘗試 | 城市 | 城市 | 其他 | 重要的 | 其他重要 | 其他_1 |
|---|---|---|---|---|---|---|---|
| 法國 | 法國 | 巴黎 | 巴黎 | 藍色的 | 019210 | 0011119 | 紅色的 |
| 西班牙 | 西班牙 | 馬德里 | 巴塞羅那 | 藍色的 | 1211 | 0019210 | 藍色的 |
| 德國 | 西班牙 | 巴塞羅那 | 巴塞羅那 | 白色的 | 019210 | 1212 | 紅色的 |
| 法國 | 英國 | 布爾多 | 倫敦 | 藍色的 | 019210 | 91021 | 紅色的 |
我必須用 NaN 填寫不重要的列(其他)的資訊,以防萬一country != ctry || city != cty。資料框結果:
| 國家 | 嘗試 | 城市 | 城市 | 其他 | 重要的 | 其他重要 | 其他_1 |
|---|---|---|---|---|---|---|---|
| 法國 | 法國 | 巴黎 | 巴黎 | 藍色的 | 019210 | 0011119 | 紅色的 |
| 西班牙 | 西班牙 | 馬德里 | 巴塞羅那 | 鈉 | 1211 | 0019210 | 鈉 |
| 德國 | 西班牙 | 巴塞羅那 | 巴塞羅那 | 鈉 | 019210 | 1212 | 鈉 |
| 法國 | 英國 | 布爾多 | 倫敦 | 鈉 | 019210 | 91021 | 鈉 |
最后我洗掉了國家和城市列。
df = df.drop(['country', 'city'], axis=1)
| 嘗試 | 城市 | 其他 | 重要的 | 其他重要 | 其他_1 |
|---|---|---|---|---|---|
| 法國 | 巴黎 | 藍色的 | 019210 | 0011119 | 紅色的 |
| 西班牙 | 巴塞羅那 | 鈉 | 1211 | 0019210 | 鈉 |
| 西班牙 | 巴塞羅那 | 鈉 | 019210 | 1212 | 鈉 |
| 英國 | 倫敦 | 鈉 | 019210 | 91021 | 鈉 |
如果我想保留為 NaN 的列可以用每個列的名稱在字串向量中表示,我將不勝感激。['other', 'other_1']
uj5u.com熱心網友回復:
按DataFrame.loc條件設定缺失值:
cols = ['other','other_1']
df.loc[df.country.ne(df.ctry) | df.city.ne(df.cty), cols] = np.nan
df = df.drop(['country', 'city'], axis=1)
洗掉列的解決方案country, city使用DataFrame.pop:
cols = ['other','other_1']
df.loc[df.pop('country').ne(df.ctry) | df.pop('city').ne(df.cty), cols] = np.nan
print (df)
ctry cty other important other_important other_1
0 France París blue 19210 11119 red
1 Spain Barcelona NaN 1211 19210 NaN
2 Spain Barcelona NaN 19210 1212 NaN
3 UK London NaN 19210 91021 NaN
uj5u.com熱心網友回復:
# list of columns
cols=['other', 'other_1']
# use mask to make NaN when condition is met
df[cols] = df[cols].mask(df['country'].ne(df['ctry']) | df['city'].ne(df['cty']))
# drop columns
df = df.drop(['country', 'city'], axis=1)
df
ctry cty other important other_important other_1
0 France París blue 19210 11119 red
1 Spain Barcelona NaN 1211 19210 NaN
2 Spain Barcelona NaN 19210 1212 NaN
3 UK London NaN 19210 91021 fNaN
轉載請註明出處,本文鏈接:https://www.uj5u.com/caozuo/527451.html
標籤:Python熊猫数据框
