我有一個與此資料框非常相似的資料框:
| 指數 | 日期 | 月 |
|---|---|---|
| 0 | 2019-12-1 | 12 |
| 1 | 2020-03-1 | 3 |
| 2 | 2020-07-1 | 7 |
| 3 | 2021-02-1 | 2 |
| 4 | 2021-09-1 | 9 |
我想結合最接近一組月份的所有日期。月份需要像這樣標準化:
| 幾個月 | 標準化月份 |
|---|---|
| 3、4、5 | 4 |
| 6、7、8、9 | 8 |
| 1、2、10、11、12 | 12 |
所以輸出將是:
| 指數 | 日期 | 月 |
|---|---|---|
| 0 | 2019-12-1 | 12 |
| 1 | 2020-04-1 | 4 |
| 2 | 2020-08-1 | 8 |
| 3 | 2020-12-1 | 12 |
| 4 | 2021-08-1 | 8 |
uj5u.com熱心網友回復:
您可以嘗試創建幾個月的字典,其中:
norm_month_dict = {3: 4, 4: 4, 5: 4, 6: 8, 7: 8, 8: 8, 9: 8, 1: 12, 2: 12, 10: 12, 11: 12, 12: 12}
然后使用此字典將月份值映射到它們各自的標準化月份值。
df['normalized_months'] = df.months.map(norm_month_dict)
uj5u.com熱心網友回復:
您可以遍歷 DataFrame 并使用replace來更改日期。
import pandas as pd
df = pd.DataFrame(data={'date': ["2019-12-1", "2020-03-1", "2020-07-1", "2021-02-1", "2021-09-1"],
'month': [12,3,7,2,9]})
for index, row in df.iterrows():
if (row['month'] in [3,4,5]):
df['month'][index] = 4
df["date"][index] = df["date"][0].replace(df["date"][0][5:7],"04")
elif (row['month'] in [6,7,8,9]):
df['month'][index] = 8
df["date"][index] = df["date"][0].replace(df["date"][0][5:7],"08")
else:
df['month'][index] = 12
df["date"][index] = df["date"][0].replace(df["date"][0][5:7],"12")
uj5u.com熱心網友回復:
您需要從第二個資料幀構造一個字典(假設df1和df2):
d = (
df2.assign(Months=df2['Months'].str.split(', '))
.explode('Months').astype(int)
.set_index('Months')['Normalized month'].to_dict()
)
# {3: 4, 4: 4, 5: 4, 6: 8, 7: 8, 8: 8, 9: 8, 1: 12, 2: 12, 10: 12, 11: 12, 12: 12}
然后map是值:
df1['month'] = df1['month'].map(d)
輸出:
index date month
0 0 2019-12-1 12
1 1 2020-03-1 4
2 2 2020-07-1 8
3 3 2021-02-1 12
4 4 2021-09-1 8`
轉載請註明出處,本文鏈接:https://www.uj5u.com/shujuku/435895.html
