這是我的資料集:
| Name | Dept | Project area/areas interested |
| -------- | -------- |-----------------------------------|
| Joe | Biotech | Cell culture//Bioinfo//Immunology |
| Ann | Biotech | Cell culture |
| Ben | Math | Trigonometry//Algebra |
| Keren | Biotech | Microbio |
| Alice | Physics | Optics |
這就是我想要的結果:
| Name | Dept |Cell culture|Bioinfo|Immunology|Trigonometry|Algebra|Microbio|Optics|
| -------- | -------- |------------|-------|----------|------------|-------|--------|------|
| Joe | Biotech | 1 | 1 | 1 | 0 | 0 | 0 | 0 |
| Ann | Biotech | 1 | 0 | 1 | 0 | 0 | 0 | 0 |
| Ben | Math | 0 | 0 | 0 | 1 | 1 | 0 | 0 |
| Keren | Biotech | 0 | 0 | 0 | 0 | 0 | 1 | 0 |
| Alice | Physics | 0 | 0 | 0 | 0 | 0 | 0 | 1 |
我不僅必須根據行將最后一列拆分為不同的列 - 我還必須重新拆分由“//”分隔的某些列值。并且資料框中的值必須替換為 1 或 0 (int)。我已經堅持了一段時間了(-_-;)
uj5u.com熱心網友回復:
您可以將 pandas.concat 與 pandas.get_dummies 結合使用,如下所示:
pd.concat([df[["Name", "Dept"]], df["Project area/areas interested"].str.get_dummies(sep='//')], axis=1)
轉載請註明出處,本文鏈接:https://www.uj5u.com/caozuo/527459.html
上一篇:記憶體問題使用Python將大型CSV檔案轉換為excel
下一篇:無法繪制回歸的最佳擬合線
