我有一個看起來像這樣的 Pandas 資料框。
Customer ID Customer Name Price_Val
5015 AXN 17.12
5015 AXN 2.08
5015 AXN 3.453
7315 BXN 2.22
7315 BXN 8.46567
3283 CXN 88
3283 CXN 0.4600
3283 CXN 6.46
3283 CXN
我想創建名為dec_value 的列。我希望 dec_value 列具有來自相應Price_Val列的小數位長度。
例如,我希望我的 dec_value 列應如下所示。
Customer ID Customer Name Price_Val dec_value
5015 AXN 17.12 2
5015 AXN 2.08 2
5015 AXN 3.453 3
7315 BXN 2.22 2
7315 BXN 8.4656 4
3283 CXN 88 0
3283 CXN 0.4600 4
3283 CXN 6.46 2
3283 CXN 0
我正在使用下面的代碼來完成上述作業。
i = 0
for value in df1['Price_Val']:
if value == '':
df1.loc[i, "dec_value "] = 0
else:
colval = value
k = str(colval)[::-1].find('.')
if k == -1:
df1.loc[i,"dec_value"] = 0
else:
df1.loc[i,"dec_value"] = str(colval)[::-1].find('.')
i=i 1
執行此操作的最有效方法是什么?
uj5u.com熱心網友回復:
將您的列轉換為字串、split點、rstrip零并計算字符數:
df['Price_Val'].fillna('').apply(lambda x: len(str(x).split('.')[-1].rstrip('0')))
或者
df['dec_value'] = (df['Price_Val'].fillna('').astype(str)
.str.split('.').str[-1]
.str.rstrip('0').str.len()
)
輸出:
Customer ID Customer Name Price_Val dec_value
0 5015 AXN 17.12000 2
1 5015 AXN 2.08000 2
2 5015 AXN 3.45300 3
3 7315 BXN 2.22000 2
4 7315 BXN 8.46567 5
5 3283 CXN 88.00000 0
6 3283 CXN 0.46000 2
7 3283 CXN 6.46000 2
8 3283 CXN NaN 0
或者,使用正則運算式:
df['dec_value'] = (df['Price_Val'].fillna('').astype(str)
.str.extract('\.(\d*[1-9])', expand=False)
.str.len().fillna(0, downcast='infer')
)
備選方案的時間安排(90k 行)
# apply
50.5 ms ± 913 μs per loop (mean ± std. dev. of 7 runs, 10 loops each)
# regex
83.9 ms ± 323 μs per loop (mean ± std. dev. of 7 runs, 10 loops each)
# str pipeline
115 ms ± 2.39 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)
轉載請註明出處,本文鏈接:https://www.uj5u.com/qukuanlian/369403.html
標籤:Python 熊猫 数据框 python-3.9
上一篇:計算欄位中單詞/字符的出現次數
下一篇:Pandas日期列:日期轉換問題
