本节前言

本节介绍数据的替换(replace)、条件选择(where / mask)、数值裁剪(clip)以及舍入操作。

1. replace()

将指定值替换为其他值。

基本用法

df = pd.DataFrame({'A': [1, 999, 3], 'B': ['x', 'y', 'y']})
 
# 标量 → 标量
df.replace(999, -1)
 
# 列表 → 列表
df.replace([1, 3], [100, 300])
 
# 字典映射
df.replace({'A': {999: -1}, 'B': {'x': 'X'}})
 
# 正则表达式
df.replace(r'^\d$', 'num', regex=True)

关键参数

参数说明
to_replace被替换的值、列表、字典或正则
value替换后的值
regex是否将 to_replace 视为正则
method’ffill’ / ‘bfill’ 填充替换值
limit最大替换次数

2. where()

保留满足条件的值,不满足的替换为其他值。

s = pd.Series([1, 5, 3, 8, 2])
 
s.where(s > 3)
# 0    NaN
# 1    5.0
# 2    NaN
# 3    8.0
# 4    NaN
 
s.where(s > 3, other=0)   # 不满足的条件替换为 0
 
df.where(df > 0, 0)        # DataFrame 级

理解 where

where(cond, other) 等价于:满足 cond 保留原值,否则用 other 替换。

3. mask()

where() 的反向操作:替换满足条件的值。

s.mask(s > 3)
# 0    1.0
# 1    NaN
# 2    3.0
# 3    NaN
# 4    2.0
 
s.mask(s > 3, other=-1)   # 满足条件替换为 -1
方法行为
where(cond)cond=True 保留;False 替换
mask(cond)cond=True 替换;False 保留

4. clip()

将数据裁剪到指定上下限。

s = pd.Series([-5, 1, 3, 10, 20])
 
s.clip(lower=0, upper=10)
# 0     0
# 1     1
# 2     3
# 3    10
# 4    10
 
s.clip(lower=0)     # 只设置下限
s.clip(upper=10)    # 只设置上限
 
# DataFrame 每列不同边界
df.clip(lower=df.min(), upper=df.max())

clip 的典型用途

  • 去除极端离群值。
  • 将预测值限制在合理业务范围内。

5. round() / floor() / ceil()

数值舍入操作。

s = pd.Series([1.234, 2.567, 3.5])
 
s.round(1)
# 0    1.2
# 1    2.6
# 2    3.5
 
s.floor()    # 向下取整
# 0    1.0
# 1    2.0
# 2    3.0
 
s.ceil()     # 向上取整
# 0    2.0
# 1    3.0
# 2    4.0
 
# round 支持每列不同小数位
df.round({'A': 1, 'B': 2})
方法行为等价函数
round(decimals)四舍五入到指定小数位—
floor()向下取整np.floor
ceil()向上取整np.ceil

与 DataFrame 的关系

  • DataFrame.round() 按列指定精度。
  • floor / ceil 也支持 axis 参数。

🔗 相关笔记