说明
除了
concat、merge、join等显式合并方法外,pandas 还提供了一些元素级组合、掩码式填充与原地更新的方法,用于更灵活的数据整合。
combine()
DataFrame.combine()
DataFrame.combine(
other, # 另一个 DataFrame
func, # 组合函数,接收两个 Series 并返回一个 Series
fill_value=None, # 缺失值的填充值
overwrite=True, # 为 False 时,其他中被 NaN 覆盖的位置保留原值
)说明
将两个 DataFrame 的同名列成对传给
func进行组合,返回新 DataFrame。
import pandas as pd
df1 = pd.DataFrame({'A': [1, 2], 'B': [3, 4]})
df2 = pd.DataFrame({'A': [5, 6], 'B': [7, 8]})
# 取每列最大值
print(df1.combine(df2, lambda s1, s2: s1.where(s1 > s2, s2)))
# A B
# 0 5 7
# 1 6 8
# 使用 fill_value 处理 NaN
df3 = pd.DataFrame({'A': [pd.NA, 2]})
df4 = pd.DataFrame({'A': [10, pd.NA]})
print(df3.combine(df4, lambda s1, s2: s1.fillna(0) + s2.fillna(0), fill_value=100))
# A
# 0 110
# 1 102
# overwrite=False:df2 的 NaN 不覆盖 df1 的值
df5 = pd.DataFrame({'A': [1, 2]})
df6 = pd.DataFrame({'A': [10, pd.NA]})
print(df5.combine(df6, lambda s1, s2: s1.where(s1.notna(), s2), overwrite=False))Series.combine()
Series.combine(
other, # 另一个 Series(或标量)
func, # 组合函数,接收两个标量返回一个标量
fill_value=None,
)s1 = pd.Series([1, 2, 3])
s2 = pd.Series([10, 20, 30])
print(s1.combine(s2, max))
# 0 10
# 1 20
# 2 30
print(s1.combine(s2, lambda x, y: x + y))
# 0 11
# 1 22
# 2 33combine_first()
说明
使用另一个对象中的值填充调用者中的 NaN,等价于 SQL 中的
COALESCE。
DataFrame.combine_first(other)
Series.combine_first(other)DataFrame.combine_first()
df_a = pd.DataFrame({'A': [1, pd.NA], 'B': [3, 4]})
df_b = pd.DataFrame({'A': [10, 20], 'B': [30, pd.NA]})
print(df_a.combine_first(df_b))
# A B
# 0 1.0 3.0
# 1 20.0 4.0说明
与
fillna()的区别:combine_first()会自动按索引和列对齐,而fillna()只填充相同位置的 NaN。
Series.combine_first()
s_a = pd.Series([1, pd.NA, 3], index=['a', 'b', 'c'])
s_b = pd.Series([10, 20, 30], index=['a', 'b', 'z'])
print(s_a.combine_first(s_b))
# a 1.0
# b 20.0
# c 3.0
# z 30.0update()
说明
用另一个对象的非 NaN 值原地修改当前对象,默认按索引和列对齐,不返回新对象。
DataFrame.update(
other, # 另一个 DataFrame / 可转为 DataFrame 的对象
join='left', # 对齐方式:'left'、'right'、'inner'、'outer'(仅允许 'left' 在旧版)
overwrite=True, # 为 False 时,other 中的 NaN 不覆盖原值
filter_func=None, # 函数,返回 True 的位置才更新
errors='ignore', # 'raise' 或 'ignore',控制列类型不一致时是否报错
)示例
df1 = pd.DataFrame({'A': [1, 2, 3], 'B': [4, 5, 6]})
df2 = pd.DataFrame({'A': [10, pd.NA, 30], 'B': [40, 50, 60]})
df1.update(df2)
print(df1)
# A B
# 0 10 40
# 1 2 50
# 2 30 60
# overwrite=False:df2 的 NaN 不覆盖
df3 = pd.DataFrame({'A': [1, 2]})
df4 = pd.DataFrame({'A': [10, pd.NA]})
df3.update(df4, overwrite=False)
print(df3)
# A
# 0 10
# 1 2Series.update()
s1 = pd.Series([1, 2, 3])
s2 = pd.Series([10, pd.NA])
s1.update(s2)
print(s1)
# 0 10
# 1 2
# 2 3align()
说明
将两个对象按索引对齐,返回对齐后的元组。常用于预对齐后再进行运算。
DataFrame.align(
other,
join='outer', # 'outer'、'inner'、'left'、'right'
axis=None, # None=所有轴,0=索引,1=列
level=None, # MultiIndex 层级
copy=None, # 是否复制
fill_value=None, # 对齐时填充缺失值
method=None, # 填充方法,如 'ffill'(常与 fill_value 互斥)
limit=None,
fill_axis=0,
)示例
s1 = pd.Series([1, 2], index=['a', 'b'])
s2 = pd.Series([10, 20], index=['b', 'c'])
aligned1, aligned2 = s1.align(s2, join='outer', fill_value=0)
print(aligned1)
# a 1
# b 2
# c 0
print(aligned2)
# a 0
# b 10
# c 20综合对比
| 方法 | 功能 | 对齐方式 | 是否原地 |
|---|---|---|---|
combine() | 元素级自定义组合 | 索引+列 | 否 |
combine_first() | 用 other 填充 NaN | 索引+列 | 否 |
update() | 用 other 覆盖当前值 | 索引+列 | 是 |
align() | 对齐索引并返回元组 | 索引+列 | 否 |
merge() | 按键合并 | 键列 | 否 |
concat() | 轴向拼接 | 轴对齐 | 否 |
注意
update()会直接修改原对象,如需保留原始数据请先copy()。
🔗 相关链接
- 14.2 merge() - 按键合并
- 14.1 concat() - 轴向拼接
- 六、数据清洗与预处理 - 缺失值填充与替换
- 五、数据选择与索引 - 索引对齐基础