说明
DataFrame.compare()与Series.compare()用于逐元素比较两个对象,输出两对象不同的部分,以多层列/索引呈现差异。
DataFrame.compare()
语法
DataFrame.compare(
other, # 要比较的另一个 DataFrame
align_axis=1, # 结果中差异展示的轴:1=列并排,0=行堆叠
keep_shape=False, # 是否保留所有列(不裁剪共同的相同列)
keep_equal=False, # 是否同时展示相同的值
result_names=('self', 'other'), # 差异层级的列名
)参数说明
| 参数 | 说明 | 默认 |
|---|---|---|
align_axis | 差异展示轴向:1 并排列,0 堆叠行 | 1 |
keep_shape | 为 True 时保留所有原始列,仅空值为 NaN | False |
keep_equal | 为 True 时相同值也显示 | False |
result_names | 差异列的两层名称 | ('self', 'other') |
示例
import pandas as pd
df1 = pd.DataFrame({'A': [1, 2, 3], 'B': [4, 5, 6]})
df2 = pd.DataFrame({'A': [1, 9, 3], 'B': [4, 5, 7]})
# 默认:仅显示差异行,差异列并排展示
print(df1.compare(df2))
# A B
# self other self other
# 1 2.0 9.0 NaN NaN
# 2 NaN NaN 6.0 7.0
# align_axis=0:差异按行堆叠
print(df1.compare(df2, align_axis=0))
# A B
# 1 self 2.0 NaN
# other 9.0 NaN
# 2 self NaN 6.0
# other NaN 7.0
# keep_shape=True:保留所有列
print(df1.compare(df2, keep_shape=True))
# A B
# self other self other
# 0 NaN NaN NaN NaN
# 1 2.0 9.0 NaN NaN
# 2 NaN NaN 6.0 7.0
# keep_equal=True:相同值也显示
print(df1.compare(df2, keep_equal=True))
# A B
# self other self other
# 0 1.0 1.0 4.0 4.0
# 1 2.0 9.0 5.0 5.0
# 2 3.0 3.0 6.0 7.0
# 自定义结果列名
print(df1.compare(df2, result_names=('left', 'right')))Series.compare()
语法
Series.compare(
other, # 另一个 Series
align_axis=1, # 差异展示轴向
keep_shape=False, # 是否保留所有原始索引
keep_equal=False, # 是否显示相同值
result_names=('self', 'other'),
)示例
s1 = pd.Series([1, 2, 3], index=['a', 'b', 'c'])
s2 = pd.Series([1, 9, 3], index=['a', 'b', 'c'])
print(s1.compare(s2))
# self other
# b 2.0 9.0
print(s1.compare(s2, keep_shape=True))
# self other
# a NaN NaN
# b 2.0 9.0
# c NaN NaN
print(s1.compare(s2, align_axis=0))
# b self 2.0
# other 9.0实际应用
使用场景
- 数据审核:比较两个版本的 DataFrame(如导入前后、EXCEL 修订前后)。
- 快速定位差异单元格。
- 模型验证:比较预测结果与真实结果。
# 找出两个 DataFrame 中所有不同的单元格
diff = df1.compare(df2)
print(diff.notna().any(axis=1)) # 哪些行有差异🔗 相关链接
- 14.1 concat() - 拼接后比较
- 14.7 其他组合与更新 - combine / update 等修改操作
- 六、数据清洗与预处理 - 缺失值处理与更新