说明

DataFrame.compare() 与 Series.compare() 用于逐元素比较两个对象,输出两对象不同的部分,以多层列/索引呈现差异。

DataFrame.compare()

语法

DataFrame.compare(
    other,               # 要比较的另一个 DataFrame
    align_axis=1,        # 结果中差异展示的轴:1=列并排,0=行堆叠
    keep_shape=False,    # 是否保留所有列(不裁剪共同的相同列)
    keep_equal=False,    # 是否同时展示相同的值
    result_names=('self', 'other'),  # 差异层级的列名
)

参数说明

参数说明默认
align_axis差异展示轴向:1 并排列,0 堆叠行1
keep_shape为 True 时保留所有原始列,仅空值为 NaNFalse
keep_equal为 True 时相同值也显示False
result_names差异列的两层名称('self', 'other')

示例

import pandas as pd
 
df1 = pd.DataFrame({'A': [1, 2, 3], 'B': [4, 5, 6]})
df2 = pd.DataFrame({'A': [1, 9, 3], 'B': [4, 5, 7]})
 
# 默认:仅显示差异行,差异列并排展示
print(df1.compare(df2))
#      A       B
#   self other self other
# 1  2.0   9.0  NaN   NaN
# 2  NaN   NaN  6.0   7.0
 
# align_axis=0:差异按行堆叠
print(df1.compare(df2, align_axis=0))
#        A    B
# 1 self  2.0 NaN
#   other 9.0 NaN
# 2 self  NaN 6.0
#   other NaN 7.0
 
# keep_shape=True:保留所有列
print(df1.compare(df2, keep_shape=True))
#      A       B
#   self other self other
# 0  NaN   NaN  NaN   NaN
# 1  2.0   9.0  NaN   NaN
# 2  NaN   NaN  6.0   7.0
 
# keep_equal=True:相同值也显示
print(df1.compare(df2, keep_equal=True))
#      A       B
#   self other self other
# 0  1.0   1.0  4.0   4.0
# 1  2.0   9.0  5.0   5.0
# 2  3.0   3.0  6.0   7.0
 
# 自定义结果列名
print(df1.compare(df2, result_names=('left', 'right')))

Series.compare()

语法

Series.compare(
    other,               # 另一个 Series
    align_axis=1,        # 差异展示轴向
    keep_shape=False,    # 是否保留所有原始索引
    keep_equal=False,    # 是否显示相同值
    result_names=('self', 'other'),
)

示例

s1 = pd.Series([1, 2, 3], index=['a', 'b', 'c'])
s2 = pd.Series([1, 9, 3], index=['a', 'b', 'c'])
 
print(s1.compare(s2))
#    self  other
# b   2.0    9.0
 
print(s1.compare(s2, keep_shape=True))
#    self  other
# a   NaN    NaN
# b   2.0    9.0
# c   NaN    NaN
 
print(s1.compare(s2, align_axis=0))
#   b  self    2.0
#      other   9.0

实际应用

使用场景

  • 数据审核:比较两个版本的 DataFrame(如导入前后、EXCEL 修订前后)。
  • 快速定位差异单元格。
  • 模型验证:比较预测结果与真实结果。
# 找出两个 DataFrame 中所有不同的单元格
diff = df1.compare(df2)
print(diff.notna().any(axis=1))   # 哪些行有差异

🔗 相关链接