本节介绍修改 DataFrame 结构的方法:删除、弹出、插入和截断。


1. drop()

drop() 删除指定标签的行或列。默认返回新对象,inplace=True 可原地修改。

删除行

df = pd.DataFrame({'A': [1, 2, 3], 'B': [4, 5, 6]}, index=['x', 'y', 'z'])
 
# 按行标签删除
df.drop('y')
df.drop(['x', 'z'])
 
# 按位置删除行(通过索引对象)
df.drop(df.index[0])

删除列

# 按列名删除(指定 axis)
df.drop('A', axis=1)
df.drop(['A', 'B'], axis='columns')
 
# 推荐写法:直接使用 index / columns 参数
df.drop(columns=['A'])
df.drop(index=['x', 'z'])
df.drop(index=['y'], columns=['B'])

drop() 参数

参数说明
labels要删除的标签(字符串或列表)
axis轴:0 或 'index'(行)、1 或 'columns'(列)
index指定按行删除(替代 labels+axis)
columns指定按列删除(替代 labels+axis)
levelMultiIndex 指定层级
inplace是否原地修改
errors'raise'(默认,标签不存在时报错)或 'ignore'(忽略)
# MultiIndex 按层级删除
mi_df = df.set_index(['城市', 'ID'])
mi_df.drop('北京', level='城市')
mi_df.drop(('北京', 101))   # 完整元组
 
# 不存在的标签
df.drop('k', errors='ignore')   # 不报错,返回原对象

2. pop()

pop() 弹出(删除并返回)指定列或标签,始终原地操作。

DataFrame.pop()

df = pd.DataFrame({'A': [1, 2], 'B': [3, 4], 'C': [5, 6]})
 
# 删除并返回列
col_B = df.pop('B')
print(col_B)    # 返回 Series
# 0    3
# 1    4
 
print(df)       # 原 DataFrame 中 B 列已删除
#    A  C
# 0  1  5
# 1  2  6

Series.pop()

s = pd.Series([1, 2, 3], index=['a', 'b', 'c'])
 
value = s.pop('b')
print(value)    # 2
print(s)        # a    1
                # c    3

pop 与 drop 区别

  • pop() 删除并返回被删除的值,常用于拆列保存
  • drop() 仅删除,不返回被删数据
  • pop() 只能操作单个标签
  • pop() 本身就是原地操作,无 inplace 参数
  • 标签不存在时抛出 KeyError

3. insert()

insert() 在指定位置插入新列。

df = pd.DataFrame({'A': [1, 2], 'B': [3, 4]})
 
# 在位置 1(索引 0 和 1 之间)插入新列 C
df.insert(1, 'C', [10, 20])
#    A   C  B
# 0  1  10  3
# 1  2  20  4
 
# 在末尾追加
df.insert(len(df.columns), 'D', [30, 40])
 
# 插入 Series(自动按索引对齐)
s = pd.Series([100, 200], index=[1, 0])
df.insert(0, 'E', s)

insert() 参数

参数说明
loc插入位置(0 到 len(columns) 之间)
column新列名(字符串)
value新列数据(标量、列表、Series、数组)
allow_duplicates是否允许列名重复(默认 False,重复时抛错)
# 标量填充整列
df.insert(2, '常数列', 0)
 
# 允许重复列名(不推荐)
df.insert(0, 'A', [99, 99], allow_duplicates=True)
df['A']    # 返回 DataFrame(多列同名)

注意

  • insert() 是原地操作,不返回新对象
  • 不支持用 -1 表示末尾位置,需使用 len(df.columns)
  • 插入 Series 时按索引标签对齐,而不是位置

4. truncate()

truncate() 在两个索引标签之间截取数据(包含端点)。特别适合时间序列中按日期截取。

用法

df = pd.DataFrame({'A': range(6)}, index=['a', 'b', 'c', 'd', 'e', 'f'])
 
df.truncate(before='b', after='e')
#    A
# b  1
# c  2
# d  3
# e  4

参数

参数说明
before截取起始标签(含)
after截取结束标签(含)
axis轴:0 = 行,1 = 列

时间序列截取

ts = pd.Series(
    range(10),
    index=pd.date_range('2024-01-01', periods=10, freq='D')
)
 
# 按日期范围截取(包含端点)
ts.truncate(before='2024-01-03', after='2024-01-06')
 
# 只指定 after(从开始到指定日期)
ts.truncate(after='2024-01-05')
 
# 只指定 before(从指定日期到结束)
ts.truncate(before='2024-01-08')

truncate 与切片对比

  • df[1:4]:按位置切片(不含端点)
  • df.loc['b':'e']:按标签包含端点
  • df.truncate(before='b', after='e'):按标签包含端点,但不会因标签不存在而报错(只截取存在的部分)

注意

  • truncate() 要求传入的 before 和 after 与索引可比较(可排序)
  • 若 before > after 或超出索引范围,只返回空数据,不会报错