本节介绍修改 DataFrame 结构的方法:删除、弹出、插入和截断。
1. drop()
drop() 删除指定标签的行或列。默认返回新对象,inplace=True 可原地修改。
删除行
df = pd.DataFrame({'A': [1, 2, 3], 'B': [4, 5, 6]}, index=['x', 'y', 'z'])
# 按行标签删除
df.drop('y')
df.drop(['x', 'z'])
# 按位置删除行(通过索引对象)
df.drop(df.index[0])删除列
# 按列名删除(指定 axis)
df.drop('A', axis=1)
df.drop(['A', 'B'], axis='columns')
# 推荐写法:直接使用 index / columns 参数
df.drop(columns=['A'])
df.drop(index=['x', 'z'])
df.drop(index=['y'], columns=['B'])drop() 参数
| 参数 | 说明 |
|---|---|
labels | 要删除的标签(字符串或列表) |
axis | 轴:0 或 'index'(行)、1 或 'columns'(列) |
index | 指定按行删除(替代 labels+axis) |
columns | 指定按列删除(替代 labels+axis) |
level | MultiIndex 指定层级 |
inplace | 是否原地修改 |
errors | 'raise'(默认,标签不存在时报错)或 'ignore'(忽略) |
# MultiIndex 按层级删除
mi_df = df.set_index(['城市', 'ID'])
mi_df.drop('北京', level='城市')
mi_df.drop(('北京', 101)) # 完整元组
# 不存在的标签
df.drop('k', errors='ignore') # 不报错,返回原对象2. pop()
pop() 弹出(删除并返回)指定列或标签,始终原地操作。
DataFrame.pop()
df = pd.DataFrame({'A': [1, 2], 'B': [3, 4], 'C': [5, 6]})
# 删除并返回列
col_B = df.pop('B')
print(col_B) # 返回 Series
# 0 3
# 1 4
print(df) # 原 DataFrame 中 B 列已删除
# A C
# 0 1 5
# 1 2 6Series.pop()
s = pd.Series([1, 2, 3], index=['a', 'b', 'c'])
value = s.pop('b')
print(value) # 2
print(s) # a 1
# c 3pop 与 drop 区别
pop()删除并返回被删除的值,常用于拆列保存drop()仅删除,不返回被删数据pop()只能操作单个标签pop()本身就是原地操作,无inplace参数- 标签不存在时抛出
KeyError
3. insert()
insert() 在指定位置插入新列。
df = pd.DataFrame({'A': [1, 2], 'B': [3, 4]})
# 在位置 1(索引 0 和 1 之间)插入新列 C
df.insert(1, 'C', [10, 20])
# A C B
# 0 1 10 3
# 1 2 20 4
# 在末尾追加
df.insert(len(df.columns), 'D', [30, 40])
# 插入 Series(自动按索引对齐)
s = pd.Series([100, 200], index=[1, 0])
df.insert(0, 'E', s)insert() 参数
| 参数 | 说明 |
|---|---|
loc | 插入位置(0 到 len(columns) 之间) |
column | 新列名(字符串) |
value | 新列数据(标量、列表、Series、数组) |
allow_duplicates | 是否允许列名重复(默认 False,重复时抛错) |
# 标量填充整列
df.insert(2, '常数列', 0)
# 允许重复列名(不推荐)
df.insert(0, 'A', [99, 99], allow_duplicates=True)
df['A'] # 返回 DataFrame(多列同名)注意
insert()是原地操作,不返回新对象- 不支持用
-1表示末尾位置,需使用len(df.columns)- 插入 Series 时按索引标签对齐,而不是位置
4. truncate()
truncate() 在两个索引标签之间截取数据(包含端点)。特别适合时间序列中按日期截取。
用法
df = pd.DataFrame({'A': range(6)}, index=['a', 'b', 'c', 'd', 'e', 'f'])
df.truncate(before='b', after='e')
# A
# b 1
# c 2
# d 3
# e 4参数
| 参数 | 说明 |
|---|---|
before | 截取起始标签(含) |
after | 截取结束标签(含) |
axis | 轴:0 = 行,1 = 列 |
时间序列截取
ts = pd.Series(
range(10),
index=pd.date_range('2024-01-01', periods=10, freq='D')
)
# 按日期范围截取(包含端点)
ts.truncate(before='2024-01-03', after='2024-01-06')
# 只指定 after(从开始到指定日期)
ts.truncate(after='2024-01-05')
# 只指定 before(从指定日期到结束)
ts.truncate(before='2024-01-08')truncate 与切片对比
df[1:4]:按位置切片(不含端点)df.loc['b':'e']:按标签包含端点df.truncate(before='b', after='e'):按标签包含端点,但不会因标签不存在而报错(只截取存在的部分)
注意
truncate()要求传入的before和after与索引可比较(可排序)- 若
before > after或超出索引范围,只返回空数据,不会报错