概述
分组操作的核心入口是
DataFrame.groupby()与Series.groupby(),返回DataFrameGroupBy/SeriesGroupBy对象,后续可通过聚合、变换、过滤、应用等方法操作。
1.1 DataFrame.groupby()
签名
DataFrame.groupby(by=None, axis=0, level=None, as_index=True, sort=True, group_keys=True, observed=False, dropna=True)
参数详解
| 参数 | 类型 | 默认值 | 说明 |
|---|---|---|---|
by | mapping、function、label、list of labels | None | 用于分组的键。可以是列名、Series、字典、函数或它们的列表 |
axis | int | 0 | 分组沿哪个轴(0 行,1 列),pandas 2.x 建议使用 axis=0 |
level | int、level name、或两者的序列 | None | 当索引为 MultiIndex 时,按指定层级分组 |
as_index | bool | True | 分组键是否作为最终结果的索引;False 时作为列保留 |
sort | bool | True | 是否对分组键排序 |
group_keys | bool | True | 调用 apply() 时是否在结果中添加分组键 |
observed | bool | False | 对 Categorical 分组时,是否仅使用实际出现的类别(True)或所有类别(False) |
dropna | bool | True | 是否丢弃分组键为 NaN 的组 |
pandas 2.3 注意
axis=1按列分组属于弃用行为,官方建议使用df.T.groupby(...)替代。
基础示例
import pandas as pd
df = pd.DataFrame({
"部门": ["A", "A", "B", "B", "C"],
"性别": ["男", "女", "男", "女", "男"],
"薪资": [8000, 9500, 12000, 11000, 9000],
})
# 按单列分组
df.groupby("部门")
# 按多列分组
df.groupby(["部门", "性别"])
# by 参数传入 Series 或列表
df.groupby(by=df["部门"])
# by 参数传入字典:按行映射到组
df.groupby(by={0: "A组", 1: "A组", 2: "B组", 3: "B组", 4: "C组"})
# by 参数传入函数:按行索引规则分组
df.groupby(by=lambda idx: "偶" if idx % 2 == 0 else "奇")
# as_index=False:分组键保留为列
df.groupby("部门", as_index=False)["薪资"].sum()
# sort=False:保持原始分组键出现顺序
df.groupby("部门", sort=False)["薪资"].sum()
# dropna=False:保留 NaN 分组
df2 = pd.DataFrame({"键": ["A", None, "B", None], "值": [1, 2, 3, 4]})
df2.groupby("键", dropna=False)["值"].sum()1.2 Series.groupby()
Series 分组与 DataFrame 类似,但结果往往直接是 Series 聚合结果。
s = pd.Series([1, 2, 3, 4, 5], index=["a", "b", "c", "d", "e"])
s.groupby(["x", "y", "x", "y", "x"]).sum()
# 输出:
# x 9
# y 6
# dtype: int64
# 与 DataFrame 分组结合
df.groupby("部门")["薪资"].mean()1.3 按列、按列表、按函数、按索引层级分组
| 分组方式 | 说明 | 示例 |
|---|---|---|
| 按列(label) | 按 DataFrame 中的列名分组 | df.groupby("部门") |
| 按列表(list) | 按多个列名或外部列表分组 | df.groupby(["部门", "性别"]) |
| 按函数(function) | 对索引或行内容调用函数后分组 | df.groupby(lambda x: x % 2) |
| 按索引层级(level) | 对 MultiIndex 指定层级分组 | df.groupby(level="年份") |
按索引层级分组示例
import numpy as np
df_mi = pd.DataFrame(
np.arange(12).reshape(6, 2),
index=pd.MultiIndex.from_product(
[["2023", "2024"], ["Q1", "Q2", "Q3"]],
names=["年份", "季度"],
),
columns=["销量", "库存"],
)
# 按一级索引分组
df_mi.groupby(level="年份").sum()
# 按多个索引层级分组
df_mi.groupby(level=["年份", "季度"]).sum()
# 同时使用 by 和 level
df_mi.groupby(by="销量", level="年份").sum()1.4 Grouper 对象
pd.Grouper用于在groupby()中提供更精细的分组规格,尤其适用于时间频率分组和跨列/索引混合分组。
常用参数
| 参数 | 说明 |
|---|---|
key | 要分组的列名或索引名 |
level | 索引层级 |
freq | 时间分组频率(如 'D'、'M'、'Q') |
closed | 时间区间端点闭合方式('left' / 'right') |
label | 聚合标签取区间左/右端点 |
sort | 是否排序 |
origin | 时间对齐起点 |
offset | 时间偏移 |
示例
# 按时间频率分组:每天 → 每月
ts_df = pd.DataFrame(
{"值": range(10)},
index=pd.date_range("2024-01-01", periods=10, freq="D"),
)
ts_df.groupby(pd.Grouper(freq="M")).sum()
# key + freq 指定时间列
sales = pd.DataFrame({
"日期": pd.date_range("2024-01-01", periods=100, freq="D"),
"金额": np.random.randint(100, 500, 100),
})
sales.groupby(pd.Grouper(key="日期", freq="W"))["金额"].sum()
# 同时使用多个 Grouper(按季度 + 地区)
df.groupby([pd.Grouper(key="日期", freq="Q"), "地区"]).sum()与
resample()的关系
pd.Grouper(freq=...)等价于先设置索引为时间列再调用resample(),但使用key时不改变原 DataFrame 索引,更加灵活。
相关笔记
- 上一节:分组与聚合 章节总览
- 下一节:分组与聚合-2-迭代与分组
- 延伸:特殊分组