本章介绍创建分类数据的多种方式,以及
CategoricalDtype的核心参数。
1. 使用 pd.Categorical()
最直接的构造方法:
import pandas as pd
cat = pd.Categorical(['a', 'b', 'c', 'a'])
print(cat)
# ['a', 'b', 'c', 'a']
# Categories (3, object): ['a', 'b', 'c']
print(cat.categories) # Index(['a', 'b', 'c'])
print(cat.codes) # [0 1 2 0]指定 categories 与 ordered
cat = pd.Categorical(
['低', '高', '中'],
categories=['低', '中', '高'],
ordered=True
)
print(cat)
# ['低', '高', '中']
# Categories (3, object): ['低' < '中' < '高']参数详解
| 参数 | 说明 |
|---|---|
values | 数据源(列表、Series、数组) |
categories | 类别的唯一值列表;若不指定,则自动从数据中推断 |
ordered | 是否有序(默认 False) |
注意
若数据中出现
categories之外的取值,不会报错,但会标记为缺失(NaN)。
2. 使用 CategoricalDtype
CategoricalDtype 是可复用的分类类型,适合同时创建多列或用于 astype。
from pandas import CategoricalDtype
cat_type = CategoricalDtype(
categories=['低', '中', '高'],
ordered=True
)
# 在 astype 中使用
s = pd.Series(['高', '低', '中']).astype(cat_type)
print(s.dtype) # category
print(s.cat.ordered) # True参数
| 参数 | 说明 |
|---|---|
categories | 类别列表(或 None 自动推断) |
ordered | 是否有序(默认 False) |
优点
将
CategoricalDtype存储在变量中,可在多个 DataFrame 中复用,保证类别结构一致。
3. 使用 astype(‘category’)
将普通列快速转为分类类型,最简单的方式:
df = pd.DataFrame({'城市': ['北京', '上海', '北京', '广州']})
df['城市'] = df['城市'].astype('category')
print(df.dtypes)
# 城市 category
# dtype: object
# 查看类别
print(df['城市'].cat.categories)
# Index(['上海', '北京', '广州'], dtype='object')
# 查看编码
print(df['城市'].cat.codes)
# 0 1
# 1 0
# 2 1
# 3 2注意
直接
astype('category')时,categories自动从数据去重产生,且ordered=False。若需指定顺序,使用CategoricalDtype。
4. 使用 pd.cut() 与 pd.qcut()
将连续数值离散化为分类数据。
pd.cut()
等宽分箱:
ages = pd.Series([18, 25, 45, 60, 80])
bins = pd.cut(ages, bins=3)
print(bins)
# 0 (17.986, 38.667]
# 1 (17.986, 38.667]
# 2 (38.667, 59.333]
# 3 (59.333, 80.0]
# 4 (59.333, 80.0]
# Categories (3, interval[float64, right]): [(17.986, 38.667] < (38.667, 59.333] < (59.333, 80.0]]自定义边界与标签:
bins = pd.cut(
ages,
bins=[0, 18, 35, 60, 100],
labels=['未成年', '青年', '中年', '老年'],
ordered=True
)
print(bins)
# 0 青年
# 1 青年
# 2 中年
# 3 老年
# 4 老年
# Categories (4, object): ['未成年' < '青年' < '中年' < '老年']pd.qcut()
等频分箱(按分位数):
data = pd.Series(range(100))
q = pd.qcut(data, q=4)
print(q.value_counts())5. 从其他类型转换
# 数值列转分类
s = pd.Series([1, 2, 2, 3])
s.astype('category')
# 字符串列转分类
s = pd.Series(['a', 'b', 'a'])
s.astype('category')
# 布尔列转分类
s = pd.Series([True, False, True])
s.astype('category')从 DataFrame 的多个列创建
df = pd.DataFrame({'A': ['x', 'y'], 'B': [1, 2]})
df = df.astype({'A': 'category', 'B': 'category'})6. 从 records / list 创建
# 从字典列表
records = [{'等级': '高', '城市': '北京'}, {'等级': '低', '城市': '上海'}]
df = pd.DataFrame(records).astype({'等级': 'category', '城市': 'category'})小结
pd.Categorical()直接创建CategoricalDtype定义可复用类型astype('category')快速转换pd.cut()/pd.qcut()数值分箱- 多种来源均可转换