本章介绍创建分类数据的多种方式,以及 CategoricalDtype 的核心参数。


1. 使用 pd.Categorical()

最直接的构造方法:

import pandas as pd
 
cat = pd.Categorical(['a', 'b', 'c', 'a'])
print(cat)
# ['a', 'b', 'c', 'a']
# Categories (3, object): ['a', 'b', 'c']
 
print(cat.categories)   # Index(['a', 'b', 'c'])
print(cat.codes)        # [0 1 2 0]

指定 categories 与 ordered

cat = pd.Categorical(
    ['低', '高', '中'],
    categories=['低', '中', '高'],
    ordered=True
)
print(cat)
# ['低', '高', '中']
# Categories (3, object): ['低' < '中' < '高']

参数详解

参数说明
values数据源(列表、Series、数组)
categories类别的唯一值列表;若不指定,则自动从数据中推断
ordered是否有序(默认 False)

注意

若数据中出现 categories 之外的取值,不会报错,但会标记为缺失(NaN)。


2. 使用 CategoricalDtype

CategoricalDtype 是可复用的分类类型,适合同时创建多列或用于 astype。

from pandas import CategoricalDtype
 
cat_type = CategoricalDtype(
    categories=['低', '中', '高'],
    ordered=True
)
 
# 在 astype 中使用
s = pd.Series(['高', '低', '中']).astype(cat_type)
print(s.dtype)   # category
print(s.cat.ordered)  # True

参数

参数说明
categories类别列表(或 None 自动推断)
ordered是否有序(默认 False)

优点

将 CategoricalDtype 存储在变量中,可在多个 DataFrame 中复用,保证类别结构一致。


3. 使用 astype(‘category’)

将普通列快速转为分类类型,最简单的方式:

df = pd.DataFrame({'城市': ['北京', '上海', '北京', '广州']})
 
df['城市'] = df['城市'].astype('category')
print(df.dtypes)
# 城市    category
# dtype: object
 
# 查看类别
print(df['城市'].cat.categories)
# Index(['上海', '北京', '广州'], dtype='object')
 
# 查看编码
print(df['城市'].cat.codes)
# 0    1
# 1    0
# 2    1
# 3    2

注意

直接 astype('category') 时,categories 自动从数据去重产生,且 ordered=False。若需指定顺序,使用 CategoricalDtype。


4. 使用 pd.cut() 与 pd.qcut()

将连续数值离散化为分类数据。

pd.cut()

等宽分箱:

ages = pd.Series([18, 25, 45, 60, 80])
 
bins = pd.cut(ages, bins=3)
print(bins)
# 0    (17.986, 38.667]
# 1    (17.986, 38.667]
# 2    (38.667, 59.333]
# 3    (59.333, 80.0]
# 4    (59.333, 80.0]
# Categories (3, interval[float64, right]): [(17.986, 38.667] < (38.667, 59.333] < (59.333, 80.0]]

自定义边界与标签:

bins = pd.cut(
    ages,
    bins=[0, 18, 35, 60, 100],
    labels=['未成年', '青年', '中年', '老年'],
    ordered=True
)
print(bins)
# 0    青年
# 1    青年
# 2    中年
# 3    老年
# 4    老年
# Categories (4, object): ['未成年' < '青年' < '中年' < '老年']

pd.qcut()

等频分箱(按分位数):

data = pd.Series(range(100))
q = pd.qcut(data, q=4)
print(q.value_counts())

5. 从其他类型转换

# 数值列转分类
s = pd.Series([1, 2, 2, 3])
s.astype('category')
 
# 字符串列转分类
s = pd.Series(['a', 'b', 'a'])
s.astype('category')
 
# 布尔列转分类
s = pd.Series([True, False, True])
s.astype('category')

从 DataFrame 的多个列创建

df = pd.DataFrame({'A': ['x', 'y'], 'B': [1, 2]})
df = df.astype({'A': 'category', 'B': 'category'})

6. 从 records / list 创建

# 从字典列表
records = [{'等级': '高', '城市': '北京'}, {'等级': '低', '城市': '上海'}]
df = pd.DataFrame(records).astype({'等级': 'category', '城市': 'category'})

小结

  • pd.Categorical() 直接创建
  • CategoricalDtype 定义可复用类型
  • astype('category') 快速转换
  • pd.cut() / pd.qcut() 数值分箱
  • 多种来源均可转换