本节前言
正确的数据类型是高效分析的前提。本节涵盖强制转换、自动推断以及数值 / 时间 / 分类转换工具。
1. astype()
强制转换数据类型。
df = pd.DataFrame({'A': ['1', '2', '3'], 'B': [1.5, 2.5, 3.5]})
df['A'].astype(int)
# 0 1
# 1 2
# 2 3
df['A'].astype('int64')
df['B'].astype('int32')
# 每列不同 dtype
df.astype({'A': int, 'B': 'float32'})
# 转换为可空类型
df['A'].astype('Int64') # 可空整数
df['A'].astype('string') # 字符串类型
df['A'].astype('category') # 分类类型注意
astype(int)不能处理 NaN,需使用astype('Int64')。astype默认返回新对象,不修改原数据。
2. convert_dtypes()
自动推断为最佳 dtype(倾向使用可空扩展类型)。
df2 = pd.DataFrame({
'A': [1, 2, None], # 自动 → Int64
'B': [1.5, None, 3.5], # 自动 → Float64
'C': ['x', 'y', None], # 自动 → string
'D': [True, False, None] # 自动 → boolean
})
df2.convert_dtypes()| 原 dtype | 转换后 |
|---|---|
| int64 含缺失 | Int64 |
| float64 含缺失 | Float64 |
| object 字符串 | string |
| bool 含缺失 | boolean |
3. infer_objects()
将 object 类型尽量推断为更具体的类型,但不会改变缺失值。
df3 = pd.DataFrame({'A': [1, 2, 3]}, dtype=object)
df3.dtypes # object
df3.infer_objects().dtypes # int64| 方法 | 作用 |
|---|---|
astype() | 显式指定目标 dtype |
convert_dtypes() | 自动推断为最合适的可空 dtype |
infer_objects() | 在保留数据不变的前提下放宽 object 推断 |
4. pd.to_numeric()
将参数转换为数值类型。
s = pd.Series(['1', '2.5', '3.0'])
pd.to_numeric(s)
# 0 1.0
# 1 2.5
# 2 3.0
# errors='coerce':无法转换的变为 NaN
pd.to_numeric(pd.Series(['1', 'abc']), errors='coerce')
# 0 1.0
# 1 NaN
# 转换为无符号 / 可空类型
pd.to_numeric(s, dtype='uint8') # pandas 2.x 支持
pd.to_numeric(s, downcast='integer')5. pd.to_datetime()
将参数转换为时间戳类型。
# 字符串转时间
pd.to_datetime('2024-01-15')
# Series 批量转换
pd.to_datetime(pd.Series(['2024-01-01', '2024-01-02']))
# 统一格式
pd.to_datetime(['15/01/2024', '16/01/2024'], format='%d/%m/%Y')
# 时间戳(单位秒)
pd.to_datetime([1700000000], unit='s')
# errors='coerce' 容错
pd.to_datetime(['2024-01-01', 'invalid'], errors='coerce')
# 推断格式
pd.to_datetime(['2024-01-01', '2024-01-02'], infer_datetime_format=True)format 推荐
大数据场景下务必指定
format参数,可将解析速度提升数倍。
6. pd.to_timedelta()
将参数转换为时间差类型。
pd.to_timedelta('1 day')
# Timedelta('1 days 00:00:00')
pd.to_timedelta(['1 days', '2 hours', '3 minutes'])
pd.to_timedelta(10, unit='hours')
# Timedelta('10 hours 00:00:00')
pd.to_timedelta(np.arange(3), unit='days')7. pd.Categorical()
创建分类数据类型。
cat = pd.Categorical(
['a', 'b', 'c', 'a'],
categories=['a', 'b', 'c', 'd'],
ordered=True
)
cat.categories # Index(['a', 'b', 'c', 'd'], dtype='object')
cat.codes # array([0, 1, 2, 0], dtype=int8)
cat.ordered # True
# 或通过 astype 转换
s = pd.Series(['low', 'high', 'medium']).astype('category')相关链接
Categorical 的更多操作见 八、分类数据(如大纲扩展笔记)。
8. pd.cut() / pd.qcut()
将连续数值离散化为区间。
pd.cut()
按等宽区间切分。
ages = pd.Series([18, 25, 30, 45, 60, 70])
pd.cut(ages, bins=3)
# 0 (17.958, 35.333]
# 1 (17.958, 35.333]
# 2 (17.958, 35.333]
# 3 (35.333, 52.667]
# 4 (52.667, 70.0]
# 5 (52.667, 70.0]
# 自定义边界
pd.cut(ages, bins=[0, 18, 35, 60, 100], labels=['少年', '青年', '中年', '老年'])pd.qcut()
按分位数等频切分。
pd.qcut(ages, q=4) # 四等分位
pd.qcut(ages, q=[0, .25, .5, .75, 1.])| 函数 | 切分方式 | 适用场景 |
|---|---|---|
pd.cut() | 等宽切分 | 已知边界、均匀分布数据 |
pd.qcut() | 等频切分 | 偏态分布、需要每箱样本量相同 |