dtypes(数据类型)是 pandas 存储数据的内部格式。pandas 在 NumPy 类型基础上扩展了可空类型和自定义扩展类型,以增强缺失值支持与表达能力。
1. 数值型
整数类型
| dtype | 说明 |
|---|
int8 | 8 位有符号整数,范围 -128 ~ 127 |
int16 | 16 位有符号整数 |
int32 | 32 位有符号整数 |
int64 | 64 位有符号整数(默认) |
uint8 | 8 位无符号整数,范围 0 ~ 255 |
uint16 | 16 位无符号整数 |
uint32 | 32 位无符号整数 |
uint64 | 64 位无符号整数 |
import numpy as np
import pandas as pd
s = pd.Series([1, 2, 3], dtype='int8')
print(s.dtype) # int8
浮点类型
| dtype | 说明 |
|---|
float16 | 半精度浮点(内存小,精度低) |
float32 | 单精度浮点 |
float64 | 双精度浮点(默认) |
s = pd.Series([1.5, 2.5], dtype='float32')
print(s.dtype) # float32
复数类型
| dtype | 说明 |
|---|
complex64 | 实部、虚部均为 32 位浮点 |
complex128 | 实部、虚部均为 64 位浮点 |
s = pd.Series([1+2j, 3+4j], dtype='complex64')
print(s.dtype) # complex64
2. 布尔型
| dtype | 说明 |
|---|
bool | 布尔值(True / False),缺失值以 NaN 或 pd.NA 表示 |
s = pd.Series([True, False, None], dtype='bool')
print(s) # 含缺失时可能自动转换为 object
3. 对象型
| dtype | 说明 |
|---|
object | Python 对象,可容纳任意类型(字符串、混合类型、自定义对象) |
string | pandas 的 StringDtype,专用于字符串存储,支持缺失值 pd.NA |
# object 类型
s1 = pd.Series(['a', 1, None], dtype='object')
# string 类型(推荐用于纯文本)
s2 = pd.Series(['a', 'b', None], dtype='string')
print(s2.dtype) # string
4. 时间型
| dtype | 说明 |
|---|
datetime64[ns] | 时间戳,纳秒精度,无时区 |
datetime64[ns, tz] | 带时区的时间戳,如 datetime64[ns, Asia/Shanghai] |
timedelta64[ns] | 时间差,纳秒精度 |
period[D] | 周期类型,D 可替换为 M、Q、Y、H 等频率 |
# 无时区时间戳
s1 = pd.Series(pd.date_range('2024-01-01', periods=3), dtype='datetime64[ns]')
# 带时区时间戳
s2 = pd.Series(pd.date_range('2024-01-01', periods=3, tz='Asia/Shanghai'),
dtype='datetime64[ns, Asia/Shanghai]')
# 时间差
s3 = pd.Series(pd.to_timedelta(['1 day', '2 days']), dtype='timedelta64[ns]')
# 周期
s4 = pd.Series(pd.period_range('2024-01', periods=3, freq='M'), dtype='period[M]')
print(s4.dtype) # period[M]
5. 可空扩展类型
可空类型以 pd.NA 表示缺失值,解决了 NaN 在整数和布尔列中强制转换的问题。
可空整数
| dtype | 说明 |
|---|
Int8Dtype | 可空的 8 位整数 |
Int16Dtype | 可空的 16 位整数 |
Int32Dtype | 可空的 32 位整数 |
Int64Dtype | 可空的 64 位整数 |
UInt8Dtype | 可空的无符号 8 位整数 |
UInt16Dtype | 可空的无符号 16 位整数 |
UInt32Dtype | 可空的无符号 32 位整数 |
UInt64Dtype | 可空的无符号 64 位整数 |
s = pd.Series([1, 2, None], dtype='Int64')
print(s) # [1, 2, <NA>]
print(s.dtype) # Int64
可空浮点数
| dtype | 说明 |
|---|
Float32Dtype | 可空的单精度浮点 |
Float64Dtype | 可空的双精度浮点 |
s = pd.Series([1.2, None], dtype='Float32')
print(s.dtype) # Float32
可空布尔
| dtype | 说明 |
|---|
BooleanDtype | 可空布尔,支持 True / False / pd.NA |
s = pd.Series([True, None, False], dtype='boolean')
print(s.dtype) # boolean
字符串
| dtype | 说明 |
|---|
StringDtype | 内存友好的字符串存储,替代 object 类型 |
s = pd.Series(['a', None], dtype='string')
print(s.dtype) # string
6. 其他扩展类型
pandas 提供了丰富的扩展类型以处理特殊数据。
CategoricalDtype(分类类型)
- 适合有限的类别值,如“性别”“城市”。
- 底层使用整数编码,节省内存。
cat_type = pd.CategoricalDtype(categories=['低', '中', '高'], ordered=True)
s = pd.Series(['低', '高'], dtype=cat_type)
print(s.dtype) # category
SparseDtype(稀疏类型)
s = pd.Series([0, 0, 1, 0], dtype=pd.SparseDtype('int64', fill_value=0))
print(s.dtype) # Sparse[int64, 0]
IntervalDtype(区间类型)
idx = pd.interval_range(0, 5, freq=1)
s = pd.Series(idx, dtype='interval')
print(s.dtype) # interval
ArrowDtype(PyArrow 类型)
- 基于 Apache Arrow 的扩展类型,提供更强的类型系统与性能。
- 可通过
pd.arrow_dtype() 创建,或使用 convert_dtypes 配合 pyarrow 后端。
arrow_type = pd.arrow_dtype('int64')
s = pd.Series([1, 2, None], dtype=arrow_type)
print(s.dtype) # int64[pyarrow]
PeriodDtype(周期类型)
period_type = pd.PeriodDtype(freq='Q')
s = pd.Series(pd.period_range('2024Q1', periods=2, freq='Q'), dtype=period_type)
print(s.dtype) # period[Q-DEC]
7. dtype 相关常用操作
| 操作 | 方法 |
|---|
| 查看列类型 | df.dtypes |
| 查看单个类型 | df['col'].dtype |
| 强制转换 | df.astype('type') |
| 自动推断 | df.convert_dtypes() |
| 类型判断 | pd.api.types.is_numeric_dtype(df['col']) |
- 使用
object 类型进行数值运算很慢,应尽量转换为数值类型。
- 可空整数类型必须在字符串大写开头(如
Int64),否则会回退到普通 int64 并丢失缺失值。