dtypes(数据类型)是 pandas 存储数据的内部格式。pandas 在 NumPy 类型基础上扩展了可空类型和自定义扩展类型,以增强缺失值支持与表达能力。


1. 数值型

整数类型

dtype说明
int88 位有符号整数,范围 -128 ~ 127
int1616 位有符号整数
int3232 位有符号整数
int6464 位有符号整数(默认)
uint88 位无符号整数,范围 0 ~ 255
uint1616 位无符号整数
uint3232 位无符号整数
uint6464 位无符号整数
import numpy as np
import pandas as pd
 
s = pd.Series([1, 2, 3], dtype='int8')
print(s.dtype)  # int8

浮点类型

dtype说明
float16半精度浮点(内存小,精度低)
float32单精度浮点
float64双精度浮点(默认)
s = pd.Series([1.5, 2.5], dtype='float32')
print(s.dtype)  # float32

复数类型

dtype说明
complex64实部、虚部均为 32 位浮点
complex128实部、虚部均为 64 位浮点
s = pd.Series([1+2j, 3+4j], dtype='complex64')
print(s.dtype)  # complex64

2. 布尔型

dtype说明
bool布尔值(True / False),缺失值以 NaN 或 pd.NA 表示
s = pd.Series([True, False, None], dtype='bool')
print(s)  # 含缺失时可能自动转换为 object

3. 对象型

dtype说明
objectPython 对象,可容纳任意类型(字符串、混合类型、自定义对象)
stringpandas 的 StringDtype,专用于字符串存储,支持缺失值 pd.NA
# object 类型
s1 = pd.Series(['a', 1, None], dtype='object')
 
# string 类型(推荐用于纯文本)
s2 = pd.Series(['a', 'b', None], dtype='string')
print(s2.dtype)  # string

4. 时间型

dtype说明
datetime64[ns]时间戳,纳秒精度,无时区
datetime64[ns, tz]带时区的时间戳,如 datetime64[ns, Asia/Shanghai]
timedelta64[ns]时间差,纳秒精度
period[D]周期类型,D 可替换为 M、Q、Y、H 等频率
# 无时区时间戳
s1 = pd.Series(pd.date_range('2024-01-01', periods=3), dtype='datetime64[ns]')
 
# 带时区时间戳
s2 = pd.Series(pd.date_range('2024-01-01', periods=3, tz='Asia/Shanghai'),
               dtype='datetime64[ns, Asia/Shanghai]')
 
# 时间差
s3 = pd.Series(pd.to_timedelta(['1 day', '2 days']), dtype='timedelta64[ns]')
 
# 周期
s4 = pd.Series(pd.period_range('2024-01', periods=3, freq='M'), dtype='period[M]')
print(s4.dtype)  # period[M]

5. 可空扩展类型

可空类型以 pd.NA 表示缺失值,解决了 NaN 在整数和布尔列中强制转换的问题。

可空整数

dtype说明
Int8Dtype可空的 8 位整数
Int16Dtype可空的 16 位整数
Int32Dtype可空的 32 位整数
Int64Dtype可空的 64 位整数
UInt8Dtype可空的无符号 8 位整数
UInt16Dtype可空的无符号 16 位整数
UInt32Dtype可空的无符号 32 位整数
UInt64Dtype可空的无符号 64 位整数
s = pd.Series([1, 2, None], dtype='Int64')
print(s)       # [1, 2, <NA>]
print(s.dtype) # Int64

可空浮点数

dtype说明
Float32Dtype可空的单精度浮点
Float64Dtype可空的双精度浮点
s = pd.Series([1.2, None], dtype='Float32')
print(s.dtype)  # Float32

可空布尔

dtype说明
BooleanDtype可空布尔,支持 True / False / pd.NA
s = pd.Series([True, None, False], dtype='boolean')
print(s.dtype)  # boolean

字符串

dtype说明
StringDtype内存友好的字符串存储,替代 object 类型
s = pd.Series(['a', None], dtype='string')
print(s.dtype)  # string

6. 其他扩展类型

pandas 提供了丰富的扩展类型以处理特殊数据。

CategoricalDtype(分类类型)

  • 适合有限的类别值,如“性别”“城市”。
  • 底层使用整数编码,节省内存。
cat_type = pd.CategoricalDtype(categories=['低', '中', '高'], ordered=True)
s = pd.Series(['低', '高'], dtype=cat_type)
print(s.dtype)  # category

SparseDtype(稀疏类型)

  • 针对大量零值或缺失值的数据,按“稀疏”方式存储。
s = pd.Series([0, 0, 1, 0], dtype=pd.SparseDtype('int64', fill_value=0))
print(s.dtype)  # Sparse[int64, 0]

IntervalDtype(区间类型)

  • 每个元素是一个区间,如 (0, 1]。
idx = pd.interval_range(0, 5, freq=1)
s = pd.Series(idx, dtype='interval')
print(s.dtype)  # interval

ArrowDtype(PyArrow 类型)

  • 基于 Apache Arrow 的扩展类型,提供更强的类型系统与性能。
  • 可通过 pd.arrow_dtype() 创建,或使用 convert_dtypes 配合 pyarrow 后端。
arrow_type = pd.arrow_dtype('int64')
s = pd.Series([1, 2, None], dtype=arrow_type)
print(s.dtype)  # int64[pyarrow]

PeriodDtype(周期类型)

  • 表示固定频率的时间段,如月度、季度。
period_type = pd.PeriodDtype(freq='Q')
s = pd.Series(pd.period_range('2024Q1', periods=2, freq='Q'), dtype=period_type)
print(s.dtype)  # period[Q-DEC]

7. dtype 相关常用操作

操作方法
查看列类型df.dtypes
查看单个类型df['col'].dtype
强制转换df.astype('type')
自动推断df.convert_dtypes()
类型判断pd.api.types.is_numeric_dtype(df['col'])

注意

  • 使用 object 类型进行数值运算很慢,应尽量转换为数值类型。
  • 可空整数类型必须在字符串大写开头(如 Int64),否则会回退到普通 int64 并丢失缺失值。