本节前言

正确的数据类型是高效分析的前提。本节涵盖强制转换、自动推断以及数值 / 时间 / 分类转换工具。

1. astype()

强制转换数据类型。

df = pd.DataFrame({'A': ['1', '2', '3'], 'B': [1.5, 2.5, 3.5]})
 
df['A'].astype(int)
# 0    1
# 1    2
# 2    3
 
df['A'].astype('int64')
df['B'].astype('int32')
 
# 每列不同 dtype
df.astype({'A': int, 'B': 'float32'})
 
# 转换为可空类型
df['A'].astype('Int64')      # 可空整数
df['A'].astype('string')     # 字符串类型
df['A'].astype('category')   # 分类类型

注意

  • astype(int) 不能处理 NaN,需使用 astype('Int64')。
  • astype 默认返回新对象,不修改原数据。

2. convert_dtypes()

自动推断为最佳 dtype(倾向使用可空扩展类型)。

df2 = pd.DataFrame({
    'A': [1, 2, None],       # 自动 → Int64
    'B': [1.5, None, 3.5],   # 自动 → Float64
    'C': ['x', 'y', None],   # 自动 → string
    'D': [True, False, None] # 自动 → boolean
})
 
df2.convert_dtypes()
原 dtype转换后
int64 含缺失Int64
float64 含缺失Float64
object 字符串string
bool 含缺失boolean

3. infer_objects()

将 object 类型尽量推断为更具体的类型,但不会改变缺失值。

df3 = pd.DataFrame({'A': [1, 2, 3]}, dtype=object)
df3.dtypes          # object
df3.infer_objects().dtypes   # int64
方法作用
astype()显式指定目标 dtype
convert_dtypes()自动推断为最合适的可空 dtype
infer_objects()在保留数据不变的前提下放宽 object 推断

4. pd.to_numeric()

将参数转换为数值类型。

s = pd.Series(['1', '2.5', '3.0'])
pd.to_numeric(s)
# 0    1.0
# 1    2.5
# 2    3.0
 
# errors='coerce':无法转换的变为 NaN
pd.to_numeric(pd.Series(['1', 'abc']), errors='coerce')
# 0    1.0
# 1    NaN
 
# 转换为无符号 / 可空类型
pd.to_numeric(s, dtype='uint8')  # pandas 2.x 支持
pd.to_numeric(s, downcast='integer')

5. pd.to_datetime()

将参数转换为时间戳类型。

# 字符串转时间
pd.to_datetime('2024-01-15')
 
# Series 批量转换
pd.to_datetime(pd.Series(['2024-01-01', '2024-01-02']))
 
# 统一格式
pd.to_datetime(['15/01/2024', '16/01/2024'], format='%d/%m/%Y')
 
# 时间戳(单位秒)
pd.to_datetime([1700000000], unit='s')
 
# errors='coerce' 容错
pd.to_datetime(['2024-01-01', 'invalid'], errors='coerce')
 
# 推断格式
pd.to_datetime(['2024-01-01', '2024-01-02'], infer_datetime_format=True)

format 推荐

大数据场景下务必指定 format 参数,可将解析速度提升数倍。

6. pd.to_timedelta()

将参数转换为时间差类型。

pd.to_timedelta('1 day')
# Timedelta('1 days 00:00:00')
 
pd.to_timedelta(['1 days', '2 hours', '3 minutes'])
pd.to_timedelta(10, unit='hours')
# Timedelta('10 hours 00:00:00')
 
pd.to_timedelta(np.arange(3), unit='days')

7. pd.Categorical()

创建分类数据类型。

cat = pd.Categorical(
    ['a', 'b', 'c', 'a'],
    categories=['a', 'b', 'c', 'd'],
    ordered=True
)
 
cat.categories   # Index(['a', 'b', 'c', 'd'], dtype='object')
cat.codes        # array([0, 1, 2, 0], dtype=int8)
cat.ordered      # True
 
# 或通过 astype 转换
s = pd.Series(['low', 'high', 'medium']).astype('category')

相关链接

Categorical 的更多操作见 八、分类数据(如大纲扩展笔记)。

8. pd.cut() / pd.qcut()

将连续数值离散化为区间。

pd.cut()

按等宽区间切分。

ages = pd.Series([18, 25, 30, 45, 60, 70])
 
pd.cut(ages, bins=3)
# 0   (17.958, 35.333]
# 1   (17.958, 35.333]
# 2   (17.958, 35.333]
# 3   (35.333, 52.667]
# 4   (52.667, 70.0]
# 5   (52.667, 70.0]
 
# 自定义边界
pd.cut(ages, bins=[0, 18, 35, 60, 100], labels=['少年', '青年', '中年', '老年'])

pd.qcut()

按分位数等频切分。

pd.qcut(ages, q=4)   # 四等分位
pd.qcut(ages, q=[0, .25, .5, .75, 1.])
函数切分方式适用场景
pd.cut()等宽切分已知边界、均匀分布数据
pd.qcut()等频切分偏态分布、需要每箱样本量相同

🔗 相关笔记