文本数据清洗是数据预处理的重要部分。pandas 通过
.str访问器提供向量化字符串操作,并支持正则表达式,极大简化了文本处理流程。
1. 字符串访问器 .str
Series 的 .str 访问器可以对每个元素批量调用字符串方法,自动跳过缺失值。
基本用法
import pandas as pd
s = pd.Series(['Alice', 'bob', 'CHARLIE', None])
s.str.lower()
# 0 alice
# 1 bob
# 2 charlie
# 3 None注意
.str只适用于字符串列(object 或 string 类型)- 非字符串列需先转换:
df['col'].astype('string')- 缺失值在
.str方法中自动保留为NaN,不会报错
2. 大小写转换
| 方法 | 说明 | 示例 |
|---|---|---|
str.lower() | 全部小写 | 'ABC' → 'abc' |
str.upper() | 全部大写 | 'abc' → 'ABC' |
str.title() | 每个单词首字母大写 | 'hello world' → 'Hello World' |
str.capitalize() | 首字母大写,其余小写 | 'hello WORLD' → 'Hello world' |
str.swapcase() | 大小写互换 | 'aBc' → 'AbC' |
str.casefold() | 强大小写折叠(更激进) | 'Straße' → 'strasse' |
s = pd.Series(['hello world', 'GOOD MORNING', 'Python'])
s.str.title()
# 0 Hello World
# 1 Good Morning
# 2 Python3. 去除空格
| 方法 | 说明 |
|---|---|
str.strip() | 去除两端空白(包括空格、换行、制表符) |
str.lstrip() | 去除左侧空白 |
str.rstrip() | 去除右侧空白 |
s = pd.Series([' python ', '\tData\n', ' pandas '])
s.str.strip()
# 0 python
# 1 Data
# 2 pandas
# 去除指定字符
s.str.strip(' .') # 去除空格和点号
# 链式处理
s.str.strip().str.lower()注意
strip()默认去除空白(空格、\t、\n、\r等),若要去除特定字符,传入字符串参数(其中的每个字符都会被去除)。
4. 替换与删除
str.replace()
str.replace() 进行向量化字符串替换,支持正则表达式。
s = pd.Series(['a1b2', 'c3d4', 'e5f6'])
# 普通替换
s.str.replace('1', 'X')
# 0 aXb2
# 1 c3d4
# 2 e5f6
# 正则替换
s.str.replace(r'\d', 'N', regex=True)
# 0 NaNbN
# 1 NcNdN
# 2 NeNfN
# 删除数字(替换为空字符串)
s.str.replace(r'\d', '')
# 0 ab
# 1 cd
# 2 ef
# 多个替换(使用回调函数或字典?)
# 注意:str.replace 不支持字典,但可多次链式调用str.removeprefix() / str.removesuffix()(pandas 2.x)
s = pd.Series(['prefix_123', 'prefix_456'])
s.str.removeprefix('prefix_')
# 0 123
# 1 456replace() 与 .str.replace() 的区别
| 方法 | 适用范围 | 说明 |
|---|---|---|
df.replace() | DataFrame/Series | 替换整个单元格值 |
s.str.replace() | 字符串列 | 替换字符串内部的子串 |
# 整值替换
df['城市'].replace('北京', '北京市')
# 子串替换
df['城市'].str.replace('北京', 'BEIJING')5. 拆分与合并
str.split()
s = pd.Series(['2024-01-01', '2024-02-15', '2024-03-30'])
# 按分隔符拆分,返回 Series of list
s.str.split('-')
# 0 [2024, 01, 01]
# 1 [2024, 02, 15]
# 2 [2024, 03, 30]
# 展开为 DataFrame(多列)
s.str.split('-', expand=True)
# 0 1 2
# 0 2024 01 01
# 1 2024 02 15
# 2 2024 03 30
# 限制拆分次数
s.str.split('-', n=1, expand=True)
# 0 1
# 0 2024 01-01
# 1 2024 02-15
# 2 2024 03-30str.rsplit()
# 从右向左拆分
s.str.rsplit('-', n=1, expand=True)str.cat()
# 将元素连接为单个字符串
s.str.cat(sep=' | ')
# '2024-01-01 | 2024-02-15 | 2024-03-30'
# 与其他 Series 列连接
df = pd.DataFrame({'姓': ['张', '李'], '名': ['三', '四']})
df['姓名'] = df['姓'].str.cat(df['名'], sep='')
# 姓 名 姓名
# 0 张 三 张三
# 1 李 四 李四
# 连接多列
df['姓名'] = df['姓'].str.cat(df[['姓', '名']], sep='')列合并
# 多列拼接为单列
df['全名'] = df['姓'] + df['名']
df['地址'] = df['省'] + '-' + df['市'] + '-' + df['区']6. 搜索与判断
str.contains()
str.contains() 检查字符串是否包含指定模式,返回布尔 Series。
s = pd.Series(['apple', 'banana', 'cherry', None])
s.str.contains('an')
# 0 False
# 1 True
# 2 False
# 3 NaN
# 忽略大小写
s.str.contains('APPLE', case=False)
# 使用正则
s.str.contains('^a|n$', regex=True)
# 将 NaN 视为 False
s.str.contains('an', na=False)参数:pat(模式)、case(是否区分大小写)、flags(正则标志)、na(NaN 处理)、regex(是否为正则)、regex 在 pandas 2.x 中改为 regex。
str.startswith() / str.endswith()
s.str.startswith('app')
# 0 True
# 1 False
# 2 False
s.str.endswith('ry')
# 0 False
# 1 False
# 2 True布尔筛选结合
df = pd.DataFrame({'邮箱': ['a@x.com', 'b@y.org', 'c@z.net']})
df[df['邮箱'].str.contains('@')]
df[df['邮箱'].str.endswith('.com')]7. 长度与计数
str.len()
s = pd.Series(['hello', 'world!', None])
s.str.len()
# 0 5.0
# 1 6.0
# 2 NaNstr.count()
# 计算子串出现次数
s = pd.Series(['abab', 'aabb', 'ab'])
s.str.count('a')
# 0 2
# 1 2
# 2 18. 提取与抽取
str.extract()
str.extract() 使用正则表达式的捕获组提取子串,返回 DataFrame(一列一个捕获组)。
s = pd.Series(['订单号:A1001', '订单号:B2002', '其他'])
# 提取数字
s.str.extract(r'(\d+)')
# 0
# 0 1001
# 1 2002
# 2 NaN
# 多个捕获组
s.str.extract(r'([A-Z])(\d+)')
# 0 1
# 0 A 1001
# 1 B 2002
# 2 NaN NaN
# 命名捕获组(用 ?P<name>)
s.str.extract(r'(?P<字母>[A-Z])(?P<数字>\d+)')
# 字母 数字
# 0 A 1001
# 1 B 2002str.extractall()
extractall() 提取所有匹配项,返回 MultiIndex DataFrame。
s = pd.Series(['a1,b2,c3', 'x9'])
s.str.extractall(r'([a-z])(\d+)')
# 0 1
# match
# 0 0 a 1
# 1 b 2
# 2 c 3
# 1 0 x 9str.findall()
findall() 返回所有匹配项的列表。
s.str.findall(r'([a-z])(\d+)')
# 0 [(a, 1), (b, 2), (c, 3)]
# 1 [(x, 9)]extract vs findall
extract提取第一个匹配项,返回 DataFrameextractall提取所有匹配项,返回 MultiIndex DataFramefindall返回列表,适合后续自定义处理
9. 填充与对齐
str.pad()
s = pd.Series(['a', 'bb', 'ccc'])
s.str.pad(width=5, side='left', fillchar='0')
# 0 0000a
# 1 000bb
# 2 00cccstr.center()
s.str.center(width=5, fillchar='-')
# 0 --a--
# 1 -bb--
# 2 -ccc-str.zfill()
# 左侧补零
pd.Series(['1', '23', '456']).str.zfill(5)
# 0 00001
# 1 00023
# 2 0045610. 分割后展开与切片
str.slice()
s = pd.Series(['abcdef', 'ghijkl'])
s.str.slice(start=0, stop=3) # 'abc', 'ghi'
s.str.slice(start=-2) # 'ef', 'kl'str.get()
# 获取列表中的第 n 个元素
s = pd.Series(['a-b-c', 'd-e-f'])
s.str.split('-').str.get(1)
# 0 b
# 1 e11. 正则表达式常用模式
| 模式 | 含义 |
|---|---|
\d | 数字 [0-9] |
\D | 非数字 |
\w | 单词字符(字母、数字、下划线) |
\W | 非单词字符 |
\s | 空白(空格、\t、\n 等) |
\S | 非空白 |
^ | 字符串开头(在 multiline 模式下为行首) |
$ | 字符串结尾 |
[abc] | 字符集 |
[a-z0-9] | 范围字符集 |
(a|b) | 或 |
* | 0 次或多次 |
+ | 1 次或多次 |
? | 0 次或 1 次 |
{n} | 恰好 n 次 |
{n,} | 至少 n 次 |
(?P<name>...) | 命名捕获组 |
示例:常见清洗
s = pd.Series([' 张三 ', '李四-北京', '王五 2024'])
# 去除空格并转为标准格式
s.str.strip().str.replace(r'\s+', '', regex=True)
# 提取中文姓名
s.str.extract(r'([\u4e00-\u9fa5]+)')
# 提取日期
logs = pd.Series(['2024-01-01 10:30:00 INFO', '2024-02-15 12:00:00 ERROR'])
logs.str.extract(r'(\d{4}-\d{2}-\d{2})')12. 综合清洗实践
df = pd.DataFrame({'原始文本': [
' Alice, 25, Beijing ',
'Bob;30;Shanghai',
'Carol 35 Guangzhou',
None
]})
# 1. 去除首尾空格
df['文本'] = df['原始文本'].str.strip()
# 2. 统一分隔符(, 和 ; 统一为 |)
df['文本'] = df['文本'].str.replace(r'[,;]', '|', regex=True)
# 3. 按分隔符展开为三列
split = df['文本'].str.split('|', expand=True)
df[['姓名', '年龄', '城市']] = split
# 4. 清洗各列
df['姓名'] = df['姓名'].str.strip()
df['年龄'] = df['年龄'].str.strip()
df['城市'] = df['城市'].str.strip()
# 5. 类型转换
df['年龄'] = pd.to_numeric(df['年龄'], errors='coerce')
# 6. 城市规范化(去空格、统一大小写)
df['城市'] = df['城市'].str.strip().str.title()小结
.str访问器提供了所有常用字符串方法的向量化版本- 替换用
str.replace(支持正则),拆用str.split,提取用str.extract- 布尔搜索
str.contains/startswith/endswith常与df.loc配合筛选- 正则表达式是文本清洗的利器,掌握常用模式即可应对大多数场景