pandas 的
.str访问器封装了re模块的核心能力,提供提取、匹配、搜索与替换方法,可将正则表达式应用于整个 Series。
1. 正则基础回顾
| 模式 | 说明 | 示例 |
|---|---|---|
. | 任意字符(除换行) | a.c 匹配 abc、a1c |
\d | 数字 | \d+ 匹配连续数字 |
\w | 单词字符(字母、数字、下划线) | \w+ 匹配单词 |
\s | 空白 | \s+ 匹配连续空白 |
^ | 字符串开头 | ^abc |
$ | 字符串结尾 | abc$ |
[abc] | 字符集 | [a-zA-Z] |
(a|b) | 分组或 | cat|dog |
* | 0 次或多次 | ab* |
+ | 1 次或多次 | ab+ |
? | 0 次或 1 次 | ab? |
{n,m} | 次数范围 | \d{3,5} |
(?P<name>...) | 命名捕获组 | (?P<year>\d{4}) |
2. 提取:str.extract()
extract() 使用正则表达式的捕获组提取子串,返回 DataFrame,每个捕获组一列。
单捕获组
s = pd.Series(['订单号 A1001', '订单号 B2002', '无订单'])
s.str.extract(r'(A\d+)')
# 0
# 0 A1001
# 1 NaN
# 2 NaN多捕获组
s.str.extract(r'([A-Z])(\d+)')
# 0 1
# 0 A 1001
# 1 B 2002
# 2 NaN NaN命名捕获组
s.str.extract(r'(?P<字母>[A-Z])(?P<数字>\d+)')
# 字母 数字
# 0 A 1001
# 1 B 2002
# 2 NaN NaN列名自动生成
- 无名捕获组:列名为
0、1、2…- 命名捕获组:列名为名称
expand参数:若只有一个捕获组,设置expand=False可返回 Series
s.str.extract(r'([A-Z])(\d+)', expand=True) # DataFrame
s.str.extract(r'[A-Z](\d+)', expand=False) # Series3. 提取全部:str.extractall()
extractall() 找出所有匹配项,返回 MultiIndex DataFrame(第一级为原索引,第二级为匹配次数)。
s = pd.Series(['a1,b2,c3', 'x9'])
s.str.extractall(r'([a-z])(\d+)')
# 0 1
# match
# 0 0 a 1
# 1 b 2
# 2 c 3
# 1 0 x 9# 提取所有数字并堆叠
s.str.extractall(r'(\d+)')
# 0
# match
# 0 0 1
# 1 2
# 2 3
# 1 0 9extractall 返回结构
- 索引为
(原索引, 匹配序号)的 MultiIndex- 若需转换为普通 DataFrame,使用
reset_index(drop=True)
4. 搜索全部:str.findall()
findall() 返回所有匹配项的列表。
s = pd.Series(['a1,b2', 'x9y8', 'no-match'])
s.str.findall(r'[a-z]\d')
# 0 [a1, b2]
# 1 [x9, y8]
# 2 []带捕获组
s.str.findall(r'([a-z])(\d)')
# 0 [(a, 1), (b, 2)]
# 1 [(x, 9), (y, 8)]
# 2 []extractall vs findall
extractall:返回结构化的 DataFrame,便于后续分析findall:返回 Python 列表,适合进一步自定义处理
5. 匹配判断:str.match() 与 str.fullmatch()
s = pd.Series(['abc123', 'xyz789', 'abc456'])
# match:从头匹配
s.str.match('abc')
# 0 True
# 1 False
# 2 True
# fullmatch:完全匹配
s.str.fullmatch(r'[a-z]+\d+')
# 0 True
# 1 True
# 2 True6. 计数:str.count()
s = pd.Series(['apple orange apple', 'banana apple'])
s.str.count('apple')
# 0 2
# 1 17. 正则替换
s = pd.Series(['Phone: 123-456-7890', 'Tel: 987-654-3210'])
# 掩码替换:隐藏电话号码
s.str.replace(r'\d{3}-\d{3}-\d{4}', '***-***-****', regex=True)
# 0 Phone: ***-***-****
# 1 Tel: ***-***-****使用捕获组重排
dates = pd.Series(['2024-01-15', '2025-06-30'])
dates.str.replace(r'(\d{4})-(\d{2})-(\d{2})', r'\3/\2/\1', regex=True)
# 0 15/01/2024
# 1 30/06/2025使用函数动态替换
import re
s = pd.Series(['价格 100 元', '价格 250 元'])
s.str.replace(r'\d+', lambda m: f"${int(m.group()) * 1.1:.2f}", regex=True)
# 0 价格 $110.00 元
# 1 价格 $275.00 元8. 正则标志位
通过 flags 参数传入 re 模块的常量组合:
import re
s = pd.Series(['Apple', 'banana', 'Cherry'])
# 忽略大小写
s.str.contains('apple', flags=re.IGNORECASE)
# 多行模式(^ 和 $ 匹配每一行)
text = pd.Series(['line1\nline2', 'single'])
text.str.extract(r'^(\w+)$', flags=re.MULTILINE)常用标志:
| 标志 | 作用 |
|---|---|
re.IGNORECASE | 忽略大小写 |
re.MULTILINE | ^ / $ 匹配每行 |
re.DOTALL | . 匹配换行符 |
re.VERBOSE | 忽略表达式中的空白与注释 |
9. 常见正则示例汇总
| 需求 | 正则 |
|---|---|
| 提取邮箱 | ([\w.+-]+@[\w-]+\.[\w.-]+) |
| 提取手机号 | (1[3-9]\d{9}) |
| 提取日期(YYYY-MM-DD) | (\d{4}-\d{2}-\d{2}) |
| 提取中文 | ([\u4e00-\u9fa5]+) |
| 提取 URL | (https?://\S+) |
| 提取金额 | ([¥$]\d+(\.\d+)?) |
| 提取身份证号 | (\d{17}[\dXx]) |
| 清理特殊字符 | [^a-zA-Z0-9\u4e00-\u9fa5] |
s = pd.Series(['联系:zhang@test.com 电话:13812345678',
'邮箱:li@test.org,电话:13900001111'])
# 提取邮箱
s.str.extract(r'([\w.+-]+@[\w-]+\.[\w.-]+)')[0]
# 提取手机号
s.str.extract(r'(1[3-9]\d{9})')[0]
# 同时提取多个字段
s.str.extract(r'(?P<邮箱>[\w.+-]+@[\w-]+\.[\w.-]+).*?(?P<电话>1[3-9]\d{9})')小结
- 提取:
extract(第一个匹配)、extractall(全部匹配)- 搜索:
findall(返回列表)- 判断:
match(开头)、contains(包含)、fullmatch(完全)- 替换:
replace支持捕获组与函数- 标志:
flags=re.IGNORECASE等控制匹配行为