pandas 的 .str 访问器封装了 re 模块的核心能力,提供提取、匹配、搜索与替换方法,可将正则表达式应用于整个 Series。


1. 正则基础回顾

模式说明示例
.任意字符(除换行)a.c 匹配 abc、a1c
\d数字\d+ 匹配连续数字
\w单词字符(字母、数字、下划线)\w+ 匹配单词
\s空白\s+ 匹配连续空白
^字符串开头^abc
$字符串结尾abc$
[abc]字符集[a-zA-Z]
(a|b)分组或cat|dog
*0 次或多次ab*
+1 次或多次ab+
?0 次或 1 次ab?
{n,m}次数范围\d{3,5}
(?P<name>...)命名捕获组(?P<year>\d{4})

2. 提取:str.extract()

extract() 使用正则表达式的捕获组提取子串,返回 DataFrame,每个捕获组一列。

单捕获组

s = pd.Series(['订单号 A1001', '订单号 B2002', '无订单'])
 
s.str.extract(r'(A\d+)')
#        0
# 0   A1001
# 1     NaN
# 2     NaN

多捕获组

s.str.extract(r'([A-Z])(\d+)')
#       0     1
# 0    A  1001
# 1    B  2002
# 2  NaN   NaN

命名捕获组

s.str.extract(r'(?P<字母>[A-Z])(?P<数字>\d+)')
#   字母    数字
# 0    A   1001
# 1    B   2002
# 2  NaN    NaN

列名自动生成

  • 无名捕获组:列名为 0、1、2 …
  • 命名捕获组:列名为名称
  • expand 参数:若只有一个捕获组,设置 expand=False 可返回 Series
s.str.extract(r'([A-Z])(\d+)', expand=True)   # DataFrame
s.str.extract(r'[A-Z](\d+)', expand=False)    # Series

3. 提取全部:str.extractall()

extractall() 找出所有匹配项,返回 MultiIndex DataFrame(第一级为原索引,第二级为匹配次数)。

s = pd.Series(['a1,b2,c3', 'x9'])
 
s.str.extractall(r'([a-z])(\d+)')
#           0  1
#   match
# 0 0       a  1
#   1       b  2
#   2       c  3
# 1 0       x  9
# 提取所有数字并堆叠
s.str.extractall(r'(\d+)')
#           0
#   match
# 0 0       1
#   1       2
#   2       3
# 1 0       9

extractall 返回结构

  • 索引为 (原索引, 匹配序号) 的 MultiIndex
  • 若需转换为普通 DataFrame,使用 reset_index(drop=True)

4. 搜索全部:str.findall()

findall() 返回所有匹配项的列表。

s = pd.Series(['a1,b2', 'x9y8', 'no-match'])
 
s.str.findall(r'[a-z]\d')
# 0    [a1, b2]
# 1       [x9, y8]
# 2            []

带捕获组

s.str.findall(r'([a-z])(\d)')
# 0    [(a, 1), (b, 2)]
# 1       [(x, 9), (y, 8)]
# 2                  []

extractall vs findall

  • extractall:返回结构化的 DataFrame,便于后续分析
  • findall:返回 Python 列表,适合进一步自定义处理

5. 匹配判断:str.match() 与 str.fullmatch()

s = pd.Series(['abc123', 'xyz789', 'abc456'])
 
# match:从头匹配
s.str.match('abc')
# 0     True
# 1    False
# 2     True
 
# fullmatch:完全匹配
s.str.fullmatch(r'[a-z]+\d+')
# 0     True
# 1     True
# 2     True

6. 计数:str.count()

s = pd.Series(['apple orange apple', 'banana apple'])
 
s.str.count('apple')
# 0    2
# 1    1

7. 正则替换

s = pd.Series(['Phone: 123-456-7890', 'Tel: 987-654-3210'])
 
# 掩码替换:隐藏电话号码
s.str.replace(r'\d{3}-\d{3}-\d{4}', '***-***-****', regex=True)
# 0    Phone: ***-***-****
# 1    Tel: ***-***-****

使用捕获组重排

dates = pd.Series(['2024-01-15', '2025-06-30'])
dates.str.replace(r'(\d{4})-(\d{2})-(\d{2})', r'\3/\2/\1', regex=True)
# 0    15/01/2024
# 1    30/06/2025

使用函数动态替换

import re
 
s = pd.Series(['价格 100 元', '价格 250 元'])
s.str.replace(r'\d+', lambda m: f"${int(m.group()) * 1.1:.2f}", regex=True)
# 0    价格 $110.00 元
# 1    价格 $275.00 元

8. 正则标志位

通过 flags 参数传入 re 模块的常量组合:

import re
 
s = pd.Series(['Apple', 'banana', 'Cherry'])
 
# 忽略大小写
s.str.contains('apple', flags=re.IGNORECASE)
 
# 多行模式(^ 和 $ 匹配每一行)
text = pd.Series(['line1\nline2', 'single'])
text.str.extract(r'^(\w+)$', flags=re.MULTILINE)

常用标志:

标志作用
re.IGNORECASE忽略大小写
re.MULTILINE^ / $ 匹配每行
re.DOTALL. 匹配换行符
re.VERBOSE忽略表达式中的空白与注释

9. 常见正则示例汇总

需求正则
提取邮箱([\w.+-]+@[\w-]+\.[\w.-]+)
提取手机号(1[3-9]\d{9})
提取日期(YYYY-MM-DD)(\d{4}-\d{2}-\d{2})
提取中文([\u4e00-\u9fa5]+)
提取 URL(https?://\S+)
提取金额([¥$]\d+(\.\d+)?)
提取身份证号(\d{17}[\dXx])
清理特殊字符[^a-zA-Z0-9\u4e00-\u9fa5]
s = pd.Series(['联系:zhang@test.com 电话:13812345678',
               '邮箱:li@test.org,电话:13900001111'])
 
# 提取邮箱
s.str.extract(r'([\w.+-]+@[\w-]+\.[\w.-]+)')[0]
 
# 提取手机号
s.str.extract(r'(1[3-9]\d{9})')[0]
 
# 同时提取多个字段
s.str.extract(r'(?P<邮箱>[\w.+-]+@[\w-]+\.[\w.-]+).*?(?P<电话>1[3-9]\d{9})')

小结

  • 提取:extract(第一个匹配)、extractall(全部匹配)
  • 搜索:findall(返回列表)
  • 判断:match(开头)、contains(包含)、fullmatch(完全)
  • 替换:replace 支持捕获组与函数
  • 标志:flags=re.IGNORECASE 等控制匹配行为