在数据处理全流程中,字符串操作与正则表达式是清洗、提取、校验信息的核心能力。下面系统梳理 Python 内置字符串方法、格式化、及在 pandas 中的实战用法。
一、Python 字符串核心操作
1. 基础特性
- 不可变序列,索引从 0 开始,支持切片
s[start:stop:step] - 单引号、双引号、三引号(多行)均可
- 转义符
\n\t,原始字符串r"..."不转义
2. 常用方法速查表
| 分类 | 方法 | 说明 | 示例 |
|---|---|---|---|
| 大小写转换 | upper() / lower() | 全转大写/小写 | 'Abc'.lower() → 'abc' |
capitalize() | 首字母大写,其余小写 | 'aBC'.capitalize() → 'Abc' | |
title() | 每个单词首字母大写 | 'hello world'.title() → 'Hello World' | |
swapcase() | 大小写反转 | 'AbC'.swapcase() → 'aBc' | |
| 判断 | startswith(prefix) / endswith(suffix) | 是否以某字符串开头/结尾 | 'data.csv'.endswith('.csv') |
isalpha() / isdigit() / isalnum() | 是否全为字母/数字/字母+数字 | '123'.isdigit() → True | |
isspace() | 是否全为空白字符 | ' '.isspace() → True | |
islower() / isupper() | 是否全为小写/大写 | ||
isnumeric() / isdecimal() | 是否全为数字字符(更宽泛) | 'Ⅷ'.isnumeric() → True | |
| 查找 | find(sub) / rfind(sub) | 返回索引,找不到返 -1 | 'abcabc'.find('a') → 0 |
index(sub) / rindex(sub) | 同上,找不到抛 ValueError | ||
count(sub) | 子串出现次数 | 'banana'.count('a') → 3 | |
| 替换 | replace(old, new, count) | 替换子串,count 限制次数 | 'a,b,c'.replace(',','-') → 'a-b-c' |
translate(table) | 结合 str.maketrans 进行字符级替换 | 'abc'.translate(str.maketrans('a','1')) → '1bc' | |
| 拆分与合并 | split(sep, maxsplit) | 按分隔符拆成列表,默认空白 | 'a,b,c'.split(',') → ['a','b','c'] |
rsplit(sep, maxsplit) | 从右侧拆分 | 'a,b,c'.rsplit(',',1) → ['a,b','c'] | |
splitlines(keepends) | 按行拆分 | 'a\nb'.splitlines() → ['a','b'] | |
partition(sep) | 拆成三元组 (前, 分隔符, 后) | 'a,b'.partition(',') → ('a',',','b') | |
join(iterable) | 用字符串连接可迭代对象 | ','.join(['a','b']) → 'a,b' | |
| 修剪填充 | strip(chars) / lstrip / rstrip | 去除两端指定字符(默认空白) | ' hi '.strip() → 'hi' |
center(width, fillchar) | 居中对齐填充 | 'hi'.center(5,'-') → '-hi--' | |
ljust(width, fillchar) / rjust | 左/右对齐填充 | ||
zfill(width) | 右对齐,左填充 0 | '42'.zfill(5) → '00042' | |
| 编码 | encode(encoding) | 转为 bytes | '你好'.encode('utf-8') |
3. 字符串格式化(三剑客)
① %-格式化(旧式)
name = "World"
"Hello, %s!" % name # 字符串
"pi = %.2f" % 3.14159 # 保留两位小数② str.format()
"{} {} {}".format(1,2,3) # 按位置
"{1} {0}".format("a","b") # 按索引
"name: {name}".format(name="X") # 按关键字
"pi: {:.2f}".format(3.14159) # 格式化格式规范:[[fill]align][sign][#][0][width][,][.precision][type]
常用:{:<10} 左对齐宽10,{:>10} 右对齐,{:^10} 居中,{:0>5} 左补0,{:,} 千位分隔,{:.2f} 保留两位小数。
③ f-string(推荐,Python 3.6+)
name = "Tom"
age = 25
f"{name} is {age} years old."
f"圆周率:{3.14159:.2f}"
f"二进制:{age:#b}" # 带 0b 前缀
f"日期:{datetime.now():%Y-%m-%d}"支持表达式 f"{a+b}",甚至调用函数,但不能用反斜杠。
三、Pandas 中的字符串操作(Series.str 访问器)
在数据处理中,通常对整个列批量应用字符串操作,避免循环。
前提:s = df['column'].astype('string') 或保持 object 类型。
| 方法 | 对应 Python/正则 | 示例 |
|---|---|---|
str.lower() / str.upper() | 同字符串方法 | df['name'].str.lower() |
str.strip() / str.lstrip() / str.rstrip() | 去空白/字符 | df['text'].str.strip() |
str.split(pat, expand=True) | 分割,expand=True 直接拆成多列 | df['name'].str.split(' ', expand=True) |
str.replace(pat, repl, regex=True) | 替换,可用正则 | df['phone'].str.replace(r'(\d{3})\d{4}(\d{4})', r'\1****\2', regex=True) |
str.contains(pat, na=False) | 返回布尔序列 | df['email'].str.contains(r'@') |
str.extract(pat) | 提取第一个捕获分组,返回列 | df['text'].str.extract(r'(\d+)') |
str.extractall(pat) | 提取所有匹配,多级索引 | 用于一行多次匹配 |
str.findall(pat) | 返回每个单元格匹配项组成的列表 | |
str.len() | 字符串长度 | df['comment'].str.len() |
str.cat(sep) | 连接 Series | df['first'].str.cat(df['last'], sep=' ') |
str.match(pat) | 完全等同于 re.match,从开头匹配 | |
str.fullmatch(pat) | 完全匹配 | |
str.slice(start,stop) / str[] | 切片 | df['code'].str[0:3] |
str.zfill(width) | 左补0 | df['id'].str.zfill(6) |
实战清洗流水线:
# 读取 Excel 后,清洗姓名和手机号列
df['姓名'] = df['姓名'].str.strip().str.replace(r'\s+', ' ', regex=True)
df['手机'] = df['手机'].astype('string').str.extract(r'(1[3-9]\d{9})') # 提取标准手机号
df['邮箱'] = df['邮箱'].str.lower().str.strip()
# 过滤无效邮箱
df = df[df['邮箱'].str.fullmatch(r'[\w.-]+@[\w.-]+\.\w+', na=False)]四、总结
- 字符串方法处理简单、明确的格式转换,如去除空格、大小写、拆分拼接。
- 正则表达式处理复杂模式匹配、提取和验证,是文本处理的瑞士军刀。
- Pandas .str 将两种能力向量化,使数据列清洗效率极高,且能优雅处理缺失值(
na参数)。 - 最佳实践:优先使用 Python 内置字符串方法(可读性好),需要模式匹配时用正则,在 DataFrame 中始终通过
.str访问器批量操作,并预编译高频率正则模式。
掌握这套组合拳,无论是从 Excel、数据库导入的脏数据清洗,还是日常文本分析需求,你都能快速实现。