Bug description:
search() (and so findall(), finditer(), sub(), split()) can miss a match that match() finds at the same
position. It happens when the pattern starts with a scoped type-flag group, (?a:...) or (?u:...), whose flag differs
from the pattern's own:
>>> import re
>>> p = re.compile(r'(?u:\w)', re.ASCII)
>>> p.match('é')
<re.Match object; span=(0, 1), match='é'>
>>> print(p.search('é'))
None
>>> re.findall(r'(?u:\w)', 'café', re.ASCII)
['c', 'a', 'f']
>>> re.match(r'(?a:\W)', 'ß')
<re.Match object; span=(0, 1), match='ß'>
>>> print(re.search(r'(?a:\W)', 'ß'))
None
>>> re.findall(r'(?a:\W)', 'Straße!')
['!']
>>> re.findall(r'(?a:\W)|x', 'Straße!') # an alternative disables the prefilter, and the result changes
['ß', '!']
Cause. _compile_info() builds the "charset" prefilter that SRE(search) uses to skip start positions.
_get_charset_prefix() looks through leading groups and combines their flags (Lib/re/_optimizer.py:331 on main at
a57d165). It returns only the charset, though, and _compile_info() compiles that
charset with the pattern's own flags (:428). A category escape inside the group (\w, \W, \d, \D, \s, \S)
therefore gets the pattern's character type in the prefilter and the group's type in the pattern body. re.DEBUG shows
the mismatch (3.14.5):
>>> re.compile(r'(?u:\w)', re.ASCII | re.DEBUG)
SUBPATTERN None 32 0
IN
CATEGORY CATEGORY_WORD
0. INFO 7 0b100 1 1 (to 8)
in
5. CATEGORY WORD <- the prefilter: ASCII \w
7. FAILURE
8: IN 4 (to 13)
10. CATEGORY UNI_WORD <- the pattern: Unicode \w
12. FAILURE
13: SUCCESS
Matches are lost wherever the prefilter's set is narrower than the body's. That covers a Unicode class inside (?u:...)
in an ASCII pattern, and a negated class (\W, \D, \S) inside (?a:...) in a default str pattern. Where the
prefilter is wider ((?a:\w) in a str pattern), it only costs speed. Explicit sets such as [a-z] do not depend on the
type flags and are unaffected. For bytes patterns, (?L:...) goes through the same code, but I could not test it: no
single-byte locale was available.
Since 3.7. The flag combination in _get_charset_prefix() came with bpo-31690 (#3885, 3557b05), which made a,
L and u usable as scoped flags. The same lines are in Lib/re/_compiler.py:481-523 / :584 on the 3.14 and 3.15
branches.
Reproduced on: 3.12.3, 3.13.15 (the official python:3.13-slim image), 3.14.5 (python-build-standalone), all
aarch64 Linux. A 10-row probe (2 controls, 8 affected shapes) gives "search() misses what match() finds" in 8 of 10 rows
on each. On main (built from 3330712, 2026-09-28), the regression test below fails the same way: re.search(r'(?u:\w)', '\xe9', re.ASCII) returns None. In main's generated code, (?u:\w) under re.ASCII gets CATEGORY WORD in the INFO
block and CATEGORY UNI_WORD in the body, and (?a:\W) in a str pattern gets CATEGORY UNI_NOT_WORD against
CATEGORY NOT_WORD.
Suggested fix
Have _get_charset_prefix() return None when a leading group changes the type flags. The search then falls back to
the general case, which is always correct:
--- a/Lib/re/_optimizer.py
+++ b/Lib/re/_optimizer.py
@@ -321,6 +321,7 @@
return prefix, prefix_skip, False
def _get_charset_prefix(pattern, flags):
+ entry_flags = flags
while True:
if not pattern.data:
return None
@@ -331,6 +332,11 @@
flags = _combine_flags(flags, add_flags, del_flags)
if flags & SRE_FLAG_IGNORECASE and flags & SRE_FLAG_LOCALE:
return None
+ if (flags ^ entry_flags) & _parser.TYPE_FLAGS:
+ # _compile_info() compiles the charset with the pattern's flags, so a
+ # class such as \w found inside (?u:...) or (?a:...) would get the
+ # wrong character type.
+ return None
iscased = _get_iscased(flags)
if op is LITERAL:
The same hunks apply to Lib/re/_compiler.py on 3.14 and 3.15. An alternative that keeps the prefilter is to return the
effective flags along with the charset and have _compile_info() compile it with them. That touches every return in
the function, and these patterns are rare, so the smaller change seemed better. A regression test:
def test_charset_prefix_scoped_type_flags(self):
# The search() prefilter must use the type flags of the group it was taken from.
self.assertEqual(re.search(r'(?u:\w)', '\xe9', re.ASCII).group(), '\xe9')
self.assertEqual(re.search(r'(?u:\d)', '\u0663', re.ASCII).group(), '\u0663')
self.assertEqual(re.search(r'(?a:\W)', '\xdf').group(), '\xdf')
self.assertEqual(re.search(r'(?a:\S)', '\u2003').group(), '\u2003')
self.assertEqual(re.findall(r'(?a:\W)', 'Stra\xdfe!'), ['\xdf', '!'])
Checked on 3.14.5 by putting a copy of Lib/re with the patched _compiler.py first on PYTHONPATH (the run printed
re._compiler.__file__ to confirm which copy it used):
- The test above fails on the unpatched copy, with
AttributeError: search() returned None. It passes on the
patched copy.
- The 10-row probe: from 8 of 10 rows wrong to 0 of 10.
test_re from the v3.14.5 tag: 165 ran, 0 failures, 3 skipped, both unpatched and patched. test_re does not
currently exercise this bug.
- A randomized differential. 90,000 generated patterns (3 seeds × 30,000: lookaround,
\b and anchor leads, scoped
a/u/L/i groups, categories, sets, alternations, repeats; str and bytes), of which 82,891 compiled, each run on
5 random subjects (414,455 cases). The check: search(s, pos) must equal the first i >= pos where match(s, i)
succeeds. Unpatched: 8 cases broke that rule, all with a scoped type-flag group. Patched: 0.
- On main (3330712), with the fix applied to
Lib/re/_optimizer.py: the regression test fails before and passes
after; test_re 171 ran, 0 failures, 4 skipped; and the same search-vs-match rule over 320,000 generated cases
(2 seeds; lookaround / \b / anchor / \s+ leads, nested a/u/i groups, str patterns) gives 1,338 violations
with main's own _optimizer.py and 0 with the fix.
CPython versions tested on:
3.12, 3.13, 3.14, CPython main branch
Operating systems tested on:
Linux
AI assistance was used to investigate this and to draft this report.
Linked PRs
Bug description:
search()(and sofindall(),finditer(),sub(),split()) can miss a match thatmatch()finds at the sameposition. It happens when the pattern starts with a scoped type-flag group,
(?a:...)or(?u:...), whose flag differsfrom the pattern's own:
Cause.
_compile_info()builds the "charset" prefilter thatSRE(search)uses to skip start positions._get_charset_prefix()looks through leading groups and combines their flags (Lib/re/_optimizer.py:331on main ata57d165). It returns only the charset, though, and
_compile_info()compiles thatcharset with the pattern's own flags (
:428). A category escape inside the group (\w,\W,\d,\D,\s,\S)therefore gets the pattern's character type in the prefilter and the group's type in the pattern body.
re.DEBUGshowsthe mismatch (3.14.5):
Matches are lost wherever the prefilter's set is narrower than the body's. That covers a Unicode class inside
(?u:...)in an ASCII pattern, and a negated class (
\W,\D,\S) inside(?a:...)in a default str pattern. Where theprefilter is wider (
(?a:\w)in a str pattern), it only costs speed. Explicit sets such as[a-z]do not depend on thetype flags and are unaffected. For bytes patterns,
(?L:...)goes through the same code, but I could not test it: nosingle-byte locale was available.
Since 3.7. The flag combination in
_get_charset_prefix()came with bpo-31690 (#3885, 3557b05), which madea,Landuusable as scoped flags. The same lines are inLib/re/_compiler.py:481-523/:584on the 3.14 and 3.15branches.
Reproduced on: 3.12.3, 3.13.15 (the official
python:3.13-slimimage), 3.14.5 (python-build-standalone), allaarch64 Linux. A 10-row probe (2 controls, 8 affected shapes) gives "search() misses what match() finds" in 8 of 10 rows
on each. On main (built from 3330712, 2026-09-28), the regression test below fails the same way:
re.search(r'(?u:\w)', '\xe9', re.ASCII)returnsNone. In main's generated code,(?u:\w)underre.ASCIIgetsCATEGORY WORDin the INFOblock and
CATEGORY UNI_WORDin the body, and(?a:\W)in a str pattern getsCATEGORY UNI_NOT_WORDagainstCATEGORY NOT_WORD.Suggested fix
Have
_get_charset_prefix()returnNonewhen a leading group changes the type flags. The search then falls back tothe general case, which is always correct:
The same hunks apply to
Lib/re/_compiler.pyon 3.14 and 3.15. An alternative that keeps the prefilter is to return theeffective flags along with the charset and have
_compile_info()compile it with them. That touches everyreturninthe function, and these patterns are rare, so the smaller change seemed better. A regression test:
Checked on 3.14.5 by putting a copy of
Lib/rewith the patched_compiler.pyfirst onPYTHONPATH(the run printedre._compiler.__file__to confirm which copy it used):AttributeError:search()returnedNone. It passes on thepatched copy.
test_refrom the v3.14.5 tag: 165 ran, 0 failures, 3 skipped, both unpatched and patched.test_redoes notcurrently exercise this bug.
\band anchor leads, scopeda/u/L/igroups, categories, sets, alternations, repeats; str and bytes), of which 82,891 compiled, each run on5 random subjects (414,455 cases). The check:
search(s, pos)must equal the firsti >= poswherematch(s, i)succeeds. Unpatched: 8 cases broke that rule, all with a scoped type-flag group. Patched: 0.
Lib/re/_optimizer.py: the regression test fails before and passesafter;
test_re171 ran, 0 failures, 4 skipped; and the same search-vs-match rule over 320,000 generated cases(2 seeds; lookaround /
\b/ anchor /\s+leads, nesteda/u/igroups, str patterns) gives 1,338 violationswith main's own
_optimizer.pyand 0 with the fix.CPython versions tested on:
3.12, 3.13, 3.14, CPython main branch
Operating systems tested on:
Linux
AI assistance was used to investigate this and to draft this report.
Linked PRs