Skip to content

re.search() misses matches when a pattern starts with a scoped (?a:...) or (?u:...) group #158374

Description

@BwL1289

Bug description:

search() (and so findall(), finditer(), sub(), split()) can miss a match that match() finds at the same
position. It happens when the pattern starts with a scoped type-flag group, (?a:...) or (?u:...), whose flag differs
from the pattern's own:

>>> import re
>>> p = re.compile(r'(?u:\w)', re.ASCII)
>>> p.match('é')
<re.Match object; span=(0, 1), match='é'>
>>> print(p.search('é'))
None
>>> re.findall(r'(?u:\w)', 'café', re.ASCII)
['c', 'a', 'f']
>>> re.match(r'(?a:\W)', 'ß')
<re.Match object; span=(0, 1), match='ß'>
>>> print(re.search(r'(?a:\W)', 'ß'))
None
>>> re.findall(r'(?a:\W)', 'Straße!')
['!']
>>> re.findall(r'(?a:\W)|x', 'Straße!')     # an alternative disables the prefilter, and the result changes
['ß', '!']

Cause. _compile_info() builds the "charset" prefilter that SRE(search) uses to skip start positions.
_get_charset_prefix() looks through leading groups and combines their flags (Lib/re/_optimizer.py:331 on main at
a57d165). It returns only the charset, though, and _compile_info() compiles that
charset with the pattern's own flags (:428). A category escape inside the group (\w, \W, \d, \D, \s, \S)
therefore gets the pattern's character type in the prefilter and the group's type in the pattern body. re.DEBUG shows
the mismatch (3.14.5):

>>> re.compile(r'(?u:\w)', re.ASCII | re.DEBUG)
SUBPATTERN None 32 0
  IN
    CATEGORY CATEGORY_WORD

 0. INFO 7 0b100 1 1 (to 8)
      in
 5.     CATEGORY WORD          <- the prefilter: ASCII \w
 7.     FAILURE
 8: IN 4 (to 13)
10.   CATEGORY UNI_WORD        <- the pattern: Unicode \w
12.   FAILURE
13: SUCCESS

Matches are lost wherever the prefilter's set is narrower than the body's. That covers a Unicode class inside (?u:...)
in an ASCII pattern, and a negated class (\W, \D, \S) inside (?a:...) in a default str pattern. Where the
prefilter is wider ((?a:\w) in a str pattern), it only costs speed. Explicit sets such as [a-z] do not depend on the
type flags and are unaffected. For bytes patterns, (?L:...) goes through the same code, but I could not test it: no
single-byte locale was available.

Since 3.7. The flag combination in _get_charset_prefix() came with bpo-31690 (#3885, 3557b05), which made a,
L and u usable as scoped flags. The same lines are in Lib/re/_compiler.py:481-523 / :584 on the 3.14 and 3.15
branches.

Reproduced on: 3.12.3, 3.13.15 (the official python:3.13-slim image), 3.14.5 (python-build-standalone), all
aarch64 Linux. A 10-row probe (2 controls, 8 affected shapes) gives "search() misses what match() finds" in 8 of 10 rows
on each. On main (built from 3330712, 2026-09-28), the regression test below fails the same way: re.search(r'(?u:\w)', '\xe9', re.ASCII) returns None. In main's generated code, (?u:\w) under re.ASCII gets CATEGORY WORD in the INFO
block and CATEGORY UNI_WORD in the body, and (?a:\W) in a str pattern gets CATEGORY UNI_NOT_WORD against
CATEGORY NOT_WORD.

Suggested fix

Have _get_charset_prefix() return None when a leading group changes the type flags. The search then falls back to
the general case, which is always correct:

--- a/Lib/re/_optimizer.py
+++ b/Lib/re/_optimizer.py
@@ -321,6 +321,7 @@
     return prefix, prefix_skip, False
 
 def _get_charset_prefix(pattern, flags):
+    entry_flags = flags
     while True:
         if not pattern.data:
             return None
@@ -331,6 +332,11 @@
         flags = _combine_flags(flags, add_flags, del_flags)
         if flags & SRE_FLAG_IGNORECASE and flags & SRE_FLAG_LOCALE:
             return None
+        if (flags ^ entry_flags) & _parser.TYPE_FLAGS:
+            # _compile_info() compiles the charset with the pattern's flags, so a
+            # class such as \w found inside (?u:...) or (?a:...) would get the
+            # wrong character type.
+            return None
 
     iscased = _get_iscased(flags)
     if op is LITERAL:

The same hunks apply to Lib/re/_compiler.py on 3.14 and 3.15. An alternative that keeps the prefilter is to return the
effective flags along with the charset and have _compile_info() compile it with them. That touches every return in
the function, and these patterns are rare, so the smaller change seemed better. A regression test:

    def test_charset_prefix_scoped_type_flags(self):
        # The search() prefilter must use the type flags of the group it was taken from.
        self.assertEqual(re.search(r'(?u:\w)', '\xe9', re.ASCII).group(), '\xe9')
        self.assertEqual(re.search(r'(?u:\d)', '\u0663', re.ASCII).group(), '\u0663')
        self.assertEqual(re.search(r'(?a:\W)', '\xdf').group(), '\xdf')
        self.assertEqual(re.search(r'(?a:\S)', '\u2003').group(), '\u2003')
        self.assertEqual(re.findall(r'(?a:\W)', 'Stra\xdfe!'), ['\xdf', '!'])

Checked on 3.14.5 by putting a copy of Lib/re with the patched _compiler.py first on PYTHONPATH (the run printed
re._compiler.__file__ to confirm which copy it used):

  • The test above fails on the unpatched copy, with AttributeError: search() returned None. It passes on the
    patched copy.
  • The 10-row probe: from 8 of 10 rows wrong to 0 of 10.
  • test_re from the v3.14.5 tag: 165 ran, 0 failures, 3 skipped, both unpatched and patched. test_re does not
    currently exercise this bug.
  • A randomized differential. 90,000 generated patterns (3 seeds × 30,000: lookaround, \b and anchor leads, scoped
    a/u/L/i groups, categories, sets, alternations, repeats; str and bytes), of which 82,891 compiled, each run on
    5 random subjects (414,455 cases). The check: search(s, pos) must equal the first i >= pos where match(s, i)
    succeeds. Unpatched: 8 cases broke that rule, all with a scoped type-flag group. Patched: 0.
  • On main (3330712), with the fix applied to Lib/re/_optimizer.py: the regression test fails before and passes
    after; test_re 171 ran, 0 failures, 4 skipped; and the same search-vs-match rule over 320,000 generated cases
    (2 seeds; lookaround / \b / anchor / \s+ leads, nested a/u/i groups, str patterns) gives 1,338 violations
    with main's own _optimizer.py and 0 with the fix.

CPython versions tested on:

3.12, 3.13, 3.14, CPython main branch

Operating systems tested on:

Linux


AI assistance was used to investigate this and to draft this report.

Linked PRs

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions