Repository navigation
<regex>: Accelerate greedy wildcard loops - #6475
Julian Müller (muellerj3) wants to merge 1 commit into
Conversation
|
Azure Pipelines: Successfully started running 1 pipeline(s). There may be pipelines that require an authorized user to comment /azp run to run. |
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The optimization preserves existing matching semantics and is supported by focused regression tests and benchmarks.
Review effort: Balanced
Findings: None
What changed in this PR
Accelerates greedy wildcard regex repetitions by using optimized search algorithms while preserving ECMAScript and POSIX semantics.
Changes:
- Adds optimized handling for greedy dot repetitions.
- Reuses line-terminator search logic.
- Adds correctness tests and performance benchmarks.
| File | Description |
|---|---|
stl/inc/regex |
Implements optimized wildcard-loop matching. |
tests/std/tests/VSO_0000000_regex_use/test.cpp |
Tests wildcard bounds and character semantics. |
tests/std/tests/GH_000995_regex_custom_char_types/test.cpp |
Tests custom character types. |
benchmarks/src/regex_match.cpp |
Benchmarks ECMAScript and POSIX wildcard loops. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
|
ARM64 numbers could be even more impressive, as #6468 is merged |
| _Cur = _STD _Find_next_ecmascript_line_terminator<_Elem>(_Cur, _Last); | ||
| break; | ||
| case _Fl_rep_dot_posix: | ||
| _Cur = _STD _Find_unchecked(_Cur, _Last, _Elem()); |
There was a problem hiding this comment.
I think we could standardize on _Elem{}.
| if constexpr (sizeof(_Elem) == 1) { | ||
| static constexpr _Elem _Line_terminators_size_1[] = { | ||
| static_cast<_Elem>(_Meta_cr), static_cast<_Elem>(_Meta_nl)}; | ||
| return _STD find_first_of(_First, _Last, _Line_terminators_size_1, _STD end(_Line_terminators_size_1)); |
There was a problem hiding this comment.
If this case is very important we could have specialized vectorized find_first_of with hardcoded needle, it could use the best strategy. I don't expect that to be dramatically faster, but still likely there's room for improvement.
Towards #5971. This speeds up
.*and similar greedy wildcard loops, delegating the search for the first non-matching character tostd::find()orstd::find_first_of()to enable vectorization when available.While the speedups are impressive, I have to point out that we fundamentally can't beat Boost.Regex in its default mode on very long inputs because Boost.Regex makes
.match any character, including newlines or nulls. In Boost.Regex, spec-conformant behavior must be turned on explicitly.For non-random-access iterators, the new code might iterate the input up to four times now instead of once (three of which are hidden in
std::distance()andstd::advance()calls), so the change can potentially degrade performance for such iterators if iterating is expensive. If it turns out to be a practical problem, we can disable this optimization for such iterators or add optimized code for non-random-access iterators.Strictly speaking, we lack unit test coverage in
_Find_next_ecmascript_line_terminator()for theboolbranch. But this case is very esoteric and the correctness of the code is easy to verify, so I think we can live with this.Benchmark
regex_match