fix(ci): the invisible-character gate never matched anything - #58
Conversation
MEASURED 2026-08-27: this gate's pattern caught 0 OF 6 invisible-character test
cases. It has never detected an NBSP, zero-width space, BOM, soft hyphen, bidi
override or word joiner.
ROOT CAUSE: the pattern used UTF-8 BYTE sequences (\xc2\xa0) while grep -P
matches CHARACTERS. Bytes c2 a0 are ONE character U+00A0; \xc2\xa0 asks for TWO
characters, U+00C2 then U+00A0, which is never present.
grep -P '\xc2\xa0' -> miss
grep -P '\x{a0}' -> MATCH
Only \x00 worked, being single-byte in both readings.
FIXED: codepoint escapes; C0 control characters \x01-\x08,\x0B,\x0C,\x0E-\x1F
added (TAB/LF/CR excluded); and grep -a, without which grep skips any NUL-bearing
file as binary.
The C0 range matters: a stray BACKSPACE byte made a workflow unparseable in
developer-ecosystem, so it never ran, and this linter called it clean.
Canonical fix: hyperpolymath/empty-linter#70. 1 file(s) here.
VERIFIED: YAML re-parsed, and the corrected pattern was confirmed to catch a real
NBSP before the change was kept.
|
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Plus Run ID: 📒 Files selected for processing (1)
Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review. 📜 Recent review details🔇 Additional comments (2)
📝 WalkthroughSummary by CodeRabbit
WalkthroughThe invisible-character gate now uses Unicode code-point escapes, detects additional C0 controls and word joiners, and scans binary files as text. ChangesInvisible-character gate
Estimated code review effort: 1 (Trivial) | ~5 minutes Merge Risk: ⚪ Minimal · up to The workflow change is localized and no actionable merge-blocking risk remains. BOM handling is not established by the supplied evidence, but this does not currently justify delaying merge. Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Description checkExplanation The description explains the root cause, lists the workflow changes, and records verification. It does not use every template heading or include the checklist, but it contains the main required information. Full details: Linked Issues checkExplanation The workflow changes satisfy the codepoint escapes, C0 control range, and grep -a objectives in [ Full details: Docstring CoverageExplanation No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0 files. (1 skipped: 1 unsupported.)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Up to standards ✅🟢 Issues
|
There was a problem hiding this comment.
Pull Request Overview
The PR successfully addresses the non-functional invisible-character gate by switching to Unicode codepoint escapes and utilizing grep -a. While the implementation aligns with the core requirements, there is a lack of regression tests to verify the detection of specific characters like U+00A0 or U+200B. Additionally, the PR mentions 6 test cases that are not visible in the diff, making verification difficult.
Recommendations include expanding the character pattern to cover modern attack vectors (like 'Trojan Source' Bidi controls) and optimizing the CI command to ensure errors are not suppressed, which could lead to false positives.
About this PR
- No automated regression tests or test data files containing invisible characters were added to the PR. To prevent future regressions, consider adding a 'known-bad' file that contains representative invisible characters.
- The PR description mentions the gate originally caught 0 of 6 test cases. Since these test cases are not included in the repository, it is difficult to verify the fix against the failing samples.
Test suggestions
- Missing recommended test scenario: Detect Non-breaking Space (U+00A0) using codepoint escape.
- Missing recommended test scenario: Detect Zero-Width Space (U+200B) using codepoint escape.
- Missing recommended test scenario: Detect C0 control character such as Backspace (\x08).
- Missing recommended test scenario: Detect NUL byte (\x00) and ensure the file is processed (not skipped as binary).
- Missing recommended test scenario: Ensure TAB, LF, and CR do not trigger a detection.
Prompt proposal for missing tests
Consider implementing these tests if applicable:
1. Missing recommended test scenario: Detect Non-breaking Space (U+00A0) using codepoint escape.
2. Missing recommended test scenario: Detect Zero-Width Space (U+200B) using codepoint escape.
3. Missing recommended test scenario: Detect C0 control character such as Backspace (\x08).
4. Missing recommended test scenario: Detect NUL byte (\x00) and ensure the file is processed (not skipped as binary).
5. Missing recommended test scenario: Ensure TAB, LF, and CR do not trigger a detection.
TIP Improve review quality by adding custom instructions
TIP How was this review? Give us feedback
| -o -name '*.idr' -o -name '*.zig' -o -name '*.v' -o -name '*.jl' \ | ||
| -o -name '*.gleam' -o -name '*.hs' -o -name '*.ml' -o -name '*.sh' \) \ | ||
| -exec grep -Prl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null | ||
| -exec grep -aPrl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null |
There was a problem hiding this comment.
🟡 MEDIUM RISK
Suggestion: Optimize the command and remove error silencing. The -r flag is redundant when using find, and using {} + is more efficient than {} \; for large file sets. Removing 2>/dev/null ensures that if grep fails due to a syntax error in the pattern, the CI job fails rather than reporting a false success.
| -exec grep -aPrl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null | |
| -exec grep -aPl "$PATTERNS" {} + > /tmp/empty-lint-results.txt |
| # non-breaking spaces, null bytes, and other invisible Unicode in source files. | ||
| set +e | ||
| PATTERNS='\xc2\xa0|\xe2\x80\x8b|\xe2\x80\x8c|\xe2\x80\x8d|\xef\xbb\xbf|\xc2\xad|\xe2\x80\x8e|\xe2\x80\x8f|\xe2\x80\xaa|\xe2\x80\xab|\xe2\x80\xac|\xe2\x80\xad|\xe2\x80\xae|\x00' | ||
| PATTERNS='\x00|[\x01-\x08\x0B\x0C\x0E-\x1F]|\x{a0}|\x{ad}|\x{200b}|\x{200c}|\x{200d}|\x{200e}|\x{200f}|\x{202a}|\x{202b}|\x{202c}|\x{202d}|\x{202e}|\x{2060}|\x{feff}' |
There was a problem hiding this comment.
🟡 MEDIUM RISK
Suggestion: Expand the pattern to include additional problematic invisible and control characters for more comprehensive coverage. This includes U+2028/U+2029 (invisible newlines), Bidi controls (U+2066-U+2069) used in spoofing, and the DEL character (\x7f).
| PATTERNS='\x00|[\x01-\x08\x0B\x0C\x0E-\x1F]|\x{a0}|\x{ad}|\x{200b}|\x{200c}|\x{200d}|\x{200e}|\x{200f}|\x{202a}|\x{202b}|\x{202c}|\x{202d}|\x{202e}|\x{2060}|\x{feff}' | |
| PATTERNS='[\x00-\x08\x0B\x0C\x0E-\x1F\x7F]|\x{a0}|\x{ad}|\x{200b}|\x{200c}|\x{200d}|\x{200e}|\x{200f}|\x{202a}|\x{202b}|\x{202c}|\x{202d}|\x{202e}|\x{2028}|\x{2029}|\x{2066}|\x{2067}|\x{2068}|\x{2069}|\x{2060}|\x{feff}' |



Measured 2026-08-27: this gate caught 0 of 6 invisible-character test cases. It has never detected an NBSP, zero-width space, BOM, soft hyphen, bidi override or word joiner.
Root cause
The pattern used UTF-8 byte sequences (
\xc2\xa0) whilegrep -Pmatches characters. Bytesc2 a0are one character U+00A0;\xc2\xa0asks for two, U+00C2 then U+00A0 — never present.Only
\x00worked, being single-byte in both readings. The gate ran, passed, and could not see what it exists to see.Fixed
\x01-\x08,\x0B,\x0C,\x0E-\x1Fadded (TAB/LF/CR excluded)grep -a— without it grep skips any NUL-bearing file as binaryThe C0 range matters: a stray backspace byte made a workflow unparseable in
developer-ecosystem, so it never ran — and this linter called it clean.Canonical fix: hyperpolymath/empty-linter#70. 1 file(s) here.
Verified: YAML re-parsed, and the corrected pattern was confirmed to catch a real NBSP before the change was kept.