fix(ci): the invisible-character gate never matched anything - #118
Conversation
MEASURED 2026-08-27: this gate's pattern caught 0 OF 6 invisible-character test
cases. It has never detected an NBSP, zero-width space, BOM, soft hyphen, bidi
override or word joiner.
ROOT CAUSE: the pattern used UTF-8 BYTE sequences (\xc2\xa0) while grep -P
matches CHARACTERS. Bytes c2 a0 are ONE character U+00A0; \xc2\xa0 asks for TWO
characters, U+00C2 then U+00A0, which is never present.
grep -P '\xc2\xa0' -> miss
grep -P '\x{a0}' -> MATCH
Only \x00 worked, being single-byte in both readings.
FIXED: codepoint escapes; C0 control characters \x01-\x08,\x0B,\x0C,\x0E-\x1F
added (TAB/LF/CR excluded); and grep -a, without which grep skips any NUL-bearing
file as binary.
The C0 range matters: a stray BACKSPACE byte made a workflow unparseable in
developer-ecosystem, so it never ran, and this linter called it clean.
Canonical fix: hyperpolymath/empty-linter#70. 1 file(s) here.
VERIFIED: YAML re-parsed, and the corrected pattern was confirmed to catch a real
NBSP before the change was kept.
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Plus Run ID: 📒 Files selected for processing (1)
Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review. 📜 Recent review details🔇 Additional comments (3)
📝 WalkthroughSummary by CodeRabbit
WalkthroughThe workflow updates invisible-character detection. The regex uses character and Unicode code-point escapes. The scan passes ChangesInvisible-character gate
Estimated code review effort: 1 (Trivial) | ~5 minutes Merge Risk: 🟡 Moderate · up to The workflow’s invisible-character gate was corrected to use codepoint escapes and broader control coverage, but the current pattern includes a BOM escape that GNU grep 3.8 rejects before matching. That can make the CI gate fail instead of validating files, so the PR should not merge until the pattern is made compatible or the limitation is explicitly accepted. Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Linked Issues checkExplanation The pull request implements the codepoint escapes, C0 control detection, and grep -a requirements in dogfood-gate.yml. It does not show the required separate leading-BOM check, updates to stdlib/ByteDetector.affine and config.ncl, or propagation to other inlined copies. Resolution Add the separate byte-wise leading-BOM check. Update stdlib/ByteDetector.affine and config.ncl with the matching C0-control logic. Propagate the corrected pattern to all required inlined copies, then verify the targeted characters, corrupted workflows, and clean files. [ Full details: Docstring CoverageExplanation No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0 files. (1 skipped: 1 unsupported.)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Up to standards ✅🟢 Issues
|
There was a problem hiding this comment.
Pull Request Overview
The PR successfully fixes the invisible-character gate by transitioning from byte-sequence patterns to Unicode codepoint escapes and ensuring files with NUL bytes are treated as text. While the logic is sound and addresses security risks like 'Trojan Source' obfuscation, the shell command used to execute the linter is inefficient.
Codacy analysis indicates the changes are up to standards. However, the primary risk is the absence of regression test files (fixtures) containing the targeted characters, which means the gate's efficacy cannot be automatically verified in this or future PRs.
About this PR
- Although the regex patterns and flags have been updated, there are no regression test files (fixtures) included in this PR that contain the invisible characters or control codes. Adding a set of known-bad files would ensure this gate remains functional in the future and facilitates verification of the fix.
Test suggestions
- Verify detection of Non-Breaking Space (U+00A0) using codepoint escape.
- Verify detection of C0 control characters (e.g., Backspace \x08) within a source file.
- Verify that files containing NUL bytes (\x00) are scanned rather than skipped as binary.
- Verify detection of Zero-Width characters (ZWSP, ZWJ, ZWNJ).
Prompt proposal for missing tests
Consider implementing these tests if applicable:
1. Verify detection of Non-Breaking Space (U+00A0) using codepoint escape.
2. Verify detection of C0 control characters (e.g., Backspace \x08) within a source file.
3. Verify that files containing NUL bytes (\x00) are scanned rather than skipped as binary.
4. Verify detection of Zero-Width characters (ZWSP, ZWJ, ZWNJ).
TIP Improve review quality by adding custom instructions
TIP How was this review? Give us feedback
| -o -name '*.idr' -o -name '*.zig' -o -name '*.v' -o -name '*.jl' \ | ||
| -o -name '*.gleam' -o -name '*.hs' -o -name '*.ml' -o -name '*.sh' \) \ | ||
| -exec grep -Prl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null | ||
| -exec grep -aPrl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null |
There was a problem hiding this comment.
🟡 MEDIUM RISK
Suggestion: Improve execution performance and observability by batching files and removing redundant flags. The current command executes a new grep process for every file, which is inefficient. Additionally, the '-r' flag is redundant because 'find' is already performing directory traversal, and removing '2>/dev/null' allows PCRE engine errors to be visible in logs.
| -exec grep -aPrl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null | |
| -exec grep -aPl "$PATTERNS" {} + > /tmp/empty-lint-results.txt |
Measured 2026-08-27: this gate caught 0 of 6 invisible-character test cases. It has never detected an NBSP, zero-width space, BOM, soft hyphen, bidi override or word joiner.
Root cause
The pattern used UTF-8 byte sequences (
\xc2\xa0) whilegrep -Pmatches characters. Bytesc2 a0are one character U+00A0;\xc2\xa0asks for two, U+00C2 then U+00A0 — never present.Only
\x00worked, being single-byte in both readings. The gate ran, passed, and could not see what it exists to see.Fixed
\x01-\x08,\x0B,\x0C,\x0E-\x1Fadded (TAB/LF/CR excluded)grep -a— without it grep skips any NUL-bearing file as binaryThe C0 range matters: a stray backspace byte made a workflow unparseable in
developer-ecosystem, so it never ran — and this linter called it clean.Canonical fix: hyperpolymath/empty-linter#70. 1 file(s) here.
Verified: YAML re-parsed, and the corrected pattern was confirmed to catch a real NBSP before the change was kept.