fix(ci): the invisible-character gate never matched anything - #82
Conversation
MEASURED 2026-08-27: this gate's pattern caught 0 OF 6 invisible-character test
cases. It has never detected an NBSP, zero-width space, BOM, soft hyphen, bidi
override or word joiner.
ROOT CAUSE: the pattern used UTF-8 BYTE sequences (\xc2\xa0) while grep -P
matches CHARACTERS. Bytes c2 a0 are ONE character U+00A0; \xc2\xa0 asks for TWO
characters, U+00C2 then U+00A0, which is never present.
grep -P '\xc2\xa0' -> miss
grep -P '\x{a0}' -> MATCH
Only \x00 worked, being single-byte in both readings.
FIXED: codepoint escapes; C0 control characters \x01-\x08,\x0B,\x0C,\x0E-\x1F
added (TAB/LF/CR excluded); and grep -a, without which grep skips any NUL-bearing
file as binary.
The C0 range matters: a stray BACKSPACE byte made a workflow unparseable in
developer-ecosystem, so it never ran, and this linter called it clean.
Canonical fix: hyperpolymath/empty-linter#70. 1 file(s) here.
VERIFIED: YAML re-parsed, and the corrected pattern was confirmed to catch a real
NBSP before the change was kept.
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Plus Run ID: 📒 Files selected for processing (1)
Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review. 📜 Recent review details🔇 Additional comments (1)
📝 WalkthroughSummary by CodeRabbit
WalkthroughThe dogfood gate now matches invisible characters by Unicode code point, includes additional control characters and the word joiner, and scans binary files as text. ChangesInvisible-character gate
Estimated code review effort: 2 (Simple) | ~10 minutes Merge Risk: ⚪ Minimal · up to The workflow narrowly corrects invisible-character detection, including BOM handling, with no actionable merge-blocking risk remaining after normal checks and review. Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Linked Issues checkExplanation The workflow change addresses codepoint escapes, C0 controls, and grep -a. It does not include the separately required leading-BOM byte check, updates to stdlib/ByteDetector.affine and config.ncl, or corrections across the other inlined gate copies required by issue Resolution Add the leading-BOM byte-wise check, update stdlib/ByteDetector.affine and config.ncl with consistent C0-control detection, and apply the corrected pattern to all required inlined dogfood-gate.yml copies. Add verification for NBSP, zero-width space, soft hyphen, bidi override, word joiner, NUL, backspace, legitimate whitespace, clean files, and leading BOMs. Full details: Docstring CoverageExplanation No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0 files. (1 skipped: 1 unsupported.)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Up to standards ✅🟢 Issues
|
There was a problem hiding this comment.
Pull Request Overview
This PR fixes the invisible-character linter which previously failed to match Unicode characters due to incorrect usage of UTF-8 byte sequences. It successfully transitions to codepoint escapes and expands detection to C0 control characters.
While the implementation logic is correct, there is a performance optimization identified in the use of find -exec. Additionally, there are no automated tests or sample files included to verify that these patterns work as intended or to prevent future regressions. The Codacy analysis is up to standards, and no critical security vulnerabilities were found.
About this PR
- Although the manual verification confirms the fix, the PR lacks automated test cases or sample files containing the targeted characters. Without these, it is difficult to guarantee the gate remains functional as the environment or requirements evolve.
Test suggestions
- Missing recommended test scenario: Verify detection of Non-Breaking Space (U+00A0) using codepoint escape.
- Missing recommended test scenario: Verify detection of C0 control characters (e.g., Backspace \x08) to prevent CI failures.
- Missing recommended test scenario: Ensure 'grep -a' correctly processes and scans files containing NUL bytes (\x00).
- Missing recommended test scenario: Verify detection of Byte Order Mark (U+FEFF).
Prompt proposal for missing tests
Consider implementing these tests if applicable:
1. Missing recommended test scenario: Verify detection of Non-Breaking Space (U+00A0) using codepoint escape.
2. Missing recommended test scenario: Verify detection of C0 control characters (e.g., Backspace \x08) to prevent CI failures.
3. Missing recommended test scenario: Ensure 'grep -a' correctly processes and scans files containing NUL bytes (\x00).
4. Missing recommended test scenario: Verify detection of Byte Order Mark (U+FEFF).
TIP Improve review quality by adding custom instructions
TIP How was this review? Give us feedback
| -o -name '*.idr' -o -name '*.zig' -o -name '*.v' -o -name '*.jl' \ | ||
| -o -name '*.gleam' -o -name '*.hs' -o -name '*.ml' -o -name '*.sh' \) \ | ||
| -exec grep -Prl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null | ||
| -exec grep -aPrl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null |
There was a problem hiding this comment.
⚪ LOW RISK
Suggestion: The -exec ... \; syntax spawns a new process for every file, which is inefficient. Using -exec ... + allows grep to process multiple files in a single batch, improving performance. Also, the -r flag is unnecessary since find is already providing specific file paths through recursion.
| -exec grep -aPrl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null | |
| -exec grep -aPl "$PATTERNS" {} + > /tmp/empty-lint-results.txt 2>/dev/null |



Measured 2026-08-27: this gate caught 0 of 6 invisible-character test cases. It has never detected an NBSP, zero-width space, BOM, soft hyphen, bidi override or word joiner.
Root cause
The pattern used UTF-8 byte sequences (
\xc2\xa0) whilegrep -Pmatches characters. Bytesc2 a0are one character U+00A0;\xc2\xa0asks for two, U+00C2 then U+00A0 — never present.Only
\x00worked, being single-byte in both readings. The gate ran, passed, and could not see what it exists to see.Fixed
\x01-\x08,\x0B,\x0C,\x0E-\x1Fadded (TAB/LF/CR excluded)grep -a— without it grep skips any NUL-bearing file as binaryThe C0 range matters: a stray backspace byte made a workflow unparseable in
developer-ecosystem, so it never ran — and this linter called it clean.Canonical fix: hyperpolymath/empty-linter#70. 1 file(s) here.
Verified: YAML re-parsed, and the corrected pattern was confirmed to catch a real NBSP before the change was kept.