Skip to content

fix(read): treat utf-8 pdf names as text and quote pdftotext - #1238

Merged
TheGreatAxios merged 3 commits into
mainfrom
cl-9474-critic-nits
Sep 29, 2026
Merged

TheGreatAxios merged 3 commits into
mainfrom
cl-9474-critic-nits

Conversation

@TheGreatAxios

Copy link
Copy Markdown
Collaborator

Summary

  • UTF-8 files named .pdf without %PDF magic still read as text through readFileBounded and the read_file plugin
  • extractor_ready quotes the pdftotext path so spaces and metacharacters paste safely into bash
  • Unreadable diagnosis uses the absolute path, matching stream errors

Verification

  • bun run typecheck, bun run build, and bun run test pass
  • bun run check green (7808 pass, 0 fail)

Related to CL-9474

@linear-code

linear-code Bot commented Sep 29, 2026

Copy link
Copy Markdown

CL-9474

Collapsed tool rows clip to 72 characters, so a long path buried the
gap. The first sentence now names missing extractor, permission
boundary, or malformed file, and the message uses the call path.
A .pdf suffix without %PDF magic is only malformed when the first chunk
is binary. extractor_ready quotes the path so it pastes into bash.
@TheGreatAxios
TheGreatAxios merged commit c19f43c into main Sep 29, 2026
13 checks passed
@TheGreatAxios
TheGreatAxios deleted the cl-9474-critic-nits branch September 29, 2026 23:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant