Skip to content

[codex] Add real knowledge document extraction progress - #2727

Merged
kwakayama merged 2 commits into
mainfrom
codex/knowledge-ingest-real-progress
Jul 2, 2026
Merged

kwakayama merged 2 commits into
mainfrom
codex/knowledge-ingest-real-progress

Conversation

@kwakayama

@kwakayama kwakayama commented Jul 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds real extraction progress for knowledge document ingestion while keeping the existing one-document-to-one-markdown-file output contract intact.

Refs veryfront/veryfront-studio#5484.

Root Cause

The visible staging failure was a timeout, but the timeout was only the symptom. The Bosch PDF is valid and native Kreuzberg/PDFium can extract it quickly. The problematic path was the Deno/WASM extraction route, which can hang or run slowly enough that ingestion appears stalled and eventually times out without any real document progress signal.

Changes

  • Adds an optional document extraction progress callback through the compatibility extractor and knowledge parser path.
  • Uses native Kreuzberg progress extraction for Deno PDF/PPT/PPTX ingestion when progress is requested.
  • Emits real page progress for PDFs after each page is extracted by native Kreuzberg.
  • Emits real slide progress for PPTX by reading visible slide text from the OpenXML deck in presentation order.
  • Emits file-level completion progress for legacy binary PPT, where slide-level progress is not available through the native binding.
  • Keeps programmatic callers and no-logger ingests on the existing fast extraction path unless they opt into progress.
  • Falls back to the previous opaque extraction path if progress extraction fails, and logs a warning with the MIME type and failure message.
  • Includes the new worker in compiled binary and npm extension package builds.

Contract

The output shape stays the same: one source document still produces one markdown file under knowledge/; progress never splits output files or changes parser return shape.

For PDFs and legacy PPT, progress is additional metadata/logging around native or whole-file extraction. For PPTX, logger-backed progress extraction uses an OpenXML slide-text path so it can emit real per-slide progress. That path preserves presentation order for visible slide text, but it can differ from the opaque Kreuzberg path in formatting/fidelity, for example speaker notes are not included. Programmatic/no-logger callers stay on the existing path unless they explicitly request progress, and failures fall back to the previous opaque extraction behavior.

Verification

  • deno task test cli/commands/knowledge/command.test.ts
  • cd extensions/ext-document-kreuzberg && deno test --allow-all src/index.test.ts
  • env -u OTEL_EXPORTER_OTLP_ENDPOINT -u OTEL_EXPORTER_OTLP_HEADERS -u OTEL_EXPORTER_OTLP_PROTOCOL -u OTEL_LOGS_EXPORTER -u OTEL_METRICS_EXPORTER -u OTEL_RESOURCE_ATTRIBUTES -u OTEL_TRACES_EXPORTER -u GRAFANA_SERVICE_ACCOUNT_TOKEN -u GRAFANA_URL deno task verify
  • Pre-push hook passed with formatting, lint, typecheck, and unit tests under the same clean env.
  • Bosch PDF smoke: 52 real page progress events, one output file, extracted content contains Bosch/manual markers.
  • PPTX smoke: 2 real slide progress events, one output file, extracted content contains both slide texts in presentation order.

@kwakayama
kwakayama marked this pull request as ready for review July 2, 2026 14:07
@kwakayama
kwakayama requested a review from kojiwakayama as a code owner July 2, 2026 14:07
Copilot AI review requested due to automatic review settings July 2, 2026 14:07

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 6f8ecd11f8

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread cli/commands/knowledge/command.ts Outdated
slug: slugs[index],
sourceReference,
}, {
onProgress: (event) => {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Skip progress callbacks when nothing can log them

When knowledge ingest runs outside a run context, createKnowledgeIngestEventLogger() returns null, but this still passes a truthy no-op onProgress callback. In Deno, KreuzbergDocumentExtractor.extractInWorker() switches PDFs/PPT/PPTX to the new progress worker solely because options.onProgress exists, so ordinary ingests with no user-visible progress now take the slower/different page-by-page/PPTX XML path and can fail on files the existing native/WASM extractor would handle. Pass the callback only when deps.eventLogger exists.

Useful? React with 👍 / 👎.

Comment on lines +95 to +97
const slidePaths = Object.keys(zip.files)
.filter((path) => /^ppt\/slides\/slide\d+\.xml$/.test(path))
.sort((left, right) => slideNumber(left) - slideNumber(right));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve the PPTX presentation order

For PPTX decks where slides were reordered or deleted and recreated, the numeric ppt/slides/slideN.xml filenames do not define the deck order; the order comes from ppt/presentation.xml and its relationships. This progress extraction path concatenates slide text in filename order, so generated knowledge markdown can put slide content in the wrong sequence for valid decks. Use the presentation slide list and relationship targets before iterating slides.

Useful? React with 👍 / 👎.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds real-time extraction progress reporting to knowledge document ingestion without changing the existing “one source document -> one markdown file under knowledge/” output contract. This threads an optional progress callback through the compat document-extraction surface and uses a new native progress-capable worker to emit page/slide/file completion events.

Changes:

  • Introduces DocumentExtractionProgressEvent + DocumentExtractionOptions and plumbs onProgress through the knowledge parser and Kreuzberg DocumentExtractor worker path.
  • Adds a native progress extraction worker for Deno that emits per-page PDF progress and per-slide PPTX progress (with file-level completion for formats without granular progress).
  • Updates build/packaging to include the new worker and its new dependencies (pdf-lib, jszip), plus adds tests covering progress forwarding and logging.

Verification:

  • Not run in this review environment.
  • PR description reports: deno task test cli/commands/knowledge/command.test.ts, deno test --allow-all extensions/ext-document-kreuzberg/src/index.test.ts, and deno task verify in a clean env.

Reviewed changes

Copilot reviewed 16 out of 17 changed files in this pull request and generated 2 comments.

Show a summary per file
File Description
tests/unit/build/compile-binary-includes.test.ts Asserts the new native progress worker is included in compiled binaries.
src/extensions/compat/native-services.ts Adds progress event/options types and extends the DocumentExtractor.extractInWorker signature to accept options.
src/extensions/compat/index.ts Re-exports the new compat progress types.
scripts/build/npm-package-metadata.ts Marks jszip and pdf-lib as extension-owned dependencies for npm packaging.
scripts/build/compile-binary.ts Ensures deno compile includes the new worker entrypoint.
scripts/build/build-npm-extension-packages.ts Transpiles both extraction workers into the npm extension package output.
extensions/ext-document-kreuzberg/src/native-progress-extraction-worker.ts New worker that extracts PDFs page-by-page and PPTX slide-by-slide and posts progress events.
extensions/ext-document-kreuzberg/src/index.ts Routes Deno extraction through the new progress worker when onProgress is provided; adds idle/hard timeouts around worker extraction.
extensions/ext-document-kreuzberg/src/index.test.ts Adds unit tests verifying the progress-capable path is selected for PDF/PPTX in Deno.
extensions/ext-document-kreuzberg/deno.json Adds jszip/pdf-lib import mappings for the new worker.
deno.lock Adds new dependency entries (and updates other resolved versions).
deno.json Exposes veryfront/extensions/first-party-import as a public export/import mapping.
cli/shared/ensure-content-processor.ts Switches to importing importFirstPartyExtensionModule via the public veryfront/extensions/first-party-import path.
cli/shared/default-contracts.ts Same public import path update for first-party extension importing.
cli/commands/knowledge/parser.ts Adds optional progress forwarding into Kreuzberg extraction from the knowledge parser path.
cli/commands/knowledge/command.ts Logs structured extraction progress events during ingestion.
cli/commands/knowledge/command.test.ts Adds tests ensuring extraction progress events are logged and the output contract remains one markdown file per source.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +161 to +169
if (message.type === "progress") {
resetIdleTimer();
try {
await options.onProgress?.(message.event);
} catch (error) {
fail(error instanceof Error ? error : new Error(String(error)));
}
return;
}
Comment on lines +36 to +40
extractInWorker(
buffer: ArrayBuffer,
mimeType: string,
options?: { onProgress?: DocumentExtractionProgress },
): Promise<string>;
@kwakayama

Copy link
Copy Markdown
Contributor Author

Critical review — Score: 58 / 100

Good intent and clean plumbing, but the core approach introduces two latent regressions and the PR description makes a "no behavior change" claim that the code contradicts. The type/contract work, build/packaging wiring, and plumbing tests are genuinely solid; the risk is concentrated in the new extraction strategy and how it's wired into the default path.

What's good

  • Clean, well-typed contract. DocumentExtractionProgressEvent / DocumentExtractionOptions threaded coherently through compat → parser → extractor. ResolvedRunKnowledgeParsersDeps replacing Required<…> is the right call.
  • Packaging done right. New worker added to compile-binary.ts includes, build-npm-extension-packages.ts loop, npm-package-metadata.ts, plus a guard test in compile-binary-includes.test.ts. The new Worker(new URL(...)) compile-trace gotcha is handled.
  • Idle + hard timeout wrapper around the worker is a reasonable safety net.

Blocking / high-severity

1. The new page-by-page path is now the DEFAULT for every PDF/PPTX ingest — contradicting the PR description.
The description says "Keeps programmatic callers on the existing fast extraction path unless they opt into progress." The code does the opposite. In command.ts, onProgress is passed unconditionally:

}, { onProgress: (event) => { deps.eventLogger?.info(...) } });

The eventLogger?. guard is inside the callback, so the callback itself is always truthy. The extractor (index.ts) then switches to extractWithNativeProgressDeno solely on options.onProgress existing. Net effect: all PDF/PPTX CLI ingests — including runs with no logger and programmatic callers — now take the heavier path. This is exactly Codex's P2 #1, and it's correct. Gate the callback: pass onProgress only when deps.eventLogger exists.

2. No fallback on extraction failure — this can break the very PDFs it aims to fix.

if (options.onProgress && isNativeProgressMimeType(mimeType)) {
  try { return await extractWithNativeProgressDeno(...); }
  catch (error) { if (!isMissingPackageError(error)) throw error; }
}

Only a missing-package error falls back to native/WASM whole-file extraction. Any other failure — pdf-lib choking on copyPages for an unusual PDF structure, an encrypted/XFA doc, a per-page extractBytes error — propagates and fails the whole ingest, where the previous native/WASM path would have extracted the file fine. Given the stated goal is fixing a timeout on a real PDF, converting "slow but works" into "hard fail" for some corpus of PDFs is a robustness regression. Fall back to the existing path on any error, not just missing-package.

3. PPTX extraction silently diverges from Kreuzberg (quality regression + two engines for one format).
The progress path extracts PPTX via hand-rolled regex over ppt/slides/slideN.xml (<a:t> runs joined with spaces). This is a different engine from the non-progress path (Kreuzberg worker), so the content of the output .md changes depending on whether progress was requested — and per #1, progress is now always requested. The regex approach drops line/paragraph structure, ignores speaker notes (notesSlideN.xml), and won't match Kreuzberg's fidelity. The PR's "does not change parser return shape / metadata-only" framing understates this: the extracted text itself changes. At minimum, call this out explicitly; ideally keep PPTX on Kreuzberg and emit slide progress around it rather than reimplementing extraction.

Medium

4. Slide ordering by filename is not deck order (Codex P2 #2 — valid). slideN.xml numbering does not reflect presentation order after reorder/delete/recreate; order lives in ppt/presentation.xml + rels. Concatenated markdown can be out of sequence for valid decks.

5. Idle timer races the onProgress await (Copilot — valid). resetIdleTimer() runs before await options.onProgress(...), and isn't reset after. A slow callback can trip the idle timeout mid-flight and terminate a healthy worker. Clear the idle timer while awaiting the callback, reset after it resolves.

6. Performance. Page-by-page splits each PDF into N single-page docs via PDFDocument.create() + copyPages + save({ useObjectStreams: false }) + a fresh native extractBytes per page. That's O(pages) document reconstructions where there was one extraction — now on the default path (#1). For large PDFs this is a real slowdown, ironic for a change motivated by a timeout.

Low / nits

  • Correctness of the new extraction logic is untested. Tests only assert routing and event forwarding — nothing verifies page-by-page text ≈ whole-doc text, slide ordering, or the (absent) error fallback. The only correctness evidence is a manual Bosch/one-PPTX smoke. The risky part is the least-covered.
  • deno.lock churn is broad and bundled. Removes @deno/dnt / @ts-morph/* / @david/code-block-writer and bumps many unrelated deps (ai-sdk, pg-protocol, systeminformation, yargs, …). Please confirm @deno/dnt removal doesn't affect the npm build and consider isolating unrelated bumps.
  • Worker event.origin check is dead code. In a Deno module worker, event.origin is empty for parent postMessage; the origin guard provides no real protection and reads as cargo-culted web security.
  • MIME helpers duplicated between index.ts and the worker (normalizeMimeType / isPdf / isPptx).
  • DocumentKreuzbergExtensionModule.extractInWorker in parser.ts models options as only { onProgress? } while DocumentExtractionOptions also has idleTimeoutMs/hardTimeoutMs — minor drift risk (Copilot).

Bottom line

The direction (real extraction progress) is worth having and the scaffolding is well done, but as written this ships a new, less-robust, differently-behaving extraction path as the default for the exact file types in the originating bug, with no fallback and no correctness tests. Address #1 (gate on logger) and #2 (fall back on any error) before merge; resolve #3/#4 or explicitly document the PPTX behavior change. With those fixed this is a straightforward approve.

Automated critical review. Score reflects: strong plumbing/packaging, offset by a default-path behavior change contradicting the description, a robustness regression with no fallback, and untested new extraction logic.

@kwakayama
kwakayama force-pushed the codex/knowledge-ingest-real-progress branch from 6f8ecd1 to 725c461 Compare July 2, 2026 14:34
@kwakayama
kwakayama force-pushed the codex/knowledge-ingest-real-progress branch from 725c461 to 98aeece Compare July 2, 2026 14:48
@kwakayama

Copy link
Copy Markdown
Contributor Author

Re-review (commit 98aeece286) — Score: 84 / 100 (was 58)

Both blockers and both medium issues from the previous review are fixed, each with a real test. Verified against the v1→v2 delta, not just the commit message.

Confirmed fixed

  1. ✅ Progress gated on the event logger. command.ts now builds parserDeps only when deps.eventLogger exists and passes undefined otherwise, so programmatic/no-logger ingests stay on the fast path — matching the PR description at last. Covered by the new does not request extraction progress when no event logger can report it test asserting hasProgressCallback === false.
  2. ✅ Fallback on any error. The isMissingPackageError-only filter is gone; progress extraction is now genuinely opportunistic (catch { /* fall back */ }) and falls through to the previous native/WASM path. Covered by the new falls back to the previous PDF extraction path when progress extraction fails test.
  3. ✅ PPTX presentation order. pptxSlidePaths() now reads ppt/presentation.xml + _rels, orders by sldIdLst, normalizes relationship targets (including ../, absolute, and query/fragment forms), appends unreferenced slides, and falls back to filename order when the manifest is missing. The new test is a real worker integration test with a JSZip-built deck where presentation order ≠ filename order — exactly the failure case. Nice.
  4. ✅ Idle-timer race. The idle timer is now cleared while awaiting options.onProgress and reset after it resolves (guarded by !settled), so a slow callback can no longer trip the idle timeout on a healthy worker.
  5. ✅ Copilot's type-drift nit: the parser's local extractInWorker signature now uses DocumentExtractionOptions.

Remaining (why not higher)

  • PPTX engine divergence stands (previous chore: add rate limiting middleware #3). The progress path still extracts PPTX via hand-rolled <a:t> regex rather than Kreuzberg — so run-context ingests get different text than programmatic ones for the same deck (speaker notes ignored, paragraph structure flattened). The mitigations are real (only logger-backed runs take it; any failure falls back; ordering now correct), but the PR body's "Progress is additional metadata/logging only" still isn't accurate for PPTX content. Update the description, or route PPTX through Kreuzberg with slide progress around it in a follow-up.
  • Silent fallback. catch { } swallows the failure reason entirely. Two consequences: (a) a pathological doc can burn the full 10-min hard timeout in the progress worker and then silently re-run whole-file extraction — worst-case latency roughly doubles versus v1 of this PR; (b) operators get no signal that the progress path is failing in the wild, which is how quiet regressions in the regex/pdf-lib path would go unnoticed. Log the error (debug/warn) before falling back, and consider not re-trying after a hard timeout.
  • Open CodeQL alert on the worker's onmessage handler. The origin check is present but written as if (event.origin && ...) — in a dedicated module worker event.origin is "", so the guard never verifies anything (and CodeQL rightly sees no effective check). It's benign for a parent-spawned dedicated worker, but either dismiss the alert with that justification or drop/restructure the dead guard rather than leaving the alert open.
  • deno.lock churn unchanged: still bundles removal of @deno/dnt/@ts-morph/* and many unrelated bumps (ai-sdk, pg-protocol, systeminformation, yargs, …) with this feature. If npm packaging no longer needs dnt, fine — but it deserves its own commit/PR line, not a side effect of adding jszip/pdf-lib.
  • Minor: normalizeMimeType/isPdf/isPptx still duplicated between index.ts and the worker; PDF page-split correctness (page-by-page text ≈ whole-doc text) is still only smoke-verified, though the any-error fallback now bounds the blast radius.

Bottom line

This is now mergeable from my side. The response to review was exemplary — every fix came with a targeted test, including a real worker-level integration test for slide ordering. Remaining items are: fix the PR description's PPTX claim, add a log line on fallback, and resolve the CodeQL alert; none block merge except arguably the CodeQL alert if your branch protection gates on it.

Automated re-review of the force-pushed head; delta-verified against previously reviewed commit 6f8ecd11f8.

@kwakayama
kwakayama force-pushed the codex/knowledge-ingest-real-progress branch from 98aeece to a43f2ba Compare July 2, 2026 15:09
@kwakayama

Copy link
Copy Markdown
Contributor Author

Re-review (commit a43f2ba220) — Score: 90 / 100 (was 84, initially 58)

Delta-verified against the previously reviewed commit 98aeece286. Every actionable item from the last round is addressed; the v2→v3 delta is tight and focused.

Confirmed fixed since last review

  1. ✅ Fallback is no longer silent. The catch now warns with the MIME type and failure message — [ext-document-kreuzberg] native progress extraction failed; falling back to opaque extraction — routed through the extension logger (extractor is now constructed in setup(ctx) with ctx.logger, with a console.warn fallback for direct construction). The updated fallback test asserts the exact warning payload. Operators will now see progress-path failures in the wild.
  2. ✅ PR description is now accurate. The Contract section explicitly documents the PPTX divergence: OpenXML slide-text path for logger-backed progress, presentation-order preservation, formatting/fidelity differences vs Kreuzberg, speaker notes excluded, and the fallback behavior. This converts the previous misrepresentation into a documented, reviewable design tradeoff — exactly what was asked.
  3. ✅ CodeQL is green on this head — the check passes and no open code-scanning alerts are reported against the PR ref.

Remaining (non-blocking)

  • PPTX dual-engine design stands, now as a documented tradeoff rather than a surprise. Long-term, the hand-rolled <a:t>/rels regex parsing in the worker is a maintenance liability (OpenXML has more edge cases than any regex set will cover — grouped shapes, tables, alternate content blocks). If a deck matters and the regex path under-extracts, the fallback won't trigger because nothing fails — it just extracts less. Consider a follow-up to route PPTX through Kreuzberg with progress emitted around it, or at least a heuristic (e.g., suspiciously low character counts → fall back).
  • Hard-timeout-then-retry worst case remains: a pathological doc can burn the 10-min hard timeout in the progress worker and then re-run whole-file extraction. Now visible via the warning, so acceptable — but a "don't retry after hard timeout" guard would still be cheap insurance.
  • deno.lock churn still bundles @deno/dnt/@ts-morph/* removal and unrelated dependency bumps with this feature. CI is fully green (including npm install smoke and binary e2e), so it demonstrably works — but it still muddies bisects and this PR's blast radius. Worth a note in the description at minimum.
  • Micro-nits carried over: dead event.origin guard in the worker (empty origin in dedicated workers; no open alert anymore, so style only), duplicated MIME helpers between index.ts and the worker, and PDF page-split correctness still verified only by smoke test (bounded by the any-error fallback).

Bottom line

Approve. Three review rounds took this from "ships a differently-behaving, no-fallback extraction path as the silent default" to "opt-in, ordered, observable, fallback-protected, documented, and test-covered at every fix site" — with full CI green (unit, integration, binary e2e, coverage gate, CodeQL). The remaining 10 points are design-debt items (dual-engine PPTX, lock churn) suited to follow-ups, not blockers.

Automated re-review of the force-pushed head; delta-verified 98aeece286 → a43f2ba220.

@kwakayama
kwakayama merged commit 4646439 into main Jul 2, 2026
30 checks passed
@kwakayama
kwakayama deleted the codex/knowledge-ingest-real-progress branch July 2, 2026 16:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants