Skip to content

feat(ai-service): full field set including line items - #39

Merged
Fluory merged 116 commits into
mainfrom
claude/feat-ai-full-fields-22
Sep 23, 2026
Merged

Fluory merged 116 commits into
mainfrom
claude/feat-ai-full-fields-22

Conversation

@Fluory

@Fluory Fluory commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Warum

Fixes #22 · Epic #17 · gestapelt auf #38 (Basis claude/feat-user-management-30).

Arbeitsstand

Was ist passiert (Klartext)

Die KI liest jetzt alle Angaben, die der Kunde braucht: zusätzlich E-Mail, Telefon und besondere Anforderungen sowie jede Position einer Anfrage mit Menge, Einheit, Werkstoff und Maßen. Für jeden einzelnen Wert gilt weiter: „belegt" nur, wenn das zitierte Stück wirklich im Dokument steht und der Wert dazu passt. Deutsche Schreibweisen (1.250 Stk., 2,5 m, 15.10.2026) werden vereinheitlicht; steht nur eine Kalenderwoche („KW 42") da, bleibt der Liefertermin „unsicher". Eine Stahlsorte „St 37" wird nicht mehr als „Stück" gelesen, ein Datum nicht als Telefonnummer. Die Positionen werden gespeichert; ihre Anzeige in der Prüfansicht folgt mit #25. Bekannte Grenze (mit Test festgehalten): Zitiert die KI einen eingeschleusten Satz wörtlich, gilt der Wert als belegt – deshalb prüft ein Mensch jeden Wert mit seiner Fundstelle.

Plan-Pflicht (SYSTEM.md §4)

  • Kein Auslöser – keine Modulgrenze, öffentliche API, Migration, Auth/Rechte, kein Zahlungs-/Daten-/Infrapfad, höchstens zwei Module, keine Architekturvarianten, umkehrbar
  • Auslöser zutreffend – Impact Manifest ausgefüllt (Plan vor Code)

Impact Manifest

  • Betroffene Module: AI-Service (services/ai: Normalisierer, Schema/Prompt v2, Verifier, Pipeline), Vertrag contracts/ai-service.openapi.yaml, extraction (Typen, Client-Validierung, Merge, Persistenz), db (Migration 0013), review (Feldbezeichnungen), export (Limitprüfung nur auf exportierte Felder), E2E-Stub.
  • Schnittstellen / Datenänderungen: Vertrag additiv: ExtractedFields + email, phone, additional_requirements; ExtractResponse.lineItems[] (index + fünf FieldResult); neuer Reason calendar_week_only; run.schemaVersion "2", run.promptVersion extract_v2. Tabelle extracted_fields + item_index (null = Kopf), Unique (run_id, field_key, item_index) NULLS NOT DISTINCT statt (run_id, field_key); bestehende Läufe bleiben gültig.
  • Akzeptanzkriterien: alle aus feat(ai-service): full field set including line items #22 – siehe Nachweis.
  • Testplan: pytest – Normalisierer (Zahlen, Einheiten, Daten, KW, Telefon), Verifier auf Positionen, Vertrag, End-to-End mit wiedergegebener Modellantwort (3 Positionen, Injektionsversuch + gepinnte Grenze); Vitest – Merge der Positionen, Vertragsverletzungen bei Positionen, Integration: Persistenz mit item_index, Freigabe nicht durch nie exportierte Felder blockiert.
  • Verifizierte Fakten: openapi-typescript 7.13.0 erzeugt die TS-Typen aus dem neu exportierten Vertrag (Drift-Test); PG17 unterstützt UNIQUE NULLS NOT DISTINCT.
  • Offene Annahmen: Positionen werden über Dokumente hinweg nicht zusammengeführt (Duplikate entscheidet der Mensch in feat(review): line items and all formats in the review UI #25); ERP-Export sendet weiterhin nur die drei Vertragsfelder. Live-Aufruf mit verschachteltem responseSchema unverifiziert (keine Credentials); die Modellantwort des End-to-End-Tests ist handgeschrieben im Vertex-Format.
  • Nicht-Ziele: siehe oben.
  • Risiken und Rollback: Prompt-Änderung ohne Eval-Lauf (Eval-Set entsteht mit feat(evals): eval runner, 15 weighted synthetic cases and CI gate #24). Rollback der Migration 0013 ist nicht rein per Revert möglich, sobald Positionszeilen existieren: zuerst DELETE FROM app.extracted_fields WHERE item_index IS NOT NULL, dann extracted_fields_run_field_item_unique und extracted_fields_item_index_check droppen, UNIQUE (run_id, field_key) als extracted_fields_run_field_unique wiederherstellen, Spalte item_index droppen (als eigene Vorwärts-Migration, per Owner-Rolle). Positionsdaten gehen dabei verloren; Kopffelder bleiben.

Geändert

  • services/ai/** (Normalisierer, Schema/Prompt v2, Verifier, Tests, Fixtures, README), contracts/ai-service.openapi.yaml (generiert)
  • src/features/extraction/{ai-service.contract,ai-client,merge,repository,fixtures,index}.ts (+ Tests), src/features/review/review.ts, src/features/export/payload.ts
  • src/db/schema/app.ts, Migration 0013_line_items.sql
  • Tests: tests/integration/{processing,review}.test.ts, tests/e2e/ai-stub.mjs
  • Doku: docs/technical/{architecture,data-model}.md, CHANGELOG.md, services/ai/README.md

Nachweis (SYSTEM.md §11)

  • verify:changed: AI-Service ruff, ruff format, pyright 0 Fehler, pytest 327 passed / 1 skipped (Docling-Layoutmodell nicht im Cache); Vitest betroffene Dateien grün
  • verify: lokal grün (Exit 0, Node 24) gegen echtes PostgreSQL 17 + SeaweedFS – Unit 135/135, Integration 97/97, depcruise ohne Verstöße, Build, Audit; CI check (inkl. AI-Service-Schritt): siehe Checks dieses PRs
  • verify:full / E2E-Spec: Smoke lokal grün (Stub liefert Schema v2)
  • Akzeptanzkriterien feat(ai-service): full field set including line items #22: sechs Kopffelder → pytest End-to-End „multi_item_request_end_to_end" + Integration Verarbeitung; Positionen mit eigenem Status/Beleg → pytest Verifier/Positionen + Integration item_index; Schema- und Prompt-Version gespeichert → run.schemaVersion/promptVersion in extraction_runs (Integrationstest prüft „2"/extract_v2); Normalisierer + KW höchstens „unsicher" → pytest Normalisierer + End-to-End (calendar_week_only); Zitat nicht im Segment → unverified, auch für Positionen → pytest (test_injected_line_item_quantity_is_never_found u. a.); TS-Vertragstypen neu + Persistenz ohne Bruch bestehender Läufe → Drift-Test, Migration 0013, Integrationstests
  • Manueller Prüfnachweis: –
  • Frischer Review (P1 vor Ready-for-review; Architektur/API/DB immer): unabhängiger Subagent – 0 Blocker, 3 should-fix + 4 nits; should-fix und 2 nits behoben, 2 nits begründet belassen (PR-Kommentar)

Doku-Entscheidung (genau eine)

  • Keine langlebige Doku betroffen – Begründung: –
  • Doku betroffen und im selben PR aktualisiert:
    • Produktdoku (P0: README): –
    • Technische Doku: docs/technical/data-model.md, services/ai/README.md
    • Architekturkarte: docs/technical/architecture.md
    • ADR: –
    • CHANGELOG [Unreleased] (sichtbares Feature oder Verhalten – im selben PR, nie „später")

Entferntes oder Umbenanntes: Unique-Constraint extracted_fields_run_field_unique ersetzt durch extracted_fields_run_field_item_unique

Dateigrößen und neue Bausteine (SYSTEM.md §7)

Dateien über 500 Zeilen im Diff (Ausnahmen: generierter Code, Lockfiles, Fixtures, Migrationen, Schemas, Ressourcen, Doku, Konfiguration):

  • keine
  • bewusst belassen – Begründung: –
  • im selben PR nach fachlicher Verantwortung geteilt
  • Folge-Issue –

Über 800 Zeilen mit neuer Fachlogik oder über 1000 Zeilen (P1/P2): nicht betroffen

Neue Shared-Komponente, Utility-Datei, Adapter oder fachlicher Service:

  • nein
  • ja – gesucht nach: vorhandenen Normalisierern im AI-Service (grounding/); erweitert statt dupliziert

Subagent-Einsätze

  • AI-Service-Teil in eigenem Worktree (Subagent), vom Hauptagenten zusammengeführt und geprüft.
  • Frischer Review (read-only): Ergebnis und Auflösung als PR-Kommentar.

Risiken / offene Punkte


🤖 Generated with Claude Code

https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1

…ion access control

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…LS with integration proofs

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…boss queues

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…nd download routes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…-only audit, composite FK

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
- sign-up requires the invitation id (link) plus the invited e-mail – no takeover by address alone
- one company per user (unique index), deterministic membership lookup, actor from membership
- organization plugin accepts only admin/clerk roles
- configurable client-IP source for the auth rate limit; local secret refused in production
- invite page: zod input, 404 for clerks, shows the invitation link; signup needs the link
- tests: wrong/missing invitation id, foreign set-active/list-members, last admin, roles
- docs: operations (rate limit/proxy, recovery), data model, exceptions register (admin plugin)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…e rejection

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
… docs for #5

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…secret

The image sets NODE_ENV=production, so the previous check blocked the local compose stack.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…re script

Python 3.13 package requestflow_ai (src layout), docling/fastapi/google-genai
pinned, CPU-only torch via the PyTorch CPU index, dev tools ruff/pyright/pytest.
The synthetic PDF fixture is generated by scripts/make_fixtures.py (reportlab).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Test-first per ADR-0001 D8: quote normalisation (whitespace, case, NFKC,
hyphenation), German number and date formats, and the verifier rules that
turn unsupported model claims into unverified. Also adds the segment and
model-output types the tests build on; the verifier does not exist yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
normalize_text folds NFKC, soft hyphens, line-break hyphenation, typographic
dashes/quotes, whitespace and case. values parses German/ISO numbers and
DD.MM.YYYY, D.M.YY and ISO dates. verify_field only keeps or downgrades the
model's status: a quote not in the cited segment, an unknown segment, a value
inconsistent with the quote, found without evidence or value, and missing
with a value all become unverified with a reason.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…ction (red)

EML: body lines with 1-based line locators, From/Subject header segments,
RFC 2047 and quoted-printable decoding, HTML-only bodies, header newline
collapse, no attachment payloads. PDF: textline segments with page + top-left
bbox via docling-parse (model-free); the layout-pipeline test only runs with
AI_TEST_DOCLING_MODELS=1. Detection by magic bytes; .msg rejected for now.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
EML via the standard library email package (policy=default): From/Subject
header segments and one segment per non-empty body line, HTML-only bodies
reduced to text lines. docling's EMAIL backend was checked but emits
paragraphs without provenance, so it cannot give line locators.

PDF via docling: the default textlines pipeline reads docling-parse text
lines (page + top-left bbox, no ML models, no network); the opt-in layout
pipeline uses DocumentConverter (OCR and tables off) and the heron layout
model. docling is imported lazily.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…ex model client (red)

The model client is exercised through the real google-genai SDK with an
httpx MockTransport replaying recorded generateContent bodies: eu multi-region
URL, bearer auth, structured-output config, token usage, fail-closed init
(no project, no credentials, dev flag without key), no silent switch to API
key mode, and schema-invalid output. Adds the env-based Settings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…lient

Prompt extract_header_v1.md (system instruction) plus render_document, which
puts segment-id-prefixed lines between <document> delimiters and neutralises
delimiter-like tags inside the document. GeminiModelClient calls Vertex via
google-genai with response_schema = ModelExtraction, JSON mime type,
temperature 0 and no tools; schema-invalid output raises ModelOutputError.
build_model_client needs VERTEX_PROJECT and credentials, loads ADC eagerly,
refuses API-key mode for Vertex, and allows the Gemini API only with
AI_ALLOW_GEMINI_API_DEV=true plus GEMINI_API_KEY.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Pipeline: PDF and EML end to end with recorded model responses, a model date
contradicting its own quote, the prompt-injection mail (the recorded model
obeys and invents a quote -> unverified), and the no-text PDF that skips the
model. API: bearer auth (401 + WWW-Authenticate), response shape, request id
handling, 400/413/415/422/429/502 mapping without echoing input, JSON logs
with IDs only. Contract: the committed OpenAPI file equals the app's schema.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
… only company policies; probe proves it bites
…iner views, pins known schemas, checks app_rw membership and runtime user
…ivate, audited invite); plugin member endpoints closed (creator role unheld)
…cl. actions, concurrency, deactivation; docs
…active (last admin could leave unaudited, clerks could list members)
…no self-deactivation, ADR-0001 D6 amendment, docs
…ar weeks (#22)

Numbers are returned as plain decimals with a dot, units map to a small
canonical set (mm, cm, m, kg, t, pcs), e-mail is lowercased, phone kept as
written. A calendar week is accepted only as a week value and flagged so the
verifier can cap it at uncertain; a date computed from a week is rejected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…erified line items (#22)

The model schema and prompt extract_v2 add three header fields and line
items (description, quantity, unit, material, dimensions). Every item field
runs through the same grounding verifier as header fields; the index is
assigned from list order. A calendar week without a date is capped at
uncertain (reason calendar_week_only). extract_header_v1.md stays unchanged.
Recorded fixtures and existing key-set assertions grew additively.
Checkpoint before regenerating the OpenAPI contract.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
… regenerate contract (#22)

Synthetic e-mail with three positions, a calendar week, phone and an
injected quantity instruction; the hand-written Vertex response is replayed
at the HTTP boundary. The contract is regenerated with
scripts/export_openapi.py (additive: 3 header keys, LineItem, lineItems,
reason calendar_week_only). README documents the normalisation rules.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…items persisted with item_index, contract checks for items; tests

Fluory commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Frischer Review (unabhängiger Subagent, read-only, review-pr + AI/RAG-, Security- und Migrations-Regeln)

Ergebnis: 0 Blocker · 3 should-fix · 4 nits. Bestätigt (per direktem Aufruf von check_value): Komma/Punkt-Verwechslung rutscht nicht durch (1250 vs. „12,50" abgelehnt, 1.25 vs. „1.250" abgelehnt), KW-Datum nie „found", E-Mail-Grenzen korrekt, Positionsindizes vom Dienst statt vom Modell, Logs ohne Inhalt; TS-Konsumenten (Review, Export, Verarbeitung) korrekt, Migration sicher für bestehende Läufe.

# Schwere Befund Auflösung
1 should-fix Einheiten st/t zu locker: „St 37-2 Blech" (Stahlsorte) → pcs, „t=5" (Dicke) → t Einheit zählt nur direkt nach einer Zahl oder als ganze Zelle; Tests „St 37-2", „St 52", „t=5", „t 5 mm" (vorher grün, jetzt korrekt abgelehnt) + Zellen „Stk." / „kg"
2 should-fix README überschätzte den Injektionsschutz bei Positionen; Grenzfall ohne Test README präzisiert (Positionen + additional_requirements); Test test_known_limitation_line_item_quoting_the_injection_passes_grounding pinnt die Grenze
3 should-fix Rollback-Hinweis falsch (Migration entfernt einen Unique-Constraint); Plan-Pflicht-Checkboxen widersprüchlich Rollback-Plan im PR-Body; Plan-Pflicht korrigiert
4 nit Freigabe konnte an einem langen, nie exportierten Freitext scheitern ERP-Limits nur für die exportierten Felder; Integrationstest
5 nit Telefonmuster erlaubte Punkte → Datum „12.10.2026" als Telefon Punkte aus dem Muster entfernt; Tests
6 nit KW-Jahr geht verloren (KW 42 statt KW 42/2026) bleibt – Wert unverändert wie vom Modell, Status „unsicher"; Notiz für #25
7 nit Prompt nimmt nur die spezifischste Zusatzanforderung bleibt; Notiz für #25 (Review-UI)

Danach: AI-Service ruff/format/pyright grün, pytest 327 passed / 1 skipped; Vitest review/export 21/21; pnpm verify läuft.


Generated by Claude Code

@Fluory
Fluory marked this pull request as ready for review September 23, 2026 08:12
@Fluory
Fluory changed the base branch from claude/feat-user-management-30 to main September 23, 2026 10:00
@Fluory
Fluory merged commit 41f0388 into main Sep 23, 2026
7 of 8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(ai-service): full field set including line items

2 participants