diff --git a/docs/branch-review-records/fddd42a9e93d9f7e66607077eb77d2a71ad979f40926fb3f0452cd67ade9e0c8.record.md b/docs/branch-review-records/fddd42a9e93d9f7e66607077eb77d2a71ad979f40926fb3f0452cd67ade9e0c8.record.md new file mode 100644 index 0000000000..22d8164442 --- /dev/null +++ b/docs/branch-review-records/fddd42a9e93d9f7e66607077eb77d2a71ad979f40926fb3f0452cd67ade9e0c8.record.md @@ -0,0 +1 @@ +| 2026-08-18 | claude/issues-reconcile-post-phase2 | 8c13f97c31e7530be8df8b03c52ebe854ac56a82 | Dedicated ledger reconciliation, re-cut off base dda4956ff: 17 queued inbox requests applied, 4 cancellations honoured, including the Phase 2 #056 result | Self-review passed. Corrects the prior tip, which was cut against 9d832452d and left one later-arriving request pending, failing the write-discipline guard as a partial transaction. Batch now complete; inbox drained to 0 pending / 251 applied. Docs-only, no product change. | check:outstanding-issues (361 rows, no ids deleted from base dda4956ff4ad); check:ledger-write-discipline (passed dda4956ff4ad..HEAD); issues:reconcile applied 17; format | diff --git a/docs/outstanding-issues-inbox/1505ba13-1b2a-47c9-b6cc-03b060f7a878.json b/docs/outstanding-issues-inbox/applied/1505ba13-1b2a-47c9-b6cc-03b060f7a878.json similarity index 100% rename from docs/outstanding-issues-inbox/1505ba13-1b2a-47c9-b6cc-03b060f7a878.json rename to docs/outstanding-issues-inbox/applied/1505ba13-1b2a-47c9-b6cc-03b060f7a878.json diff --git a/docs/outstanding-issues-inbox/1591ee4a-ce24-4091-93ba-ac4e7819fb60.json b/docs/outstanding-issues-inbox/applied/1591ee4a-ce24-4091-93ba-ac4e7819fb60.json similarity index 100% rename from docs/outstanding-issues-inbox/1591ee4a-ce24-4091-93ba-ac4e7819fb60.json rename to docs/outstanding-issues-inbox/applied/1591ee4a-ce24-4091-93ba-ac4e7819fb60.json diff --git a/docs/outstanding-issues-inbox/1ead9bb9-5e9d-4b54-b530-87586d594f0a.json b/docs/outstanding-issues-inbox/applied/1ead9bb9-5e9d-4b54-b530-87586d594f0a.json similarity index 100% rename from docs/outstanding-issues-inbox/1ead9bb9-5e9d-4b54-b530-87586d594f0a.json rename to docs/outstanding-issues-inbox/applied/1ead9bb9-5e9d-4b54-b530-87586d594f0a.json diff --git a/docs/outstanding-issues-inbox/4302ff5b-582a-4204-94bf-d554c953900f.json b/docs/outstanding-issues-inbox/applied/4302ff5b-582a-4204-94bf-d554c953900f.json similarity index 100% rename from docs/outstanding-issues-inbox/4302ff5b-582a-4204-94bf-d554c953900f.json rename to docs/outstanding-issues-inbox/applied/4302ff5b-582a-4204-94bf-d554c953900f.json diff --git a/docs/outstanding-issues-inbox/48e960d3-36e0-4be0-af16-7959d62b1907.json b/docs/outstanding-issues-inbox/applied/48e960d3-36e0-4be0-af16-7959d62b1907.json similarity index 100% rename from docs/outstanding-issues-inbox/48e960d3-36e0-4be0-af16-7959d62b1907.json rename to docs/outstanding-issues-inbox/applied/48e960d3-36e0-4be0-af16-7959d62b1907.json diff --git a/docs/outstanding-issues-inbox/6697868a-cc70-48e3-ad6f-598ff4031110.json b/docs/outstanding-issues-inbox/applied/6697868a-cc70-48e3-ad6f-598ff4031110.json similarity index 100% rename from docs/outstanding-issues-inbox/6697868a-cc70-48e3-ad6f-598ff4031110.json rename to docs/outstanding-issues-inbox/applied/6697868a-cc70-48e3-ad6f-598ff4031110.json diff --git a/docs/outstanding-issues-inbox/779bc0e9-c28f-40ab-abda-cfdcf93ea752.json b/docs/outstanding-issues-inbox/applied/779bc0e9-c28f-40ab-abda-cfdcf93ea752.json similarity index 100% rename from docs/outstanding-issues-inbox/779bc0e9-c28f-40ab-abda-cfdcf93ea752.json rename to docs/outstanding-issues-inbox/applied/779bc0e9-c28f-40ab-abda-cfdcf93ea752.json diff --git a/docs/outstanding-issues-inbox/7c1f5850-ddcd-47da-89ad-99314d0040fe.json b/docs/outstanding-issues-inbox/applied/7c1f5850-ddcd-47da-89ad-99314d0040fe.json similarity index 100% rename from docs/outstanding-issues-inbox/7c1f5850-ddcd-47da-89ad-99314d0040fe.json rename to docs/outstanding-issues-inbox/applied/7c1f5850-ddcd-47da-89ad-99314d0040fe.json diff --git a/docs/outstanding-issues-inbox/8650e564-7835-47b3-a7e2-476f369b9eea.json b/docs/outstanding-issues-inbox/applied/8650e564-7835-47b3-a7e2-476f369b9eea.json similarity index 100% rename from docs/outstanding-issues-inbox/8650e564-7835-47b3-a7e2-476f369b9eea.json rename to docs/outstanding-issues-inbox/applied/8650e564-7835-47b3-a7e2-476f369b9eea.json diff --git a/docs/outstanding-issues-inbox/8e28c783-d655-4480-852c-bafaf0b3c07b.json b/docs/outstanding-issues-inbox/applied/8e28c783-d655-4480-852c-bafaf0b3c07b.json similarity index 100% rename from docs/outstanding-issues-inbox/8e28c783-d655-4480-852c-bafaf0b3c07b.json rename to docs/outstanding-issues-inbox/applied/8e28c783-d655-4480-852c-bafaf0b3c07b.json diff --git a/docs/outstanding-issues-inbox/993a1c72-96bd-4e48-b1a4-3ea3b603acac.json b/docs/outstanding-issues-inbox/applied/993a1c72-96bd-4e48-b1a4-3ea3b603acac.json similarity index 100% rename from docs/outstanding-issues-inbox/993a1c72-96bd-4e48-b1a4-3ea3b603acac.json rename to docs/outstanding-issues-inbox/applied/993a1c72-96bd-4e48-b1a4-3ea3b603acac.json diff --git a/docs/outstanding-issues-inbox/a1c319d6-a098-460b-8b42-6794ca6690fd.json b/docs/outstanding-issues-inbox/applied/a1c319d6-a098-460b-8b42-6794ca6690fd.json similarity index 100% rename from docs/outstanding-issues-inbox/a1c319d6-a098-460b-8b42-6794ca6690fd.json rename to docs/outstanding-issues-inbox/applied/a1c319d6-a098-460b-8b42-6794ca6690fd.json diff --git a/docs/outstanding-issues-inbox/a39af37c-b4dd-4e4b-84dc-395d86937615.json b/docs/outstanding-issues-inbox/applied/a39af37c-b4dd-4e4b-84dc-395d86937615.json similarity index 100% rename from docs/outstanding-issues-inbox/a39af37c-b4dd-4e4b-84dc-395d86937615.json rename to docs/outstanding-issues-inbox/applied/a39af37c-b4dd-4e4b-84dc-395d86937615.json diff --git a/docs/outstanding-issues-inbox/a727ac1a-1d72-41bd-88f9-76945528bc97.json b/docs/outstanding-issues-inbox/applied/a727ac1a-1d72-41bd-88f9-76945528bc97.json similarity index 100% rename from docs/outstanding-issues-inbox/a727ac1a-1d72-41bd-88f9-76945528bc97.json rename to docs/outstanding-issues-inbox/applied/a727ac1a-1d72-41bd-88f9-76945528bc97.json diff --git a/docs/outstanding-issues-inbox/d25147c8-3da0-4062-9ce0-356e59a16c63.json b/docs/outstanding-issues-inbox/applied/d25147c8-3da0-4062-9ce0-356e59a16c63.json similarity index 100% rename from docs/outstanding-issues-inbox/d25147c8-3da0-4062-9ce0-356e59a16c63.json rename to docs/outstanding-issues-inbox/applied/d25147c8-3da0-4062-9ce0-356e59a16c63.json diff --git a/docs/outstanding-issues-inbox/d958d671-362d-41ed-b72f-102357dbf156.json b/docs/outstanding-issues-inbox/applied/d958d671-362d-41ed-b72f-102357dbf156.json similarity index 100% rename from docs/outstanding-issues-inbox/d958d671-362d-41ed-b72f-102357dbf156.json rename to docs/outstanding-issues-inbox/applied/d958d671-362d-41ed-b72f-102357dbf156.json diff --git a/docs/outstanding-issues-inbox/f6798c46-7f43-4481-8e02-bcac1bf9b262.json b/docs/outstanding-issues-inbox/applied/f6798c46-7f43-4481-8e02-bcac1bf9b262.json similarity index 100% rename from docs/outstanding-issues-inbox/f6798c46-7f43-4481-8e02-bcac1bf9b262.json rename to docs/outstanding-issues-inbox/applied/f6798c46-7f43-4481-8e02-bcac1bf9b262.json diff --git a/docs/outstanding-issues.md b/docs/outstanding-issues.md index f17e5f24e5..2a9ffd12d1 100644 --- a/docs/outstanding-issues.md +++ b/docs/outstanding-issues.md @@ -93,12 +93,11 @@ removed after current-main verification; it is not missing recommended work. | 38 | `#183` | A2 | Operator — Sentry + Specialist | Next approved observability window with SENTRY_AUTH_TOKEN | 1–2 hours | Create Sentry metric alert for production DB span p95 > 500ms (`span.op:db`, environment production). **Stop:** no secret printing; blocked until token/env available. | | 39 | `#206` | A2 | Specialist — answer UI contract | With AnswerState producer work (`#207`) | 2–4 hours | `partial_retrieval` has no app-facing producer — decide RAG contract vs UI-only mapping before AnswerCard. **Stop:** no retrieval behaviour change without RAG flag. | | 40 | `#211` | A3 | High — TypeScript strictness | Dedicated migration branch | multi-PR | Plan and start `noUncheckedIndexedAccess` migration (1266 errors); highest-risk files first. **Stop:** do not flip the flag on main without a staged plan. | -| 41 | `#222` | A3 | High — headers / search chrome | During headers redesign decision | 2–4 hours | Decide whether mode-home-template / search-results-header-band are in PageHeader scope or permanently out. **Stop:** do not flatten phone composer ownership. | -| 42 | `#235` | A3 | High — design-system evidence | Next warmed local proof-shot pass | 1–2 hours | Capture missing ADOPTION.md §7 proof shots for adopted surfaces. **Stop:** not visual-baseline PNGs (`#118`); no Playwright snapshot commit. | -| 43 | `#239` | Optional | High — phone chrome | When phone orientation QA is available | 15–30 min | Manual phone rotation check for ResizeObserver-only phone chrome reserve. **Gate:** `verify:phone-chrome` still owns automated coverage. **Stop:** do not widen reserve heuristics without reproduction. | -| 44 | `#240` | Optional | High — design tokens | Next design-owner review | 15–30 min | Confirm tooltip visual hard-clip asymmetry with design owner (sr-only keeps full text). **Stop:** no product change without that confirmation. | -| 45 | `#242` | A2 | High — design-system baselines | After human review of Linux baselines | 1–2 hours | Commit approved Linux visual baselines and promote adoption not-committed → committed. **Stop:** never commit baselines from an unreviewed machine run. | -| 46 | `#248` | A2 | Operator — Supabase + Specialist | After PR #1614 symptom repair; approved live/history window | 1–2 hours | Investigate why 20260705180000 search-health indexes were missing on live despite applied history; decide if drift checks should catch this class. **Stop:** no hosted mutation without approval. | +| 41 | `#235` | A3 | High — design-system evidence | Next warmed local proof-shot pass | 1–2 hours | Capture missing ADOPTION.md §7 proof shots for adopted surfaces. **Stop:** not visual-baseline PNGs (`#118`); no Playwright snapshot commit. | +| 42 | `#239` | Optional | High — phone chrome | When phone orientation QA is available | 15–30 min | Manual phone rotation check for ResizeObserver-only phone chrome reserve. **Gate:** `verify:phone-chrome` still owns automated coverage. **Stop:** do not widen reserve heuristics without reproduction. | +| 43 | `#240` | Optional | High — design tokens | Next design-owner review | 15–30 min | Confirm tooltip visual hard-clip asymmetry with design owner (sr-only keeps full text). **Stop:** no product change without that confirmation. | +| 44 | `#242` | A2 | High — design-system baselines | After human review of Linux baselines | 1–2 hours | Commit approved Linux visual baselines and promote adoption not-committed → committed. **Stop:** never commit baselines from an unreviewed machine run. | +| 45 | `#248` | A2 | Operator — Supabase + Specialist | After PR #1614 symptom repair; approved live/history window | 1–2 hours | Investigate why 20260705180000 search-health indexes were missing on live despite applied history; decide if drift checks should catch this class. **Stop:** no hosted mutation without approval. | @@ -123,7 +122,7 @@ removed after current-main verification; it is not missing recommended work. | #001 | P2 | task | Semantic reranking still gated off | `RAG_SEMANTIC_RERANK_ENABLED=false` from PR #901. Do not enable until the provider-backed 36/36 retrieval-quality gate **and** an ambiguity-focused canary are explicitly approved and recorded. | `docs/process-hardening.md` (Semantic reranking rollout debt); PR #901 | 2026-07-21 | | #053 | P1 | task | Execute cross-border privacy/legal package | Execute OpenAI and Railway DPAs; decide ZDR and Australian data residency; obtain prompt-cache behavior in writing; review subprocessors; obtain APP 8 and APP 5/1 counsel sign-off. Do not represent the release as privacy-approved or alter final public privacy wording before sign-off. | `docs/openai-cross-border-basis.md`; `docs/privacy-impact-assessment.md` | 2026-07-24 | | #055 | P2 | task | Run one exact-SHA full release and PR gate | Before the next full-confidence release/handoff, record the candidate/PR SHA and run the local/provider release gates, Firefox/WebKit, required hosted CI, and actionable GitHub review-thread closure once. Stop at the first actionable failure and rerun only the repaired smallest gate. | `docs/launch-operator-runbook.md`; `docs/codex-review-protocol.md` | 2026-07-24 | -| #056 | P2 | task | Reconcile the existing staging migration history | `Clinical KB Staging` already exists as a healthy, empty Supabase/Railway tier with distinct secrets and no production clinical data, but it is behind the repository migration chain. In the next approved staging schema window, apply the exact missing migration chain, then re-run indexing, health, identity and data-boundary proof. Do not recreate the environment or copy production clinical documents. COUNTS REMEASURED 2026-08-15 (read-only connector query against ikoiolksxqxfxgiyqpnu): the gap is now 26, not 24, and it is widening rather than static. Staging holds 166 rows in supabase_migrations.schema_migrations with latest version 20260719055623; the repository holds 192 migration files, of which 16 - not fourteen - are dated after 20260719055623. So the composition is ten earlier history holes plus sixteen later migrations. The two-migration increase is simply new work landing on main while staging stands still, which is the expected behaviour of an unattended staging tier and not a new defect - but it does mean the apply chain grows every week this row stays open, and any runbook written against the earlier figures is already stale. Re-measure the counts at the start of the window rather than trusting these; they were accurate on 2026-08-15 and will not stay that way. | current-main staging verification; `docs/staging-setup.md`; `docs/operator-backlog.md` | 2026-07-27 | +| #056 | P2 | task | Reconcile the existing staging migration history | PHASE 2 RUN 2026-08-18 in an owner-authorized staging window. REPLAY COMPLETE AND PROVEN; check:drift now RUN and RED with 19 findings, which is the substantive result of this phase. Target Clinical KB Staging ref ikoiolksxqxfxgiyqpnu via the Supabase MCP connector, ref re-verified on every call; production sjrfecxgysukkwxsowpy was never a mutation target and the only production interaction was list_projects. GAP re-measured at the start of the window as 28, not the 26 recorded 2026-08-17: 166 staging rows against 194 repository files, being ten earlier history holes plus eighteen versions after 20260719055623. The chain was replayed in version order; staging now holds 194 rows, latest 20260814151000, zero statements IS NULL, and a two-way version diff against supabase/migrations is empty both ways. Every replayed row was read back and md5-compared against the repository file: all 28 match byte-for-byte. supabase db push was unavailable (SUPABASE_ACCESS_TOKEN absent per #183, DB password operator-only) and MCP apply_migration was rejected because it stamps connector-generated versions, which docs/staging-setup.md forbids; each migration ran verbatim through execute_sql with an explicit schema_migrations row carrying the repository version and name. No production clinical document was copied and no worker started: documents and document_chunks remain 0. DRIFT RESULT: 19 unexpected rows, exit 1. Because staging now carries the complete byte-verified chain, this is not staging staleness - it is the committed migration chain and supabase/schema.sql disagreeing, and check:drift builds its expected side from schema.sql. Decomposition: (a) seven match_* def_hash mismatches caused by SET work_mem, which pg_get_functiondef renders and def_hash does not strip; grep -c work_mem supabase/schema.sql returns 0 and only migration 20260724000000 sets it, so schema.sql is the stale side and the fix is repo-side, not a production deploy. The eighth work_mem target, match_document_table_facts_text, is absent from the drift list precisely because 20260724120000 re-created it without restating work_mem, which confirms the mechanism. (b) eight objects schema.sql declares that no migration creates or drops - five document_embedding_fields indexes, documents_status_idx, and the documents_updated_at and ingestion_jobs_updated_at set_updated_at triggers; the two triggers mean updated_at maintenance would silently not exist in any environment built from migrations alone. (c) three table column-set mismatches on document_chunks, rag_visual_eval_cases and rag_visual_eval_runs, not yet expanded per column. (d) one index def mismatch on document_chunks_content_trgm_idx, which is one of the two trigram indexes rebuilt in the 2026-08-14 production incident window and whose canonical definition should be confirmed against what was actually built. PROGRAMME CONSEQUENCE: until schema.sql and the chain are reconciled, a production drift finding cannot be assumed to mean production drifted; for the seven work_mem functions the opposite holds. This argues for reconciling schema.sql to the chain before spending a production window on Phase 3. BEARING ON #316: that row carries the work_mem explanation as an untested hypothesis about production's ten mismatched RPCs; it is now measured on staging for seven of them with no production call. It does not close Phase 1.2 - production reports ten, staging seven, and the residual (match_document_table_facts_text plus the _v2 outliers match_document_chunks_text_v2 and match_document_index_units_hybrid_v2) needs its own diffs. #316 was deliberately not updated from this session. ALSO FIXED HERE: scripts/check-drift.ts forwarded only three of five identity keys to checkSupabaseProjectConfig, so any staging URL resolved to production and was rejected as a mismatch; proven before and after against identical env (before: mismatch/production/sjrfecxgysukkwxsowpy, after: ready/staging/ikoiolksxqxfxgiyqpnu). Six clean-replay findings are recorded in docs/audit/live-drift-forensics-2026-08.md section Phase 2, none patched, no migration edited. #057 soak and rollback is now unblocked on parity grounds, though the drift reconciliation above should land first. | current-main staging verification; `docs/staging-setup.md`; `docs/operator-backlog.md` | 2026-07-27 | | #057 | P2 | task | Complete staging soak and rollback rehearsal | After #056, run the documented soak and rollback against an exact candidate; retain latency/error/rollback evidence. Stop on unsafe data, identity mismatch, or an unowned rollback decision. | `docs/launch-operator-runbook.md`; `docs/audit/capacity-review.md` | 2026-07-24 | | #011 | P3 | task | Auth DB-connection allocation is operator-only | Supabase Auth (GoTrue) is capped at ~10 absolute DB connections (Supabase perf advisor). Switch to **percentage-based** allocation in the Supabase **dashboard** before the first compute scale-up — **not settable via SQL/MCP** (operator-owned). Verify via a staging soak + an approval-gated read-only advisor re-check. | `docs/auth-connection-cap-runbook.md`; `docs/process-hardening.md` (Known follow-up debts) | 2026-07-21 | | #013 | P3 | rec | Route catalogue weight remains measurement-gated; public field INP is unavailable | Keep the payload work open, but correct the measurement state: on 2026-08-13 the official Chrome UX Report current-record API returned 404 no-data for the psychiatry.tools origin and every reviewed URL (root, Therapy, Documents search, DSM, Forms, Services, Specifiers, and Formulation). Field INP is therefore unavailable because the site does not meet CrUX eligibility/coverage, not unverified and not a pass. The production app deliberately has no browser Sentry bundle, so adding RUM would expand the approved telemetry and privacy envelope. Next: continue lab LCP/TBT and interaction traces; request a separate privacy/operator decision before adding browser RUM, or recheck CrUX after traffic eligibility changes. Do not block route payload fixes waiting for a field dataset that does not exist, and do not infer an INP pass from absence. | session 2026-08-13 official CrUX current-record queries; https://cruxvis.withgoogle.com/ | 2026-07-21 | @@ -159,7 +158,6 @@ removed after current-main verification; it is not missing recommended work. | #195 | P3 | task | M1: Repo-host hardening (branch protection and required checks) | **Outcome:** GitHub branch-protection rulesets and required checks match audit §8 / maturity M1. **Next:** maintainer GitHub UI work; not a repo-file change. Record evidence in the ledger when done. **Stop:** agents must not weaken required checks. | docs/maturity-backlog-workorders.md M1; #086 | 2026-07-31 | | #206 | P2 | task | AnswerState partial_retrieval has no app-facing producer | VERIFIED CORRECT 2026-08-12 — re-checked against merged main during the full ledger sweep and left unchanged: `partial_retrieval` is declared in src/lib/answer-state-types.ts:63 and handled in answer-clipboard.ts:75, but nothing in src/app or the retrieval path produces it — still no app-facing producer, as the row says. Do not synthesise it from candidate counts. This stamp exists so a later reader can tell "checked and still true" from "never looked at"; the two were indistinguishable before. PR-E step 0 found nothing in the client payload names which expected sources were unavailable (retrievalDiagnostics = candidate counts; conflictsOrGaps = prose). RetrievalStateBanner supports the state but PR-J adoption can only emit ready/stale_evidence/source_only. Next action: decide whether a separate RAG contract PR should add a named missing-source signal (governance preflight + RAG impact line + offline eval); until then do not synthesise the state from counts. Pinned by tests/answer-state-contract.test.ts and SPEC 13 / COMPONENTS 2. | PR-E step 0, session 2026-08-02 | 2026-08-02 | | #211 | P3 | task | Plan and start the noUncheckedIndexedAccess migration | **DEPRIORITISED 2026-08-12 (yield review against current main), and that judgment still holds** — each site is a local judgment, no open ledger row traces a defect to unchecked indexed access, and the diff conflicts with every open PR. Do it in scoped batches after the clinical and CI-trust work. This update carries that conclusion forward rather than replacing it; what has changed is that the batches now exist on paper and the count was wrong. **RE-MEASURED AND PLANNED 2026-08-14 in PR #1944.** The staged plan is docs/no-unchecked-indexed-access-migration-plan.md; the migration has NOT started and tsconfig.json is unchanged, so this row stays open and stays deprioritised. Measured against main at d47aa6d rather than reusing the 2026-08-02 figure: **1,445 errors across 269 files, up from 1,266**. The drift is itself a finding — the flag is off, so nothing stops new unchecked indexing landing, and any plan built on the stale count under-scopes. The measurement also reshapes the job in a way that supports doing it in batches: tests/ (713) plus design-scratch mockups (237) are two-thirds of the population and carry no production consequence, so the genuinely risky remainder is about 500 errors, not 1,445. Shape is 71 percent TS2532/TS18048, which a guard fixes; the 368 TS2345/TS2322 need a real decision about what the absent case means. Hot spots unchanged and confirmed: answer-verification.ts (41), rag-extractive-answer.ts (23), worker/main.ts (23), evidence.ts (19). Six stages, cheapest first, each flagged mechanical or manual with its own gate. Key constraint the plan records: noUncheckedIndexedAccess is a whole-project option and narrowing include does not isolate a directory, because TypeScript still reports errors in every transitively imported file — so the flag flips exactly once in the final PR and intermediate stages are verified by a baseline ratchet in the shape of scripts/design-system-contract-baseline.json. Stage 6 touches src/lib/rag/**, so the plan writes out the flag-before-editing, RAG impact line, and live-canary obligations. Stop unchanged: do not flip the flag on main ahead of the final stage. | session 2026-08-02 /ledger sweep — docs/review-findings-2026-08-02.md | 2026-08-02 | -| #222 | P3 | task | Headers surface only partially converged in PR-J: mode-home-template and search-results-header-band untouched | VERIFIED CORRECT 2026-08-12 — re-checked against merged main and left open: Still unconverged: src/components/mode-home-template.tsx defines ModeHomeStatusNotice locally (:232) and imports neither PageHeader nor the DS EmptyState; search-results-header-band.tsx is likewise untouched. Note the adjacency — in-flight PR #1842 delegates ModeHomeStatusNotice to the DS EmptyState under #221, which is a different conversion from the PageHeader question this row asks. Re-check after #1842 merges. Builder A converged DsmPageHeader, InformationPageHeader and InformationPageBreadcrumbs onto PageHeader plus Breadcrumb, and declined two files with reasons. mode-home-template.tsx ModeHomeHero is a centred display hero on the fluid text-hero token and is the slot the in-flow phone composer sits in, so converging it onto a left-aligned PageHeader is a redesign of 13 mode homes that collides with the one-composer-per-page contract. search-results-header-band.tsx is a results spine carrying status, counts and filters, not a page-title stack, so its pin tests/search-results-header-band.dom.test.tsx remains unflipped. Both are defensible; both leave the headers surface partially adopted. Next action: decide whether either is in scope at all, or record them as permanently out of the PageHeader vocabulary. Found during PR-J adoption, 2026-08-03. | session 2026-08-03 (PR-J Wave 5, Builder A) | 2026-08-02 | | #231 | P1 | issue | Generation fallbacks no longer stick in answer cache; lithium generation quality still falls back safely | PARTIAL 2026-08-12: This PR fixes the clinically consequential stale-fallback path: every answer whose routing or degraded reason contains generation_fallback is excluded from rag_response_cache. Offline evidence: 96 focused answer-route tests and 574 RAG fixture/contract tests passed. Approved live baseline/final canaries preserved 36/36 document and content recall at 1.0 with zero per-case reciprocal-rank regressions; the final 44-case answer gate had zero citation or numeric-grounding failures. A budget extension was tested and rejected: four cache-bypassed 'Lithium dosing?' probes remained grounded, cited safe extractive fallbacks at 35-40 second candidate budgets; the decisive 40-second probe completed generation in 25.272 seconds and 27.237 seconds total with route_deadline_exceeded=false, but failed generation quality. Therefore OPENAI_ANSWER_TIMEOUT_MS and the route budget are not the current residual binding cause. INSTRUMENT NOW EXISTS 2026-08-14: the "Next: instrument" half of this row is done. Commit a3bc4da adds scripts/probe-generation-quality.ts — one cache-bypassed live answer reporting the structured generation_quality_gate_reasons, provider-backed, refusing demo mode, never caching or logging the probe. The same commit adjudicates PR #1861: superseded for phase 1, close recommended, with the numeric-retry half deferred to phase 2 pending probe evidence. So do not review #1861 as though it were the live fix, and do not re-implement the probe. Next: run scripts/probe-generation-quality.ts in an environment that has OPENAI and Supabase credentials — it is blocked in offline containers, which is why it has not been run yet — then make a separate bounded output-quality fix with an offline fixture and live canary. Stop: do not increase route/provider timeouts or cache any generation fallback. INCIDENT ADDENDUM 2026-08-14 (later the same day): rung-2 evidence was then measured live - supabase_rpc_latency_ms 31610 on a semantic query (route budget 25000 starved generation), caused by the #316 dropped trigram indexes; after their owner-approved restore, 1535 (text fast path) / 8519 (hybrid). Pre-generation latency was the binding residual cause of semantic-query source-only fallbacks in that window; evidence in docs/audit/live-drift-forensics-2026-08.md. S1 (A1 phase 2) must re-verify generation_quality_gate:* dominance on healthy latency (run the probe with node --env-file=.env.local, which the probe does not load itself) before choosing a code mitigation rung. The route-budget stop condition stands unchanged. | sessions 2026-08-14: instrument adjudication + live incident probes (owner-authorized Supabase connector) | 2026-08-04 | | #235 | P3 | task | ADOPTION.md section 7 proof shots exist for only four of the adopted surfaces | CLOSURE ATTEMPTED AND REJECTED 2026-08-14 — read this before closing again. PR #1940 queued a `done` for this row citing ADOPTION.md section 7.1's per-surface executable-evidence table; the closure was cancelled on review with the reason "executable evidence does not replace the requested desktop and phone proof shots". The cancellation is correct, and the trap is worth naming: section 7.1 opens with "This PR records executable evidence RATHER THAN committing image baselines", so the very section that looks like the evidence says in its first line that it is not. A test that proves a component is mounted is not a picture of the surface, and this row asks for the picture. IN FLIGHT note retired: PR #1842 merged, so the do-not-start warning no longer applies. The requirement is unchanged. The adoption contract asks for a proof shot per adopted surface. Wave 5 captured four - DSM header, settings rows, patient panel, answer surface - and none for the forms fold, the catalogue and docs surfaces, the headers convergence, or the empty states adopted since. Section 7 therefore reads as complete while most of the adoption is unevidenced, which matters because the proof shot is what a later reader uses to tell an intended restyle from a regression (the #229 DSM eyebrow was almost rediscovered as a defect for exactly this reason). Next action: capture the missing shots against a warmed local server (npm run ensure) and attach them to section 7. Cheap and mechanical - no gate, no provider access. Stop: this is not the visual-baseline harness (#118) - do not commit Playwright snapshot PNGs or flip that job to blocking. Stop: do not close this row on unit, DOM or contract evidence of any kind. | session 2026-08-04 (DS V2 Wave 5 close-out capture) | 2026-08-04 | | #239 | P3 | rec | Manual phone rotation check for ResizeObserver-only phone chrome reserve | PR #1616 phone overlay reserve publishes only from ResizeObserver quiet-window deliveries. Desktop↔phone and late-mount recovery are covered; orientation that does not change stack height is a narrower trigger. Next: rotate a physical phone on a chrome-overlay route and confirm --phone-overlay-chrome-h updates. | PR #1616 review findings; session 2026-08-05 | 2026-08-05 | @@ -186,12 +184,12 @@ removed after current-main verification; it is not missing recommended work. | #312 | P3 | issue | check:playwright-browser-revision reporting OK does NOT mean browsers are installed — and installing the matching revision is a cheap first option | Progress 2026-08-15: PR #1965 landed on main (commit 3ec6116) and closes the false-OK gap for the pinned Chromium revision by resolving the effective cache and requiring a launchable binary. Keep this issue open: the unscoped test:e2e and release matrix also require Firefox and WebKit, and the check does not yet report whether their locked revisions are installed. Next: enumerate required and installed revisions for all browser families, with a Chromium-only-cache regression; project-scoped --project=chromium runs may continue to require Chromium alone. | session 2026-08-12; scripts/playwright-browser-preflight.mjs:127-152; scripts/run-playwright.mjs:50-53; #290 close-out; archived #255 | 2026-08-12 | | #314 | P2 | issue | Ship compact compressed registry projections and verify live transfer | UPDATE 2026-08-16: the provider-free implementation is now on current main via commit ca788d41e: view=summary/search projections and gzip are shipped in repository code. Remaining scope is external-only: deploy an exact authorized SHA, verify /api/registry/records headers and transfer bytes on that deployment, then rerun live LCP. Stop: do not close from local payload measurements, and do not deploy or query production without explicit target authorization. | Commit ca788d41e on current main; original latency audit evidence | 2026-08-13 | | #315 | P3 | rec | If the ui-smoke scroll-hide flake (archived #290) recurs, start from the reporter-stranding mechanism — and treat the old regression window as unconfirmed | Independent verification on 2026-08-13 (second session, fresh cloud container, pinned Chromium 1234 installed per #312) measured the archived #290 flake at BOTH ends of its recorded window and corrects the archive's causal story: the bad SHA 9ab3b73ad itself passed 16 recorded executions — reproducer isolated --repeat-each=5 (5 passed, ~1.0s each), one full tests/ui-smoke.spec.ts --project=chromium run (98 tests passed, 2.5m, 0 flaky), and reproducer x10 under deliberate CPU contention (6 busy-loop processes on 4 cores, run times 1.2-1.5s: 10 passed). Current main a76f280 also 5/5. So the recovery was NOT drift — the exact commit that measured 2/5-3/5 failures passes cleanly here — and the e8adde1b9..9ab3b73a window is unconfirmed; the failure was specific to the original machine's environment/load profile. Recorded as a comment on PR #1884 (issuecomment-5272932999). On recurrence, do not re-bisect first: test the stranding mechanism. computeScrollHideUpdate (src/components/clinical-dashboard/use-hide-on-scroll.ts) re-evaluates only on scroll/resize events, and its viewportHeightChanged / maxOffset-range-change guards deliberately zero accumulated down-travel (contract-asserted in tests/use-hide-on-scroll.test.ts) — so geometry churn consuming the final steps of a gesture strands the not-hidden state permanently until the next event, matching the recorded ~11.5s toHaveAttribute timeout signature (the assertion DOES auto-retry for 10s; the attribute genuinely never flips). Fastest confirmation: a diagnostic page.on('console') trace logging which guard fires per evaluation. The window itself was one PR (#1744 mode-routing, true merge a503c22) whose net diff touched no scroll-hide code — content-bisect axes, if ever needed: tests/ vs src/ split, use-home-mode-seed/use-last-app-mode neutralized, prefetchModeDestination reverted, positional heading click restored to a settle wait. Stop: any guard change is a behaviour change to protected phone chrome — needs a failing trace first, never speculatively; do not weaken the assertion or tap targets. | session 2026-08-13; PR #1884 comment; archived #290; #312 | 2026-08-13 | -| #316 | P1 | issue | Live DB has 20 currently missing repo-defined indexes and 10 retrieval RPC bodies diverge; weekly live-drift has been red since 2026-07-26 with no routing | Combined 2026-08-14 update, superseding the two partial requests cancelled in this same batch. PHASE 0 CLOSED including the forced-dispatch proof its definition of done required: live-drift dispatched on main (Actions run 31813064485) failed at the drift step, the always() capture step still ran, the migration-history step correctly skipped, and the separate drift-routing job then created issue #1963 "Live drift check failing" carrying the label, run URL, job result, trigger and the full findings block. Routing is now also covered offline by tests/live-drift-workflow.test.ts, mutation-verified. INCIDENT REPAIR, owner-approved in-session: the two retrieval-critical indexes documents_title_trgm_idx and document_chunks_content_trgm_idx were restored with CREATE INDEX CONCURRENTLY plus ANALYZE, both indisvalid and indisready at 648 kB and 68 MB, re-verified afterwards by an independent read-only query. Before and after supabase_rpc_latency_ms 31610 to 1535 on the text fast path and 8519 hybrid, with match_document_chunks_text_v2 at 14 ms. No repo schema change was needed because the definitions were already codified. CORRECTED FIGURES measured 2026-08-14, superseding the 2026-08-09 numbers this row was opened with: 10 match_* def_hash mismatches (unchanged), 20 missing_live indexes rather than 21, and the same 2 unexpected_live. ATTRIBUTION STILL OPEN: migration 20260705180000 recorded 14 executed statements so it was not mark-applied, and the 20260804110240 guard validates four other indexes and never checks this pair, so it gives no existence bound for 2026-08-04. The drop window is therefore 2026-07-05 to 2026-08-02 and the dashboard audit-history pairing remains owner action; #248 stays open. NEXT: Phase 3 RPC reconciliation before Phase 4, per the plan's ordering that the change which can alter clinical answers precedes the ones that only speed them up. Evidence: docs/audit/live-drift-forensics-2026-08.md. OFFLINE ANALYSIS 2026-08-17 (no live call made; Supabase MCP was unauthenticated in that session). THIS ROW'S 'NEXT: Phase 3' IS NOT EXECUTABLE AS WRITTEN - correcting it is the point of this update. Phase 3's prerequisite in docs/database-remediation-playbook.md is the Phase 1 dossier with '10 diffs classified', and docs/audit/live-drift-forensics-2026-08.md section 1.2 records all ten as UNCLASSIFIED with 'per-function diff hunks still pending'. Phase 3's own pasted prompt says any UNCLASSIFIED entry stays untouched and is escalated to the owner. So Phase 3 currently has ZERO executable entries and a production window would accomplish nothing. THE ACTUAL NEXT STEP IS A READ-ONLY WINDOW to finish Phase 1.2, which is far cheaper and lower risk than the production window this row's wording implies. Do not request a production window first. TESTABLE HYPOTHESIS PREPARED OFFLINE that should shorten that read-only window: def_hash is computed in migration 20260706200000 as md5 of pg_get_functiondef with block comments, line comments and all whitespace stripped - it does NOT strip SET attributes, which pg_get_functiondef renders. Migration 20260724000000_optimize_rpc_work_mem.sql applies ALTER FUNCTION ... SET work_mem = '64MB' to exactly EIGHT of the ten UNCLASSIFIED functions (match_document_chunks_hybrid, match_document_embedding_fields_hybrid, match_document_index_units_hybrid, match_document_memory_cards_hybrid, match_document_memory_cards_hybrid_v2, match_document_chunks_text, match_document_lookup_chunks_text, match_document_table_facts_text). So one read-only query - does live carry SET work_mem on those eight - plausibly classifies 8 of 10 in a single step, and if live lacks it they are repo-ahead but NON-BEHAVIOURAL for answer content (a planner memory setting affects latency, not ordering or content), so they would not need the full eval-canary pair Phase 3's repo-ahead rule assumes. This is a HYPOTHESIS, not a finding: live state was never read. The two remaining outliers, match_document_chunks_text_v2 and match_document_index_units_hybrid_v2, are absent from that migration and need their own diffs; their canonical bodies are in 20260717160000_optimize_owner_public_retrieval.sql, 20260713020000_owner_plus_public_retrieval.sql and 20260717162000_bound_versioned_retrieval_match_count.sql. DISCARDED REASONING, recorded so nobody repeats it: 'the drift manifest never mentions work_mem' is NOT evidence for or against the hypothesis, because supabase/drift-manifest.json stores only signature, def_hash and acl per function and never stores body text at all. Flag before editing: this whole surface is protected RAG retrieval, so any actual change needs the RAG-surface flag, the PR RAG impact line, and per-RPC approval as the playbook already specifies. | PR #1968 review against docs/audit/live-drift-forensics-2026-08.md, 2026-08-15 | 2026-08-13 | +| #316 | P1 | issue | Live DB has 20 currently missing repo-defined indexes and 10 retrieval RPC bodies diverge; weekly live-drift has been red since 2026-07-26 with no routing | PHASE 1.2 COMPLETE 2026-08-18 (read-only connector window, four SELECT statements, zero writes; project ref verified as sjrfecxgysukkwxsowpy before the first query). ALL TEN RPC def_hash mismatches are now CLASSIFIED and every one is attribute-only: the live pg_get_functiondef carries a SET work_mem clause that supabase/schema.sql (the manifest source, replayed by scripts/generate-drift-manifest.ts) does not, and stripping exactly that one line from the live text reproduces the manifest hash byte-for-byte for 10/10 using the exact 20260706200000 rule (md5 of pg_get_functiondef with block comments, line comments and whitespace stripped). Bodies, signatures, return shapes, plan_cache_mode/search_path clauses and ACLs are identical to the repo. ZERO body divergences, ZERO repo-ahead, ZERO UNCLASSIFIED. The PR #2017 hypothesis is confirmed for the eight AND extends to the two _v2 outliers, which also carry live-only work_mem. Split by direction: (a) MIRROR-STALE x4 - match_document_chunks_text, match_document_lookup_chunks_text, match_document_memory_cards_hybrid, match_document_memory_cards_hybrid_v2 all live 64MB = migration 20260724000000; only schema.sql omits the clause; remedy is repo-only (add the clause to schema.sql, npm run drift:manifest), no hosted change - PHASE 3 MAY EXECUTE THESE NOW. (b) LIVE-AHEAD ATTRIBUTE-ONLY x6 - match_document_chunks_hybrid, match_document_embedding_fields_hybrid, match_document_index_units_hybrid, match_document_index_units_hybrid_v2 are 128MB on live (no recorded migration sets 128MB; 20260724000000 records 64MB for the first three and nothing for the v2); match_document_chunks_text_v2 is 64MB on live with no migration ever setting it; match_document_table_facts_text is 64MB on live although the recorded chain drops it (20260724120000 recreates the function after 20260724000000; live proconfig order search_path,plan_cache_mode,work_mem proves an ALTER re-applied afterwards outside recorded history). Remedy per plan is codify-as-live: new migration ALTER FUNCTION ... SET work_mem = ordered after every recreate, plus schema.sql mirror and regenerated manifest, PR body 'RAG impact: no retrieval behaviour change - codifying already-live attribute'; the migration marks applied against an already-matching state so no hosted change. OWNER DECISIONS FLAGGED, NOT ASSERTED: (1) confirm 128MB on the four is intended, or standardise to the recorded 64MB - that direction IS a hosted change and should carry a before/after latency measurement; (2) canary exemption - work_mem is planner memory, it changes plans and latency not the ORDER BY/LIMIT result set (only rows with exactly equal sort keys could reorder), so the recommendation is that codify-as-live proceeds without an eval-canary and any hosted value change is confirmed by the Phase 5 EXPLAIN re-run rather than an eval dispatch; owner to grant. Query 4 of the session confirmed 20260724000000 is the only recorded migration mentioning work_mem (8 statements). Playbook trap-list correction recorded in the forensics file (20260724120000 DOES contain a create or replace function at line 9); playbook not edited. NEXT: Phase 2 staging parity (#056, running concurrently) then Phase 3 with the classifications above; Phase 3 needs no eval-canary approval from this dossier. Residual open question is provenance of the 128MB/_v2 settings, which pairs with the section 1.1 dashboard audit-history owner action. | Phase 1.2 read-only Supabase connector session 2026-08-18 (ref sjrfecxgysukkwxsowpy verified; SELECT only), evidence in docs/audit/live-drift-forensics-2026-08.md section 1.2 | 2026-08-13 | | #317 | P2 | task | Verify registry-backed service records preserve facet metadata | #1878 introduced the services filter-contract tree and #1882 later merged the identical tree, so no merge-conflict audit is required. Current main uses ServiceRecord.catalogPayload.tags and fixture coverage verifies 219 records. Add focused offline tests that recordToRow and rowToServiceRecord preserve all six tag dimensions and degrade safely when payloads are malformed or absent. Do not add a second facets carrier unless a failing test proves the current contract inadequate. | PR #1921 review; #1878/#1882 tree comparison; service-facets.ts; registry-records.ts | 2026-08-13 | | #318 | P1 | task | The medication interaction lexicon has never been clinically reviewed and its sign-off block is empty | docs/medication-interaction-lexicon-review.md is generated by npm run medications:lexicon-report and expands every lexicon term to the catalogue drugs it resolves to, with how many CRITICAL/HIGH rows depend on it, sorted by severe usage. It is marked UNREVIEWED and its sign-off table is unfilled, so every red and amber drug-drug interaction alert is currently an unvalidated mapping over source-backed text. The wording shown to a clinician is always verbatim catalogue prose; what is unreviewed is which drugs a phrase like 'NSAIDs' or 'CNS depressants' was taken to mean. Next: a clinician reads the term table top-down and fills in the sign-off block. Stop: do not treat check:medication-lexicon-report passing as review - that check only proves the sheet describes the current lexicon, not that the mappings are correct. WORKLIST PREPARED 2026-08-15 (PR #1991): docs/medication-lexicon-review-worklist.md gives the top ten terms by severe usage (236 of 390 severe firings, 61 percent) with resolved drug sets and six prioritised questions. TWO OF THE THREE DEFECTS THIS ROW CITES WERE ALREADY CLOSED: the ARB/Carbapenem substring match is fixed and guarded, and lithium is reachable (9 rows / 9 severe). The divergent Warfarin pair remains and is worse than stated - warfarin-vka and warfarin-anticoagulant carry 3 interaction rows each with ZERO in common, so which record is opened changes which warnings appear. TWO FIXES LANDED 2026-08-17, owner-approved, both mechanical rather than clinical. (1) DEAD SLUG: the tcas selector listed slug 'dothiepin' but the catalogue keys the drug as 'dosulepin' (same drug, current INN), so the slug matched zero records and Dosulepin - whose own record flags Toxicity in OD FATAL and Anticholinergic HIGH - fired none of the term's 20 CRITICAL/HIGH rows. Fixed; restoring the author's evident intent, corroborated by the catalogue already filing it subclass TCA. Measured effect after regenerating data/medication-interaction-index.json: 22 rows now name dosulepin as a counterparty, 20 of them CRITICAL/HIGH, up from 0 via this term; aggregate resolution is unchanged (523 rows, 362 resolved, 161 unresolved, 423 with a catalogue target) because those rows already resolved through other TCAs, so this widens counterparties inside already-resolved rows rather than resolving new ones. Durable guard added: the coverage test now fails on ANY selector slug or denySlug that resolves to no catalogue record. The pre-existing test only required a TERM to resolve to some drug, so tcas stayed green on five of its six slugs - that is exactly how this shipped. (2) THE REVIEW INSTRUMENT'S TWO BLIND SPOTS: missedClassMembers() in scripts/build-medication-lexicon-report.ts skipped any surface stem shorter than four characters, which made the check unable to fire at all for tcas and arbs (ppis was rescued by its long surface 'proton pump inhibitors'), and it read only class and subclass, never tag. So the sheet's printed 'Checks that ran and found nothing' line was false for two terms - a printed clean result that could not have found anything is worse than no line, because it retires the question. Both closed: the floor is now 3, the shortest stem any real surface produces, and the haystack includes tag. The sheet now raises the Celecoxib/Parecoxib coxib gap itself (2 flagged, up from 1). Design note recorded because the first attempt was wrong: the fix originally matched short acronyms as whole tokens, and mutation testing showed that branch did no protective work - the leading word boundary already stops 'arb' reaching inside 'Carbapenem' - while it would newly MISS a subclass spelled 'TCAs', a regression in the dangerous direction. It is a plain prefix match, pinned by a pluralised-subclass test. STILL OPEN AND STILL YOURS: the sign-off block is untouched and the sheet is still UNREVIEWED, which is the only thing that closes this row. Five clinical questions remain with their mappings deliberately unchanged - nsaids excluding Celecoxib/Parecoxib across 38 severe rows (now auto-flagged); maois excluding Moclobemide across 17 severe rows, which the sheet still CANNOT surface because Moclobemide's tag is also RIMA and RIMA/MAOI are synonyms in pharmacology but unrelated as strings; opioids including Loperamide across 35 severe rows in the false-alert direction; acei and arbs resolving to one drug each, which is catalogue coverage rather than a narrow selector (ramipril, lisinopril, irbesartan, telmisartan, valsartan are absent from the catalogue entirely); and anticoagulants including three antiplatelets while deliberately excluding Aspirin on identical class metadata. Also worth its own row: src/lib/medication-interaction-lexicon.ts alone classifies clinicalRisk FALSE under classifyPullRequestFiles, and only the generated data/medication-interaction-index.json makes a lexicon PR clinical-risk - so a lexicon edit that changes which drugs a CRITICAL phrase resolves to would skip the governance preflight if the index were not regenerated in the same PR. | PR #1923; docs/medication-interaction-lexicon-review.md; docs/samd-classification-medication-considerations.md | 2026-08-13 | | #320 | P3 | task | Crop-to-page overlay remains unbuilt; bbox already reaches viewer state at runtime but is untyped, unvalidated, and unused | **Outcome:** selecting an indexed table or diagram can highlight its region on the PDF page, or the capability is deliberately retired — either way it stops living only in a plan document. **Detail:** this is the one Phase 3 capability never built (docs/plans/document-viewer-redesign-plan.md, Phase 3 table, 'Out of scope'). It had no ledger row until now, which is how work disappears between sessions: the plan doc marks it out of scope and nothing in durable memory says it remains owed. **The data path is partially live, not dropped.** src/lib/document-detail.ts SELECTs bbox alongside the other image columns, and withImageTableMetadata spreads every selected field except metadata. bbox therefore survives the runtime response and reaches DocumentViewer's image state. The gap is static and behavioural: DocumentDetailImage in src/lib/document-detail-contract.ts does not declare bbox, ImageRow in src/components/document-viewer/types.ts aliases that contract, no normalisation validates the stored value, and no viewer code renders it. Verified against exact PR head 2ac0f48a820be62947112efbb5d0845a702dad8e on 2026-08-13. **Shape of the work, in order:** (1) establish the ingestion coordinate space and stored shape, add a normalised bbox field to DocumentDetailImage, and add a focused loader or route-serialization test proving bbox survives with the promised shape. Do not change the selected-field mapping unless that test demonstrates an actual loss. (2) Only then draw the highlight over the rendered page when a figure is selected, accounting for the virtualized page column, the per-page raster scale from resolveViewportScale, and rotation. **Why it was scoped out rather than overlooked:** the contract and normalisation work has a wider blast radius than the component-only Phase 3 diff, and crop geometry quality from ingestion is separate debt — the redesign plan's residual-risk section says not to block viewer UX on perfect crops. **Stop:** do not land the typed-contract and normalisation half inside a viewer-only PR; it changes what the document-detail API promises and needs its own review and governance preflight. Do not render raw, unvalidated bbox values — a highlight over the wrong region of a clinical source is worse than no highlight. | session 2026-08-13 document-viewer remaining-work inventory; docs/plans/document-viewer-redesign-plan.md Phase 3 table; src/lib/document-detail.ts bbox projection | 2026-08-13 | -| #321 | P3 | task | Four follow-up groups cover nine controls after #291 | Six controls in the differential comparison page stay coupled to its planned rewrite and pinned density test. The filmstrip Page unknown control is a later mechanical change. DocumentViewer needs its persistent access reason split from transient loading before classification. The pin-limit control remains a capacity-state judgement. These are four source groups and nine controls, not four controls. | PR #1778 body; verified against main 2d27039 | 2026-08-14 | -| #322 | P1 | issue | Two catalogue records are both named Warfarin and share no interaction rows, so which one a clinician opens changes the warnings | RE-GRADED TO P1 AND CORRECTED 2026-08-17 on traced evidence, not inference. The original claim that the two records carry "three interaction rows each with ZERO in common" is FACTUALLY WRONG and understates the defect. Traced in data/medication-interaction-index.json: rows 0 (Pharmacokinetic, CYP2C9 inhibitors - amiodarone/fluconazole/metronidazole/cotrimoxazole) and 2 (Dietary, vitamin K) ARE substantively shared by both records. The real defect is COMPLEMENTARY INCOMPLETENESS, and it lands precisely on psychiatric drugs. warfarin-vka carries a CRITICAL CYP2C9 INDUCERS row (carbamazepine, St John's Wort - INR collapses toward 1.0, stroke risk) which warfarin-anticoagulant does NOT have. warfarin-anticoagulant carries a HIGH bleeding row for NSAIDs/Aspirin/SSRIs with twelve counterparties including six SSRIs (citalopram, escitalopram, fluoxetine, fluvoxamine, paroxetine, sertraline) which warfarin-vka does NOT have. So neither record is complete on its own, and each omits a psychiatrically central interaction the other holds. Resolution is strictly per-slug with no union anywhere: src/lib/medication-interactions.ts:207 and :241 read INDEX.bySlug[slug], and INDEX.names maps slug to display name, not name to slugs. Practical consequence for this app's actual user: a psychiatrist checking warfarin against sertraline sees the bleeding warning only if they happened to open warfarin-anticoagulant, and sees NOTHING if they opened warfarin-vka; checking warfarin against carbamazepine is the exact mirror. Nothing on screen distinguishes the two records - both display as "Warfarin". Secondary artefact worth fixing in the same pass: each record lists the OTHER as a counterparty (warfarin-vka's row-0 counterparties include warfarin-anticoagulant), so the duplicate is being treated as a drug that interacts with itself. Scope of this finding: it establishes the STRUCTURAL inconsistency only. Whether the interaction content of either record is clinically correct is not assessed here and is not an agent's call. Next: a clinician decides whether to merge the two records, delete one, or relabel them, and confirms the merged interaction set is complete - merging is the obvious candidate precisely because the two sets are complementary. Stop: do not patch the catalogue data automatically; this is clinical content. Grouping: treat alongside #318 as the clinical-content cluster - both are unvalidated medication-interaction surfaces in a prescribing tool. | PR #1923; docs/medication-interaction-lexicon-review.md flag section; tests/medication-interaction-lexicon-coverage.test.ts | 2026-08-13 | +| #321 | P3 | task | Four follow-up groups cover nine controls after #291 | PARTIAL 18 August 2026. Of the four follow-up groups: (1) the filmstrip 'Page unknown' control is FIXED — document-image-filmstrip.tsx converted its data-driven disabled state from native disabled to aria-disabled=true + ignoreUnavailableActivation + an sr-only reason, per docs/wiring-conventions.md's stated-reason pattern (settles this one control from #291's follow-up list); tests/document-image-filmstrip.dom.test.tsx gained a focused case (aria-disabled, not natively disabled, accessible description, click is a no-op), vitest run: 3 passed. The other three groups are unchanged and still not single-PR-sized: the six differential comparison page controls remain coupled to its own planned rewrite and pinned density test; DocumentViewer's persistent-access-reason/transient-loading split is a classification design decision, not yet made; the pin-limit control remains a capacity-state judgement call. Stays open for those three. | PR #1778 body; verified against main 2d27039 | 2026-08-14 | +| #322 | P1 | issue | Two catalogue records are both named Warfarin and share no interaction rows, so which one a clinician opens changes the warnings | PR #2069 (gemini/clinical-medication-graph-dedup) is open and targets this row. As of 2026-08-18 it reconciles both Warfarin catalogue records to carry the same interaction row set (the row's core ask), but its own new coverage test (tests/medication-interaction-lexicon-coverage.test.ts) still fails: the two duplicate records don't cross-resolve each other by name. That's a content/authoring decision (cross-link vs. merge-to-one-canonical-record) left for clinical/authoring review, not yet fixed. Stays open pending that PR landing correctly — flagging so a future session doesn't open a duplicate PR for the same dedup (see #292). | PR #1923; docs/medication-interaction-lexicon-review.md flag section; tests/medication-interaction-lexicon-coverage.test.ts | 2026-08-13 | | #323 | P2 | task | 35 of 328 catalogue medications sit outside the resolved interaction graph, so the tool can never warn about them | Measured 2026-08-13 from data/medication-interaction-index.json using both endpoints of every row with a resolved counterparty: 35 of the catalogue's 328 medications sit outside the resolved interaction graph. They are concentrated in aperients (8), antibiotics (5), antidiabetics (4) and vitamins (3); psychiatry-relevant examples include topiramate and zolpidem. The former 127 count considered only inbound counterparty references and wrongly labelled source-only drugs such as celecoxib unreachable even though their own rows emit alerts. This is primarily CORPUS coverage: widening it requires authoring an interaction row or making existing source content machine-resolvable with clinical review, not indiscriminately widening lexicon selectors. PR #1923 closed the safety half - evaluateMedicationInteractions now reports unreachableCounterparties, composeMedicationVerdict treats it as incomplete so green is unreachable, and MedicationInteractionBlock names the uncovered drugs and says the absence of a warning is not evidence of safety. The generated list by class is the 'What this tool can never warn about' section of docs/medication-interaction-lexicon-review.md and refreshes with the report. Next: prioritise clinically relevant gaps on the prescribing surface. Stop: do not close this by loosening the matcher; that reintroduces the false-positive class (Sodium content, Vitamin K, hyperkalaemia prose) that was deliberately rejected. | PR #1923; docs/medication-interaction-lexicon-review.md coverage section; src/lib/medication-interactions.ts UNREACHABLE_SLUGS | 2026-08-13 | | #324 | P1 | rec | No gate detects a merged PR whose content is silently reverted by a later merge resolution | **Outcome:** the file-level merge-loss detector is delivered; one authoritative row now tracks its remaining operational decision. **Delivered:** PR #1944 added scripts/audit-merge-loss.mjs through npm run audit:merge-loss and focused tests. It compares every changed file in a bounded main-history window with the landing commit's first parent, then reports possible reverts for human review. The implementation independently rediscovered the acf78bf casualties, including the #1803 token-retirement loss, and deliberately remains advisory because blob equality cannot distinguish a deliberate revert from an accidental merge-resolution loss. **Remaining:** decide whether it runs after merges or on a schedule, who triages positive findings, and whether the separate branch-versus-squash inbox-request-loss case should be a second detector or a mode of the same tool. A scheduled or required check without a named human disposition path would become ignorable noise. **Stop:** do not reimplement the delivered script, and do not make either detector blocking or auto-close findings until that ownership decision exists. SIGNAL-TO-NOISE CHARACTERISED AND TWO FIXES LANDED 2026-08-15 (owner-approved in session; still advisory, still unscheduled, still not blocking). (1) DEFECT FOUND AND FIXED: treeEntryReader split ls-tree output on the literal two-character sequence backslash-t rather than a tab, so the tree entry kept the filename. Same-path comparisons were unaffected, which is why the tool still found real losses, but isReconciliationMove compares an inbox path against its applied/ path, so the exemption could never match. Measured at 8069188 over 14 days: 51 findings / 255 flagged files / filesExempted 0, versus 11 findings / 66 flagged files / 189 exempted after a one-character fix - the exemption the script's own docstring says exists to stop inbox noise burying the genuine #1803 signal had been dead since it was written. Root cause of the escape: every test injected entryAt directly, so bare entries compared equal whether or not the path was stripped. Closed permanently by extracting parseTreeEntry as an exported pure function and testing it against real ls-tree output; mutation-verified (reintroducing backslash-t fails 3 tests plus the self-test). (2) MECHANISM CLASSIFIER ADDED: classifyRemoval walks the commits touching each flagged file between the landing and the ref, oldest first, takes the first whose tree entry already equals the pre-landing entry, and reports whether that commit was a merge (accidental) or single-parent (usually deliberate, and its subject says why). This is what makes the report triageable: over the window, 14 of 66 flagged files were merge-resolution removals with 13 from the single documented bad merge acf78bf, while all 52 others had explanatory single-parent subjects such as 'Re-land the --shadow-tight retirement', 'rework the viewer for phone and PWA reading' and 'ci: speed iteration without weakening gates'. Merge-resolution findings now sort first; unknown is reported rather than guessed. Mutation-verified in three directions (tab bug, newest-first walk, unknown-as-deliberate). GENUINE STILL-UNREPAIRED LOSSES, re-verified against main after it advanced past 8069188: #1800 fuzzy catalogue wiring is absent from therapies.ts, specifiers.ts and factsheets-data.ts AND all three of its tests carry zero fuzzy assertions so nothing can go red (tracked by #330); #1804's removal of UniversalSearchAlsoMatches from forms mode is reverted so the component is back at forms-search-results-page.tsx lines 44 and 894 with its guard assertions reverted, APPARENTLY UNTRACKED; #1796's ALLOWED_NODE_MAJOR_VERSIONS [24, 26] allowance is gone so worker/validate-runtime.ts still hard-codes nodeMajor() !== 24, APPARENTLY UNTRACKED; #1803 and #1807 lost design-system doc status rows while their code landed, so docs and code disagree. NEXT - the three decisions this row exists for are still open and are deliberately NOT implemented: (a) schedule, recommended weekly on a 14-day window rather than post-merge, because a post-merge trigger fires roughly 380 times per 14 days here and at merge time the loss has not happened yet; (b) triage owner, recommended routing to a pinned issue reusing the live-drift routing already covered by tests/live-drift-workflow.test.ts, with one named human, and not a required check; (c) recommended ONE tool with a --mode flag rather than a second detector, since the inbox case shares the window, landing enumeration and tree-entry comparison and differs only in paths and exemptions. Also recommended: the phantom-SHA class (a ledger record asserting a fix at 720e7027, an object that does not exist) is a DIFFERENT family - a ledger assertion with no landed content, checkable with git cat-file -e - and should get its own row rather than being folded into this tool. Stop unchanged: do not make either detector blocking or auto-close findings until the ownership decision exists. | session 2026-08-13 blob sweep; PR #1944 audit implementation and tests; PR #1937 inbox-loss case; consolidated by PR #1956 review follow-up | 2026-08-13 | | #325 | P3 | rec | A queued update request can silently clobber a row that changed after the request was written | **Outcome:** the inbox cannot apply a stale rewrite over someone else's newer content without anyone noticing. **Detail:** the inbox intake fixed ID allocation — ids are assigned at reconciliation, so two branches can no longer collide on a number, which was the sharper of the two hazards. It does not address content staleness. An 'update' request carries a full replacement '--detail' string written against whatever the author read at queue time; reconciliation applies it verbatim. If the target row changed on main between queueing and reconciling, the newer content is overwritten with no signal. The multiple-pending-mutations guard does not catch this: it fires only when two requests target the same id, not when one request is simply old. **Live near-miss, 2026-08-13:** a document-viewer ledger pass was drafted against a base four days stale, and its '#215' restatement was composed from that stale reading. It was caught only because the author re-read every row against current main before queueing — a discipline, not a gate. The same pass had already had to discard a directly-allocated '#295' because main had since claimed it; that half is now structurally impossible, this half is not. **Next:** consider fingerprinting the target row at queue time — the request schema is versioned ('version: 1'), so a 'baseRow' hash could be added to add/update/done payloads and compared at reconcile, refusing (or requiring an explicit override) when the row moved underneath. Weigh against just documenting the re-read discipline: this costs a schema bump plus writer, reconcile and self-test changes, and the failure needs a multi-day-stale base to bite. **Stop:** do not make reconciliation merge or three-way-diff detail text — a replacement that silently becomes a merge is harder to reason about than one that refuses. | session 2026-08-13 document-viewer ledger truth pass, PR #1930; scripts/ledger-inbox.mjs request schema | 2026-08-13 | @@ -201,7 +199,6 @@ removed after current-main verification; it is not missing recommended work. | #329 | P2 | issue | All live mobile routes breach LCP; shared CSS delivery and JavaScript are the current bottleneck | PR #1927 is merged and deployed to Railway production at exact SHA f2abf5baf3f449a1803bedef9dc107f30b70db93. Three-sample live medians on that SHA are Documents 3374 ms, DSM 3961 ms, Forms 3507 ms, root 3819 ms, Therapy 3422 ms, and Services 3793 ms; desktop LCP is 580-679 ms and mobile CLS remains within the rule. The production CSS split is retained and reduced four canonical medians modestly, but every mobile route still breaches 2500 ms. Root trace attribution is now concrete: TTFB 283 ms, LCP render delay 3449 ms, the 46,724-byte transferred shared stylesheet completes at 3644 ms under the throttled critical-request contention, total main-thread work is 1785 ms, script evaluation is 1030 ms, and shared chunk 8322 alone consumes 870 ms CPU. This is separate from canonical #117, which continues to track the unresolved Therapy catalogue payload and per-field safety decision. Next: split the 4,251-line global stylesheet by route ownership and reduce the shared search-shell/root client boundary before repeating the same bounded live matrix. Therapy field safety review remains required for search/pathways. INP remains unverified because Lighthouse does not measure it and no usable CrUX result exists. Stop: do not strip clinical fields, weaken the Lighthouse budget, refresh a passing baseline to hide latency, or claim an INP pass. | PR #1927; Railway deployments 1224ed55-210d-443b-94e5-20f87475468c and 810cc8b3-e39a-493f-b18f-8c63d150d53f; live Web Vitals runs 31719448766 and 31719451951; PR #1933 review | 2026-08-13 | | #332 | P3 | task | Three mode-nav icon glyphs sit at 17px, off the --spacing-icon-* scale, and no gate flags them | Split out of #275 rather than folded into its badge-box token. mode-nav/mode-nav.tsx:64 and :214 and mode-nav/nav-slot-ink.tsx:44 size their with h-[1.0625rem] w-[1.0625rem] — 17px against an icon scale of 12/14/16/20/24 (--spacing-icon-xs..xl in the globals.css @theme block). #275 counted these among its five files because they share the badge's number, but they are a different role: the badge is a text-bearing box sized around its own --text-2xs numeral, these are glyphs. They are now the only consumers of that value, since the badge moved to --spacing-search-band-badge. Nothing gates this: check-icon-scale.mjs enforces only the retired 4.5 (18px) half-step and its header states it deliberately does NOT flag arbitrary h-[Nrem], because non-icon boxes legitimately use that form. So this is unguarded and will not self-report. Why it was not just fixed: snapping to size-icon-md (16px) or size-icon-lg (20px) visibly changes nav chrome at every breakpoint, and 17px is close enough to 16 that the choice looks arbitrary without seeing it rendered — a design call, not a token swap. Next: get a Chromium look at mode-nav at phone and desktop widths with the icon at 16 and at 20, pick one, then migrate all three together. If 17px turns out to be deliberate, say so in a comment at the call site and consider whether check:icon-scale should flag off-scale arbitrary icon sizes on -typed elements specifically, which would have surfaced this. Stop: do not add a 17px step to --spacing-icon-* to make the problem go away — that token block's own comment argues against widening the scale off the 4px grid, and it would sanction the drift rather than resolve it. | session 2026-08-14; split from #275; check-icon-scale.mjs header | 2026-08-14 | | #334 | P3 | issue | Claude Code web containers can ship Node 22 with no node_modules, so npm ci fails engine-strict before any work starts | Hit 2026-08-14 at the start of a Claude Code on the web session, and it blocks a session completely until worked around, so it is worth recording even though the cause is the container image rather than this repo. The container provided /opt/node20, /opt/node21 and /opt/node22 with node22 on PATH, no nvm, and no node_modules in either the primary checkout or a fresh worktree. package.json requires node >=24.15.0 <25 with engine-strict, so 'npm ci --include=dev' aborts immediately with 'notsup Required: {node: >=24.15.0 <25, npm: 11.x} Actual: {npm: 10.9.7, node: v22.22.2}'. Nothing in the repo can fix this from inside, because the failure happens before any repo script can run — .nvmrc correctly says 24 and is simply not consulted, and there is no nvm for it to drive. Workaround used, which took about a minute and is safe: fetch the current 24.x from the nodejs.org dist index, untar to /opt/node24, and prefix subsequent commands with 'export PATH=/opt/node24/bin:/opt/node24/bin:/root/.local/bin:/root/.cargo/bin:/usr/local/go/bin:/opt/node22/bin:/opt/maven/bin:/opt/gradle/bin:/opt/rbenv/bin:/root/.bun/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin'. Everything downstream then behaved normally — npm ci, the full unit suite, build, and the Playwright-free gates all passed. Worth knowing that this is a DIFFERENT surface from the Codex Cloud provisioning path: scripts/setup-codex-cloud.sh and scripts/setup-codex-worktree.mjs cover Codex, and docs/codex-cloud.md is explicit that Cloud mirrors the tracked toolchain, but neither runs for a Claude Code web session, so that hardening does not carry over. Next: decide whether this deserves repo-side help at all. Options are a short note in the AGENTS.md or CLAUDE.md orientation telling an agent to install Node 24 to /opt/node24 and re-export PATH rather than concluding the environment is broken, or a small bootstrap script equivalent to the Codex ones that a web session can run first. Prefer the note: a bootstrap script that downloads a runtime is a bigger surface than the problem. Stop: do not relax the engines range, drop engine-strict, or pass --force to get npm ci through — the Node 24 floor is enforced deliberately in several places (preinstall, check:runtime, scripts/dev-free-port.mjs) and loosening it to accommodate a bad container would disable a real guard. | session 2026-08-14; Claude Code web container for PR #1942 | 2026-08-14 | -| #336 | P3 | rec | Decide whether responsive breakpoint windows get named tokens, or stay raw min-[]/max-[] everywhere | Split out of #275 rather than guessed at. The repo defines ZERO --breakpoint-* tokens, and at least nine sites hand-write the arbitrary form: min-[414px]:max-[429px] at clinical-dashboard/result-filter-control.tsx:231, plus max-[359px] (search-heading-mockups, differentials/diagnosis-map-panel.tsx:1036, clinical-dashboard/account-setup-dialog.tsx:98) and max-[389px] (factsheets/factsheets-search-page.tsx:176, clinical-dashboard/search-results-header-band.tsx:532, factsheets-compact-view-mockups). #275 asked for the 414-429 window to be tokenised alongside the badge box; that was deliberately NOT done, because naming one window while eight peers stay raw reintroduces exactly the one-call-site drift #275 exists to stop, just on a different axis. This is a real decision with two defensible answers and it should be made once, for all of them. (a) Stay raw and say so in docs/design-system/GATES.md: the values are per-device band edges carrying measured justifications in their own comments, they are not a scale, and a Tailwind 4 --breakpoint-* entry adds BOTH the min and max variant to every utility in the build for a single consumer. (b) Name them: Tailwind 4 --breakpoint- generates : and max-:, so the 414-429 window needs two entries (414px and 430px, since max-[429px] is inclusive and max- is exclusive), and 359/389 would want their own. Note the mockup hits are design scratch and out of scope for any gate. Next: pick (a) or (b), record it in GATES.md section 3 so the next session does not re-derive it, and only then migrate. Stop: do not migrate one window ahead of the decision. | session 2026-08-14; split from #275 during the design-token relands PR | 2026-08-14 | | #337 | P3 | rec | npm run format in an uninstalled worktree runs a different Prettier than the lockfile pins and manufactures false drift | MEASURED 2026-08-14 in a Claude-on-web container during PR #1943, by running the commands rather than reasoning about them. The repo pins prettier ^3.9.6 in package.json with 3.9.6 in package-lock.json, but the container had no node_modules, so 'npm run format' (prettier --write .) resolved Prettier through npx and got 3.8.1. The older Prettier disagreed with files that are correctly formatted under the pinned version and REWROTE 31 files nobody had touched, including src/lib/rag/rag-cache.ts, src/lib/rag/rag-provider.ts, src/lib/openai.ts, src/lib/types.ts, tests/route-reachability.test.ts and several docs. Committing that output would have turned a docs-only PR into one classifyPullRequestFiles scores as ragRanking and clinicalRisk, pulling in a Clinical Governance Preflight and a RAG impact line for changes that were pure formatting noise, and would have collided with four sibling sessions working the same tree. Proof it was an artifact and not real drift: 'npx prettier@3.9.6 --check' on the same files returns 'All matched files use Prettier code style!' -- main is clean. This is the same failure class as archived row #087 (never act on a knip finding from a worktree that has not been installed) but strictly worse, because knip only reports while format WRITES, and the false result arrives already applied to the working tree. Next: make the version explicit rather than incidental -- either pin the binary in the format and format:changed scripts, or fail closed when the resolved Prettier version does not match the lockfile, so the command cannot silently run the wrong one. A pre-push guard already reconstructs an exact-lock environment for this reason (scripts/guard-push.mjs), so the precedent for refusing to trust an unpinned local Prettier exists. Stop: do not commit the output of npm run format from a worktree that has not been installed, and do not conclude formatting drift exists on main without re-checking under the pinned version. | session 2026-08-14 PR #1943; package.json ^3.9.6; package-lock.json 3.9.6; npx prettier --version 3.8.1 vs npx prettier@3.9.6 | 2026-08-14 | | #338 | P3 | issue | The visual ISSUES-LIST.html register cannot be refreshed from any non-Windows session, so it drifts silently as work moves to cloud sessions | **Outcome:** either the rendered register is refreshable from any session that can reconcile, or it is retired and the Markdown ledger is the only artifact. **Detail, observed 2026-08-14 during the reconciliation in PR #1956.** `.claude/skills/issues/SKILL.md` refreshes the register by invoking `refresh-issues-list.ps1` under the operator's Windows `.codex\scripts` directory and writing `ISSUES-LIST.html` into their OneDrive folder — both absolute Windows paths. A Linux, container, or Codex/Claude Cloud session can run `npm run issues:reconcile` perfectly well (it did: 35 requests, write-discipline verified) but cannot run the refresh and cannot even check how stale the artifact is. The skill already handles this correctly for a single run — it says a stale visual artifact must not invalidate a valid canonical transaction, which is the right call — so this is not a correctness bug. The problem is cumulative: every cloud reconciliation widens the gap, and nothing measures it, so a reader opening the HTML has no way to tell whether it is an hour or a month behind. **Why it is P3 and not higher:** `docs/outstanding-issues.md` is the canonical rendered source and is always current; only the convenience artifact drifts. **Next, cheapest first:** decide whether the register is still wanted. If yes, the smallest fix is a stamp rather than a port — have the refresh write the reconciliation commit SHA into the HTML so staleness is visible at a glance, and have reconcile print a reminder naming the commit that needs it. A full cross-platform port (a Node renderer under `scripts/`) is the larger option and probably only worth it if the register is load-bearing for someone. If nobody reads it, retiring it and deleting that skill section is cheaper than either. **Stop:** do not improvise a substitute renderer or hand-write the HTML from a cloud session — an artifact that looks refreshed but was produced by a different generator is worse than one that is visibly stale. | PR #1956 reconciliation; .claude/skills/issues/SKILL.md refresh section; session 2026-08-14 | 2026-08-14 | | #339 | P2 | task | Favourites Continue and Recent are driven by hard-coded demo timestamps; real saved items have no last-opened data | Surfaced while shipping #164 (PR #1983), which made both surfaces prominent. src/components/clinical-dashboard/favourites-command-library-page.tsx derives 'most recently used' from lastUsedScore(item.lastUsed), and item.lastUsed comes from lastUsedByItemId — a hard-coded five-entry literal keyed to demo slugs ('Today 08:44', 'Yesterday 16:12', ...). Anything else, including every real registry favourite, falls back to the literal string 'Saved', which lastUsedScore buckets at 1000. pinnedItemIds is likewise a hard-coded two-item Set. The consequence after #164: for a signed-in user with real favourites, the Continue card and the Recent panel are effectively arbitrary — every item ties at the same score and the order is whatever the source array happened to be. Note that recentQueries in the shell is search-query history, not viewed-item history, so it cannot back this. Next: add a per-favourite last-opened timestamp. Cheapest is a client-side recents store keyed by favourite id written on open; the durable version is a column on the account favourites record so it survives a device change, which is a schema plus /api/account/favourites change and needs the usual migration review. Either way, pinning should stop being a hard-coded id set. Stop: do not fabricate a timestamp at render time from anything other than a recorded open event — an invented 'last used' on a clinical reference list is worse than an honest absence. | session 2026-08-15; PR #1983; favourites-command-library-page.tsx lastUsedByItemId/pinnedItemIds | 2026-08-15 | @@ -223,6 +220,9 @@ removed after current-main verification; it is not missing recommended work. | #Q5JHBJ | P2 | task | Deploy the 20260818090000 schema_drift_snapshot v2 history probe (Phase 6.1) in an approved production window after Phase 4, triage the first unguarded no-statements report, and author fail-fast guard migrations for the pre-contract 2026-07-01..02 and 2026-07-12 rows | Repo side of remediation plan Phase 6 landed (probe migration, guard-migration contract in docs/database-drift-detection.md + AGENTS.md, tests/migration-history-guards.test.ts, tests/search-health-index-coverage.test.ts + supabase/search-health-unmonitored-indexes.json). The migration is NOT deployed: it needs the owner-approved production migration deploy window (plan approval map, Phase 6.1, after Phase 4). Expected first live run: the ~9 remaining section 1.1 versions plus the 2026-07-12 batch (20260712165915..20260712173000) are reported as unguarded no_statements findings because no repo-provable guard exists; each needs a validation guard migration (20260804110240 pattern) + a migration_history allowlist entry, not a bare allowlist. Also decide the 8 monitor-candidate indexes in the ratchet file by a required_indexes migration (Phase 4.4). Do not touch #316/#056 for this; those rows are owned by the Phase 1.2 / Phase 2 sessions. | Phase 6 worker chat 2026-08-18; docs/database-remediation-plan.md section 6; docs/audit/live-drift-forensics-2026-08.md Phase 6 | 2026-08-17 | | #43SSS0 | P3 | rec | Three spring easing tokens in globals.css are dead: zero var() references and zero utility usage | --spring-tight, --spring-bouncy and --spring-gentle (src/app/globals.css:222-224) are declared in the @theme block but have no var() consumer in any stylesheet and no generated-utility consumer in src/. Tailwind v4.3.3 tree-shakes unused theme variables, so they never reach the compiled CSS — they are source noise, not shipped weight. Found while confirming (during PR #2046) that --animate-answer-ecg survives that same tree-shaking because it IS referenced via var() from the project's own CSS; --ease-spring is the working precedent for that pattern. Next: delete the three tokens, or wire them to the motion surfaces they were intended for. | PR #2046 phone/PWA answer-progress animation defect, 2026-08-17 | 2026-08-17 | | #75JA0P | P2 | issue | Playwright runs the whole suite with reducedMotion:"reduce", so no gate reflects the default user configuration | playwright.config.ts:61 sets contextOptions: { reducedMotion: "reduce" } suite-wide, and every motion assertion has to opt out per-test via page.emulateMedia({ reducedMotion: "no-preference" }). That inversion is why three consecutive PRs (#1974, #1989, #1995) shipped green while a physical iPhone with OS Reduce Motion on showed a frozen, blank answer-progress panel: the suite never exercised the reported configuration. PR #2046 added tests/ui-phone-motion.spec.ts to cover that one surface, but the suite-wide default remains inverted for every other motion behaviour. Next: decide whether the suite default should be no-preference with reduce opted into per-test (the safer direction), or keep the current default and add a contract test that fails when a motion assertion has no explicit emulateMedia call. | PR #2046 phone/PWA answer-progress animation defect, 2026-08-17 | 2026-08-17 | +| #1PN5BM | P3 | issue | H5a residual: whether a constant similarity of 1 may contribute to a confidence label is still open, and after G1 it lives only in the hazard doc | Packet G1 (PR #2053, merged 2026-08-17) implemented owner decision Option B: buildDocumentSummaryResults now stamps similarity_origin "document_context" on document-summary rows, deriveConfidence is unchanged, and document summaries still reach "high". That closed the LEGIBILITY half of the H5a live residual -- the fabricated 1.0 is no longer indistinguishable from a perfect cosine at any surface that reads a row. It did NOT answer the underlying governance question: may a score nobody measured contribute to the confidence label a clinician reads at all? Option B was chosen because tagging has no measured safety cost while Option A (tag as synthetic_text, capping summaries at "medium") is a label downgrade without measured gain -- so the question was deferred deliberately, not resolved. The paired question row #J912J9 is being closed by G1, so once that closure reconciles this knowledge survives only in docs/clinical-hazard-analysis.md H5a and not in the queue anyone reads. NEXT: no action required unless a measured signal appears; if it does, the tag is what makes the fix cheap -- any future gate can now discriminate the document-summary route without re-deriving provenance. Guard rails already in place: tests/rag-score.test.ts pins the discriminating pair (two document_context citations >= 0.82 -> "high"; the identical scores tagged synthetic_text -> "medium"), so a silent change in either direction goes red. | Packet G1 session 2026-08-17 (PR #2053); docs/clinical-hazard-analysis.md H5a; closes-with #J912J9 | 2026-08-18 | +| #SDQSFD | P2 | rec | ci-change-scope rag_eval_changed regex misses src/lib/rag/** (post-#994 layout), so a src/lib/rag-only PR skips eval:rag:adversarial:offline and the RAG eval CI job | scripts/ci-change-scope.mjs:290 matches only src/lib/rag.ts and src/lib/rag-*.ts (the pre-#994 layout); src/lib/rag/rag.ts, src/lib/rag/rag-answer-instructions.ts, src/lib/rag/answer-composition.ts do not set rag_eval_changed=true. verify-pr-local.mjs:123-124 then selects only check:rag:fixtures and ci.yml:413-425 skips the safety/RAG eval job. Packet S2 (2026-08-18) was covered only because it also touched tests/answer-*.test.ts and scripts/fixtures/*. Fix: add a src/lib/rag/ prefix (or /^src\/lib\/rag\//) to ragEvalPatterns with a scope test proving src/lib/rag/rag.ts alone trips rag_eval_changed; workflow/policy scope, own PR (operational-risk classifier), never bundled with a RAG behaviour change. | packet S2 review, docs/rag-improvement/HANDOVER.md | 2026-08-18 | +| #DREDWA | P3 | rec | Ledger writer self-tests use only legacy numeric ids, which is why a Crockford-id lookup bug survived the ULID migration unnoticed | Fixed in PR #2053: issueRowFingerprint (scripts/check-outstanding-issues.mjs) matched only /^#(\d+)$/ and keyed on entry.number, which is null on ULID-backed rows, so it returned null for every Crockford display locator -- and ledger-inbox.mjs reads a null fingerprint as "no such row" and refuses. npm run issues:done and issues:update were therefore unusable for EVERY row minted since the ULID migration, failing with "#J912J9 is not in Open items" about a row plainly in Open items. It surfaced only because a session happened to need to close two Crockford-id rows. ROOT CAUSE OF THE SURVIVAL, not of the bug: the self-tests and fixtures in scripts/outstanding-issues.mjs and tests/outstanding-issues-writer.test.ts exercise the writer almost entirely with legacy #005/#006/#013-style ids, so no test ever drove a Crockford id through the fingerprint path. A second, subtler trap sits in the same area and is now pinned but not generally guarded: Crockford's alphabet includes 0-9, so a ULID-derived locator can be ENTIRELY digits (the writer test's own id is #041061) and is indistinguishable from a legacy id by pattern -- branching on id shape rather than resolving against the table silently misses exactly those rows, which is how the first attempt at the fix still returned null. NEXT: add a fixture row with a ULID/Crockford id (ideally an all-digit one) to the shared ledger test fixtures and drive every writer entry point -- addIssue, resolveIssue, updateIssue, issueRowFingerprint, and the ledger-inbox done/update/reconcile paths -- through both id generations, so the next lookup left behind by an id-scheme change fails a test instead of a user's command. | Packet G1 session 2026-08-17 (PR #2053), discovered while queueing the G1 closures | 2026-08-18 | ## Resolved / archive @@ -491,3 +491,5 @@ Move resolved rows here with the resolution date and a one-line outcome. Keep th | #0MSNT8 | task | Governance Option B decided: tag document-summary rows with similarity_origin 'document_context', keep the confidence label — implement as packet G1 | Implemented as packet G1 on branch claude/g1-rag-document-context-qn9ubx. Every element of the queued Option B scope landed: document_context added to the similarity_origin union in src/lib/types.ts and to src/lib/answer-stream-contract.ts (as an allow-set, so the union and the stream validator cannot drift), stamped in buildDocumentSummaryResults (src/lib/rag/rag-row-contracts.ts), deriveConfidence left unchanged, rag.ts synthetic_similarity_count left counting only synthetic_text, and docs/clinical-hazard-analysis.md H5a updated to mark the decision implemented. Four discriminating pins, each mutation-checked against the change it exists to catch: summary rows carry the tag (tests/rag-retrieval-row-contract.test.ts); two document_context citations at >= 0.82 still yield high while the identical scores tagged synthetic_text still yield medium, and the telemetry counter ignores the new value (tests/rag-score.test.ts); the stream validator accepts every declared union member with a compile-time exhaustiveness guard (tests/answer-incremental-delivery.test.ts). Gates: vitest 58/58 across the five touched suites, full unit 642 files / 6879 passed, eval:rag:offline 36 golden cases / 603 tests, check:rag:fixtures, verify:pr-local all 18 stages green. No canary per the decision. HANDOVER status row updated. Closes the paired governance question row #J912J9. | 2026-08-17 | | #6BG9X2 | task | R2 + R3: claim-support strictness rejects verbatim-faithful guideline restatements (directive normativity; topic-overlap dilution) — packet S1c | Fixed in PR #2052 (packet S1c): R2 descriptive-norm disjunct in normativeDirectiveActions (digit-anchored, descriptiveContext-guarded, adversarial negatives for care-record prose, unrelated imperatives, and incidental norm-adjacent action words) and R3 adjacent atom-free-segment topic lending confined to the overlap clause with all other gates single-segment. Measured offline before loosening: 87 -> 78 sole-overlap rejections (46 -> 42 unique), zero protective fixture flips across the 613-test offline corpus; discriminating negatives pin cross-bullet dose mis-binding, non-adjacent synthesis, and alien-topic claims. Post-merge canary pair owner-approved. Note: request file emitted via the repo inbox schema because issues:done rejects ULID display ids (issueRowFingerprint is numeric-only). | 2026-08-17 | | #J912J9 | issue | Decide whether a fabricated similarity of 1 on document-summary rows may earn the high confidence label a clinician reads | Answered and implemented (packet G1, PR for branch claude/g1-rag-document-context-qn9ubx). Owner decided Option B on 2026-08-17: document-summary rows keep the high confidence label, and the fabricated similarity gets its own provenance value rather than being folded into synthetic_text. Landed: document_context added to the similarity_origin union (src/lib/types.ts) and to the streamed-preview client-source validator (src/lib/answer-stream-contract.ts), and stamped in buildDocumentSummaryResults (src/lib/rag/rag-row-contracts.ts). Per the decision deriveConfidence (src/lib/rag/rag-answer-support.ts) is unchanged and still excludes only synthetic_text, and rag.ts synthetic_similarity_count still counts only synthetic_text; both are pinned by discriminating tests that go red on the rejected Option A fold. Option A (tag as synthetic_text so summaries cap at medium) recorded as rejected. docs/clinical-hazard-analysis.md H5a marks the decision implemented and names the residual: the tag closes the legibility gap, not the deeper question of whether a constant 1.0 should contribute to a confidence label, but any future gate can now discriminate the route without re-deriving provenance. No retrieval behaviour change; no canary. | 2026-08-17 | +| #222 | task | Headers surface only partially converged in PR-J: mode-home-template and search-results-header-band untouched | DECIDED AND RECORDED 18 August 2026 as DECISIONS.md C7. mode-home-template.tsx and search-results-header-band.tsx are permanently declared outside the PageHeader vocabulary: the mode-home hero is a centred display hero the in-flow composer sits beneath (redesign risk + one-composer-per-page collision), and the results-header band is a status/count/filter spine, not a title stack, pinned by tests/search-results-header-band.dom.test.tsx. Noted separately in the same decision: ModeHomeStatusNotice already converged onto the DS EmptyState via PR #1842 (#221) — a different, already-closed conversion from the PageHeader question this row asked. Docs-only change. | 2026-08-18 | +| #336 | rec | Decide whether responsive breakpoint windows get named tokens, or stay raw min-[]/max-[] everywhere | DECIDED AND RECORDED 18 August 2026. Chose (a): responsive breakpoint windows stay raw min-[…]/max-[…] everywhere; no --breakpoint-* @theme tokens added. Decision and full rationale recorded in docs/design-system/GATES.md §3 (prohibition-table row) and new §3b (18 Aug 2026), covering all nine current call sites (result-filter-control.tsx, search-heading-mockups.tsx x3, diagnosis-map-panel.tsx, factsheets-search-page.tsx, search-results-header-band.tsx, factsheets-compact-view-mockups.tsx). Docs-only change; no migration performed, per the row's own stop rule. | 2026-08-18 |