Skip to content

Whole-blog rebuild: post header, ink code blocks, wide covers, responsive mobile covers (2608 2.2b) - #494

Merged
pftg merged 5 commits into
masterfrom
blog-rebuild-v2
Aug 20, 2026
Merged

pftg merged 5 commits into
masterfrom
blog-rebuild-v2

Conversation

@pftg

@pftg pftg commented Aug 20, 2026

Copy link
Copy Markdown
Member

Whole-blog rebuild (Paul: "rebuild the whole blog, we will measure the whole blog")

Per-phase measurement gates retired — the blog is rebuilt in full here and
ONE 28-day session-weighted read runs against the rebuilt whole afterwards
(protocol: docs/projects/2608-.../40.01-blog-engagement-baseline.md).

Post pages

  • Header per the Rescue Room prototype: date · reading time above the
    title → title → description as a proper dek → byline (slimmed to author)
  • Cover breaks wide of the text column: 900px article, 680px measure
  • Code blocks on the system ink (--rr-ink-900) instead of Chroma's
    Dracula gray — the page's dark moments now come from one palette
  • Measure fixed in px, not ch: ch is font-relative per element, so
    a shared 68ch gave the 12px meta ~480px, the H1 ~1900px and the prose a
    third width — three measures on one page
  • Dek suppressed when it just repeats the opening paragraph (dev.to
    imports: 529 of 607 posts), on two tells — trailing ... and a normalised
    prefix match against the body

Lists (index + tag pages)

  • Mobile covers restored and responsive: img-cropped.html grew
    mobileWidth + mobileSizes; rows and the feature stack covers
    full-width above the text. Mobile is where the humans are — 28% of clicks
    on 6% of impressions (GSC).
  • Feature cover is now the first mobile visual, so it loads eager.

Scope safety

All new rules are under .post-article, which only blog single.html
emits. Course chapters share .blog, single-post.css and
blog-single.css but never that class — Codex verified nothing here reaches
them (C3 owns course visuals).

Review trail

codex:codex-rescue returned BLOCK on a real regression I introduced —
stacked covers declared a fixed 430px slot while rendering full-width to
860px, so tablets got sources at ~half the pixels needed. Fixed with vw-based
mobileSizes, plus its three P2s (cover sizes overstated by border-box
padding, the dek false negative, the lazy feature cover). All verified in
built output.

Gates

bin/hugo-build green (dev + production), marketing ratchet green, blog
suite 69 tests / 0 failures, determinism proven, macOS baselines re-recorded
and reviewed via bin/record-baselines (kept 6 / restored 48 on the final
pass). Linux legs recording via CI dispatch on this branch.

🤖 Generated with Claude Code

pftg and others added 2 commits August 20, 2026 18:45
…mobile covers (2608 2.2b)

Post template (scoped .post-article - course/single.html shares the bundle
and C3 owns course visuals, nothing here reaches chapters):
- date · reading time above the title, description as dek, byline slimmed
  to author (date/time moved up)
- cover breaks wide of the text column (900px article, 680px measure)
- code blocks on the system ink (--rr-ink-900) over Chroma dracula's gray -
  !important beats the inline background noClasses emits
- FIXED measure in px, not ch: ch is font-relative PER ELEMENT, so a shared
  68ch gave the 12px meta ~480px, the H1 ~1900px and the prose a third
  width - three measures on one page
- one more legacy-cascade fight (title centering vs style.css heading
  margins) documented and scoped; same Phase 1a death list as #0066d6
- dek suppressed on dev.to imports (description auto-truncated from the
  opening paragraph, tell: trailing '...') - it printed the first sentence
  twice back to back; caught in baseline review

Lists: img-cropped.html grows mobileWidth (mobile srcset 860w instead of
the 160w that forced covers hidden); rows + feature stack covers full-width
above text on mobile - where the humans are (GSC: 28% of clicks on 6% of
impressions).

Baselines: 29 blog baselines re-recorded via bin/record-baselines' first
at-scale run (69 tests, 0 failures, kept 29 / restored 44 non-blog
rewrites); special posts re-recorded after the dek fix.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ic, eager feature

codex verdict was BLOCK on a real quality regression I introduced, plus
three P2s - all four fixed:

- P1: stacked mobile covers declared a FIXED 430px slot while the CSS
  renders them full-width to 860px, so tablets got sources ~half the pixels
  they needed (860px/DPR2 needed 1568w, got 860w). img-cropped.html now
  takes mobileSizes (a vw expression) alongside mobileWidth: rows declare
  calc(100vw - 76px) with a 784px max, the feature calc(100vw - 116px)
  with 744px. Verified in built output, both expressions correct.
- P2: post cover sizes said 900px, but global border-box makes the 900px
  article 864px of content - now (max-width:900px) calc(100vw - 36px), 864px.
- P2: the dek ... heuristic missed a post whose description matched the
  opening paragraph verbatim. Second tell added (normalised 60-char prefix
  match against .Plain) rather than patching the one file. All three cases
  verified: false-negative suppressed, hand-written lead kept, ... import
  suppressed.
- P2: the feature cover is the first mobile visual and was still lazy -
  img-cropped grew a loading param, feature passes eager.

Codex also confirmed course isolation holds (post-article appears only in
blog single.html + blog-single.css; course chapters emit .blog/.single-content
but never .post-article) and no CLS risk from the width/height attrs.

Baselines: full blog re-record, 69 tests 0 failures, 6 baselines moved
(the mobile covers + the dek-suppressed inline-style post).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Aug 20, 2026 •

Copy link
Copy Markdown
Contributor

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 11692340-7f3f-4c3e-b947-70080055bcaa


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

pftg and others added 2 commits August 20, 2026 18:57
Three concepts updated, not just a log entry:
- architecture/cover-image-pipeline.md: the new img-cropped mobileWidth /
  mobileSizes / loading params, and WHY mobileWidth alone is insufficient
  (a fixed px slot under-serves one end of the 320-784px range; the failure
  mode is a silent quality regression no test catches).
- architecture/blog-list-page.md: the three shared partials that stopped
  index/tag drift, the .Kind branch for the taxonomy root, and the three
  traps (dev-disabled term kinds vs primary navigation, the created_at date
  fallback, string-valued tags 500ing every term page).
- build/test-gates.md: bin/record-baselines replaces the hand-rolled
  FORCE_SCREENSHOT_UPDATE dance, with its glob semantics and --linux rule.

Conformant under --strict.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

The bundle-root indexes are what a cold session reads first (progressive
disclosure), so a concept that grew scope without its one-liner following
is invisible. blog-list-page now advertises tag pages + shared partials +
traps; cover-image-pipeline the responsive params; test-gates the
record-baselines wrapper.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@pftg
pftg merged commit e1fa540 into master Aug 20, 2026
6 checks passed
@pftg
pftg deleted the blog-rebuild-v2 branch August 20, 2026 17:16
pftg added a commit that referenced this pull request Aug 20, 2026
The pass queued on #516 could not run while the session sat on #518, where
this bundle state did not exist. #516 has merged, so it runs now.

**Two corrections to workflows/site-redesign-rollout.md, written four passes
ago and both inherited from the plan doc without independent verification:**

* The engagement figure. The concept repeated "25.2% scroll / 26.3s vs a
  32.9-40.3% site average". Measured: one 3-day Clarity window of five, and
  the lowest. The windows swing 2.9x (29.89/51.13/75.11/50.91/25.56) and
  session-weighted across 743 bot-filtered sessions the blog sits at 44.31% /
  34.97s - at or above the average it was said to trail. The low window also
  straddles the 08-20 deploy, so the clean pre-ship baseline is 08-06->08-17:
  451 sessions, 56.4% / 40.1s. Blog-first still holds on a better fact - GSC
  puts the blog at 77% of the site's entire Google traffic.

* The course coupling. The concept repeated 20.01's "2.2 couples the course
  page". True of the FILE, false of the SELECTORS: course/single.html:55
  renders class="single-content" with no .post-article, and all 15 styled
  rules in pages/blog-single.css are .post-article-prefixed. DECOUPLED - 2.3
  need not follow 2.2. The genuinely shared file is single-post.css, which
  also drives bin/generate-template-pdfs.

Both failures share a cause, now named as a rule in that concept: **check
phase status against GIT, not the plan table.** Phases 2.1 and 2.2 had already
shipped (#487 and #494, both 2026-08-20) while the plan still listed them
pending, and a status answer was given from the table. A plan records what was
decided; only the tree records what shipped.

**Added:**

* build/test-gates.md - a skip_area mask blinds a gate STRUCTURALLY where
  tolerance blinds it statistically. All four blog-index screenshots mask
  .post-feature, which IS the feature slot, so the index content area has
  never been visually gated at any tolerance, and two phases shipped through
  that hole. Also: local gates are the merge authority while CI is unreliable,
  with the resulting Linux-red debt stated rather than hidden; and quote the
  `[snap_diff] N screenshots compared` count, since a suite that compared
  nothing also prints "0 failures".

* workflows/analytics-access.md - /blog/ fires no scroll_depth at all
  (page/analytics.html:72 gates on .IsPage, false for list pages), so GA4
  cannot see the blog index and Clarity is the only instrument that can. Plus
  the 3-day-window trap: session-weight across every window, and check whether
  a window straddles a deploy.

* workflows/review-swarm.md - non-colliding agents can still collide with an
  unmerged branch, and never switch branches under a running agent (it
  silently changes files it is mid-read of and nothing errors).

okf validate --strict: conformant, no warnings on any edited concept.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
pftg added a commit that referenced this pull request Aug 20, 2026
Every finding verified against the tree before acting. All eight valid.

**P1 - the one that got published.** `.okf/workflows/analytics-access.md`
claimed Clarity's per-page numbers "contradict its own aggregate for an
identical window and page-set" (~9% vs 25.56%, ~3x). Not the same page-set:
~9% was session-weighted over the TOP TEN pages, 25.56% covers EVERY /blog/
page. The omitted long tail can account for the entire gap. No disagreement was
demonstrated and per-post analysis was never ruled out - it needs the full page
rows. Corrected in the concept AND in 40.01, kept as a worked near-miss because
the shape recurs: the API returns a top-N subset by default and the aggregate
on request, so comparing them is the most available mistake to make. The
section's own rule is "state the denominator".

**P1 - alias inventory would have broken live CSS.** 20.03 omitted
blog-list.css:78 and named vibe-code-rescue.css, which has ZERO var(--rr-*)
references. Verified inventory now in the spec: blog-list.css (9 lines),
single-post.css:434-491 (6), blog-single.css (3). single-post.css carries CTA
and tag colour/background declarations reaching the course bundle, so deleting
the aliases in 1a.4 on the old list would have broken blog AND course.

**P1 - headline baseline was contaminated.** Deploy time now confirmed: #487 at
17:35 and #494 at 20:16 on 2026-08-20, both inside the 08-18->20 window. The
34.97s/743-session headline is superseded by the clean 08-06->08-17 window (451
sessions, 56.4% / 40.1s), with the 12-vs-28-day length mismatch recorded as a
follow-on rather than papered over.

**P1 - Linux baselines.** Codex is right that CLAUDE.md:148 requires both legs
before a PR. That is knowingly overridden (Paul 2026-08-21, CI unreliable). The
override and its cost - master's Linux job red until one batched dispatch - are
now stated in the spec, with an explicit instruction to do the dispatch before
merge if CI is healthy when the phase runs.

**P2 fixes:** the analytics gate cannot use `eq .Section "blog"` for tag pages
(hugo.toml:37-41 rewrites the term PERMALINK; the taxonomy is `tag = "tags"`,
so .Section is `tags`) - needs an explicit term/taxonomy predicate; inline
!important H1 styles remain at themes/beaver/layouts/list.html:51,70 and a
stylesheet rule cannot override them; the three !importants in blog-single.css
must NOT be probed for removal - 20.02:69-81 records that they fight legacy
heading-margin rules, not the retired anchor rule, and removing them restores a
title-alignment regression; R4 marked done and linked to 40.01.

Gates: okf validate --strict conformant, no warnings on edited concepts;
bin/hugo-build clean. Docs + bundle only, no code or CSS touched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
pftg added a commit that referenced this pull request Aug 20, 2026
…rywhere

No P1s this round. Three findings were defects in the OKF bundle itself.

**My own concept contradicted itself.** site-redesign-rollout.md quoted ~145
GSC clicks at line 65 and the corrected 105 at line 79, and said "live phase
status lives in the plan doc" a few lines after establishing that the plan went
stale and status must come from git. Both fixed: neither document is a status
source, and the figure is now 105 throughout.

**The mask warning was overstated.** The masks hide the listing ROWS and the
FEATURE SLOT, not "the entire content area" - lead, filters, CTA and pagination
stay covered - and the post template has 24 dedicated baselines, so Phase 2.2
was never unguarded. Narrowed to what is true: Phase 2.1's rows and feature slot
went unseen.

**The GA4 scroll claim was too absolute.** Only the CUSTOM 25/50/75/90 milestones
are lost to the .IsPage gate; if enhanced measurement is on, the built-in
`scroll` (90%) still fires. Not verified either way here, so the concept now
says so rather than asserting GA4 sees nothing - discarding a usable signal
because a doc overstated a gap is its own error.

**Merge time is not deploy time.** The baseline treated #487/#494 merge
timestamps as proof the window was contaminated. GitHub Pages publishes on a
separate run that can lag or fail. Downgraded to CONTAMINATED-PENDING-
CONFIRMATION with the restore condition stated.

**Arithmetic:** 12 days to 28 needs 16 more, not 12 - "four more 3-day pulls"
reaches 24. Corrected in both places it appeared.

**Tag pages do not share all of 2.1.** They have no feature slot (it lives
behind a first-page guard in blog/list.html), so verifying them against the full
scope list returns a false negative.

**Scope recount:** seven edits across four files, not six across three - 3.7 was
added in review and the summary never caught up, which would let an executor
skip the taxonomy cleanup.

**Baseline churn was describing already-shipped work.** Marked historical; the
residual work in 4b requires ZERO visual delta, and a moved baseline there is a
regression to investigate, not one to accept.

**Swept the corrected metrics through the canonical summaries** - the project
README and 20.01 itself both still presented 25.2%/26.3s as current. A cold
session reads those first.

okf validate --strict conformant; bin/hugo-build clean. Docs + bundle only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
pftg added a commit that referenced this pull request Aug 20, 2026
…519)

* Phase 2 specs + pre-ship baseline: both blog phases already shipped

Three artifacts from a parallel spec/measurement pass. The headline is a
premise inversion that changes what Phase 2 work remains.

**Phases 2.1 and 2.2 already shipped on master**, both on 2026-08-20:
#487 (17:35, blog index/tags/posts restyle) and #494 (20:16, whole-blog
rebuild). Verified with `git merge-base --is-ancestor e1fa540 origin/master`.
The plan's Phase 2 table still lists them as pending rows; 20.03 and 20.04 are
therefore specs-of-record plus residual punch-lists, not to-do lists.

**The engagement number that justified blog-first does not survive recomputation.**
The plan cites 25.2% scroll / 26.3s against a 32.9-40.3% site average. That is
ONE 3-day Clarity window of five, and the lowest; the windows swing 2.9x
(29.89 / 51.13 / 75.11 / 50.91 / 25.56%). Session-weighted over all 743
bot-filtered sessions the blog sits at 44.31% scroll / 34.97s - at or above the
average it was said to trail. That window also straddles the 08-20 deploy, so
the clean pre-ship baseline is 08-06 -> 08-17: 451 sessions, 56.4% / 40.1s.

What does hold up strategically: GSC shows the blog at 105 clicks / 28d,
**77% of the entire site's Google traffic**.

**The course coupling was pointed at the wrong file.** 20.01's "2.2 note" is
true of the file and false of the selectors: `course/single.html:55` renders
`class="single-content"` with no `.post-article`, and all 15 styled rules in
`pages/blog-single.css` are `.post-article`-prefixed. Only two selectors reach
course. DECOUPLED - 2.3 need not follow 2.2. The genuinely shared file is
`single-post.css`.

Two blindnesses found and verified, both of which explain why nobody noticed
the phases had shipped:

* `/blog/` fires no `scroll_depth` at all - `page/analytics.html:72` gates on
  `.IsPage`, false for list pages. GA4 cannot see the blog index.
* All four blog/index screenshots mask `.blog-post` AND `.post-feature`
  (`desktop_site_test.rb:34,42`, `mobile_site_test.rb:25,33`). `.post-feature`
  IS the feature slot. The visual gate has never covered the index's content
  area, at any tolerance.

Gaps are recorded as gaps, not estimated: per-post scroll depth is unobtainable
(Clarity per-page 0-2% contradicts its own aggregate 25.56% for the identical
window, ~3x), no pre-ship GA4 scroll_depth exists (it shipped WITH the rebuild),
and no conversion metric exists for the window.

Open decision for Paul: the cover shipped article-bleed at 900px, not the
plan's full-bleed. Recommendation is to keep 900px - a 100vw break-out risks
horizontal body scroll across 624 post dirs and invalidates the
`sizes="...864px"` on the LCP image.

Docs only - no code, templates, or CSS touched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* OKF: run the deferred pass, and correct two of my own claims

The pass queued on #516 could not run while the session sat on #518, where
this bundle state did not exist. #516 has merged, so it runs now.

**Two corrections to workflows/site-redesign-rollout.md, written four passes
ago and both inherited from the plan doc without independent verification:**

* The engagement figure. The concept repeated "25.2% scroll / 26.3s vs a
  32.9-40.3% site average". Measured: one 3-day Clarity window of five, and
  the lowest. The windows swing 2.9x (29.89/51.13/75.11/50.91/25.56) and
  session-weighted across 743 bot-filtered sessions the blog sits at 44.31% /
  34.97s - at or above the average it was said to trail. The low window also
  straddles the 08-20 deploy, so the clean pre-ship baseline is 08-06->08-17:
  451 sessions, 56.4% / 40.1s. Blog-first still holds on a better fact - GSC
  puts the blog at 77% of the site's entire Google traffic.

* The course coupling. The concept repeated 20.01's "2.2 couples the course
  page". True of the FILE, false of the SELECTORS: course/single.html:55
  renders class="single-content" with no .post-article, and all 15 styled
  rules in pages/blog-single.css are .post-article-prefixed. DECOUPLED - 2.3
  need not follow 2.2. The genuinely shared file is single-post.css, which
  also drives bin/generate-template-pdfs.

Both failures share a cause, now named as a rule in that concept: **check
phase status against GIT, not the plan table.** Phases 2.1 and 2.2 had already
shipped (#487 and #494, both 2026-08-20) while the plan still listed them
pending, and a status answer was given from the table. A plan records what was
decided; only the tree records what shipped.

**Added:**

* build/test-gates.md - a skip_area mask blinds a gate STRUCTURALLY where
  tolerance blinds it statistically. All four blog-index screenshots mask
  .post-feature, which IS the feature slot, so the index content area has
  never been visually gated at any tolerance, and two phases shipped through
  that hole. Also: local gates are the merge authority while CI is unreliable,
  with the resulting Linux-red debt stated rather than hidden; and quote the
  `[snap_diff] N screenshots compared` count, since a suite that compared
  nothing also prints "0 failures".

* workflows/analytics-access.md - /blog/ fires no scroll_depth at all
  (page/analytics.html:72 gates on .IsPage, false for list pages), so GA4
  cannot see the blog index and Clarity is the only instrument that can. Plus
  the 3-day-window trap: session-weight across every window, and check whether
  a window straddles a deploy.

* workflows/review-swarm.md - non-colliding agents can still collide with an
  unmerged branch, and never switch branches under a running agent (it
  silently changes files it is mid-read of and nothing errors).

okf validate --strict: conformant, no warnings on any edited concept.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Fix all 8 Codex findings; one had already reached the OKF bundle

Every finding verified against the tree before acting. All eight valid.

**P1 - the one that got published.** `.okf/workflows/analytics-access.md`
claimed Clarity's per-page numbers "contradict its own aggregate for an
identical window and page-set" (~9% vs 25.56%, ~3x). Not the same page-set:
~9% was session-weighted over the TOP TEN pages, 25.56% covers EVERY /blog/
page. The omitted long tail can account for the entire gap. No disagreement was
demonstrated and per-post analysis was never ruled out - it needs the full page
rows. Corrected in the concept AND in 40.01, kept as a worked near-miss because
the shape recurs: the API returns a top-N subset by default and the aggregate
on request, so comparing them is the most available mistake to make. The
section's own rule is "state the denominator".

**P1 - alias inventory would have broken live CSS.** 20.03 omitted
blog-list.css:78 and named vibe-code-rescue.css, which has ZERO var(--rr-*)
references. Verified inventory now in the spec: blog-list.css (9 lines),
single-post.css:434-491 (6), blog-single.css (3). single-post.css carries CTA
and tag colour/background declarations reaching the course bundle, so deleting
the aliases in 1a.4 on the old list would have broken blog AND course.

**P1 - headline baseline was contaminated.** Deploy time now confirmed: #487 at
17:35 and #494 at 20:16 on 2026-08-20, both inside the 08-18->20 window. The
34.97s/743-session headline is superseded by the clean 08-06->08-17 window (451
sessions, 56.4% / 40.1s), with the 12-vs-28-day length mismatch recorded as a
follow-on rather than papered over.

**P1 - Linux baselines.** Codex is right that CLAUDE.md:148 requires both legs
before a PR. That is knowingly overridden (Paul 2026-08-21, CI unreliable). The
override and its cost - master's Linux job red until one batched dispatch - are
now stated in the spec, with an explicit instruction to do the dispatch before
merge if CI is healthy when the phase runs.

**P2 fixes:** the analytics gate cannot use `eq .Section "blog"` for tag pages
(hugo.toml:37-41 rewrites the term PERMALINK; the taxonomy is `tag = "tags"`,
so .Section is `tags`) - needs an explicit term/taxonomy predicate; inline
!important H1 styles remain at themes/beaver/layouts/list.html:51,70 and a
stylesheet rule cannot override them; the three !importants in blog-single.css
must NOT be probed for removal - 20.02:69-81 records that they fight legacy
heading-margin rules, not the retired anchor rule, and removing them restores a
title-alignment regression; R4 marked done and linked to 40.01.

Gates: okf validate --strict conformant, no warnings on edited concepts;
bin/hugo-build clean. Docs + bundle only, no code or CSS touched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* OKF: lift the Hugo permalink trap out of the spec, and correct a shipped deprecation

The Codex round fixed the 2608 specs. One finding was a durable CODE fact left
in a project doc, where nobody doing template work would find it.

* architecture/hugo-site.md — **a permalink rewrite does NOT change .Section or
  .Kind.** The taxonomy is `tag = "tags"`; only [permalinks.term] rewrites the
  URL. A page served at /blog/tags/rails/ still has .Section == "tags", so every
  `eq .Section "blog"` condition MISSES tag pages while reading as though it
  covers them — the URL says blog, the page object does not. The proposed
  analytics gate was written exactly this way and would have shipped
  instrumentation that skipped the pages it named.

* architecture/blog-list-page.md — the same drift in a second form. That concept
  already records index and tag templates drifting apart in MARKUP, fixed with
  shared partials. Unifying markup did not unify PREDICATES: a .Section guard
  added anywhere still covers one and skips the other. Also records the inline
  !important H1 styles still at list.html:51,70.

* design/site-palette.md — two corrections. --color-primary no longer "dies in
  Phase 1a.2"; it is GONE as of #518, and the seven surviving matches in the CSS
  are comments recording what each rule replaced — a loose grep reads them as
  survival. And the --rr-* alias deprecation was missing from the concept
  entirely: it now names the three live consumers and the rule that matters,
  **verify by grep at deletion time, never against a written inventory.** That
  inventory was wrong twice in one review, and single-post.css belongs to the
  COURSE bundle, so deleting the aliases early breaks blog and course together.

okf validate --strict: conformant, no warnings on any edited concept.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Close the cover deviation: keep 900px article-bleed

Decided under Paul's standing "do not wait for me" grant rather than parked.

20.01's Phase 2 table says "full-bleed"; the implementation shipped
article-bleed at 900px. Keeping 900px:

* A 100vw break-out inside a centred column is the classic source of
  horizontal body scroll, which CLAUDE.md forbids outright - across 624 post
  directories at once.
* The sizes attribute (single.html:99, :114) is written for an 864px render
  box. Full-bleed makes every one of those wrong, so this is not a CSS-only
  change but CSS plus a srcset/sizes revision on the LCP image.
* The stated purpose - one confident visual in the fold, wider than the prose
  - is already met at 900px against a 680px measure.

Reversible in one max-width plus a sizes revision, which is why it did not
need a person. The TABLE wording should be corrected to "article-bleed", not
the code chased to match the table.

bin/hugo-build clean. Docs only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Fix all 7 second-round Codex findings

**P1 - R2 still directed an executor to probe the !important removal.** I fixed
the S4a discussion last round and missed the residual that repeats it. Deleted,
with the reason inline: those declarations fight LEGACY HEADING-MARGIN rules,
not the retired anchor rule (20.02:69-81), and removing them restores the
title-alignment regression. R2 is now comment hygiene only.

**Superseded engagement figure was still the stated rationale in BOTH specs.**
25.2% / 26.3s / 219 is one 3-day window of five, the lowest, straddling the
08-20 deploy; clean pre-ship is 56.4% / 40.1s over 451 sessions, and GSC is 105
blog clicks, not 145. "Visitors leave in the first quarter" is unsupported and
is now retracted in both specs. What survives is better: those 105 clicks are
77% of the site's entire Google traffic.

**Contaminated window was still named as the active comparator.** The headline
was struck last round but the surrounding prose still quoted 743 sessions /
44.31%. Replaced with a table that makes the clean 451-session figures primary
and secondary and marks the 743 numbers as audit-only.

**The disproven population claim survived in a second place.** Gap 1 still said
per-page and aggregate cover "the same window and page-set" and concluded
protocol step 3 cannot run. Corrected: top-ten vs all-pages, so step 3 remains
EXECUTABLE and the open task is retrieving all rows. Second time this round a
fix landed in one location and missed its duplicate; swept for every corrected
claim before committing this time.

**The 4.2x bot-gap multiplier is withdrawn.** GA4 Organic Search includes Bing
and DDG; the GSC figure is Google only. Not equivalent populations, so the
multiplier is overstated. analytics-access.md:122-125 prescribes the correct
comparison and it was not run. What stands without it: Direct is 8,598 sessions,
91% of the total, which is not plausible human direct navigation.

**The "site-redesign-rollout.md does not exist" claims are closed in all three
places.** It exists and governs both specs; it read as absent only because the
specs were drafted from a worktree on an unmerged branch predating it. Recorded
generalisably: a missing-file conclusion from inside a worktree is a branch
question first.

**"blog-list.css is their last consumer" corrected** - it is one of three, and
that summary contradicted the verified inventory later in the same file. Also
fixed the line list (nine lines, not the six claimed, and 203 not 204).

bin/hugo-build clean. Docs only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Fix all 9 third-round Codex findings; sweep the corrected metrics everywhere

No P1s this round. Three findings were defects in the OKF bundle itself.

**My own concept contradicted itself.** site-redesign-rollout.md quoted ~145
GSC clicks at line 65 and the corrected 105 at line 79, and said "live phase
status lives in the plan doc" a few lines after establishing that the plan went
stale and status must come from git. Both fixed: neither document is a status
source, and the figure is now 105 throughout.

**The mask warning was overstated.** The masks hide the listing ROWS and the
FEATURE SLOT, not "the entire content area" - lead, filters, CTA and pagination
stay covered - and the post template has 24 dedicated baselines, so Phase 2.2
was never unguarded. Narrowed to what is true: Phase 2.1's rows and feature slot
went unseen.

**The GA4 scroll claim was too absolute.** Only the CUSTOM 25/50/75/90 milestones
are lost to the .IsPage gate; if enhanced measurement is on, the built-in
`scroll` (90%) still fires. Not verified either way here, so the concept now
says so rather than asserting GA4 sees nothing - discarding a usable signal
because a doc overstated a gap is its own error.

**Merge time is not deploy time.** The baseline treated #487/#494 merge
timestamps as proof the window was contaminated. GitHub Pages publishes on a
separate run that can lag or fail. Downgraded to CONTAMINATED-PENDING-
CONFIRMATION with the restore condition stated.

**Arithmetic:** 12 days to 28 needs 16 more, not 12 - "four more 3-day pulls"
reaches 24. Corrected in both places it appeared.

**Tag pages do not share all of 2.1.** They have no feature slot (it lives
behind a first-page guard in blog/list.html), so verifying them against the full
scope list returns a false negative.

**Scope recount:** seven edits across four files, not six across three - 3.7 was
added in review and the summary never caught up, which would let an executor
skip the taxonomy cleanup.

**Baseline churn was describing already-shipped work.** Marked historical; the
residual work in 4b requires ZERO visual delta, and a moved baseline there is a
regression to investigate, not one to accept.

**Swept the corrected metrics through the canonical summaries** - the project
README and 20.01 itself both still presented 25.2%/26.3s as current. A cold
session reads those first.

okf validate --strict conformant; bin/hugo-build clean. Docs + bundle only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
pftg added a commit that referenced this pull request Aug 20, 2026
…utes

A Codex review found a P1 that made the whole SEO effort a no-op, plus a
bug in my own generator. Both verified in source before fixing.

1. THE REPAIRS WERE NOT DURABLE. lib/sync/post.rb:39-45 re-pulls title
   and description from dev.to unless the frontmatter carries
   `seo_override: true`. Its own comment says "the 10-min sync cron
   clobbers any locally-edited SEO snippet". None of the 141 files I
   edited had the flag, so every rebuilt description and the hand-written
   SimpleCov title would have been reverted within ten minutes of merge -
   silently, with the commits still in history looking like they worked.

   All 140 dev.to-backed files now set it, verified by script: 140 with
   the flag, 0 missing. The SimpleCov post gets it too, since its title
   and description were both hand-written.

2. MY GENERATOR ATE UNDERSCORES. clean() ran gsub(/[*_`]/, "") to strip
   markdown emphasis, which also stripped underscores inside identifiers.
   `stringify_keys` shipped as "stringifykeys" - destroying the exact
   keyword that post targets, in the description Google reads.

   Emphasis is now stripped only when the markers actually wrap
   something, and inline code is unwrapped rather than deleted. Blast
   radius measured rather than assumed: exactly 1 of 141 descriptions
   contained an underscore, so one post was mangled, not many. All 141
   were restored to their pre-commit state and regenerated from source
   with the corrected cleaner - 140 rebuilt, 93 left for a human.

Two factual corrections in the new posts, both verified in the gems:

- R2 claimed langchainrb gives token counts "on every response". The
  Hugging Face, llama.cpp and Replicate response classes never override
  those methods, so they inherit base_response and raise
  NotImplementedError. Now scoped to the mainstream adapters, with the
  exceptions named.

- R3 said a retried POST could duplicate a row write or an email. It
  cannot. chat.rb:233 runs handle_tool_calls only after a response
  returns, and the retry middleware wraps the HTTP request below that -
  so a lost response can double-charge the completion, but the first
  attempt never reached the tool. Duplicate side effects come from
  retrying the job around the call. The post now says so.

Local macOS suite (not CI): 34 runs, 87 assertions, 0 failures, and 53
screenshots compared with no failures. That also answers the CI
Screenshot Tests failure - footer and CTA regions on homepage/services
do not reproduce locally, and the linux baselines date from #494 with
design work landed since (#503 tokens, #508 palette). Drift, not this
diff.

bin/hugo-build 8/8.
pftg added a commit that referenced this pull request Aug 20, 2026
… entries (#520)

* fix(seo): both page-1 posts were serving Google a truncated SERP entry

Two posts rank page-one and take almost no clicks: `solid queue vs
sidekiq` at position 5.2 with 0.86% CTR (116 impressions, 1 click), and
`simplecov` at 10.3 with 0.87% (115 impressions, 1 click). Expected CTR
at position 5 is roughly 6%.

I assumed by analogy to the Kamal post that this was a weak title, then
read the RENDERED output instead of the source, and the real defect was
mechanical.

`layouts/partials/seo/enhanced-meta-tags.html` truncates a blog title to
45 chars (:14), appends " | JetThoughts Blog", then truncates the whole
string at 60 with an ellipsis (:8, :22). A title over 45 therefore gets
cut twice and lands in the SERP trailing an ellipsis AFTER the brand:

  before: Solid Queue vs Sidekiq: Complete Comparison | JetThoughts…
  after:  Solid Queue vs Sidekiq: When Each Wins | JetThoughts Blog

  before: How we configure Simplecov for our Ruby on | JetThoughts…
  after:  How to Configure SimpleCov in Rails | JetThoughts Blog

That is also WHY the pipeline's "title <= 45 chars" rule exists - a rule
I had been following without knowing the mechanism.

Descriptions were broken independently. Solid Queue's ran 162 chars and
lost "included." to the 160 cap. SimpleCov's ended in a literal "..." in
the source file - a dev.to import artifact, since dev.to truncates
descriptions near 100 chars and the importer carried the ellipsis into
our meta tag verbatim.

Both descriptions now answer the query instead of listing the contents,
and both render whole.

Content untouched. Solid Queue is 2,161 words and did not need it;
SimpleCov is 561 words and DOES look thin against SimpleCov's own docs
at position 10 - flagged, not fixed here, because that is a rewrite
rather than a snippet fix.

Also in this commit: the 2026-08-20 P0 gate override recorded in 20.09,
so a later session reads the content sprint as a deliberate call rather
than a gate nobody checked.

Verified in rendered HTML, not source: bin/hugo-build 8/8 validators,
zero ellipsis in either title or description.

* fix(seo): stop truncating post titles twice - 454 of 614 shipped an ellipsis

Paul authorised dropping the brand suffix if it helped. It does, and the
underlying bug was arithmetic the template contradicted itself on.

`enhanced-meta-tags.html` cut a post title to 45, appended
" | JetThoughts Blog" (19 chars) for 64, then cut the result again at 60.
45 + 19 = 64 > 60 ALWAYS, so any title of 42+ characters was truncated
twice and reached the SERP trailing an ellipsis after the brand name:

  How we configure Simplecov for our Ruby on | JetThoughts…
  Solid Queue vs Sidekiq: Complete Comparison | JetThoughts…
  Insights from Zapier's CTO on Managing Remote…

The template's own two constants could never both hold. Its effective
safe length was 41 chars, not the 45 the code implied - which is also
why the blog pipeline's "title <= 45" rule exists, a rule I had been
following without knowing the mechanism.

Measured before the change: 454 of 614 posts (74%) carried titles over
45 chars. After: of 690 rendered post titles, exactly ONE still contains
an ellipsis - "Why AI Hasn't Blown Our Minds…Yet", where it is deliberate
punctuation.

Fix is structural rather than numeric: truncate EXACTLY once. Post titles
take the full 60 with no brand suffix; the second length guard moved
inside the non-blog branch, where the suffix is still appended and still
needs it. Google appends the site name itself when it wants one.

Diagnosis note, because the wrong answer was convincing: reading the
source suggested the blog branch already handled this. Rendered output
disagreed. Two probe markers in the template proved the second truncate
was firing on top of the first - the branch was right and ran twice.
Source-reading would have shipped a no-op.

Gates: bin/hugo-build 8/8 validators. `bin/qtest --changed` reports no
visual-affecting changes, correctly - title and meta tags move no pixels.
og:title and twitter:title verified to follow the same clean string.

* fix(seo): rebuild 141 meta descriptions the dev.to import left mid-sentence

234 of 615 posts served Google a description ending in a bare '...'.
The cause is structural: 529 posts (86% of the blog) came from dev.to,
which truncates its description near 100 chars, and the importer copied
that ellipsis into our frontmatter verbatim.

The full sentence is almost always still in the post body, so these are
rebuilt from the body's first real prose - headings, code fences, images
and list markers skipped.

STRICT on purpose: a description is only written when whole sentences
fit the 158-char budget. My first attempt fell back to a word-boundary
cut and I inspected the output before shipping it - roughly a quarter
ended in dangling phrases:

  "...and create not just a good headline, but a catchy one? No matter
   what your content type is, and if you're either writing a small"
  "...to get the results you envisioned for your"

That is worse than the ellipsis it replaced, because nothing signals to
the reader that the text was cut. So the rule is now sentence-boundary
or skip.

Cost of the stricter rule: 141 rebuilt instead of 220, and 92 left for a
human rather than auto-filled badly. Verified mechanically across all
141 written: zero end in '...', zero end without terminal punctuation.

Remaining 92 need a written description - they are posts whose opening
prose has no complete sentence inside the budget. Not attempted here.

bin/hugo-build: 8/8 validators.

* feat(content): R2 - RubyLLM vs langchainrb for Rails

Queue row R2. Researched by unpacking both gems, not from memory:
ruby_llm 1.16.0, langchainrb 0.19.5, langchainrb_rails 0.1.12.

The thesis came out of the file lists rather than a feature survey. One
gem ships chat/agent/embedding/cost/model-registry; the other ships
vectorsearch/chunker/loader/output_parsers/evals. They are not competing
implementations of one job, so the decision is "do you retrieve from your
own documents", not "which is better". That also satisfies the plan's
constraint that R2 must LINK the LangChain guides rather than cannibalize
them.

Three critics plus a cold-eyes gate ran. What they caught:

- PROVIDER COUNT WAS WRONG. I wrote 10 for ruby_llm; it registers 13
  (lib/ruby_llm.rb:104-116). My `ls providers/ | head -20` truncated the
  listing and I read the cut-off result as complete - the same error as
  an undersized grep window earlier in the session. The correct number
  is better for the post: 13 vs 13 is a tie, so the row settles nothing,
  which is the point. The old draft warned against choosing on provider
  count while printing a wrong one.

- THE STABILITY ROW WAS A CHEAP SHOT. "past 1.0 vs pre-1.0" implied
  quiet upgrades. ruby_llm ships upgrade_to_v1_7/v1_9/v1_10/v1_14
  generators and an acts_as_legacy shim - four schema migrations inside
  minor releases. Now symmetric: budget for migrations on either.

- I DISMISSED A GEM I HAD NOT OPENED. The draft waved at
  langchainrb_rails 0.1.12 with "that version tells you what to expect".
  Opening it proved me wrong: four generators (pgvector, pinecone,
  chroma, prompt), an ActiveRecord hook, a Railtie. The truth argues the
  thesis better than the smirk did - ruby_llm's generators scaffold a
  conversation, langchainrb_rails' scaffold a vector store.

- COST CLAIM OVERSTATED. langchainrb does ship prompt/completion/total
  token counts; what it lacks is price and the cache/thinking split.
  "You supply the price table" replaces "you fly blind".

- FIVE BANNED DEFINITIONAL CONSTRUCTIONS, one slogany flip, one negative
  parallelism, one "teams add last and wish they had added first"
  chiasmus. All removed; sweep now returns zero.

Rejected three critic rewrites that would have fabricated evidence - an
invented client bill going "$340 to $1,900", a count of "two of the last
four apps", an OpenAI format change "found in production". None
happened. Used the one real published incident instead (nine schemas,
1549 green tests, VCR matching on method and URI only) and checked my
citation against the source post, which caught me inflating it to "nine
agent pipelines" when it was nine schemas in one pipeline.

Model id in the sample is claude-sonnet-4-5, not gpt-4o. gpt-4o still
resolves in the registry, but the gem's own README uses a current model
and this post argues that models get retired underneath you.

Gates: bin/hugo-build 8/8. check-post-visuals back at floor 72 (the
decision tree, which routes rather than restating the table). Mermaid
pre-rendered, 499.9px viewBox = 9.36px at 390px against a 9px floor.
Rendered scroll gate at 390px: zero console errors, no page overflow.

Cover generated and verified in rendered output - cold-eyes caught that
the missing cover.png was falling back SILENTLY to the site default
og-default.jpg with no build error. og:image now resolves to the real
file.

Dropped the Raw HTTP column from the table: at 390px it rendered 520px
wide in a 464px container with overflow-x visible, so the whole column
was clipped and unreachable. Its cells all read "you write it" and it
has its own section. Table now fits exactly at 464.

Noted, not fixed: that clipping is a theme-level issue, not specific to
this post - wide tables have no scroll container. Needs its own change
plus a visual regression run.

* feat(content): R3 rescoped - what RubyLLM retries, and for how long

The queue row R3 read "rate limits, token budgets, retries, streaming
into Turbo". Audited against content/blog/ before drafting and three of
those four were already owned by posts that shipped AFTER the row was
groomed:

  token budgets  -> cost-optimization-llm-applications-token-management
  streaming      -> fibers-async-ruby-llm-streaming-rails
  rate limits    -> same post, "Rate limiting the upstream calls"

Writing it as specified would have cannibalised two posts - the same
collision that killed R9 and redirected the Kamal work. Only retries
were unclaimed, so the post is retry semantics. Verdict recorded in
20.09 and the row retired.

Researched by unpacking ruby_llm 1.16.0 and faraday-retry 2.4.0. I was
wrong twice before reading them, which is the reason the post exists:

1. Guessed the 0.1s interval was too fast for a rate limit. Wrong -
   faraday-retry reads Retry-After AND the rate-limit reset header and
   takes the larger, so a 429 gets the provider's number.
2. Then assumed a sane ceiling on that wait. Wrong - max_interval
   defaults to Float::MAX (middleware.rb:55) and ruby_llm never sets it,
   so the guard that would abandon an over-long wait never fires.

The other half is connection.rb:111 adding :post back to Faraday's
IDEMPOTENT_METHODS, which excludes POST by design. Correct for chat
completions, dangerous for a tool call with a side effect.

Fact-checker verified all eight core claims against source and caught:

- retries vs attempts off-by-one (max: 3 is three RETRIES, four attempts)
- request_timeout IS exposed and defaults to 300s. My draft said you
  cannot configure your way out; the real worst case stacks four
  attempts x 300s against up to three Retry-After waits, so well over
  half an hour rather than the five minutes one Retry-After suggests
- "would never retry anything at all" was false - a GET for the model
  list still retries. Narrowed to "a single completion"
- Faraday::RetriableResponse is listed but inert, because retry_statuses
  is never set. Source-true, behaviour-false - now labelled dormant
- the OpenAI source link implied provider header behaviour I had not
  verified. Replaced with a note telling the reader to check their own
  provider's docs

Cold-eyes caught the worst one: I cited our own pool-exhaustion post as
evidence for what happens when a model call is NOT in a job. That
incident happened INSIDE a job. It now reads "a job is not a free pass
either", which also resolves a contradiction with "inside a Sidekiq job
that is fine" three paragraphs earlier. It also pulled the model-
retirement claim back to what the source post actually supports,
matching the correction shipped in #509.

Gates: bin/hugo-build 8/8. check-post-visuals at floor 72 (decision
diagram, 257px viewBox - 12px text renders true-size at 390px, clear of
the 9px floor). Rendered scroll gate at 390px: zero console errors, no
page overflow, table fits at 464. og:image resolves to the post's own
cover derivative, verified in built HTML rather than assumed.

Ran two critics rather than four - the fact-checker and cold-eyes, which
between them caught every shipping defect on R2 while the style critics
caught style. Stating the reduction rather than implying full coverage.

* style(content): plainer English in R2 and R3, and drop the padded source lists

Paul asked for plainer English and less AI-feel. The voice guide's first
gate is exactly this - a sentence the reader has to decode has already
failed, no matter how it scores on the mechanical checks.

The sentences that needed it were the ones carrying a metaphor where a
plain word would do, or two ideas welded together:

  "Both gems brought their own centre of gravity into Rails"
    -> "Each gem brought the thing it is good at into Rails"
  "Providers are a tie at thirteen each, so that row settles nothing"
    -> "Both ship thirteen providers, so that row will not help you choose"
  "Neither gem promises you a quiet upgrade"
    -> "Neither one gives you upgrades for free"
  "That failure is quieter than it sounds"
    -> "That kind of change is easy to miss"
  "A second gem constructing its own requests doubles the surface"
    -> "twice as many places for that to hide"
  "What you keep writing yourself is the boring layer"
    -> "What you write yourself is the dull but necessary part"
  "Source-true, behaviour-inert"
    -> "It is in the code, but nothing reaches it"

Also removed both Sources blocks. Paul flagged them as redundant and the
slop critic had said the same thing earlier - six links, every one a
first-party vendor page the reader can find from the gem name. A trailing
bibliography reads as generated; thoughtbot links where the claim lives.
Each post now ends with one line naming the two things actually read and
telling the reader to check their own installed versions.

Facts unchanged: 13/13 providers, the 1549-test incident, Float::MAX, the
0.1/0.2/0.4 schedule, every version number. bin/hugo-build 8/8.

Tension worth noting for later: #510 standardised 20 posts onto a
"## Sources" heading. That was about naming lists consistently, not about
whether they should exist. This commit is the other half - do not pad one
with generic vendor links just because the heading is there.

* docs(voice): the plain-English rule needed a tech-stream example

Section 0 had one worked example, from a first-person LinkedIn post,
where the defects were borrowed drama and inverted causality. Today's
pass on two tech posts hit a different family, and it is the one that
recurs in the Rails stream: an abstraction standing where a plain word
fits.

Added the seven-row table of what shipped vs what it became - 'centre of
gravity', 'settles nothing', 'quiet upgrade', 'doubles the surface',
'Source-true, behaviour-inert'. Every one of those passed banned-word,
em-dash and slop checks. They fail the only test that matters: the
reader has to translate before they can use the sentence.

Named the tell so it is greppable in review: a noun phrase doing a
verb's job. 'Centre of gravity', 'the surface', 'the boring layer' all
gesture at a shape instead of saying what happens. The fix is the
concrete verb, after which the metaphor is unnecessary.

Also added a 'padded source list' row to the structural-patterns table.
Six first-party vendor links shipped on a draft today; Paul flagged them
as redundant and a slop critic had called the shape a generated-text
tell earlier in the same session. Worth recording because #510 had just
standardised 20 posts onto a '## Sources' heading - making a section
cheap to add is what makes padding it easy, so the naming convention and
this rule have to travel together.

* fix(seo): the description repairs would have been reverted in ten minutes

A Codex review found a P1 that made the whole SEO effort a no-op, plus a
bug in my own generator. Both verified in source before fixing.

1. THE REPAIRS WERE NOT DURABLE. lib/sync/post.rb:39-45 re-pulls title
   and description from dev.to unless the frontmatter carries
   `seo_override: true`. Its own comment says "the 10-min sync cron
   clobbers any locally-edited SEO snippet". None of the 141 files I
   edited had the flag, so every rebuilt description and the hand-written
   SimpleCov title would have been reverted within ten minutes of merge -
   silently, with the commits still in history looking like they worked.

   All 140 dev.to-backed files now set it, verified by script: 140 with
   the flag, 0 missing. The SimpleCov post gets it too, since its title
   and description were both hand-written.

2. MY GENERATOR ATE UNDERSCORES. clean() ran gsub(/[*_`]/, "") to strip
   markdown emphasis, which also stripped underscores inside identifiers.
   `stringify_keys` shipped as "stringifykeys" - destroying the exact
   keyword that post targets, in the description Google reads.

   Emphasis is now stripped only when the markers actually wrap
   something, and inline code is unwrapped rather than deleted. Blast
   radius measured rather than assumed: exactly 1 of 141 descriptions
   contained an underscore, so one post was mangled, not many. All 141
   were restored to their pre-commit state and regenerated from source
   with the corrected cleaner - 140 rebuilt, 93 left for a human.

Two factual corrections in the new posts, both verified in the gems:

- R2 claimed langchainrb gives token counts "on every response". The
  Hugging Face, llama.cpp and Replicate response classes never override
  those methods, so they inherit base_response and raise
  NotImplementedError. Now scoped to the mainstream adapters, with the
  exceptions named.

- R3 said a retried POST could duplicate a row write or an email. It
  cannot. chat.rb:233 runs handle_tool_calls only after a response
  returns, and the retry middleware wraps the HTTP request below that -
  so a lost response can double-charge the completion, but the first
  attempt never reached the tool. Duplicate side effects come from
  retrying the job around the call. The post now says so.

Local macOS suite (not CI): 34 runs, 87 assertions, 0 failures, and 53
screenshots compared with no failures. That also answers the CI
Screenshot Tests failure - footer and CTA regions on homepage/services
do not reproduce locally, and the linux baselines date from #494 with
design work landed since (#503 tokens, #508 palette). Drift, not this
diff.

bin/hugo-build 8/8.
pftg added a commit that referenced this pull request Aug 20, 2026
…state integrity fixes (#522)

* OKF: lift what three review rounds taught into concepts

#519 took three Codex rounds (8, 7, 9 findings, none cosmetic). The individual
fixes shipped with that PR. What belongs in the bundle is why a docs-only
change needed three passes.

**workflows/review-swarm.md — a correction is an edit, and it can break what
the file already held.** Round three's findings were largely defects that
rounds one and two INTRODUCED. Adding "phase status comes from git" left a
sentence four lines away still routing to the plan doc. Adding a corrected
click figure left the superseded one earlier in the same file. The correcting
mindset asks "am I right here now" and does not look sideways at the invariants
the document already carried — so after correcting a claim, re-read the WHOLE
file, not the paragraph.

Two companions to it:

* **Sweep a corrected metric through the canonical summaries.** A figure was
  fixed in two specs and a measurement record while the project README and the
  plan's own justification still presented the superseded value as current —
  and those are what a cold session reads FIRST. A number lives in more places
  than the document that owns it.
* **Budget more than one review round for docs.** Prose has no compiler and no
  test; the only gate is a reader checking claims against the tree. Across three
  rounds: a residual that would have reintroduced a known regression, an alias
  inventory that would have broken live CSS, an arithmetic error in a
  measurement plan, an overstated bot multiplier. One CLEAN round is the signal
  to stop; one round is not.

**workflows/analytics-access.md — cut a measurement window on the DEPLOY, not
the MERGE.** The first version of that correction used #487/#494 merge
timestamps as proof a Clarity window was contaminated. A merge is not a
release: Pages publishes from a separate workflow run that can lag, fail, or be
re-run. This cuts both ways — it can condemn a usable window as easily as bless
a contaminated one. Read the deployment record; failing that, mark the window
contaminated-pending-confirmation with the restore condition written down.

okf validate --strict conformant; bin/hugo-build clean. Bundle only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Fix both Codex findings, including 14 future-dated timestamps

**Every OKF timestamp written today was ~2h in the future.** I stamped LOCAL
time with a `Z` suffix: actual UTC was 2026-08-20T23:18 while the session clock
showed 2026-08-21T01:18 (+0200). Fourteen stamps across eight concepts.

That is not cosmetic in this bundle specifically: `.okf/index.md` resolves
concurrent-edit conflicts by taking the LATER timestamp, so a future-dated
stamp silently outranks a genuinely newer edit from a parallel session. All
fourteen corrected to real UTC, and the rule is now recorded next to the
conflict rule it undermines: take the value from `date -u`, never compose it
from the displayed date.

**The deploy conclusion was still asserted as fact in the same paragraph that
documents it as unconfirmed.** The added rule says cut the window on the deploy
and admits only merge timestamps were read; four lines above, the prose still
stated the low window contained the ship. Now conditional, with the restore
condition explicit: if the deploy landed after 08-20, that window is clean and
its 292 sessions come back.

Worth naming: that second defect is precisely what the rule added one commit
earlier describes — "a correction is an edit, and it can break what the file
already held." I committed it inside the rule about it, which is the strongest
evidence for the rule that could exist.

okf validate --strict conformant; bin/hugo-build clean. Bundle only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Finish both integrity fixes properly: measured timestamps, deploy state swept

Both findings were my own fixes left half-done.

**Timestamps: the sweep caught one class and missed its sibling.** I grepped
`2026-08-21T*` and corrected those, but stamps written EARLIER in the session
were also local-time-with-a-Z and already carried an 08-20 date (23:45, 23:50),
so the grep never saw them - and they sat ahead of HEAD.

More importantly the converted values were never MEASURED. I derived them by
subtracting two hours from times I had invented, and as review noted the
originating offsets are not all +0200 (commit 8fa4149 carries +0300). Inventing
more precise-looking numbers would repeat the original error.

So every stamp this session authored is normalised to one MEASURED value from
`date -u` (2026-08-20T23:28:00Z), and the log records that per-edit precision
was not recoverable. Monotonic, never future-dated, no invented precision. Other
sessions' stamps (the sonnet-5 entry at 2026-08-20T00:00Z) are untouched.

**Deploy state: I made the concept conditional and left every summary that
feeds it asserting the opposite.** 40.01 said "the deploy time is now
confirmed", the README and 20.01 called the window clean, and both phase specs
plus site-redesign-rollout repeated it - while analytics-access said no
deployment record had been read. Seven locations swept to
provisional/pending-confirmation with the restore condition stated: if the
deploy landed after 08-20, the 292 excluded sessions come back.

That is exactly the rule committed two commits ago - "sweep a corrected metric
through the canonical summaries, not just its home" - violated in the next
change. Recorded rather than quietly fixed, because the pattern is the point:
correcting the owning document feels like completing the correction, and it is
not.

okf validate --strict conformant; bin/hugo-build clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Convert the timestamps instead of flattening them; correct the sweep count

Two fixes, both to my own repair.

**Flattening destroyed what `verified:` exists for.** Normalising every affected
stamp to the sweep time fixed the future-dating and turned three distinct
verification events in test-gates.md into three identical entries, losing their
order. The field records a HISTORY of checks; overwriting it with "now" falsely
dates past verifications to the correction run and undermines the very
conflict rule the sweep was protecting.

Converted properly instead. This session ran at +0200 (verified: local 01:18 ==
UTC 23:18), so each stamp it wrote as local-with-Z converts by -2h and the
distinct values come back: test-gates keeps 21:45 / 22:00 / 22:50 in order,
seo-meta-tags 21:50, hugo-site / blog-list-page / site-palette 23:20, ci-gates
22:00, rollout 22:30. Only the three concepts actually edited in this change
carry the measured 23:28. The rule now says convert with the offset the stamp
was written at, or mark it unknown - never overwrite history with now.

**The count was wrong.** "Fourteen across eight" was itself an unmeasured
assertion in durable guidance. Measured: 18 timestamp literals across 9
concepts (11 .okf/ files touched, less index.md and log.md, which are reserved
rather than concepts). Corrected.

okf validate --strict conformant; no stamp ahead of the verified clock.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Anchor every timestamp to a verifiable commit; make the deploy state conditional throughout

**P1: -2h was the wrong conversion, and no conversion could have been right.**
Git records this repo's commits at +03:00 while the shell reports +0200 - 86c2c91
committed at 02:11:35+03:00 = 23:11:35Z, so a stamp of 23:20Z inside it claimed
to postdate its own containing commit. But the deeper problem is that the
original local times were never MEASURED; I wrote round numbers. No offset
recovers truth from an invented value.

So the stamps are no longer converted at all - each is anchored to a VERIFIABLE
event: the commit in which it landed. seo-meta-tags and ci-gates take
2026-08-20T22:27:35Z (#516). hugo-site, blog-list-page, site-palette and
test-gates take 23:11:35Z (#519). test-gates' verified entries split correctly -
`git show 8fa4149:.okf/build/test-gates.md` shows which two existed at #516, so
those carry 22:27:35Z and the later one 23:11:35Z, preserving order. The three
concepts edited in this change carry a measured `date -u` value. Every stamp is
now defensible by a command anyone can re-run.

**P2: the conditional state stopped one level short.** analytics-access and the
top of 40.01 said pending-confirmation while the labels downstream still read
"contaminated", "CLEAN", and "straddles the 08-20 ship" as fact - so a reader
following the summaries would discard a possibly-valid window regardless. Six
downstream labels made conditional, and the held-out row now states the
restore condition rather than being struck through as though settled.

okf validate --strict conformant; bin/hugo-build clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Restore distinct provenance and finish the conditional sweep

**Third time collapsing distinct events, now fixed with recovered times.** The
two opus-5 verified entries in test-gates.md originated in separate commits -
67bdd78 (2026-08-21T00:43:35+03:00 = 21:43:35Z) and 6c4b7ab
(2026-08-20T23:47:30+02:00 = 21:47:30Z) - and I had assigned both the #516
squash time, erasing their order for the second time in this branch. Both are
now their own commit's time, recoverable by anyone with `git log -1 --format=%cI`.

Worth stating plainly: I wrote the rule "convert with the originating offset or
mark unknown, never overwrite history with now" and then violated it in the same
patch, twice. The pull toward a single tidy value is strong precisely because it
LOOKS like consistency.

**The conditional sweep reached one more file.** 20.03 still said the window
"straddles the 2026-08-20 deploy" and called 08-06→17 clean, so a cold session
reading the 2.1 spec would discard a possibly-valid window as settled fact. Now
provisional with the restore condition.

Deliberately NOT changed: `.okf/log.md` still records what was believed at the
time. That file's stated contract is "records what changed, not what is true" -
rewriting dated history to match current belief would make it useless as an
audit trail. The concepts carry current truth; the log carries the sequence.

okf validate --strict conformant; bin/hugo-build clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant