Content: RubyLLM agent evals + debugging posts (R7, R8) - #499
Conversation
Agent-eval-specific companion to the monitoring post: bounded schemas as the first check, a per-iteration jsonb audit trail on the run's metrics row, funnel_summary, golden sets grown from flagged runs, and eval replays against a local qwen3:0.6b. SEO: title 30 chars, description 150 chars (og identical), primary keyword in first 100 words, 5 internal links, 7 external citations. First-hand material sanitized from a production talent-matching pipeline (shapes and lessons only - no prompts, model IDs, or score bands). Reviewer verdicts: - Tech fact-check (core-reviewer): found 3 blocking errors in the draft - a false "no Rails engine for LLM testing" claim (ruby_llm-evals exists), an inverted lazy-block claim, and a config snippet that could not route to the local model. All three rewritten against the gem source at ruby_llm-1.16.0 and the ruby_llm-evals README. - Cold-eyes 9-check gate (core-reviewer, fresh context): "PUBLISH-READY" after fixing an unsourced universal claim and two rhetorical flips. - 4-eyes pre-commit (core-reviewer): "VERDICT: APPROVED FOR COMMIT" - all three blockers verified closed, including an empirical check that the published initializer resolves (Scorer < BaseAgent => model=qwen3:0.6b provider=ollama temp=0.0). Two claims were cut rather than published because the sources did not support them: "answers in milliseconds" (no benchmark) and "landed in the very next commit" (26 commits sit between the two). Gates: bin/hugo-build green, bin/check-post-visuals exit 0, mermaid pre-rendered to SVG at 14.45px effective font on a 390px viewport (floor 9px), scroll gate clean at 1280x800 and 390x844 with zero console errors and zero 404s. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Independent codex pass with a runnable probe found two blockers in the
evals post.
1. The post told readers a block handed to `temperature` or `model` is
"dropped without a warning" and "does nothing". True for temperature;
false and dangerous for model. `def model(model_id = nil, **options)`
assigns `@chat_kwargs = options`, so the block form leaves model_id
nil and wipes the inherited model and provider - the agent then runs
on config.default_model, a hosted model where the reader believed a
local one was pinned. Measured against the post's own BaseAgent:
inherited {provider: "ollama", model: "qwen3:0.6b"} became {}.
The sentence now separates the harmless case from the destructive one.
2. The citation linked crmne/ruby_llm/blob/v1.16.0/... which 404s - that
repo's tags carry no `v` prefix. Corrected to /blob/1.16.0/, verified
200. Note this is per-repo: socketry/falcon does use v0.57.0.
Also dropped "and this column landed the next day" - the same
unverifiable interval as the "very next commit" claim cut pre-commit,
in softer wording. "The metrics row came first" carries the point.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro Plus Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Third and last post of the RubyLLM x Rails queue. Its thesis is what a green Rails suite actually asserts when the thing under test lives on someone else's server, told through three incidents from the talent-matching pipeline we run. Source claims verified against installed gems rather than paraphrased, which is the failure mode that produced six wrong claims in this cluster today: - VCR DEFAULT_MATCHERS = [:method, :uri] at request_matcher_registry.rb line 9 (vcr 6.4.0) - verbatim - ruby_llm chat.rb:115 duck-types to_json_schema - verbatim - pin_connection! at connection_pool.rb:366, active_connection? at :419 returning connection_lease.connection, with the two caveats the Rails docs state in the comment directly above it. The draft cited :414; corrected to :419. - commit f95c0d3ec exists with the exact message quoted (the squash of f303304da) - both schematist links resolve Cold-eyes gate applied nine checks and fixed five things in place. The one that mattered: the connection-pool section re-derived the outage that multi-agent-llm-rails-rubyllm owns as its centerpiece. Compressed to a single sentence plus an explicit hand-off link - this post owns the TEST, the sibling owns the outage. Also cut a slogany flip and a paragraph that restated its own H2, softened "half the context window" to "a shorter context window" (fake precision about an unnamed model), removed a manufactured same-morning coincidence between a provider action and a deploy bug, and closed an evidence gap where the prose claimed four key-scrubbing filters while the code block showed two. The mermaid was redrawn from a five-node straight line restating one sentence into a four-node convergence: two different request bodies collapsing onto one match key. Render-gated at 1280x800 and 390x844 - 18px labels measure 15.8px effective at mobile against the 9px floor, 420px tall so it is not a wall, and it scores 4/4 on the visual gate. Gates: hugo-build green, check-post-visuals at floor, zero console errors, 21/21 requests 200, zero em dashes, zero banned words, title 33 chars, description 154. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
R8 added: Its thesis: what a green Rails suite actually asserts when the thing under test lives on someone else's server. Three incidents — nine schemas swapped with 1549 tests still green, a provider retiring the model five features ran on, and a connection lease that no busy-count could see. Every source claim verified against installed gems, not paraphrased (the failure mode behind six wrong claims in this cluster today):
Cold-eyes gate fixed five things. The one that mattered: the connection-pool section re-derived the outage that Diagram redrawn from a five-node straight line restating one sentence into a four-node convergence (two request bodies → one match key). Render-gated at both viewports: 15.8px effective at mobile vs the 9px floor, 420px tall, 4/4 on the visual criteria. Gates: hugo-build green · visuals ratchet at floor · zero console errors · 21/21 requests 200 · zero em dashes · title 33 / description 154. One judgment call flagged rather than hidden: the model-retirement incident names no provider and links no deprecation notice, so it's the one claim a reader can't corroborate. It's framed as an observation chain (ran |
Codex follow-up on the correction itself. 'The two fail differently,
and the difference matters' announced what the next two sentences show;
cut. And 'leaves model_id as nil, so the method assigns an empty hash'
implied conditionality - the assignment is unconditional, which is why
model(provider: 'x') { } would keep the provider. Reworded, and the
two cases split into separate paragraphs so the if-X-then-Y structure
is visible rather than buried in a 684-char block.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A codex pass RAN the test instead of reading it and found the post's central payoff broken: `assert_nil ...active_connection?` fails before the regression it claims to pin is even introduced, so a reader pasting it would conclude their own code has the bug. Cause, unconditional and in Rails itself. test_fixtures.rb:201-204 calls pin_connection! and then lease_connection on every fixture pool, which populates the lease and marks it sticky. Since active_connection? is literally `connection_lease.connection`, it returns a live adapter from setup, before the test body runs a line. Sticky also means later with_connection calls skip their own release, so the lease never clears on its own. Measured: transactional + clean call FAILS; transactional + fiber FAILS; non-transactional + clean passes; non-transactional + re-added wrapper correctly detected. Fix, verified 4/4 green including rollback integrity: drop the fixture's lease immediately before exercising the subject. The transaction stays open, only the lease clears. Added that line to the snippet and a fourth caveat to the prose naming test_fixtures.rb:201-204 as the reason - the Rails docs list three caveats for this method and not this one, which is exactly why it cost an afternoon. Everything else in the section survives and is confirmed correct: pinning really does flatten busy-counts, :366 and :419 are right, and both documented caveats hold. Corrected the bare `connection_pool.rb` path to connection_adapters/abstract/connection_pool.rb. Also from the same pass: - August 16 was doing double duty. This post dated the model incident to it while multi-agent-llm-rails-rubyllm dates the pool outage to the same day, and R8's third incident IS that outage - so two of three incidents silently collided on one date across the cluster. This one is now "a separate morning". - Softened the retirement claim to what was actually observed. A deprecation notice in the answer body is not what a retired model returns (that is model_not_found), so the post now says retired or quietly substituted and notes the effect was identical. - allow_http_connections_when_no_cassette = false is VCR's default; the post framed it as a habit that holds up. Now says so and gives the real reason to write it down. Verified unchanged by the same pass, recorded so it is not re-checked: all three schematist payload claims (bare Draft 2020-12, $schema and title inside the strict schema, name collapsing to "response") are correct through ruby_llm's real normalize_schema_payload, and the VCR.configure block is valid - filter_sensitive_data aliases define_cassette_placeholder and its block does receive the interaction. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…o claim (#509) Found by reading the codex report for #499 after the fact. Its DO NOT MERGE was never about the connection test - codex confirmed that fix by execution, matching my own harness (passes clean, fails when with_connection is restored, rollback intact, test_fixtures.rb:201-204 cited accurately). The blocker was a P1 I never checked: line 19 still read "a retired model took out five features at once", a definite claim, while the body at line 74 says "whether the model was being retired or quietly substituted, the effect was the same". The softening landed in the prose and not in the metadata, so the published page kept making the claim the softening existed to remove. Now: "a hosted model stopped answering the way it used to", which is the body's own wording and asserts only what we observed. description/og_description say models "can be retired" - a general risk the post supports rather than a claim about this incident - so they are left alone. Lesson for the merge gate: I verified the executable claim and the date collision, then merged without diffing frontmatter against the body it summarises. Metadata is published copy and needs the same claims check as prose. Noted while committing, not fixed here: content/blog/debugging-rubyllm-agents-rails matches a broad "debug" .gitignore rule. The tracked files edit fine, but a NEW file added to that bundle would be silently skipped. The negation pattern that fixes it is still uncommitted in another worktree. bin/hugo-build: 8/8 validators.
Sync for 2026-08-20's merged work (#499, #501, #506, #509, #510). All three land in workflows/blog-pipeline.md because they are gates, not background - per the bundle's own rule that a lesson mattering six weeks later belongs in a concept rather than the log. 1. Technical claims must be executed, not read. Every wrong technical claim shipped today came from reading source and inferring behaviour. The Kamal guide asserted a stale traefik: key sits in deploy.yml doing nothing; kamal config answers 'unknown key: traefik'. The bad inference came from Validator::Configuration#allow_extensions? => true, which governs YAML extension keys and not arbitrary config keys. 2. Frontmatter is published copy. #509 shipped live with a body softened to 'whether retired or quietly substituted' while twitter_description still asserted 'a retired model took out five features at once'. 3. Citation lists use ## Sources. 4 posts already used it; 16 had drifted across three forms, and the bold variant survived the first survey because my grep matched only the other two. Concept timestamp + verified stamp added honestly (claude/opus-5, today). workflows/index.md entry widened to name the three gates so a cold session finds them without opening the file. Validated: uv run okf_validate.py .okf --strict -> conformant, 0 errors. Warning count unchanged at 64 (measured against a clean tree first, so the number is a baseline rather than a claim).
docs(okf): three publishing gates learned from claims we shipped wrong Sync for 2026-08-20's merged work (#499, #501, #506, #509, #510). All three land in workflows/blog-pipeline.md because they are gates, not background - per the bundle's own rule that a lesson mattering six weeks later belongs in a concept rather than the log. 1. Technical claims must be executed, not read. Every wrong technical claim shipped today came from reading source and inferring behaviour. The Kamal guide asserted a stale traefik: key sits in deploy.yml doing nothing; kamal config answers 'unknown key: traefik'. The bad inference came from Validator::Configuration#allow_extensions? => true, which governs YAML extension keys and not arbitrary config keys. 2. Frontmatter is published copy. #509 shipped live with a body softened to 'whether retired or quietly substituted' while twitter_description still asserted 'a retired model took out five features at once'. 3. Citation lists use ## Sources. 4 posts already used it; 16 had drifted across three forms, and the bold variant survived the first survey because my grep matched only the other two. Concept timestamp + verified stamp added honestly (claude/opus-5, today). workflows/index.md entry widened to name the three gates so a cold session finds them without opening the file. Validated: uv run okf_validate.py .okf --strict -> conformant, 0 errors. Warning count unchanged at 64 (measured against a clean tree first, so the number is a baseline rather than a claim).
New post:
evaluating-rubyllm-agents-rails— how to evaluate RubyLLM agents with bounded schemas, a per-iteration audit trail, and replayable evals on a local model.Written by the batch pipeline (3-critic panel + 4-eyes + cold-eyes), then corrected by an independent codex pass that ran a probe against the installed gem rather than reading it.
Caught pre-commit by the panel: a false thesis ("no Rails engine for LLM testing" —
ruby_llm-evalsexists), a config snippet that would raiseModelNotFoundErroras written, an artifact that didn't match the real schema, and two unsupported claims.Caught by the codex pass, fixed in the last commit:
temperatureormodelis harmlessly ignored. True for temperature; false for model —def model(model_id = nil, **options)assigns@chat_kwargs = options, so the block form wipes the inherited model and provider and the agent falls back toconfig.default_model. Measured:{provider: "ollama", model: "qwen3:0.6b"}became{}. The two cases are now separated.crmne/ruby_llm/blob/v1.16.0/..., which 404s — that repo's tags carry novprefix. Fixed and verified 200. (Per-repo:socketry/falcondoes usev0.57.0.)Verified clean by that pass: the ollama config, the
ModelNotFoundErrormechanism (confirmed by execution), the lazy-block list, inheritance behaviour, theruby_llm-evalscharacterization, the 523MB figure, all internal links, the mermaid pre-render hash, and voice/canon.Gates:
bin/hugo-buildgreen with all 8 validators ·check-post-visualsexit 0 · scroll gate at 1280x800 and 390x844 with zero console errors and no overflow · mermaid at 14.45px effective vs the 9px floor.Note: gates ran before the theme rebuild (#494) landed; worth a re-check of the rendered page after rebase if anything looks off.
🤖 Generated with Claude Code