Skip to content

Content: RubyLLM agent evals + debugging posts (R7, R8) - #499

Merged
pftg merged 5 commits into
masterfrom
rubyllm-evals-debugging-batch
Aug 20, 2026
Merged

pftg merged 5 commits into
masterfrom
rubyllm-evals-debugging-batch

Conversation

@pftg

@pftg pftg commented Aug 20, 2026

Copy link
Copy Markdown
Member

New post: evaluating-rubyllm-agents-rails — how to evaluate RubyLLM agents with bounded schemas, a per-iteration audit trail, and replayable evals on a local model.

Written by the batch pipeline (3-critic panel + 4-eyes + cold-eyes), then corrected by an independent codex pass that ran a probe against the installed gem rather than reading it.

Caught pre-commit by the panel: a false thesis ("no Rails engine for LLM testing" — ruby_llm-evals exists), a config snippet that would raise ModelNotFoundError as written, an artifact that didn't match the real schema, and two unsupported claims.

Caught by the codex pass, fixed in the last commit:

  • The post said a block handed to temperature or model is harmlessly ignored. True for temperature; false for model — def model(model_id = nil, **options) assigns @chat_kwargs = options, so the block form wipes the inherited model and provider and the agent falls back to config.default_model. Measured: {provider: "ollama", model: "qwen3:0.6b"} became {}. The two cases are now separated.
  • The citation linked crmne/ruby_llm/blob/v1.16.0/..., which 404s — that repo's tags carry no v prefix. Fixed and verified 200. (Per-repo: socketry/falcon does use v0.57.0.)
  • Dropped "and this column landed the next day" — the same unverifiable interval as a claim already cut, in softer wording.

Verified clean by that pass: the ollama config, the ModelNotFoundError mechanism (confirmed by execution), the lazy-block list, inheritance behaviour, the ruby_llm-evals characterization, the 523MB figure, all internal links, the mermaid pre-render hash, and voice/canon.

Gates: bin/hugo-build green with all 8 validators · check-post-visuals exit 0 · scroll gate at 1280x800 and 390x844 with zero console errors and no overflow · mermaid at 14.45px effective vs the 9px floor.

Note: gates ran before the theme rebuild (#494) landed; worth a re-check of the rendered page after rebase if anything looks off.

🤖 Generated with Claude Code

pftg and others added 2 commits August 20, 2026 19:21
Agent-eval-specific companion to the monitoring post: bounded schemas as
the first check, a per-iteration jsonb audit trail on the run's metrics
row, funnel_summary, golden sets grown from flagged runs, and eval
replays against a local qwen3:0.6b.

SEO: title 30 chars, description 150 chars (og identical), primary
keyword in first 100 words, 5 internal links, 7 external citations.

First-hand material sanitized from a production talent-matching pipeline
(shapes and lessons only - no prompts, model IDs, or score bands).

Reviewer verdicts:
- Tech fact-check (core-reviewer): found 3 blocking errors in the draft -
  a false "no Rails engine for LLM testing" claim (ruby_llm-evals exists),
  an inverted lazy-block claim, and a config snippet that could not route
  to the local model. All three rewritten against the gem source at
  ruby_llm-1.16.0 and the ruby_llm-evals README.
- Cold-eyes 9-check gate (core-reviewer, fresh context): "PUBLISH-READY"
  after fixing an unsourced universal claim and two rhetorical flips.
- 4-eyes pre-commit (core-reviewer): "VERDICT: APPROVED FOR COMMIT" -
  all three blockers verified closed, including an empirical check that
  the published initializer resolves (Scorer < BaseAgent =>
  model=qwen3:0.6b provider=ollama temp=0.0).

Two claims were cut rather than published because the sources did not
support them: "answers in milliseconds" (no benchmark) and "landed in the
very next commit" (26 commits sit between the two).

Gates: bin/hugo-build green, bin/check-post-visuals exit 0, mermaid
pre-rendered to SVG at 14.45px effective font on a 390px viewport
(floor 9px), scroll gate clean at 1280x800 and 390x844 with zero
console errors and zero 404s.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Independent codex pass with a runnable probe found two blockers in the
evals post.

1. The post told readers a block handed to `temperature` or `model` is
   "dropped without a warning" and "does nothing". True for temperature;
   false and dangerous for model. `def model(model_id = nil, **options)`
   assigns `@chat_kwargs = options`, so the block form leaves model_id
   nil and wipes the inherited model and provider - the agent then runs
   on config.default_model, a hosted model where the reader believed a
   local one was pinned. Measured against the post's own BaseAgent:
   inherited {provider: "ollama", model: "qwen3:0.6b"} became {}.
   The sentence now separates the harmless case from the destructive one.

2. The citation linked crmne/ruby_llm/blob/v1.16.0/... which 404s - that
   repo's tags carry no `v` prefix. Corrected to /blob/1.16.0/, verified
   200. Note this is per-repo: socketry/falcon does use v0.57.0.

Also dropped "and this column landed the next day" - the same
unverifiable interval as the "very next commit" claim cut pre-commit,
in softer wording. "The metrics row came first" carries the point.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Aug 20, 2026 •

Copy link
Copy Markdown
Contributor

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: f47feac5-9416-4aac-a027-c2f5e0f4ac69


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Third and last post of the RubyLLM x Rails queue. Its thesis is what a
green Rails suite actually asserts when the thing under test lives on
someone else's server, told through three incidents from the
talent-matching pipeline we run.

Source claims verified against installed gems rather than paraphrased,
which is the failure mode that produced six wrong claims in this
cluster today:
- VCR DEFAULT_MATCHERS = [:method, :uri] at request_matcher_registry.rb
  line 9 (vcr 6.4.0) - verbatim
- ruby_llm chat.rb:115 duck-types to_json_schema - verbatim
- pin_connection! at connection_pool.rb:366, active_connection? at :419
  returning connection_lease.connection, with the two caveats the Rails
  docs state in the comment directly above it. The draft cited :414;
  corrected to :419.
- commit f95c0d3ec exists with the exact message quoted (the squash of
  f303304da)
- both schematist links resolve

Cold-eyes gate applied nine checks and fixed five things in place. The
one that mattered: the connection-pool section re-derived the outage
that multi-agent-llm-rails-rubyllm owns as its centerpiece. Compressed
to a single sentence plus an explicit hand-off link - this post owns the
TEST, the sibling owns the outage. Also cut a slogany flip and a
paragraph that restated its own H2, softened "half the context window"
to "a shorter context window" (fake precision about an unnamed model),
removed a manufactured same-morning coincidence between a provider
action and a deploy bug, and closed an evidence gap where the prose
claimed four key-scrubbing filters while the code block showed two.

The mermaid was redrawn from a five-node straight line restating one
sentence into a four-node convergence: two different request bodies
collapsing onto one match key. Render-gated at 1280x800 and 390x844 -
18px labels measure 15.8px effective at mobile against the 9px floor,
420px tall so it is not a wall, and it scores 4/4 on the visual gate.

Gates: hugo-build green, check-post-visuals at floor, zero console
errors, 21/21 requests 200, zero em dashes, zero banned words,
title 33 chars, description 154.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@pftg pftg changed the title Content: RubyLLM agent evals post (R7) Content: RubyLLM agent evals + debugging posts (R7, R8) Aug 20, 2026
@pftg

pftg commented Aug 20, 2026

Copy link
Copy Markdown
Member Author

R8 added: debugging-rubyllm-agents-rails — completes the RubyLLM × Rails queue (R1/R4/R5 shipped earlier today, R7 above, R8 now).

Its thesis: what a green Rails suite actually asserts when the thing under test lives on someone else's server. Three incidents — nine schemas swapped with 1549 tests still green, a provider retiring the model five features ran on, and a connection lease that no busy-count could see.

Every source claim verified against installed gems, not paraphrased (the failure mode behind six wrong claims in this cluster today):

  • DEFAULT_MATCHERS = [:method, :uri] at request_matcher_registry.rb:9 (vcr 6.4.0) — verbatim
  • chat.rb:115 duck-types to_json_schema — verbatim
  • pin_connection! at :366, active_connection? at :419 returning connection_lease.connection — the draft cited :414, corrected
  • f95c0d3ec exists with the quoted message; both schematist links 200

Cold-eyes gate fixed five things. The one that mattered: the connection-pool section re-derived the outage that multi-agent-llm-rails-rubyllm owns as its centerpiece. Now one sentence plus a hand-off link — this post owns the test, the sibling owns the outage. Also removed a manufactured same-morning coincidence between a provider action and a deploy bug, and closed an evidence gap where the prose claimed four key-scrubbing filters while the code block showed two.

Diagram redrawn from a five-node straight line restating one sentence into a four-node convergence (two request bodies → one match key). Render-gated at both viewports: 15.8px effective at mobile vs the 9px floor, 420px tall, 4/4 on the visual criteria.

Gates: hugo-build green · visuals ratchet at floor · zero console errors · 21/21 requests 200 · zero em dashes · title 33 / description 154.

One judgment call flagged rather than hidden: the model-retirement incident names no provider and links no deprecation notice, so it's the one claim a reader can't corroborate. It's framed as an observation chain (ran .ask("say ok"), got a deprecation notice, refreshed the registry) rather than a bare assertion. Naming the model would make it airtight if you'd rather.

pftg and others added 2 commits August 20, 2026 21:00
Codex follow-up on the correction itself. 'The two fail differently,
and the difference matters' announced what the next two sentences show;
cut. And 'leaves model_id as nil, so the method assigns an empty hash'
implied conditionality - the assignment is unconditional, which is why
model(provider: 'x') { } would keep the provider. Reworded, and the
two cases split into separate paragraphs so the if-X-then-Y structure
is visible rather than buried in a 684-char block.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A codex pass RAN the test instead of reading it and found the post's
central payoff broken: `assert_nil ...active_connection?` fails before
the regression it claims to pin is even introduced, so a reader pasting
it would conclude their own code has the bug.

Cause, unconditional and in Rails itself. test_fixtures.rb:201-204 calls
pin_connection! and then lease_connection on every fixture pool, which
populates the lease and marks it sticky. Since active_connection? is
literally `connection_lease.connection`, it returns a live adapter from
setup, before the test body runs a line. Sticky also means later
with_connection calls skip their own release, so the lease never clears
on its own.

Measured: transactional + clean call FAILS; transactional + fiber FAILS;
non-transactional + clean passes; non-transactional + re-added wrapper
correctly detected.

Fix, verified 4/4 green including rollback integrity: drop the fixture's
lease immediately before exercising the subject. The transaction stays
open, only the lease clears. Added that line to the snippet and a fourth
caveat to the prose naming test_fixtures.rb:201-204 as the reason - the
Rails docs list three caveats for this method and not this one, which is
exactly why it cost an afternoon.

Everything else in the section survives and is confirmed correct:
pinning really does flatten busy-counts, :366 and :419 are right, and
both documented caveats hold. Corrected the bare `connection_pool.rb`
path to connection_adapters/abstract/connection_pool.rb.

Also from the same pass:
- August 16 was doing double duty. This post dated the model incident to
  it while multi-agent-llm-rails-rubyllm dates the pool outage to the
  same day, and R8's third incident IS that outage - so two of three
  incidents silently collided on one date across the cluster. This one
  is now "a separate morning".
- Softened the retirement claim to what was actually observed. A
  deprecation notice in the answer body is not what a retired model
  returns (that is model_not_found), so the post now says retired or
  quietly substituted and notes the effect was identical.
- allow_http_connections_when_no_cassette = false is VCR's default; the
  post framed it as a habit that holds up. Now says so and gives the
  real reason to write it down.

Verified unchanged by the same pass, recorded so it is not re-checked:
all three schematist payload claims (bare Draft 2020-12, $schema and
title inside the strict schema, name collapsing to "response") are
correct through ruby_llm's real normalize_schema_payload, and the
VCR.configure block is valid - filter_sensitive_data aliases
define_cassette_placeholder and its block does receive the interaction.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@pftg
pftg merged commit 1dfbd72 into master Aug 20, 2026
8 of 9 checks passed
@pftg
pftg deleted the rubyllm-evals-debugging-batch branch August 20, 2026 20:04
pftg added a commit that referenced this pull request Aug 20, 2026
…o claim (#509)

Found by reading the codex report for #499 after the fact. Its DO NOT
MERGE was never about the connection test - codex confirmed that fix by
execution, matching my own harness (passes clean, fails when
with_connection is restored, rollback intact, test_fixtures.rb:201-204
cited accurately).

The blocker was a P1 I never checked: line 19 still read "a retired model
took out five features at once", a definite claim, while the body at line
74 says "whether the model was being retired or quietly substituted, the
effect was the same". The softening landed in the prose and not in the
metadata, so the published page kept making the claim the softening
existed to remove.

Now: "a hosted model stopped answering the way it used to", which is the
body's own wording and asserts only what we observed.

description/og_description say models "can be retired" - a general risk
the post supports rather than a claim about this incident - so they are
left alone.

Lesson for the merge gate: I verified the executable claim and the date
collision, then merged without diffing frontmatter against the body it
summarises. Metadata is published copy and needs the same claims check as
prose.

Noted while committing, not fixed here: content/blog/debugging-rubyllm-agents-rails
matches a broad "debug" .gitignore rule. The tracked files edit fine, but
a NEW file added to that bundle would be silently skipped. The negation
pattern that fixes it is still uncommitted in another worktree.

bin/hugo-build: 8/8 validators.
pftg added a commit that referenced this pull request Aug 20, 2026
Sync for 2026-08-20's merged work (#499, #501, #506, #509, #510). All
three land in workflows/blog-pipeline.md because they are gates, not
background - per the bundle's own rule that a lesson mattering six weeks
later belongs in a concept rather than the log.

1. Technical claims must be executed, not read. Every wrong technical
   claim shipped today came from reading source and inferring behaviour.
   The Kamal guide asserted a stale traefik: key sits in deploy.yml doing
   nothing; kamal config answers 'unknown key: traefik'. The bad
   inference came from Validator::Configuration#allow_extensions? => true,
   which governs YAML extension keys and not arbitrary config keys.

2. Frontmatter is published copy. #509 shipped live with a body softened
   to 'whether retired or quietly substituted' while twitter_description
   still asserted 'a retired model took out five features at once'.

3. Citation lists use ## Sources. 4 posts already used it; 16 had drifted
   across three forms, and the bold variant survived the first survey
   because my grep matched only the other two.

Concept timestamp + verified stamp added honestly (claude/opus-5, today).
workflows/index.md entry widened to name the three gates so a cold
session finds them without opening the file.

Validated: uv run okf_validate.py .okf --strict -> conformant, 0 errors.
Warning count unchanged at 64 (measured against a clean tree first, so
the number is a baseline rather than a claim).
pftg added a commit that referenced this pull request Aug 20, 2026
docs(okf): three publishing gates learned from claims we shipped wrong

Sync for 2026-08-20's merged work (#499, #501, #506, #509, #510). All
three land in workflows/blog-pipeline.md because they are gates, not
background - per the bundle's own rule that a lesson mattering six weeks
later belongs in a concept rather than the log.

1. Technical claims must be executed, not read. Every wrong technical
   claim shipped today came from reading source and inferring behaviour.
   The Kamal guide asserted a stale traefik: key sits in deploy.yml doing
   nothing; kamal config answers 'unknown key: traefik'. The bad
   inference came from Validator::Configuration#allow_extensions? => true,
   which governs YAML extension keys and not arbitrary config keys.

2. Frontmatter is published copy. #509 shipped live with a body softened
   to 'whether retired or quietly substituted' while twitter_description
   still asserted 'a retired model took out five features at once'.

3. Citation lists use ## Sources. 4 posts already used it; 16 had drifted
   across three forms, and the bold variant survived the first survey
   because my grep matched only the other two.

Concept timestamp + verified stamp added honestly (claude/opus-5, today).
workflows/index.md entry widened to name the three gates so a cold
session finds them without opening the file.

Validated: uv run okf_validate.py .okf --strict -> conformant, 0 errors.
Warning count unchanged at 64 (measured against a clean tree first, so
the number is a baseline rather than a claim).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant