Skip to content

Falcon benchmarks: say what the numbers can and cannot support - #640

Merged
pftg merged 1 commit into
masterfrom
blog/falcon-benchmark-methodology
Aug 28, 2026
Merged

pftg merged 1 commit into
masterfrom
blog/falcon-benchmark-methodology

Conversation

@pftg

@pftg pftg commented Aug 28, 2026

Copy link
Copy Markdown
Member

Content-only, one post — our #2 organic page (31 clicks / 2,898 impressions via the alias).

Verdict: not the same defect class, and here is why that matters

After removing fabricated figures from two posts today, the obvious move was to delete these tables too. That would have been wrong.

Two things argue the numbers are real:

  1. Internally coherent. Falcon shows 102ms against a forced 100ms I/O delay — exactly what a working async server does. Puma's degradation is proportional to its worker-slot count (4 workers × 5 threads = 20 slots) against 400 concurrent users.
  2. The hello-world table shows Agoo beating Falcon on every column — 7,000 req/s vs 6,000, less memory, less CPU. Nobody fabricates a competitor beating their own recommendation.

Deleting numbers on suspicion is as bad as inventing them.

What is actually wrong

Defect Evidence
No load generator named 0 hits for wrk/ab/k6/oha/autocannon anywhere in 5,375 words
No date recorded 0 hits
App under test not described —
Action Cable listed as a web server It is a Rails component that runs on one

The sibling post falcon-web-server-production-tuning-benchmarks names its tooling and has a Sources section. The standard exists in this repo; this post missed it.

The fix

Numbers reframed as indicative rather than reproducible, with the gap named explicitly rather than papered over, and readers pointed at the tuning post for figures that carry a method.

The Agoo row is now called out in prose as running against us — leaving it unremarked reads as an oversight rather than a disclosure.

Action Cable → Action Cable on Puma, which makes the row mean something.

Why not just delete the tables

They are the post's substance and there is no positive evidence they are false. The honest repair is to state what they can support. That is a different judgement from the Solid Queue table, which was arithmetically impossible, and the 450/120/73 triple, which appeared verbatim across two unrelated stacks.

Gates

bin/hugo-build clean · bin/rake test:links 0 errors · post in the crawl · 0 em dashes

🤖 Generated with Claude Code

https://claude.ai/code/session_016PUkwFTsiv7EB2DYKogbpg

Audited the tables on our #2 organic page after the fabricated-figure sweep.
Verdict: NOT the same defect class, and worth saying why rather than deleting
on suspicion.

Two things argue the numbers are real. They are internally coherent - Falcon at
102ms against a forced 100ms I/O delay is exactly what a working async server
does, and the Puma row degrades in proportion to its worker-slot count under
400 concurrent users. And the hello-world table shows Agoo beating Falcon on
every column: more requests per second, less memory, less CPU. Nobody fabricates
a competitor beating their own recommendation.

What is actually wrong is different. No load generator is named anywhere in the
post, no date is recorded, and the app under test is not described - so nobody
can re-run these and check us. Meanwhile the sibling tuning post names its
tooling and carries a Sources section, which shows the standard exists and this
post missed it.

Fixed by saying so: the numbers are now framed as indicative rather than
reproducible, with the gap named explicitly and readers pointed at the tuning
post for figures that have a method attached. The Agoo row is called out in
prose as running against us, because leaving it unremarked reads as an oversight
rather than a disclosure.

Also corrected a category error: Action Cable was listed as a web server
alongside Puma and Falcon. It is a Rails component that runs ON a server, so the
row now reads "Action Cable on Puma", which makes the comparison mean something.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016PUkwFTsiv7EB2DYKogbpg
@pftg
pftg merged commit 10f0d13 into master Aug 28, 2026
@pftg
pftg deleted the blog/falcon-benchmark-methodology branch August 28, 2026 13:44
@coderabbitai

coderabbitai Bot commented Aug 28, 2026

Copy link
Copy Markdown
Contributor

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 0906e062-3f84-4520-ad70-626863dc3bdd


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant