Routes every query to the cheapest model that can handle it: three local tiers
on an RTX 5090 (free), three Claude cloud tiers above them (metered), with an
optional answer-then-verify judge that catches bad local answers and escalates.
Born from a research question — "could a distributed volunteer network replace
centralized LLM serving?" — whose honest answer (see simulations/) was that
the winning cost architecture is exactly this: local-first, escalate rarely,
frontier as a scalpel.
| File | What it is |
|---|---|
cascade.py |
The router. Six tiers, difficulty classifier, judge mode, cost tracking. |
cascade_code.py |
Coding agent powered by the tiers — it builds and improves this project (including its own source). |
bench.py |
Benchmark: plain vs judge mode, local-only, graded answers. |
PLAN.md |
Build plan with phase status and open findings. |
| Tier | Model | Serves | Cost |
|---|---|---|---|
| T1 | qwen3:8b | trivial (difficulty 1) | free |
| T2 / T2c | qwen3:32b / qwen3-coder:30b | standard / code (2) | free |
| T3 | qwen3.6 + thinking | hard (3) | free |
| T4 | claude-haiku-4-5 | easy tiers when Ollama is down | $1/$5 per Mtok |
| T5 | claude-sonnet-5 | expert (4) | $2/$10 (intro) |
| T6 | claude-opus-5 | frontier (5), server-side refusal fallbacks | $5/$25 |
python cascade.py # REPL (!up !stats !tier N !quit)
python cascade.py "question" # one-shot auto-routed
python cascade.py --judge "question" # verify local answers, escalate on fail
python cascade.py --tier 6 "question" # force a tier
python cascade_code.py --dir . --steps 30 "task description" # coding agent
python bench.py # plain-vs-judge benchmark (local only)
Cloud tiers activate automatically when credentials exist (ANTHROPIC_API_KEY
or an ant auth login profile). Without them, routing caps at T3 with a note.
- VRAM: Qwen modelfiles default to 32K context, which pushes the 32B to
55GB and spills onto CPU (6 tok/s).
cascade.pysets explicit per-tiernum_ctxso everything stays 100% GPU (60-190 tok/s). Keep it that way. - claude.ai subscription ≠ API credits. The cloud tiers bill prepaid API credits at platform.claude.com, a separate wallet from any Claude plan.
- Router decisions log to
router_log.jsonl(gitignored data; future benchmark/tuning input).
The research phase that produced this design (2026-07-30, one session): Monte Carlo and analytic models answering, in order —
sim.py— volunteer-swarm serving architectures. Verdict: never split model layers across the WAN; draft locally, verify remotely.train_sim.py— distributed training feasibility (DiLoCo islands, memory limits, RL rollouts as the swarm's real niche).cost_sim.py— distributed company vs centralized cost. Verdict: ~2-3x serving at best, training worse; sunk consumer hardware is the only edge.frontier_sim.py— when a fixed GPT-5-level bar becomes affordable (~$28M rented burst 2027, ~$10M 2028 at 3x/yr efficiency).security_sim.py— audit economics: serving fraud deterred at 0.1% overhead; training poisoning managed at 5-10%.strategy_sim.py— final ranking of cost levers. Cascade routing ~8-10x beats distribution ~1.1-3x; hence this repo.
Each is standalone: python simulations/<name>.py prints its tables.
One phase at a time (see PLAN.md). cascade-code implements; a reviewer reads every diff before commit. Logs are data — don't delete them casually.