Skip to content

Latest commit

 

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Cascade — a local-first, auto-tiering LLM system

Routes every query to the cheapest model that can handle it: three local tiers on an RTX 5090 (free), three Claude cloud tiers above them (metered), with an optional answer-then-verify judge that catches bad local answers and escalates. Born from a research question — "could a distributed volunteer network replace centralized LLM serving?" — whose honest answer (see simulations/) was that the winning cost architecture is exactly this: local-first, escalate rarely, frontier as a scalpel.

The system

File What it is
cascade.py The router. Six tiers, difficulty classifier, judge mode, cost tracking.
cascade_code.py Coding agent powered by the tiers — it builds and improves this project (including its own source).
bench.py Benchmark: plain vs judge mode, local-only, graded answers.
PLAN.md Build plan with phase status and open findings.

Tiers

Tier Model Serves Cost
T1 qwen3:8b trivial (difficulty 1) free
T2 / T2c qwen3:32b / qwen3-coder:30b standard / code (2) free
T3 qwen3.6 + thinking hard (3) free
T4 claude-haiku-4-5 easy tiers when Ollama is down $1/$5 per Mtok
T5 claude-sonnet-5 expert (4) $2/$10 (intro)
T6 claude-opus-5 frontier (5), server-side refusal fallbacks $5/$25

Usage

python cascade.py                          # REPL (!up !stats !tier N !quit)
python cascade.py "question"               # one-shot auto-routed
python cascade.py --judge "question"       # verify local answers, escalate on fail
python cascade.py --tier 6 "question"      # force a tier

python cascade_code.py --dir . --steps 30 "task description"   # coding agent
python bench.py                            # plain-vs-judge benchmark (local only)

Cloud tiers activate automatically when credentials exist (ANTHROPIC_API_KEY or an ant auth login profile). Without them, routing caps at T3 with a note.

Notes that save you a debugging session

  • VRAM: Qwen modelfiles default to 32K context, which pushes the 32B to 55GB and spills onto CPU (6 tok/s). cascade.py sets explicit per-tier num_ctx so everything stays 100% GPU (60-190 tok/s). Keep it that way.
  • claude.ai subscription ≠ API credits. The cloud tiers bill prepaid API credits at platform.claude.com, a separate wallet from any Claude plan.
  • Router decisions log to router_log.jsonl (gitignored data; future benchmark/tuning input).

simulations/

The research phase that produced this design (2026-07-30, one session): Monte Carlo and analytic models answering, in order —

  1. sim.py — volunteer-swarm serving architectures. Verdict: never split model layers across the WAN; draft locally, verify remotely.
  2. train_sim.py — distributed training feasibility (DiLoCo islands, memory limits, RL rollouts as the swarm's real niche).
  3. cost_sim.py — distributed company vs centralized cost. Verdict: ~2-3x serving at best, training worse; sunk consumer hardware is the only edge.
  4. frontier_sim.py — when a fixed GPT-5-level bar becomes affordable (~$28M rented burst 2027, ~$10M 2028 at 3x/yr efficiency).
  5. security_sim.py — audit economics: serving fraud deterred at 0.1% overhead; training poisoning managed at 5-10%.
  6. strategy_sim.py — final ranking of cost levers. Cascade routing ~8-10x beats distribution ~1.1-3x; hence this repo.

Each is standalone: python simulations/<name>.py prints its tables.

Working agreements

One phase at a time (see PLAN.md). cascade-code implements; a reviewer reads every diff before commit. Logs are data — don't delete them casually.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages