Skip to content

Count what each provider actually serves, and name the model that answered each turn - #122

Closed
ParallelEntrepreneur wants to merge 9 commits into
mainfrom
colonizer/issue-39-5082fc9d
Closed

ParallelEntrepreneur wants to merge 9 commits into
mainfrom
colonizer/issue-39-5082fc9d

Conversation

@ParallelEntrepreneur

Copy link
Copy Markdown
Collaborator

Closes #39.

A user wired a local DeepSeek server into Colonizer, saw Settings report it Reachable · 1 ms · 1 model, and concluded the local model was never used. Nothing was broken. The provider was pointed at subagent_model, and colonies almost never spawn subagents, so it legitimately saw almost no traffic. The gateway was the only component that knew anything, and it kept only live gauges — in-flight and queued, both zero whenever nobody is mid-request, which is almost always. There was no cumulative record anywhere, so "is this thing ever used?" had no answer in the UI.

This adds the record, and makes the UI say what it knows and no more.

What changed

Cumulative counters per provider (ask 1). gateway.rs gains a ProviderUsage — requests, failures, fallbacks, duration_ms, last_request_at — backed by atomics in UsageCounters, with a Timed RAII guard for wall-clock. They persist to <data_dir>/provider-usage.json, are seeded back on startup, flushed by a 5 s background task when dirty and once on shutdown, and forgotten when a provider is deleted. A missing or corrupt file degrades to defaults rather than failing to boot. Settings renders them per provider card; GET /api/providers returns them.

"Configured but unused" (ask 2). providers::used_by() reports which of model / subagent_model / background_model resolve to a given provider, across the global agent env and every org override. An empty list means configured but not referenced by any model setting. When a provider is wired only off the main path — the exact shape of this issue — the card says so instead of looking broken.

Which model served a turn (ask 3). sessionStream.ts diffs each turn_end event's cumulative model_usage map against the previous one to name the models that served that turn, shown in the chat turn footer (lead model, +N when several, full list in the tooltip) on both the complete and failed lines.

Where to point a local model (ask 4). Answered in prose, not by changing a default. README, the Claude Code module README and docs/protocol.md now say plainly that the orchestrator model carries nearly all of a colony's traffic and that subagent routing rarely sees any — "To put real traffic on your own hardware, point the orchestrator model at it." The supporting observation (four recent colonies of 413–2,095 events spawned 0, 0, 0 and 1 subagents) is presented as a sample, not a law. No default was changed; if reviewers would rather move the default, that is a separate and more invasive call than this issue needs.

Honesty decisions worth reviewing

This is a PR about a UI that misled someone, so where a counter can't support a claim, the wording backs off:

  • requests includes queue-rejected attempts (a real request against the provider, and failures must stay a subset), but duration_ms excludes queue time — it is timed from slot acquisition. Otherwise a saturated max_concurrent: 1 provider that served nothing would display "10m 0s total" of pure queue wait. Both facts are stated in the struct doc and in protocol.md.
  • fallbacks is a prediction, not an observation. The gateway answers with a fallback and the colony's router retries on Claude; that retry never returns through the gateway. The card says "N of them got a Claude fallback" and the tooltip says outright that it is a prediction.
  • A failure part-way through a streamed body is not counted. Counting it on the anthropic wire would mean threading counters through the streaming type, which this change does not do. It is counted on the openai wire, where the error return already had the counter in scope. All three descriptions — struct doc, protocol.md, UI tooltip — now say exactly this.
  • An absent usage is not "Never used". usage and used_by are optional in types.ts because older Motherships don't send them. A missing object renders "No usage recorded — this Mothership doesn't report usage"; "Never used" is reserved for counters that genuinely read zero. The same applies to used_by: the wiring note is skipped rather than guessed.
  • Counters begin at zero on upgrade, so the card is explicit that it counts "since it first kept tally".

Verification

No CI exists in this repo, so everything below was run locally.

  • cargo test --workspace: 107 passed, 0 failed, run twice with identical results (the new tests are async and multi-threaded). Includes new tests for every outcome a request can have, restart persistence with a corrupt-file fallback, provider deletion forgetting its tally, an openai body that fails after the headers, a queue timeout counting a request and a failure but no duration, and concurrent writers always leaving a parseable file.
  • cargo clippy --workspace --all-targets: 18 warnings, exactly the pre-existing baseline, zero new — cross-checked against the diff's added-line ranges.
  • web: npm run build (tsc --noEmit && vite build) green.
  • modules/agents/claude-code: 58/58 pass, including the model_usage emission test the web reducer depends on. services/telemetry: 8/8 pass.
  • The reducer was validated by executing the real module against eight scenarios (first turn, absent model_usage, full replay, post-resume totals shrink, stale seq, unchanged counts, NaN/Infinity/string garbage, mixed shrink-and-grow) rather than by reading it. It never emits a negative delta or invents a model.
  • The dirty-flag fix was mutation-tested: its assertion fails against the previous any-short-circuit form.
  • cargo fmt was deliberately not run. This repo has no rustfmt.toml and rustfmt disagrees with the house style in every file of both crates; new code matches its surroundings by hand.

Look closely at

  • write_usage (gateway.rs). Three callers write this file — the flush loop, the shutdown flush, and forget_usage. They previously shared one .tmp path, so interleaved writes could publish invalid JSON, which on next boot would silently reset every provider's tally and report "Never used" for a busy provider — the very lie this issue is about. Each call now uses a unique tmp name and cleans up after a failed write. A residual window remains and is documented in the comment: renames can still land out of order, so a tally can be briefly stale, but each published snapshot is complete and parseable.
  • Dirty flags are cleared before the write, so a failed write drops that batch until the next change. The in-memory tally stays correct; the flush protocol was deliberately not restructured.
  • Timed's lifetime. Moving the timer after the permit must not stop it covering the streamed body. It is still moved into the same Guards tuple and dropped when the stream ends — worth a second pair of eyes.
  • The queue-timeout test pins the contract, not the handler. App is only constructible in main.rs and there is no test precedent for building one, so the test replays the counting sequence against a real exhausted semaphore instead of driving proxy. A revert of the two-line Timed move would be caught by review, not by that test.
  • Known cosmetic gap: if a provider is deleted in the window between provider lookup and counter acquisition, a zero-valued ghost entry can persist in the JSON, and a request in flight during deletion loses its last ticks. Left alone as accounting noise with no panic risk.
  • Pre-existing, not introduced here: a client that survives a colony resume without reloading keeps a stale lastSeq while the event log restarts at seq 1; the whole chat thread misrenders in that state, and the new models line inherits it.
  • Invariant to preserve: modelsOfTurn attributing a first-seen turn_end's whole total is correct only because replay is always complete. If a replay tail-limit is ever added, the first replayed turn would silently absorb the colony's all-time usage.

🤖 Generated by Colonizer in a microVM

…settlers into a crew (#71)

Implements the Settler Showcase design (claude.ai/design): the ant, the card and
parallel settlers.

The ant (AntAvatar.tsx, index.css) is now a pixel-art worker in side profile,
26 by 20 pixels with crisp edges and a joint for every moving part. The ground,
thought pixels, scribble and lens move a pixel at a time.
- working: a tripod gait, carrying what the tool is (a lens to search or read, a
  leaf to edit, a block to build or install, a flag to test).
- thinking: looks up while a thought fills in, dot by dot.
- writing: nods over a line it writes and rewrites.
- done: hops once and rests with a check.
- paused: sits with its antennae drooped.
- A failed tool result makes it stumble over a pebble, then carry on.
- Each role wears its own mark: a Scout's pack, a Builder's hard hat, an
  Inspector's monocle, and so on.
- Reduced motion freezes every state in a readable pose.

The card (SettlerCard.tsx) keeps the settler's name, task and one status line.
It adds a step count and a trail along its foot that moves while the settler
works, and the frame turns green when the settler is done. Collapsed, a done
card shows the first sentence of its report. Opened, it shows the report as
markdown, then the steps: in plain words with the command or path each acted
on, or the raw transcript in the technical view.

Grouping: settlers the orchestrator sends out together are one crew. Their
events interleave while they work, and each settler now keeps one card for the
whole burst instead of splitting into "continued" cards every time another
settler speaks. A crew's cards hang off one rail under a strip of its ants on a
shared trail ("3 settlers · 2 at work · 1 done"), and their loops are staggered
so they never move in lockstep.

Departures from the design, where it showed data colonies don't send:
- A done card doesn't show a duration, because the stream doesn't record when
  a settler finished.
- A failed step says "didn't work", not "didn't work · retried", because
  nothing says it was retried.

The mock's colony now sends a crew of three (Scout, Tester, Inspector) with a
failing test run, so ?mock=1 shows every part of this.
…e top of the README (#72)

Installing was explained only in the README, about 200 lines down, and the site's
homepage showed `scripts/install.sh` with no word that it has to run inside a clone.
docs/install.md is now the one place for it, mirrored to colonizer.dev/docs/install
alongside vision, architecture and protocol. It covers:

- what a machine needs, including Claude Code on Linux;
- the build-and-run commands, and the two installer options;
- the first run, and what the installer does on a Mac;
- where settings, credentials, colonies and the app live;
- how to update.

The README's "Run it" moves up to right after the introduction: the four commands
and a pointer to the guide. The introduction no longer says Linux only, since an
Apple Silicon Mac has run a colony end to end (#36).

colonizer.dev/docs/install only exists once the website re-syncs its docs after this
merges.
A version tag now publishes a GitHub release with a prebuilt app for Linux x86_64
and one for Apple Silicon, plus an installer:

    curl -fsSL https://colonizer.dev/install.sh | sh

Nothing on the machine needs Rust or Node.js to install it.

.github/workflows/release.yml
- linux-binaries: builds colonizer-agentd and rtk for both colony architectures,
  and the Linux harness, as static musl binaries. They are built inside
  rust:1-alpine with `docker run`, because GitHub's JavaScript actions don't run
  in Alpine containers. A static harness starts on any glibc; one built on the
  Ubuntu 24.04 runner would not start on Debian 12 or Ubuntu 22.04.
- bundle: runs scripts/install.sh --bundle on ubuntu-24.04 and macos-15, then a
  smoke test. The app starts from the unpacked archive, finds its assets there,
  reports agentd, the web UI and the claude-code module, and serves the UI.
- release: runs for tags only. The tag must match crates/colonizer's version. It
  writes SHA256SUMS and publishes the archives with install.sh.

A release contains no Anthropic code. The Claude Agent SDK is "all rights reserved",
and its optional platform packages are Claude Code itself (224 MB).
- The bundle installs the module with --omit=optional. The runner already passes
  pathToClaudeCodeExecutable, so those packages were never used.
- scripts/record-fetch-at-install.mjs takes the SDK out and writes its lockfile
  tarball URL and sha256, after checking the tarball against package-lock.json's
  integrity.
- scripts/install-release.sh fetches the SDK from the npm registry against that
  record. On a Mac it also fetches the linux-arm64 Claude Code build from
  Anthropic's stable channel, checked against the manifest. macOS's own plutil
  reads the manifest, so no Node.js is needed. The Workflow's Package step fails
  if the SDK is found in a bundle.

scripts/install-release.sh (published as install.sh)
- Refuses platforms it can't serve, including Linux without read-write /dev/kvm.
- Checks the archive against SHA256SUMS.
- Swaps ~/.local/share/colonizer/app in whole, and links ~/.local/bin/colonizer.
- Running it again updates in place. It reuses a Claude Code binary that still
  matches the manifest, and never touches settings or colonies.
- Accepts COLONIZER_VERSION, COLONIZER_APP and --pull-image.
- Everything runs inside main(), so a truncated download runs nothing.

scripts/install.sh --bundle skips the KVM check and the Claude Code fetch, and
uses binaries from $COLONIZER_PREBUILT instead of building them.
build-agentd.sh and build-rtk.sh accept COLONIZER_BUILD_HERE=1, to build inside
an existing rust:1-alpine instead of a microVM. rtk's source is then unpacked
outside the repository, because cargo refused to build it as a stray member of
the harness workspace (the local test below caught that).

Tested on an Apple Silicon Mac:
- In a rust:1-alpine microVM with COLONIZER_BUILD_HERE=1: agentd, rtk and the
  harness built as static aarch64 musl binaries. The static harness started in a
  Debian (node:24-bookworm) microVM and served /api/status.
- `install.sh --bundle` with those guest binaries: a 129 MB app, a 44 MB archive,
  and no Agent SDK inside.
- `curl … | sh` against the archive served locally, into an empty HOME, took 41 s.
  It fetched the SDK (sha256 checked) and Claude Code 2.1.267 (manifest checked).
- The installed app reported its assets in the new location: agentd, web,
  claude-code, the guest Claude binary, and msb 0.6.18 from the relocated vendor
  directory. It served the UI.
- Running the installer again reused Claude Code, left no app.new or app.old,
  and kept the data directories.
- The installed colonizer-agentd, rtk and claude-guest each ran `--version`
  inside node:24-bookworm.
- The runner's 58 tests passed against an SDK restored from the npm tarball
  without optional packages.

On GitHub (this pull request's run): both linux-binaries jobs built in rust:1-alpine, and both
bundles passed the smoke test from their unpacked archives (97 MB for linux-x86_64, with headscale and
tailscale; 43 MB for darwin-arm64).

Not tested yet:
- Installing the linux-x86_64 release on a Linux machine.
- Launching a colony from an installed release.
- colonizer.dev/install.sh, which the website adds once a release exists. Until
  then, install from the release URL directly.
docs/install.md opens with "Install a release":
- `curl -fsSL https://colonizer.dev/install.sh | sh`;
- what the installer checks and where it installs;
- the two things it fetches from Anthropic's own channels, since a release
  carries no Anthropic code;
- updating by running it again;
- how to install a particular version, or pull the colony image up front.

The build steps are now "Build from source". "What you need" separates what a
release needs (git, gh, curl, tar) from what a source build adds (Node.js 20+,
Rust 1.88+). "Where things live" and "Updating" cover both ways. The README's
"Run it" leads with the one command, with the source build under it.

v0.1.0 was installed from the published release into an empty HOME on an Apple
Silicon Mac (52 s): it fetched the SDK and Claude Code 2.1.267, started, and
reported agentd, the web UI, the claude-code module, the guest Claude binary
and msb 0.6.18. The Linux release is smoke-tested in CI but hasn't been installed
on a Linux machine yet. colonizer.dev/install.sh is added by the website.
…, and publish them on release (#76)

Reserves the names on crates.io and puts the source there. The crates are not an
install: `cargo install` gives the mothership binary without microsandbox, the
in-VM daemon, the agent module or the web UI. Each crate's README says so and
points to `curl -fsSL https://colonizer.dev/install.sh | sh`. The site and docs
don't mention cargo.

- The harness package is renamed `colonizer-harness`, because `colonizer` on
  crates.io belongs to an unrelated project. Its binary is still `colonizer`, so
  the command, target/release/colonizer and the release archives don't change.
  scripts/install.sh and release.yml build it with `-p colonizer-harness`.
- The Headroom bundle pins move from vendor/vendor.lock into
  crates/colonizer/headroom.lock, because `include_str!` reached outside the crate
  and `cargo publish` packages only the crate. fetch-vendor.sh loses its skip for
  kind `bundle`, which no longer occurs. docs/protocol.md and
  headroom-bundle.yml name the new file.
- Both crates are at 0.1.1 and have a crates.io README, a homepage and a readme
  field.
- release.yml: the release job checks the tag against both crate versions. A new
  crates job, tags only and after the release, publishes whichever of
  colonizer-agentd and colonizer-harness crates.io doesn't have yet. It uses
  Trusted Publishing (rust-lang/crates-io-auth-action v1.0.5, pinned, with
  id-token: write), so no token is stored. crates.io won't create a crate that
  way ("Trusted Publishing tokens do not support creating new crates. Publish the
  crate manually, first"), so the first version is published by hand, and later
  tags publish on their own.

Checked:
- `cargo test --workspace` passes.
- `cargo publish --dry-run --locked` packages and verifies both crates. The
  harness package includes headroom.lock and README.md.
- actionlint passes on release.yml.
…m to 0.1.2 (#77)

The crates.io pages now open like the repository: a banner, a badge row, then
what the crate is.

- assets/banners/banner.html draws both banners in colonizer.dev's colours and
  type: a mono eyebrow, a headline, a lead, a foot line and a diagram leading to
  the command.
  - colonizer-harness: "The mothership.", with the colony graph from the site,
    ending in `colonizer`.
  - colonizer-agentd: "Inside every colony.", showing a colony microVM with agentd
    between Claude Code, the event log and the terminal, and the private mesh up
    to the mothership.
- scripts/render-crate-banners.sh renders them at 2560×800 with headless Chrome
  into assets/banners/<crate>.png. The READMEs link them by their
  github.com/ghraw URL on main, because crates.io renders a README
  outside the repository.
- crates/colonizer/README.md and crates/colonizer-agentd/README.md:
  - banner and badges: crates.io version, latest release, docs, MIT;
  - what each crate is, with the install command, and why the crate alone isn't
    an install;
  - for the harness, what Colonizer does and the two crates;
  - for agentd, what it does and its four endpoints.
  The claims come from the README and docs/protocol.md §3.
- Both crates are at 0.1.2: crates.io shows a README only with a new version.
  Tagging v0.1.2 publishes them through Trusted Publishing.

`cargo publish --dry-run` packages and verifies both crates.
…laiming no runtime downloads (#78)

docs/vision.md defined a settler as "the agent module working inside a colony".
Since #70 and #71, the colony chat also calls the subagents that module sends out
settlers, named for their role (Scout, Builder, Inspector…), and colonizer.dev is
getting a page about colonies and settlers. The vocabulary now covers both: the
agent module is a colony's first settler, and the subagents it sends out are
settlers too, each with a role.

Principle 4 said "No runtime downloads". That hasn't been true for a while:
microsandbox pulls a colony image the first time a colony needs it, and
switching Headroom on downloads its bundle. docs/install.md already says so.
The principle now keeps what holds (your hardware, its own dependencies, no
cloud account) and names the two downloads after install.
…itches it on (#79)

colonizer.dev/live will show a dot for every area, about 25 km across, where a
mothership is online, lit while its colonies work. This adds both ends.

The receiver, services/telemetry, is a Cloudflare Worker with a D1 table,
deployed at telemetry.colonizer.dev:
- POST /v1/heartbeat takes {install_id, version, platform, colonies} and ignores
  any other field. The location is Cloudflare's own estimate for the request,
  snapped to a 25 km grid cell; only the cell is stored. The install id is kept
  as a SHA-256 hash. The IP address is used, in memory, for rate limiting only.
  {install_id, online: false} deletes the row.
- GET /v1/presence is the public view: totals and counts per cell, never per
  mothership. Online means a heartbeat in the last 12 minutes.
- Rows older than an hour are pruned during heartbeats. A cron trigger would
  need a workers.dev subdomain on the account, and this needs nothing.
- The privacy-relevant logic is in src/presence.js with node tests.

The mothership, crates/colonizer/src/telemetry.rs:
- Off until the user answers. The answer and a random v4 install id live in
  <config>/telemetry.json. The id is created on switching on and forgotten on
  switching off, so two periods on the map can't be linked.
- While on, a heartbeat every 5 minutes (the service may name 60 s to 1 h), and
  within a minute of the colony count changing. Switching off and shutting down
  send online: false, so the dot goes at once.
- DO_NOT_TRACK or COLONIZER_TELEMETRY=off keep it off whatever Settings says;
  COLONIZER_TELEMETRY_URL points it elsewhere.
- GET/PUT /api/telemetry return the status and the exact next heartbeat.

The web UI asks once, after GitHub and Claude are connected, with "Show on the
map" / "No thanks" / "What is sent". Settings has a Live map section with the
switch and the heartbeat as JSON.

docs/telemetry.md says what is sent, what is kept and for how long, what is
public, and how to keep it off; the README's trust model, configuration and
development sections and docs/protocol.md point to it.

The crates are bumped to 0.1.3, the first release with the live map.
…wered each turn

Refs #39

Co-Authored-By: Colonizer <noreply@colonizer.dev>
@ParallelEntrepreneur

Copy link
Copy Markdown
Collaborator Author

Superseded by the PR above, which is this commit cherry-picked onto current main.

This branch was cut from the history main had before it was rewritten, so eight of its nine commits are already on main under different SHAs and git saw two lineages for the same content. Cherry-picking the one real commit left three trivial conflicts, all resolved there.

It also surfaced a break no textual merge would have flagged: Gateway::new() grew a data_dir argument here, and main.rs gained a test helper on main after this branch was cut that still called it with none. Cherry-picked cleanly, then failed to compile. Fixed in the new PR.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

No way to tell whether a model provider is ever used

1 participant