Skip to content

Bound free inference spending and isolate internal endpoints - #107

Merged
mrmps merged 3 commits into
mainfrom
codex/inference-spending
Sep 22, 2026
Merged

mrmps merged 3 commits into
mainfrom
codex/inference-spending

Conversation

@mrmps

@mrmps mrmps commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner

Bound public inference spending before provider work and keep expensive classification available to funded workspaces.

Before (Pricing · desktop) After (Pricing · desktop)
Before After
Before (Pricing · mobile) After (Pricing · mobile)
Before After
Before (Public navigation) After (Public navigation)
Before After
  • Free traffic shares atomic $100/day, $0.50/IP/day, $0.01/request and four concurrent requests/IP limits. IPv6 /64s share an allowance; retries, aliases and unfunded keys cannot bypass it.
  • Funded requests reserve workspace credit once, check each provider attempt locally, and settle actual token usage after responding. Uncertain spend remains held and triggers operator alerts.
  • Cache Spur reputation checks within a 45,000-lookup rolling 32-day ceiling. Known anonymous proxy infrastructure requires a funded key. Chat and skill-review routes require a separate internal credential and disappear from public discovery.
  • Return actionable errors and publish the limits across pricing, API docs and discovery. Idempotency keys prevent duplicate execution with 409; responses are not cached.

Free smart mode supports requests that fit its allowance; large prompts/batches need a funded key. Existing classification-count quotas still apply. Paid requests bypass Spur and the free coordinator, while existing workspace quota checks remain. Provider timeouts retain their allowance; unresolved billing requires reconciliation rather than an automatic refund.

The CI artifact spending-e2e records behavior through the production Vite build, workerd and SQLite Durable Objects using explicit provider fixtures. Separate HTTP scenarios use concurrent native PostgreSQL connections to verify shared balances across keys. Live-provider checks run after deployment.

@coderabbitai

coderabbitai Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 1ba5f20a-d697-4437-b14a-7c7b1bea3944

📥 Commits

Reviewing files that changed from the base of the PR and between a68b524 and e13eec9.

📒 Files selected for processing (36)
  • .github/workflows/check.yml
  • .github/workflows/deploy.yml
  • e2e/spending.mjs
  • migrations/postgres/0012_spending_idempotency.sql
  • package.json
  • src/agents.ts
  • src/alerts.ts
  • src/cost.ts
  • src/docs.ts
  • src/features/billing/plans.tsx
  • src/home.ts
  • src/http/classification.ts
  • src/http/mcp.ts
  • src/http/spending-classification.ts
  • src/index.ts
  • src/jev.ts
  • src/openapi.ts
  • src/pages.ts
  • src/pricingui.ts
  • src/retail-rates.json
  • src/server.ts
  • src/skills.ts
  • src/spending/free-budget.ts
  • src/spending/index.ts
  • src/spending/permit.ts
  • src/spending/policy.ts
  • src/typesafe-compat.ts
  • src/wellknown.ts
  • test/chat.test.ts
  • test/skills.test.ts
  • tests/plans-prices.test.ts
  • tests/public-pricing.test.ts
  • tests/retail-rates.test.ts
  • tests/spending.e2e.test.ts
  • wrangler.example.toml
  • wrangler.local.toml
 ________________________________________________________________
< Fully armed and operationally intelligent code reviewer bunny. >
 ----------------------------------------------------------------
  \
   \   \
        \ /\
        ( )
      .( o ).
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@mrmps
mrmps merged commit bed15c9 into main Sep 22, 2026
2 of 3 checks passed
mrmps added a commit that referenced this pull request Sep 24, 2026
* Make admin analytics dense and searchable on one page

* Fix payment reconciliation during overdue flags, refunds, and provider failures (#100)

* Add TypeSafe SDK-compatible API (#102)

* Reduce classification latency with west-region Laya and fewer account round trips (#103)

* Cut Laya fast latency below 150ms

* Retain Oregon placement and remove redundant account plan lookup

* Apply workspace keys to TypeSafe SDK requests (#104)

* Serve Laya on Beam and add Kev, retiring the Modal deployment (#105)

Beam serves the same System One protocol as TypeSafe, so Laya stops being a
bespoke HTTP client and becomes a third transport in src/jev.ts: one packer,
one retry policy, one validator, one meter. Kev joins it as `model: "kev"`.

Two Beam limits have no analogue at TypeSafe. Every model refuses more than
32 named questions per request, and the contexts are small — Laya answers
within 512 tokens of state plus one question, Kev packs 8,192. Both reject an
overflow rather than truncating, so a context refusal is translated to
max_tokens_exceeded and runJevBatches halves the batch and recovers. Laya's
16-item ceiling is empirical: Beam accepted 20 items of ordinary support text
and refused 25.

Behaviour preserved deliberately:
  - Lane quotas (fast 60/min, bulk 1,000/min) are a product decision about
    shared capacity, not a property of Modal, so callers keep their limits.
  - Beam requests are not retried. The old Modal client did not retry either,
    and a lane quota counts attempts, so retrying would spend a caller's
    budget on a model that will not clear inside a backoff. Retries stay a
    Jev-only behaviour, now expressed as Backend.attempts.
  - The quota-refusal abort from #103 still cancels in-flight inference.

Behaviour that changed, visibly:
  - A result is labelled jev/laya or jev/kev. Beam does not report which
    checkpoint answered, so laya-0.3.4-<checkpoint>-<lane> is gone, and with
    it the timing fields only Modal could populate (backendMs, headersMs).
  - Account analytics records a real provider cost for these models instead
    of marking spend unknown, because Beam reports token counts.
  - Cold-start 503s cannot happen: there is no pool of ours to start.

Cost: usage-priced at $0.021 per 1M input tokens with no idle charge, against
$1,168 a month for the warm Modal lane alone after the us-west multiplier.

BEAM_API_KEY is the only credential either model needs; LAYA_FAST_URL,
LAYA_BULK_URL, LAYA_MODAL_KEY and LAYA_MODAL_SECRET are removed.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Describe the Beam transport in the repository instructions (#106)

AGENTS.md told contributors that Jev's transports live in src/jev.ts and said
nothing about Laya or Kev. Both now go through the same module, and the two
rules that are easy to get wrong — Beam's 32-question ceiling and the fact
that a Beam request is never retried because a lane quota counts attempts —
were only discoverable by reading the code.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Bound free inference spending and isolate internal endpoints (#107)

* Bound public inference spending and protect internal endpoints

* Verify spending boundaries and finish paid admission and recovery alerts

* Demonstrate live provider billing through funded admission

* Keep anonymous live smart checks within the request allowance (#108)

* Simplify billing to input tokens plus Smart escalations (#109)

* Price base input at Jev's rate and Smart reviews per escalation

* Expose customer usage headers to browser API clients

* Fix agent-reported batch recovery and spending error reporting (#110)

* Let paid requests settle into a bounded negative balance (#111)

* Allow bounded paid overdrafts and preserve completed Smart results

* Describe the hold without implying prepaid coverage is required

* Add agent testimonials to feedback (#112)

* Route documents over 32,000 characters to chunklaya

An input over MAX_CHARS used to be refused with input_too_long. When
CHUNKLAYA_URL and CHUNKLAYA_TOKEN are set and CHUNKLAYA_ENABLED is "true",
it is now answered by chunklaya, our own long-document service (Laya behind
a chunk-and-index harness, github.com/myxamediyar/chunklaya, serve/), and
the result is labelled chunklaya/multilingual. Nothing at or under 32,000
characters changes; an explicit model "jev" keeps its ceiling; model
"chunklaya" selects it for shorter text. Unconfigured, the old 400 stands.

The service speaks System One, so it is a fourth Backend in src/jev.ts
through the bearer transport Beam already uses, with the URL and token read
from the environment, one document per request, a 60 s deadline, and no
per-token cost (the pod is billed by the hour). Its refusals are reported
as chunklaya_input, chunklaya_busy and chunklaya_unavailable in our own
words, and never fall back to Jev or the LLM chain. Ceilings: 4,000,000
characters per input, 20 inputs per request, no smart tier.

Dimensions pack against the backend's limits so every question about a
document travels in one request. Billing prices the model at zero in the
rate card, the spending table and the reservation bound. Docs, OpenAPI, MCP
and the CLI render the new model and codes from the same constants.

The deploy workflow passes CHUNKLAYA_URL and CHUNKLAYA_TOKEN to the Worker
when both exist as repository secrets, and leaves them out otherwise.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Price the chunklaya route as chunklaya under a spending permit

postBeam sent every request through providerFetch under Beam's name, so a
request routed to chunklaya was priced as beam:chunklaya/multilingual,
which the price table does not have, and was refused as unpriced_model.
The provider the function already receives is now the one it names.

test/chunklaya.test.ts runs the route under a Permit, as production does,
and asserts it is priced as chunklaya at zero.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Let the deploy's anonymous probe pass through the reputation gate

The post-deploy check curls one anonymous classification from the hosted
runner. The free-traffic reputation gate can refuse a runner's IP with 403
proxy_requires_payment, which is the Worker answering with its own gate,
not a broken deploy; the last two deploys were marked failed by it, and
`curl --fail` discarded the body that would have said so.

The probe now prints the status and body, passes with a notice on the
gate's own 403, still requires "spam" on a 200, and fails on anything
else. The funded end-to-end step that followed it runs again.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Give operator inference a separate bounded daily allowance

* Limit anonymous traffic by label set (#116)

* Limit anonymous traffic by label set

* Preserve public quota documentation contract

* Close cross-endpoint label allowance bypasses

* Add paginated admin view of all retained label sets (#117)

* Exclude obsolete raw keys from the label registry (#118)

* Index collected label names for the admin catalog (#119)

* Expose classification token usage and customer pricing (#121)

* Classify long documents with Jev screening and fixed context pricing (#122)

* Implement Jev long-context screening with fixed context pricing

* Budget selected evidence against serialized Jev requests

* Give long-context Jev calls bounded longer deadlines (#123)

* Open POST /v1/systemone to images through the dgemma service

A System One body that names model "dgemma" or carries an images array is
forwarded to the image-capable DiffusionGemma service (vLLM's structured-read
mode, vllm-project/vllm#57250), which speaks the same contract. Every other
body still goes to TypeSafe. Neither answers for the other: a refused body is
400 dgemma_input with the service's reason, a saturated service 429
dgemma_busy, a down or unconfigured one 503 dgemma_unavailable, and images
sent under another model 400 images_unsupported.

Images are data URLs (PNG, JPEG, WebP, GIF), at most 4 and 900,000 base64
characters together. The service is reached through the DGEMMA_URL and
DGEMMA_TOKEN Worker secrets when DGEMMA_ENABLED is "true"; the deploy
workflow passes the pair through when both repository secrets are set. The
route is priced as dgemma at zero for workspace permits, since the service is
billed by the hour. OpenAPI, the docs and AGENTS.md describe the field, the
model and the codes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Enable paid 10M-token long-context classification jobs (#125)

* Add paid ten-million-token long-context jobs

* Allow browser preflight for long-context uploads

* Tighten the dgemma door after review

Route on the presence of the images field, so an empty array is refused
rather than forwarded as an unknown key; require an https service address
and reject redirects; keep the 60-second deadline through the body read;
accept a 200 only when it reports this model, an entry for every question
asked and a usage count; and scope the zero-rate trial to System One, where
an images field selects the service as surely as its name.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Let the System One response schema carry a skipped dgemma answer

A question the service skips under ask_if is answered null; the published
response schema now allows that beside the answer objects.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Keep the dgemma request's redirects manual

The Workers runtime has no redirect "error" mode and throws on the option,
so every request to the image-capable service failed before it was sent and
the route answered dgemma_unavailable. Redirects are now "manual": a 3xx
comes back as a response and is answered as an outage, which is what the
runtime's own message recommends. A test pins the option and the mapping.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Support time-limited complimentary Pro grants

* Start complimentary Pro year on verified signup

* Accept whole 10M-token documents in one upload (#127)

* Add URL classification with metered Context.dev scraping (#128)

* Restore public chat with classification tools (#130)

* Restore public chat with classification tools

* Keep non-chat private routes protected

* Keep chat subtitle readable with conversation controls

* Measure chat usage in the unified analytics dashboard

* Align chart units and verify dashboard navigation and export

* Count every chat classification tool variant

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: myxamediyar <mukhamediyar@berkeley.edu>
Co-authored-by: Mukhamediyar <86684638+myxamediyar@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant