Bound free inference spending and isolate internal endpoints - #107
Merged
Merged
Conversation
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Note Currently processing new changes in this PR. This may take a few minutes, please wait... ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (36)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
mrmps
added a commit
that referenced
this pull request
Sep 24, 2026
* Make admin analytics dense and searchable on one page * Fix payment reconciliation during overdue flags, refunds, and provider failures (#100) * Add TypeSafe SDK-compatible API (#102) * Reduce classification latency with west-region Laya and fewer account round trips (#103) * Cut Laya fast latency below 150ms * Retain Oregon placement and remove redundant account plan lookup * Apply workspace keys to TypeSafe SDK requests (#104) * Serve Laya on Beam and add Kev, retiring the Modal deployment (#105) Beam serves the same System One protocol as TypeSafe, so Laya stops being a bespoke HTTP client and becomes a third transport in src/jev.ts: one packer, one retry policy, one validator, one meter. Kev joins it as `model: "kev"`. Two Beam limits have no analogue at TypeSafe. Every model refuses more than 32 named questions per request, and the contexts are small — Laya answers within 512 tokens of state plus one question, Kev packs 8,192. Both reject an overflow rather than truncating, so a context refusal is translated to max_tokens_exceeded and runJevBatches halves the batch and recovers. Laya's 16-item ceiling is empirical: Beam accepted 20 items of ordinary support text and refused 25. Behaviour preserved deliberately: - Lane quotas (fast 60/min, bulk 1,000/min) are a product decision about shared capacity, not a property of Modal, so callers keep their limits. - Beam requests are not retried. The old Modal client did not retry either, and a lane quota counts attempts, so retrying would spend a caller's budget on a model that will not clear inside a backoff. Retries stay a Jev-only behaviour, now expressed as Backend.attempts. - The quota-refusal abort from #103 still cancels in-flight inference. Behaviour that changed, visibly: - A result is labelled jev/laya or jev/kev. Beam does not report which checkpoint answered, so laya-0.3.4-<checkpoint>-<lane> is gone, and with it the timing fields only Modal could populate (backendMs, headersMs). - Account analytics records a real provider cost for these models instead of marking spend unknown, because Beam reports token counts. - Cold-start 503s cannot happen: there is no pool of ours to start. Cost: usage-priced at $0.021 per 1M input tokens with no idle charge, against $1,168 a month for the warm Modal lane alone after the us-west multiplier. BEAM_API_KEY is the only credential either model needs; LAYA_FAST_URL, LAYA_BULK_URL, LAYA_MODAL_KEY and LAYA_MODAL_SECRET are removed. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * Describe the Beam transport in the repository instructions (#106) AGENTS.md told contributors that Jev's transports live in src/jev.ts and said nothing about Laya or Kev. Both now go through the same module, and the two rules that are easy to get wrong — Beam's 32-question ceiling and the fact that a Beam request is never retried because a lane quota counts attempts — were only discoverable by reading the code. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * Bound free inference spending and isolate internal endpoints (#107) * Bound public inference spending and protect internal endpoints * Verify spending boundaries and finish paid admission and recovery alerts * Demonstrate live provider billing through funded admission * Keep anonymous live smart checks within the request allowance (#108) * Simplify billing to input tokens plus Smart escalations (#109) * Price base input at Jev's rate and Smart reviews per escalation * Expose customer usage headers to browser API clients * Fix agent-reported batch recovery and spending error reporting (#110) * Let paid requests settle into a bounded negative balance (#111) * Allow bounded paid overdrafts and preserve completed Smart results * Describe the hold without implying prepaid coverage is required * Add agent testimonials to feedback (#112) * Route documents over 32,000 characters to chunklaya An input over MAX_CHARS used to be refused with input_too_long. When CHUNKLAYA_URL and CHUNKLAYA_TOKEN are set and CHUNKLAYA_ENABLED is "true", it is now answered by chunklaya, our own long-document service (Laya behind a chunk-and-index harness, github.com/myxamediyar/chunklaya, serve/), and the result is labelled chunklaya/multilingual. Nothing at or under 32,000 characters changes; an explicit model "jev" keeps its ceiling; model "chunklaya" selects it for shorter text. Unconfigured, the old 400 stands. The service speaks System One, so it is a fourth Backend in src/jev.ts through the bearer transport Beam already uses, with the URL and token read from the environment, one document per request, a 60 s deadline, and no per-token cost (the pod is billed by the hour). Its refusals are reported as chunklaya_input, chunklaya_busy and chunklaya_unavailable in our own words, and never fall back to Jev or the LLM chain. Ceilings: 4,000,000 characters per input, 20 inputs per request, no smart tier. Dimensions pack against the backend's limits so every question about a document travels in one request. Billing prices the model at zero in the rate card, the spending table and the reservation bound. Docs, OpenAPI, MCP and the CLI render the new model and codes from the same constants. The deploy workflow passes CHUNKLAYA_URL and CHUNKLAYA_TOKEN to the Worker when both exist as repository secrets, and leaves them out otherwise. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Price the chunklaya route as chunklaya under a spending permit postBeam sent every request through providerFetch under Beam's name, so a request routed to chunklaya was priced as beam:chunklaya/multilingual, which the price table does not have, and was refused as unpriced_model. The provider the function already receives is now the one it names. test/chunklaya.test.ts runs the route under a Permit, as production does, and asserts it is priced as chunklaya at zero. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Let the deploy's anonymous probe pass through the reputation gate The post-deploy check curls one anonymous classification from the hosted runner. The free-traffic reputation gate can refuse a runner's IP with 403 proxy_requires_payment, which is the Worker answering with its own gate, not a broken deploy; the last two deploys were marked failed by it, and `curl --fail` discarded the body that would have said so. The probe now prints the status and body, passes with a notice on the gate's own 403, still requires "spam" on a 200, and fails on anything else. The funded end-to-end step that followed it runs again. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Give operator inference a separate bounded daily allowance * Limit anonymous traffic by label set (#116) * Limit anonymous traffic by label set * Preserve public quota documentation contract * Close cross-endpoint label allowance bypasses * Add paginated admin view of all retained label sets (#117) * Exclude obsolete raw keys from the label registry (#118) * Index collected label names for the admin catalog (#119) * Expose classification token usage and customer pricing (#121) * Classify long documents with Jev screening and fixed context pricing (#122) * Implement Jev long-context screening with fixed context pricing * Budget selected evidence against serialized Jev requests * Give long-context Jev calls bounded longer deadlines (#123) * Open POST /v1/systemone to images through the dgemma service A System One body that names model "dgemma" or carries an images array is forwarded to the image-capable DiffusionGemma service (vLLM's structured-read mode, vllm-project/vllm#57250), which speaks the same contract. Every other body still goes to TypeSafe. Neither answers for the other: a refused body is 400 dgemma_input with the service's reason, a saturated service 429 dgemma_busy, a down or unconfigured one 503 dgemma_unavailable, and images sent under another model 400 images_unsupported. Images are data URLs (PNG, JPEG, WebP, GIF), at most 4 and 900,000 base64 characters together. The service is reached through the DGEMMA_URL and DGEMMA_TOKEN Worker secrets when DGEMMA_ENABLED is "true"; the deploy workflow passes the pair through when both repository secrets are set. The route is priced as dgemma at zero for workspace permits, since the service is billed by the hour. OpenAPI, the docs and AGENTS.md describe the field, the model and the codes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Enable paid 10M-token long-context classification jobs (#125) * Add paid ten-million-token long-context jobs * Allow browser preflight for long-context uploads * Tighten the dgemma door after review Route on the presence of the images field, so an empty array is refused rather than forwarded as an unknown key; require an https service address and reject redirects; keep the 60-second deadline through the body read; accept a 200 only when it reports this model, an entry for every question asked and a usage count; and scope the zero-rate trial to System One, where an images field selects the service as surely as its name. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Let the System One response schema carry a skipped dgemma answer A question the service skips under ask_if is answered null; the published response schema now allows that beside the answer objects. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Keep the dgemma request's redirects manual The Workers runtime has no redirect "error" mode and throws on the option, so every request to the image-capable service failed before it was sent and the route answered dgemma_unavailable. Redirects are now "manual": a 3xx comes back as a response and is answered as an outage, which is what the runtime's own message recommends. A test pins the option and the mapping. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Support time-limited complimentary Pro grants * Start complimentary Pro year on verified signup * Accept whole 10M-token documents in one upload (#127) * Add URL classification with metered Context.dev scraping (#128) * Restore public chat with classification tools (#130) * Restore public chat with classification tools * Keep non-chat private routes protected * Keep chat subtitle readable with conversation controls * Measure chat usage in the unified analytics dashboard * Align chart units and verify dashboard navigation and export * Count every chat classification tool variant --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: myxamediyar <mukhamediyar@berkeley.edu> Co-authored-by: Mukhamediyar <86684638+myxamediyar@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Bound public inference spending before provider work and keep expensive classification available to funded workspaces.
Free smart mode supports requests that fit its allowance; large prompts/batches need a funded key. Existing classification-count quotas still apply. Paid requests bypass Spur and the free coordinator, while existing workspace quota checks remain. Provider timeouts retain their allowance; unresolved billing requires reconciliation rather than an automatic refund.
The CI artifact
spending-e2erecords behavior through the production Vite build, workerd and SQLite Durable Objects using explicit provider fixtures. Separate HTTP scenarios use concurrent native PostgreSQL connections to verify shared balances across keys. Live-provider checks run after deployment.