Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 5 additions & 6 deletions docs/postgres-setup.md
Original file line number Diff line number Diff line change
Expand Up @@ -101,12 +101,11 @@ connection strings into commands that will remain in shell history. Migrations
are transactional, ordered, and checksum-verified; editing an already-applied
migration is rejected.

The production template targets `aws:us-west-2`, supported by Wrangler 4.122's
placement schema. Cloudflare runs fetch handlers in a nearby Cloudflare data
center, not inside AWS itself. Placement does not move an existing Neon project
or Durable Object, and is not a residency guarantee. Confirm the Neon primary
and upstream latency before activating the production migration. See
[Cloudflare placement](https://developers.cloudflare.com/workers/configuration/placement/).
The production Worker targets Oregon (`aws:us-west-2`) for fetch execution.
This is not a data-residency guarantee and does not move an existing Neon
database or Durable Object. Verify the active database's region separately;
Worker placement alone does not establish full colocation.
See [Cloudflare placement](https://developers.cloudflare.com/workers/configuration/placement/).

Do not enable the new account offering until retail token rates, Autumn events,
legacy Pro handling, reconciliation, and end-to-end production checks are ready.
Expand Down
16 changes: 8 additions & 8 deletions inference/laya/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,15 +21,15 @@ One single-label decision is one question. Multi-label asks one question per lab

Text ≤2,000 characters, 2–16 labels ≤100 characters each, instructions ≤400 characters. The combined model input must also fit the selected checkpoint's context (512 English, 1,024 multilingual); long labels/instructions may hit the smaller question budget. Reject instead of truncate.

Modal app `classifier-laya-router-trial` has global compute placement (no west-coast pin), one ingress in `us-east`, and private proxy authentication. This is not a GPU replica in every region, and does not promise 40–70ms worldwide. Pools are isolated. Large flashes are shed, not absorbed into an unbounded backlog. Accepted payloads are only held in memory, never stored as durable jobs.
Modal app `classifier-laya-router-trial-west` colocates its compute and ingress in `us-west`, near west-coast Worker execution, and uses private proxy authentication. This avoids Modal's former east-to-global inter-region hop, but it is not a GPU replica in every region and does not promise the same client round-trip worldwide. Pools are isolated. Large flashes are shed, not absorbed into an unbounded backlog. Accepted payloads are only held in memory, never stored as durable jobs.

## Cost

At [Modal list prices](https://modal.com/pricing), checked 2026-09-20, approximate allocation cost is $0.80/hour L4 + 2 × $0.0473/hour CPU + 4 × $0.008/hour memory = **$0.9266/hour per lane**.
At [Modal list prices](https://modal.com/pricing), checked 2026-09-20, base allocation is $0.80/hour L4 + 2 × $0.0473/hour CPU + 4 × $0.008/hour memory. The narrow-region 1.75× multiplier makes the colocated deployment approximately **$1.6216/hour per lane**.

- Warm fast: about **$22.24/day or $667 per 30-day month**.
- Bulk: about **$0.93 per allocated hour**, including startup/idle tails; another ~$667/month if continuously active.
- Both continuously active: about **$44.48/day or $1,334/month**. Fleet caps are not a hard dollar budget; CPU overage, other infrastructure, reviews, taxes and credits are separate.
- Warm fast: about **$38.92/day or $1,168 per 30-day month**.
- Bulk: about **$1.62 per allocated hour**, including startup/idle tails; another ~$1,168/month if continuously active.
- Both continuously active: about **$77.83/day or $2,335/month**. Fleet caps are not a hard dollar budget; CPU overage, other infrastructure, reviews, taxes and credits are separate.

Laya's explicit retail token rate is zero during the trial. Smart reviews retain existing prices. Provider-token spend is not Modal GPU hosting spend: inspect Modal usage for the actual bill. Account analytics marks combined provider cost unknown when Modal is involved.

Expand All @@ -54,7 +54,7 @@ Worker configuration lives in `wrangler.example.toml`; production deploys from m
To pause new requests, set `LAYA_ENABLED = "false"` in the canonical config and deploy. **Disabling the Worker route does not stop the warm GPU bill.** To stop both trial pools immediately:

```sh
uvx --from modal modal app stop classifier-laya-router-trial
uvx --from modal modal app stop classifier-laya-router-trial-west
```

Existing Jev remains available. Redeploy `deploy.py` and re-enable the Worker flag to resume. Do not scale up to hide overload without revisiting costs.
Expand All @@ -63,13 +63,13 @@ Watch model labels `laya-0.3.4-<checkpoint>-<lane>` in existing classifier analy

## Quota admission rollout

`QuotaCoordinator` keeps each caller's existing fast/smart tier counters and fast/bulk Laya counters in one object. A warm Laya request checks and writes both quotas in one durable operation. Jev still shares its tier allowance with Laya, and a Laya lane still shares its allowance across tiers. Decision and question costs remain separate. A tier refusal spends nothing; a lane refusal retains the tier debit. Combined Laya admission fails closed if either counter is unavailable; Jev retains its existing fail-open behavior.
`QuotaCoordinator` keeps each caller's existing fast/smart tier counters and fast/bulk Laya counters in one object. A warm Laya request checks and writes both quotas in one durable operation. Anonymous fast calls first pass a local Cloudflare burst shield, then exact admission overlaps inference; an exact refusal aborts the in-flight Modal request and is still authoritative. Jev still shares its tier allowance with Laya, and a Laya lane still shares its allowance across tiers. Decision and question costs remain separate. A tier refusal spends nothing; a lane refusal retains the tier debit. Combined Laya admission fails closed if either counter is unavailable; Jev retains its existing fail-open behavior.

Deploy the transfer-aware `RateLimiter`, coordinator binding and migration with `QUOTA_COORDINATOR_ENABLED = "false"` first. After that deployment completes, enable the flag in a second deployment. On first use, each old object freezes its counters and forwards later requests. The coordinator imports the frozen snapshot without resetting the minute or day allowance. Interrupted imports retry the same snapshot. Only opaque Durable Object IDs, scope names and counters are stored. Migration adds latency on the first call, not every call.

Rollback by disabling `QUOTA_COORDINATOR_ENABLED` while retaining the new classes, bindings and forwarding code. **Do not roll back to code predating the transfer protocol:** its old counters are frozen and no longer authoritative. The disabled path continues to follow transferred counters. Neither successful admission nor forwarding uses unconfirmed storage writes.

Run `node inference/laya/latency.mjs <output.json>` before and after deployment from the same client. It records 20 sequential synthetic requests, separates the first call, and reports end-to-end, Worker, quota, Modal and backend timings. Do not equate backend compute time or the Worker-to-Modal span with client latency. The main Worker is in Oregon for account database access; the Modal ingress is in Virginia. Moving every request east would also move database-dependent account traffic, so placement must be evaluated per path.
Run `node inference/laya/latency.mjs <output.json>` before and after deployment from the same client. It records 20 sequential synthetic requests, separates the first call, and reports end-to-end, Worker, quota, Modal and backend timings. Do not equate backend compute time or the Worker-to-Modal span with client latency. The Worker targets Oregon (`aws:us-west-2`); Modal ingress and compute use `us-west`. Existing Durable Objects and databases are not relocated by these settings. Measure other client regions independently.

## Evidence

Expand Down
7 changes: 5 additions & 2 deletions inference/laya/deploy.py
Original file line number Diff line number Diff line change
Expand Up @@ -18,9 +18,12 @@ def download():
.env({"USE_TF": "0", "LAYA_REVISION": REVISION, "HF_HUB_OFFLINE": "1", "TRANSFORMERS_OFFLINE": "1"})
.add_local_file(root / "adapter.py", "/root/adapter.py")
.add_local_file(root / "runtime.py", "/root/runtime.py"))
app = modal.App("classifier-laya-router-trial")
app = modal.App("classifier-laya-router-trial-west")
resources = dict(image=image, gpu="L4", cpu=2, memory=4096,
compute_region=None, routing_region="us-east", unauthenticated=False,
# Keep Modal's ingress and container together near west-coast
# traffic: an east-coast ingress adds a continent-scale hop
# that dwarfs ~30 ms inference.
compute_region="us-west", routing_region="us-west", unauthenticated=False,
max_containers=1, target_concurrency=1, startup_timeout=240)


Expand Down
2 changes: 1 addition & 1 deletion inference/laya/smoke.py
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ async def main():
report = {}
async with httpx.AsyncClient(headers=headers, timeout=30) as client:
for lane in ("fast", "bulk"):
url = f"https://miryaboy--classifier-laya-router-trial-{lane}.us-east.modal.direct"
url = f"https://miryaboy--classifier-laya-router-trial-west-{lane}.us-west.modal.direct"
started = time.monotonic()
while time.monotonic() - started < 240:
try:
Expand Down
2 changes: 1 addition & 1 deletion src/docs.ts
Original file line number Diff line number Diff line change
Expand Up @@ -95,7 +95,7 @@ LAYA TRIAL
Fast stays warm. Bulk starts on demand and can return 503 while starting.
On 429 or 503, respect Retry-After and use bounded retries with backoff.
The repository CLI retries Laya for up to three minutes per batch.
Global GPU placement is not a replica in every region or a latency promise.
One regional GPU deployment is not a replica in every region or a latency promise.
Accepted work is held only in memory; there is no durable batch-job service.
Overload never silently switches the model or processing lane.

Expand Down
3 changes: 1 addition & 2 deletions src/http/classification.ts
Original file line number Diff line number Diff line change
Expand Up @@ -34,7 +34,6 @@ export async function accountClassification(request: Request, env: AppEnv & Part
type: `${source} · ${body.dimensions ? "Dimensions" : body.multi ? "Multi-label" : "Single-label"}`, meteringMode: "tokens",
});
if (!reservation) throw new AppError(401, "Missing account credential.");
const plan = await env.APP_DB.prepare("SELECT billing_plan FROM app_accounts WHERE id=?").bind(accountId).first<{ billing_plan: string }>();
const meter = newMeter();
let admissionError: AppError | undefined;
let reservationQueue = Promise.resolve();
Expand Down Expand Up @@ -67,7 +66,7 @@ export async function accountClassification(request: Request, env: AppEnv & Part
let response: Response;
try {
response = await worker.fetch(new Request(request.url, { method: "POST", headers: request.headers, body: text }), env as Env, ctx,
{ account: { id: accountId, multiplier: plan?.billing_plan === "free" ? 1 : 10 }, meter });
{ account: { id: accountId, multiplier: reservation.billingPlan === "free" ? 1 : 10 }, meter });
await reservationQueue;
if (admissionError) throw admissionError;
} catch (error) {
Expand Down
52 changes: 40 additions & 12 deletions src/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -930,6 +930,14 @@ async function escalate(env: Env, inputs: string[], labels: string[], instructio
* fast-tier answers under a smart-tier label.
*/
type LayaRequestTiming = LayaTiming & { runMs?: number };
type LayaRun = Promise<Awaited<ReturnType<typeof runLaya>>>;

function startLaya(env: Env, plan: LayaPlan, meter: Meter | undefined, timing: LayaRequestTiming | undefined, signal?: AbortSignal): LayaRun {
const started = performance.now();
return runLaya(env, plan, meter, timing, signal).finally(() => {
if (timing) timing.runMs = performance.now() - started;
});
}

async function classifyMany(
env: Env,
Expand All @@ -941,16 +949,15 @@ async function classifyMany(
meter?: Meter,
layaPlan?: LayaPlan,
layaTiming?: LayaRequestTiming,
layaRun?: LayaRun,
): Promise<{ results: Result[]; escalationFailed: number }> {
const keys = jevKeys(env);
if (keys || layaPlan) {
const started = Date.now();
let jev: Awaited<ReturnType<typeof jevClassify>> | null = null;
try {
if (layaPlan) {
const runStarted = performance.now();
jev = await runLaya(env, layaPlan, meter, layaTiming);
if (layaTiming) layaTiming.runMs = performance.now() - runStarted;
jev = await (layaRun ?? startLaya(env, layaPlan, meter, layaTiming));
} else jev = await jevClassify(keys!, inputs, labels, instructions, !!multi, meter);
} catch (e) {
if (layaPlan) throw e;
Expand Down Expand Up @@ -995,14 +1002,12 @@ async function classifyMany(
}

/** Keep each field independent, including smart escalation and the bounded LLM fallback. */
async function classifyMatrix(env: Env, inputs: string[], dimensions: Dimension[], batches: DimensionBatch[], tier: Tier, instructions: string | undefined, meter: Meter, layaPlan?: LayaPlan, layaTiming?: LayaRequestTiming) {
async function classifyMatrix(env: Env, inputs: string[], dimensions: Dimension[], batches: DimensionBatch[], tier: Tier, instructions: string | undefined, meter: Meter, layaPlan?: LayaPlan, layaTiming?: LayaRequestTiming, layaRun?: LayaRun) {
const started = Date.now();
let jev: Awaited<ReturnType<typeof classifyDimensions>> | undefined;
const keys = jevKeys(env);
if (layaPlan) {
const runStarted = performance.now();
const flat = await runLaya(env, layaPlan, meter, layaTiming);
if (layaTiming) layaTiming.runMs = performance.now() - runStarted;
const flat = await (layaRun ?? startLaya(env, layaPlan, meter, layaTiming));
jev = inputs.map((_, i) => flat.slice(i * dimensions.length, (i + 1) * dimensions.length));
} else if (keys) {
try { jev = await classifyDimensions(keys, batches, meter); }
Expand Down Expand Up @@ -1987,28 +1992,51 @@ const worker = {
if (!enterprise && decisions > rpm) return fail(`Maximum ${rpm} ${dimensions ? "decisions" : "inputs"} per ${account ? "account" : "public"} ${tier} request; split the batch to fit the per-minute quota`, 400, dimensions ? "too_many_decisions" : "too_many_inputs");
const regularQuotaStarted = performance.now();
const combinedQuota = !!layaPlan && env.QUOTA_COORDINATOR_ENABLED === "true" && !!env.QUOTAS;
let layaRun: LayaRun | undefined;
let layaAbort: AbortController | undefined;
if (combinedQuota && layaTiming && processing === "fast" && env.LAYA_FAST_ADMISSION) {
try {
const key = env.LIMITER.idFromName(`laya:fast:${quotaOwner}`).toString();
const checks = await Promise.all(Array.from({length:layaPlan!.cost}, () => env.LAYA_FAST_ADMISSION!.limit({key})));
if (checks.some(check => !check.success))
return fail("Laya trial limit reached; retry after the window resets",429,"laya_rate_limit",{"retry-after":"1"});
} catch {
return fail("Laya admission control is temporarily unavailable",503,"laya_unavailable",{"retry-after":"1"});
}
layaAbort = new AbortController();
layaRun = startLaya(env, layaPlan!, meter, layaTiming, layaAbort.signal);
// Exact admission can still reject before this promise is awaited.
void layaRun.catch(() => {});
}
let gate: AdmissionResult;
if (combinedQuota) {
const quotas: Omit<Quota,"id">[] = enterprise ? [] : [{scope:tier,cost:decisions,limit:rpm,daily:TIERS[tier].daily * multiplier}];
quotas.push({scope:`laya:${processing}`,cost:layaPlan!.cost,limit:LAYA_LIMITS[processing].rpm,daily:LAYA_LIMITS[processing].daily});
try {
const response = await admit({LIMITER:env.LIMITER,QUOTAS:env.QUOTAS!},quotaOwner,quotas,!!layaTiming);
// The coordinator returns stage timings without the diagnostic
// storage.sync(); its output gate still preserves counter durability.
const response = await admit({LIMITER:env.LIMITER,QUOTAS:env.QUOTAS!},quotaOwner,quotas,false);
readQuotaTiming(response, layaTiming ? regularQuotaTiming : undefined);
gate = await response.json() as AdmissionResult;
if (typeof gate?.limited !== "boolean" || !Number.isFinite(gate.remaining) ||
(gate.limited ? !["tier","lane"].includes(gate.limitedBy!) : !Number.isFinite(gate.laneRemaining)))
throw new Error("Invalid quota admission response");
if (gate.limited && gate.limitedBy === "lane") return fail("Laya trial limit reached; retry after the window resets",429,
gate.scope === "day" ? "rate_limit_day" : "laya_rate_limit", {"retry-after":String(gate.resetIn ?? 60)});
if (gate.limited && gate.limitedBy === "lane") {
layaAbort?.abort();
return fail("Laya trial limit reached; retry after the window resets",429,
gate.scope === "day" ? "rate_limit_day" : "laya_rate_limit", {"retry-after":String(gate.resetIn ?? 60)});
}
layaRemaining = gate.laneRemaining ?? -1;
} catch {
layaAbort?.abort();
return fail("Laya admission control is temporarily unavailable",503,"laya_unavailable", {"retry-after":"1"});
}
} else gate = enterprise
? { limited: false, remaining: -1 }
: await limited(env, tier, quotaOwner, decisions, multiplier, layaTiming ? regularQuotaTiming : undefined);
regularQuotaMs = performance.now() - regularQuotaStarted;
if (gate.limited) {
layaAbort?.abort();
const perDay = gate.scope === "day";
// The moment someone runs out of room is the moment to say where more is.
// A free caller is told the plan that lifts this exact limit and gets its
Expand Down Expand Up @@ -2053,12 +2081,12 @@ const worker = {
let escalationFailed = 0;
try {
if (dimensions) {
const r = await classifyMatrix(env, inputs, dimensions, dimensionBatches, tier, instructions, meter, layaPlan, layaTiming);
const r = await classifyMatrix(env, inputs, dimensions, dimensionBatches, tier, instructions, meter, layaPlan, layaTiming, layaRun);
matrix = r.results;
results = matrix.flat();
escalationFailed = r.escalationFailed;
fallbackDecisions = r.fallbackDecisions;
} else ({ results, escalationFailed } = await classifyMany(env, inputs, labels, tier, instructions, multi, meter, layaPlan, layaTiming));
} else ({ results, escalationFailed } = await classifyMany(env, inputs, labels, tier, instructions, multi, meter, layaPlan, layaTiming, layaRun));
} catch (e) {
if (e instanceof LayaError) return fail(e.message, e.status,
e.status === 400 ? "laya_input" : e.status === 429 ? "laya_rate_limit" : "laya_unavailable",
Expand Down
7 changes: 5 additions & 2 deletions src/laya.ts
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ export type LayaEnv = {
LAYA_MODAL_KEY?: string;
LAYA_MODAL_SECRET?: string;
LAYA_ENABLED?: string;
LAYA_FAST_ADMISSION?: RateLimit;
LIMITER: DurableObjectNamespace;
};
export const LAYA_LIMITS = {
Expand Down Expand Up @@ -84,7 +85,7 @@ const probability = (v: unknown): v is number => typeof v === "number" && Number

export type LayaTiming = { fetchMs?: number; headersMs?: number; backendMs?: number };

export async function runLaya(env: LayaEnv, plan: LayaPlan, meter?: Meter, timing?: LayaTiming): Promise<JevResult[]> {
export async function runLaya(env: LayaEnv, plan: LayaPlan, meter?: Meter, timing?: LayaTiming, signal?: AbortSignal): Promise<JevResult[]> {
const url = plan.processing === "fast" ? env.LAYA_FAST_URL : env.LAYA_BULK_URL;
if (env.LAYA_ENABLED !== "true" || !url || !env.LAYA_MODAL_KEY || !env.LAYA_MODAL_SECRET)
throw new LayaError("Laya trial is currently unavailable", 503);
Expand All @@ -100,7 +101,9 @@ export async function runLaya(env: LayaEnv, plan: LayaPlan, meter?: Meter, timin
try {
response = await fetch(url + "/predict", { method: "POST",
headers: { "content-type": "application/json", "Modal-Key": env.LAYA_MODAL_KEY, "Modal-Secret": env.LAYA_MODAL_SECRET },
body: JSON.stringify({ batch }), signal: AbortSignal.timeout(Math.min(15_000, deadline - Date.now())) });
body: JSON.stringify({ batch }), signal: signal
? AbortSignal.any([signal, AbortSignal.timeout(Math.min(15_000, deadline - Date.now()))])
: AbortSignal.timeout(Math.min(15_000, deadline - Date.now())) });
} catch { throw new LayaError("Laya timed out or could not be reached", 503, 5); }
if (timing) timing.headersMs = (timing.headersMs ?? 0) + performance.now() - started;
if (response.status === 429 || response.status === 503)
Expand Down
Loading
Loading