Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions .github/workflows/check.yml
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,10 @@ jobs:
run: bun test tests/spending.e2e.test.ts
env:
POSTGRES_TEST_URL: postgres://postgres:spending-test@localhost:5432/postgres
- name: Verify document admission under concurrent PostgreSQL requests
run: bun e2e/document-admission.ts
env:
POSTGRES_TEST_URL: postgres://postgres:spending-test@localhost:5432/postgres
- name: Render secret-free build configuration
env:
CLOUDFLARE_ACCOUNT_ID: "00000000000000000000000000000000"
Expand All @@ -52,6 +56,8 @@ jobs:
run: npm run test:e2e:long-context-billing
- name: Verify ten-million-token job and settlement
run: npm run test:e2e:long-context-job
- name: Verify one whole ten-million-token document upload
run: npm run test:e2e:whole-document
- name: Save long-context verification evidence
uses: actions/upload-artifact@v4
with:
Expand All @@ -60,6 +66,8 @@ jobs:
captures/long-context-e2e.json
captures/long-context-billing.json
captures/long-context-job.json
captures/whole-document.json
captures/document-admission.json
- name: Save classification usage evidence
uses: actions/upload-artifact@v4
with:
Expand Down
3 changes: 3 additions & 0 deletions .github/workflows/deploy.yml
Original file line number Diff line number Diff line change
Expand Up @@ -104,6 +104,8 @@ jobs:
run: bun e2e/complimentary-pro.ts
- name: Verify ten-million-token job and settlement
run: npm run test:e2e:long-context-job
- name: Verify one whole ten-million-token document upload
run: npm run test:e2e:whole-document
- name: Save long-context verification evidence
uses: actions/upload-artifact@v4
with:
Expand All @@ -112,6 +114,7 @@ jobs:
captures/long-context-e2e.json
captures/long-context-billing.json
captures/long-context-job.json
captures/whole-document.json
- name: Save complimentary Pro verification evidence
uses: actions/upload-artifact@v4
with:
Expand Down
2 changes: 1 addition & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,7 @@ the plain text (`curl classifier.dev`), the HTML and the Markdown never drift.
- Docs are plain text with UPPERCASE headings (`src/docs.ts`, `src/pages.ts`); `renderDoc` turns them into HTML and `toMarkdown` into Markdown.
- Discovery files (`/.well-known/*`, sitemap, robots, auth.md) are generated in `src/wellknown.ts` from `SITE` and the MCP tool table — edit the source, never a served file.
- The MCP servers (`src/mcp.ts`) are stateless Streamable HTTP; tools call the API through `worker.fetch` so limits and logging are shared.
- Long-context jobs in `src/long-context-job.ts` accept ordered bounded parts through a SQLite Durable Object. It holds a maximum-price workspace reservation, retains only selected evidence until completion/cancellation/24-hour expiry, and settles the original part-token total once. Keep `wrangler.example.toml`, the generated docs and the job E2E in sync with this contract.
- Whole-document uploads use `POST /v1/classify`: `src/document-upload.ts` stream-parses JSON or UTF-8 text up to 10M tokens/100 MB, preserving exact cl100k counts across internal fragment boundaries. `src/long-context-job.ts` stores source privately while queued, deletes it as SQLite Durable Object alarms screen it, then judges and settles automatically. Reserve the actual document price after upload. Delete remaining source/evidence on failure, cancellation or 24-hour expiry. The old manual-part API remains supported. Keep generated docs and `e2e/whole-document.mjs` in sync; test through the built Worker, real Durable Objects and PostgreSQL ledger.
- Never commit secrets; `.secrets.env`, `.dev.vars` are ignored. `eval/data/` is ignored except the summary copied to `src/vs-jev.json`.
- Measured numbers on the site come from `eval/`; do not type numbers in by hand.
- Jev is asked through Vercel's AI Gateway first when `AI_GATEWAY_API_KEY`
Expand Down
24 changes: 14 additions & 10 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -226,7 +226,7 @@ provider context check. Eligible chunks can be omitted when that budget fills;
This does not guarantee full-document final reading or universal accuracy.

Use Fast with a workspace key backed by paid balance or an active paid
subscription; anonymous access and signup credit do not qualify. Per request:
subscription; anonymous access and signup credit do not qualify. Synchronous limits:
250,000 original context tokens summed across inputs, 20 documents, 32 decisions,
and a 1 MB body. Retail is $0.084/M original `cl100k_base` context tokens,
counted once across inputs regardless of dimensions and actual inference usage.
Expand All @@ -236,15 +236,19 @@ tokenization or chunking, bounding tokenizer work on pathological inputs.
No eligible evidence returns `422 long_context_no_evidence` without charge.
Explicit `model: "chunklaya"` keeps its separate legacy opt-in behavior.

For up to 10 million input tokens, funded workspaces use a long-context job:
create a UUID job with a `max_tokens` ceiling, upload ordered parts of at most
50,000 tokens and 1 MB each, then call `finish`. The server screens every part
without retaining the full source and keeps a bounded set of selected excerpts
until completion or 24-hour expiry. The final Jev call reads at most 20,000
selected evidence tokens per document. `max_tokens` reserves workspace credit;
successful completion charges the actual part-token sum at the same $0.084/M
rate. Cancellation, expiry and no eligible evidence refund the reservation.
The full API sequence and resume/status route are in the generated docs.
Send one whole document of up to 10 million input tokens (100 MB) to
`POST /v1/classify` with a funded workspace key. Use the normal JSON input and
labels, or send `text/plain` with a `labels` query parameter. Large uploads
return `202` with a `status_url`; `Prefer: respond-async` also works for smaller
documents. The server splits, screens, retries and finishes automatically.
The repository CLI handles upload and polling with
`node cli/classify.js a,b --document document.txt --json`.

The exact whole-document token count sets the reservation and $0.084/M charge.
Source text is stored privately while queued and deleted as it is screened.
Final Jev reads at most 20,000 selected evidence tokens. Completion, failure,
cancellation and 24-hour expiry delete remaining source and evidence; failed
or canceled work is refunded. Results remain available until expiry.

## Analytics

Expand Down
14 changes: 14 additions & 0 deletions cli/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,19 @@ One line per input, in input order: `label`, `confidence`, `text`, tab-separated
everything the API returns; `--quiet` gives labels only; `--count` gives a
histogram.

## One whole document

With a funded workspace key in `CLASSIFY_API_KEY`, upload a UTF-8 document of
up to 10 million tokens (100 MB):

classify renewal,cancellation --document agreement.txt --json

The CLI streams the file once and waits for the result. The server handles
chunking, screening, retries and final judgment. The price is $0.084 per
million original tokens; 10M tokens costs $0.84. Failed jobs are refunded.
`CLASSIFY_JOB_TIMEOUT` controls how long to poll (seconds, default 86400).
Interrupting the CLI does not cancel accepted work; stderr includes its job ID.

## Built for agents

The confidence is calibrated (on a six-way emotion set, answers at ≥ 0.9 were
Expand Down Expand Up @@ -50,6 +63,7 @@ resumes; a daily quota stops immediately. Errors go to stderr with exit code 1.
-k, --max <n> at most n labels (implies --multi)
-s, --smart re-ask uncertain answers of a reasoning model
-i, --instructions <text> extra criteria
--document <file> one whole UTF-8 document (funded workspace)
-r, --review <t> print only inputs with confidence below t
-c, --count label histogram instead of rows
-j, --json NDJSON output
Expand Down
77 changes: 74 additions & 3 deletions cli/classify.js
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,8 @@
// tab-separated and greppable, or NDJSON with --json. Errors go to stderr with
// exit 1; nothing else ever does.

import { readFileSync, writeFileSync, mkdirSync, statSync, realpathSync } from "node:fs";
import { createReadStream, readFileSync, writeFileSync, mkdirSync, statSync, realpathSync } from "node:fs";
import { randomUUID } from "node:crypto";
import { homedir } from "node:os";
import { join } from "node:path";
import { fileURLToPath } from "node:url";
Expand All @@ -23,6 +24,7 @@ const HELP = `classify ${PKG.version} — sort text into your own labels, with a

USAGE
classify <labels> "<text>" one text, prints the label
classify <labels> --document file.txt one whole document, uploads and waits
classify <labels> < items.txt one line per input, in order
cat items.jsonl | classify <labels> --field title

Expand All @@ -42,9 +44,10 @@ OPTIONS
-k, --max <n> at most n labels (implies --multi)
-s, --smart re-ask uncertain answers of a reasoning model (slower)
--model laya opt into the automatically routed Laya trial (default: jev)
--model chunklaya the long-document model; also automatic past 32,000 characters
--model chunklaya opt into the legacy long-document model
--processing bulk Laya bulk lane; default fast is one decision per call
-i, --instructions <text> extra criteria: "judge only the service, ignore the food"
--document <file> upload one UTF-8 document, up to 10M tokens (paid workspace)
-r, --review <t> print only inputs with confidence below t
-c, --count print a label histogram instead of rows
-j, --json NDJSON output with every field the API returns
Expand Down Expand Up @@ -86,7 +89,7 @@ Docs: https://classifier.dev Agent skill: npx skills add https://classifier.de
function parseArgs(argv) {
const o = { labels: null, text: null, multi: false, max: null, smart: false, instructions: "",
review: null, count: false, json: false, quiet: false, field: "text", id: null,
endpoint: ENDPOINT, model: "jev", processing: "fast", apiKey: process.env.CLASSIFY_API_KEY || process.env.CLASSIFIER_API_KEY || "", help: false, version: false };
endpoint: ENDPOINT, model: "jev", processing: "fast", document: null, apiKey: process.env.CLASSIFY_API_KEY || process.env.CLASSIFIER_API_KEY || "", help: false, version: false };
const positional = [];
for (let i = 0; i < argv.length; i++) {
const a = argv[i];
Expand All @@ -109,6 +112,7 @@ function parseArgs(argv) {
else if (a === "--id") o.id = next();
else if (a === "--endpoint") o.endpoint = next();
else if (a === "--api-key") o.apiKey = next();
else if (a === "--document") o.document = next();
else if (a === "--model") o.model = next();
else if (a === "--processing") o.processing = next();
else if (a.startsWith("--") && a.includes("=")) { argv.splice(i + 1, 0, a.slice(a.indexOf("=") + 1)); argv[i] = a.slice(0, a.indexOf("=")); i--; }
Expand Down Expand Up @@ -205,6 +209,7 @@ async function post(o, inputs) {
continue;
}
const payload = await res.json().catch(() => ({}));
if (res.status === 202) return checkResults(o, await waitDocument(o, payload), inputs.length);
if (res.ok) return checkResults(o, payload, inputs.length);
last = payload.error || `HTTP ${res.status}`;
// A daily quota cannot recover during a normal CLI run. Keep the API's
Expand All @@ -224,6 +229,62 @@ async function post(o, inputs) {
throw new Error(last);
}

async function waitDocument(o, accepted) {
const url = new URL(accepted.status_url, o.endpoint);
if (url.origin !== new URL(o.endpoint).origin) throw new Error("The API returned an invalid job URL.");
if (!process.env.CLASSIFY_NO_PROGRESS) process.stderr.write(`classify: document accepted; waiting for ${accepted.id}\n`);
const deadline = Date.now() + (Number(process.env.CLASSIFY_JOB_TIMEOUT) || 86400) * 1000;
let failures = 0;
while (Date.now() < deadline) {
let response;
try {
response = await fetch(url, { headers: { authorization: `Bearer ${o.apiKey}` }, signal: AbortSignal.timeout(TIMEOUT_MS) });
} catch {
if (++failures >= 5) throw new Error(`Polling interrupted. Your job continues at ${url}`);
await sleep(2000); continue;
}
const job = await response.json().catch(() => ({}));
if (response.ok && job.status === "finished") return job.result;
if (job.status === "failed") throw new Error(job.error?.message || "Document classification failed; the reservation was refunded.");
if (!response.ok && response.status !== 429 && response.status < 500)
throw new Error(job.error || `Unable to read job: HTTP ${response.status}`);
if (!response.ok && ++failures >= 5) throw new Error(`Polling interrupted. Your job continues at ${url}`);
if (response.ok) failures = 0;
await sleep(Math.min(30, Math.max(1, Number(response.headers.get("retry-after")) || 2)) * 1000);
}
throw new Error(`Stopped waiting. Check your job at ${url}`);
}

async function classifyDocument(o) {
if (!o.apiKey) throw new Error("--document requires a funded workspace API key (CLASSIFY_API_KEY).");
if (o.text !== null || o.smart || o.model !== "jev" || o.max !== null)
throw new Error("--document accepts one file with Jev fast, optional --multi and --instructions.");
if (!statSync(o.document).isFile()) throw new Error("--document must name a UTF-8 text file.");
const url = new URL(o.endpoint);
for (const label of o.labels) url.searchParams.append("label", label);
if (o.instructions) url.searchParams.set("instructions", o.instructions);
if (o.multi) url.searchParams.set("multi", "true");
const id = randomUUID();
for (let attempt = 0; attempt < ATTEMPTS; attempt++) {
const stream = createReadStream(o.document);
try {
const response = await fetch(url, { method: "POST", duplex: "half", body: stream,
headers: { authorization: `Bearer ${o.apiKey}`, "content-type": "text/plain; charset=utf-8",
"idempotency-key": id, prefer: "respond-async" }, signal: AbortSignal.timeout(TIMEOUT_MS) });
const payload = await response.json().catch(() => ({}));
if (response.status === 202) return await waitDocument(o, payload);
if (response.status < 500 && response.status !== 409 && response.status !== 429)
throw new Error(payload.error || `Document upload failed: HTTP ${response.status}`);
if (attempt + 1 === ATTEMPTS) throw new Error(payload.error || "Document upload is unavailable.");
} catch (error) {
if (error instanceof TypeError || error.name === "TimeoutError") {
if (attempt + 1 === ATTEMPTS) throw new Error(`Upload interrupted. Retry with Idempotency-Key ${id}, or check /v1/long-context/jobs/${id}/status.`);
} else throw error;
} finally { stream.destroy(); }
await sleep(1000 * 2 ** attempt);
}
}

const ATTEMPTS = 5;

let hinted = false;
Expand Down Expand Up @@ -439,6 +500,16 @@ export async function main(argv) {
if (o.labels.length > 100) fail("at most 100 labels");
try { batchSize(o); } catch (e) { fail(e.message); }

if (o.document) {
const payload = await classifyDocument(o);
const results = checkResults(o, payload, 1);
if (o.count) process.stdout.write(formatCount(o, results).join("\n") + "\n");
else if (kept(o, results[0])) process.stdout.write((o.json
? JSON.stringify({ file: o.document, ...results[0], usage: payload.usage, pricing: payload.pricing })
: labelOf(o, results[0])) + "\n");
return 0;
}

const hint = updateHint();

let items;
Expand Down
60 changes: 60 additions & 0 deletions e2e/document-admission.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,60 @@
import assert from "node:assert/strict";
import { SQL } from "bun";
import { mkdir, readdir, readFile, writeFile } from "node:fs/promises";
import { postgresDatabase, hashToken, type AppEnv } from "../src/server/db";
import { authorizeAndReserve } from "../src/server/usage";
import { refundTokenReservation } from "../src/server/token-ledger";

// Failure contract: different concurrent document IDs must not overdraw; the
// same admission ID must debit once even when recovery reads race. Refunds
// restore one hold. Real PostgreSQL connections, migrations and production SQL.
assert.ok(process.env.POSTGRES_TEST_URL, "Set POSTGRES_TEST_URL to a local test database.");
const sql = new SQL(process.env.POSTGRES_TEST_URL!, { max: 20 });
const schema = `document_admission_${crypto.randomUUID().replaceAll("-", "")}`;
const token = "classifier_agent_document_admission_fixture";
const env = { APP_ACCOUNTS_ENABLED: "true", APP_DB: postgresDatabase(async (queries, transaction) =>
sql.begin(transaction ? "ISOLATION LEVEL SERIALIZABLE" : "ISOLATION LEVEL READ COMMITTED", async connection => {
await connection.unsafe(`SET LOCAL search_path TO ${schema}`);
const results = [];
for (const query of queries) {
const rows = await connection.unsafe(query.sql, query.params);
results.push({ results: Array.from(rows), meta: { changes: rows.count } });
}
return results;
})) } as AppEnv;
const report: unknown[] = [];
try {
await sql.unsafe(`CREATE SCHEMA ${schema}`);
await sql.begin(async connection => {
await connection.unsafe(`SET LOCAL search_path TO ${schema}`);
for (const file of (await readdir("migrations/postgres")).filter(p => p.endsWith(".sql")).sort())
await connection.unsafe(await readFile(`migrations/postgres/${file}`, "utf8")).simple();
});
await env.APP_DB.prepare("INSERT INTO app_accounts(id,email,name,balance,paid_balance,reset_at,created_at,period_start) VALUES('fixture','fixture@example.com','Fixture',100,100,now()+interval '1 day',now(),now())").run();
await env.APP_DB.prepare("INSERT INTO app_agents(id,account_id,name,client,status,token_hash,prefix,created_at) VALUES('fixture','fixture','Fixture','API','connected',?,'fixture',now())")
.bind(await hashToken(token)).run();
const reserve = (id: string) => authorizeAndReserve(new Request("http://fixture/", {
headers: { authorization: `Bearer ${token}` },
}), env, 60, 1, { meteringMode: "tokens", reservationId: id });
for (const sameId of [false, true]) {
const shared = crypto.randomUUID();
const attempts = await Promise.allSettled(Array.from({ length: 20 }, () => reserve(sameId ? shared : crypto.randomUUID())));
const accepted = attempts.flatMap(a => a.status === "fulfilled" && a.value ? [a.value] : []);
assert.equal(new Set(accepted.map(a => a.id)).size, 1);
if (sameId) assert.equal(accepted.length, 20, JSON.stringify(attempts.filter(a => a.status === "rejected").map(a => String(a.reason))));
else assert.equal(accepted.length, 1);
const account = await env.APP_DB.prepare("SELECT balance::integer AS balance,paid_balance::integer AS paid FROM app_accounts WHERE id='fixture'").first();
assert.deepEqual(account, { balance: 40, paid: 40 });
const agent = await env.APP_DB.prepare("SELECT used::integer AS used FROM app_agents WHERE id='fixture'").first();
assert.deepEqual(agent, { used: 60 });
await Promise.all(Array.from({ length: 5 }, () => refundTokenReservation(env.APP_DB, accepted[0].id)));
assert.deepEqual(await env.APP_DB.prepare("SELECT balance::integer AS balance FROM app_accounts WHERE id='fixture'").first(), { balance: 100 });
report.push({ sameId, attempts: 20, accepted: accepted.length, distinctHolds: 1, account, agent, refundedBalance: 100 });
}
await mkdir("captures", { recursive: true });
await writeFile("captures/document-admission.json", JSON.stringify({ runtime: "PostgreSQL concurrent connections", results: report }, null, 2));
console.log("Document admission concurrency passed; captures/document-admission.json");
} finally {
await sql.unsafe(`DROP SCHEMA IF EXISTS ${schema} CASCADE`);
await sql.close();
}
Loading
Loading