Typed decisions on the Apple Neural Engine. 1.4 ms each, with 94.6% of the graph executing on the ANE, from a 22.6M-parameter model with no fine-tuning.
On CLINC150 — 1,300 human-written utterances collected independently of this project, 300 in-scope plus 1,000 genuinely out-of-scope — that gets 93.3% accuracy, ECE 0.041, and 0.994 AUROC at spotting out-of-scope input.
It is a System-1 decision engine: typed decisions over a sentence encoder, paired with a second tier that only runs when the first is not sure.
tier 1 MiniLM-L6 (22.6M) on the ANE 1.4 ms answers most inputs
tier 2 Apple Foundation Model (~3B) 1212 ms answers what tier 1 escalates
Tier 1 is not smart. It is cheap enough to run on everything, and — once calibrated — honest enough to say when it should not decide. That is the whole idea; the accuracy lives in tier 2.
Proof of concept. Tier 2 is served by fm serve, the Apple Foundation Models
CLI built into macOS 27 at /usr/bin/fm — there is nothing external to install
for it. fmserve.py talks to it over stdlib HTTP only.
The ANE is easy to claim and hard to verify — a Core ML model that silently falls back to CPU still returns the right answer, just slower. Four things, each measured by a script in this repo.
It lands there. where.py asks Core ML's own compute planner for the
per-op device assignment rather than inferring it from timing: 157 of 166 ops
(94.6%) on the MLNeuralEngineComputeDevice. The nine on CPU are five
fp32↔fp16 casts, the embedding gather — a table lookup is not an ANE
operation — and three mask ops (greater_equal, add, select). Every
matmul, softmax and layernorm is on the ANE, including 21 of the 22 adds.
fp16 costs almost nothing. The usual objection to the ANE is precision, so
drift.py runs the fp32 torch path against the fp16 Core ML path end to end,
anchors and queries both, at two label counts:
| identical decision | broke a correct answer | |
|---|---|---|
| CLINC, K=10, 1,300 rows | 300/300 in-scope | 0 |
| MASSIVE, K=60, 2,974 rows | 2,971/2,974 | 1 |
The one loss is "please turn lights off", iot_hue_lightoff →
iot_hue_lighton, reported at p=0.489 on both paths — a coin flip in fp32
too. Every CLINC disagreement is an out-of-scope item, which has no correct
label to lose.
What 1.4 ms buys. Running on everything, rather than on what you can afford. And because the cost is per input rather than per decision — one encode, then a dot product per question — three typed questions about one ticket cost 1.13 ms, of which the three decisions themselves are 25 µs:
three questions, three separate calls 3.46 ms
three questions, one State 1.13 ms
the three questions alone, vector precomputed 0.025 ms
Under load, serve.py does 708 req/s at 16 clients on a laptop, p50
1.21 ms — against decider-2b's published 431 req/s at 64 clients on an NVIDIA
GH200. Different work per decision and not a benchmark either side agreed to,
but the shape of the difference is the point.
That 1.4 ms is the saturated figure. Like every Core ML backend, the
encoder goes cold between calls: a 1 ms gap costs 44%, and an application
deciding something every few hundred milliseconds should expect ~3 ms. The ANE
pays the smallest penalty of the three, so its lead widens when calls are
sporadic — 2.8× over CPU and 2.7× over GPU at a 200 ms gap. warmup.py has the
table.
What is not measured: energy. laya-coreml reports 0.154 J per decision on
the ANE against 0.429 J on a compiled GPU path — 2.78× the energy for only
1.39× the latency — which suggests the ANE's real argument is power, not
speed, and that this README makes the weaker half of the case. Attributing
joules to a 1.4 ms operation honestly is harder than it sounds, so it is
recorded as open in NEXT_STEPS.md rather than guessed at.
▶ demos/bluesky/ANE-jetstreaming.mp4 — 24 seconds of every link posted to Bluesky being sorted as it arrives.
systemANE · the Bluesky firehose, sorted on the Neural Engine
─────────────────────────────────────────────────────────────────────
00:22 posts 1078 48.9/s cards 178 8.1/s shown 78 refused 100 56%
encode 3.15 ms (warm 1.44) capacity 695 posts/s warm using 2.5% of one thread
─────────────────────────────────────────────────────────────────────
news 0.53 new [penncapital-star.com] Trump likes data centers. Congress...
music 0.53 rule [youtube.com ] All Star - Smash Mouth (SKA PUNK Cover)
news 0.34 new [propublica.org ] The FBI Anti-Corruption Squad Was Circling...
rule means the link's domain is in a lookup table. new means the encoder
placed it from the card's title alone — which is 81.6% of the stream, since
a 25-minute capture turned up 2,997 distinct domains and 66% of them appeared
exactly once. Anchors are built from domains that do carry a rule, so the
labels are free and nobody on this project wrote them.
Tested by splitting on domain, never on post — centroids from the Guardian, Globo, AP and Spiegel, scored on the NYT, BBC, Reuters and Le Monde: 0.821 accuracy, 0.704 macro-F1 over 823 held-out cards from sites never seen in training, against a 0.599 majority baseline.
The errors are on screen too, and they are the ones that number predicts — a
book listing, "Continental Drift by Mai-Linh Hong", goes to music, because
"title by person" is how a song is credited. Roughly one row in five is wrong.
It is also the case this engine was built for rather than a flattering one: in
~/Downloads a two-line host rule placed 99% of files and the model was
decoration. A long tail is what gives an encoder a job, and this is one.
demos/bluesky/README.md has the harvest, the rules and what went wrong.
python3.12 -m venv .venv && .venv/bin/pip install -r requirements.txt
.venv/bin/python build.py # convert MiniLM -> encoder.mlpackage (once)
.venv/bin/python where.py # confirm it lands on the ANE
fm serve & # tier 2, in another shell (port 1976)
.venv/bin/python cascade.py # the two-tier democascade.py --no-tier2 runs tier 1 alone, and needs no Foundation Model at
all. Python 3.12 — coremltools has no 3.14 wheels. macOS 27 on Apple silicon:
the ANE for tier 1, the built-in Foundation Model for tier 2.
cascade.py, verbatim, with fm serve running:
ticket tier1 p marg sim why tier2 ms
------------------------------------------------------------------------------------------------
I was charged twice for last month billing 0.95 0.89 0.41
can't log in after the update, it says... auth 1.00 1.00 0.59
the app crashes every time I open the... bug 1.00 1.00 0.54
would be great if you supported dark mode feature 1.00 0.99 0.32 structural feature 2477
where is my package, it was supposed t... shipping 1.00 1.00 0.54
my subscription renewed but I cancelle... billing 0.99 0.98 0.31
2FA codes never arrive on my phone auth 1.00 1.00 0.39
the checkout page throws a 500 when I ... billing 0.94 0.89 0.28
it broke bug 0.75 0.52 0.27 ambiguous bug 575
I need help with my account auth 0.99 0.97 0.44
this isn't working auth 0.76 0.65 0.22 ambiguous bug 585
something is wrong bug 0.96 0.94 0.31
can someone call me feature 0.65 0.47 0.14 out-of-scope (none)
what are your office hours bug 0.44 0.13 0.07 out-of-scope (none)
I'd like to speak to a manager feature 0.91 0.88 0.17 out-of-scope (none)
happy holidays everyone shipping 0.46 0.21 0.08 out-of-scope (none)
tier 1 answered 9/16, refused 4/16 (out of scope), escalated 3/16
tier 1 latency: mean 1.42 ms
tier 2 latency: mean 1212 ms (856x tier 1)
cascade total 3660 ms vs 19399 ms if every ticket went to the LLM (5.3x saving)
Three things in that table are the point:
- The clear tickets are routed in ~1.4 ms each. No LLM.
- The four out-of-scope inputs are refused, not guessed at. They are not escalated either: tier 2 is bound to the same five labels, so it would also be forced to pick a wrong one. Refusing is cheaper and more correct.
- Escalation fixed a wrong answer. "this isn't working" was heading for
authat p=0.76; the margin was under threshold, tier 2 was asked, and it came backbug.
Thresholds come from calibration_tickets.json, fitted by
calibrate_tickets.py on tickets.VALIDATION — 72 tickets disjoint from the
sixteen above, so this is a held-out test and not a restatement of the fit.
Those 72 are hand-written by the author, though, which is why the headline
number comes from CLINC150 instead.
enc = Encoder()
route = Choice(enc, ROUTES) # option descriptions embedded once
d = route("I was charged twice") # ~1.4 ms
# Decision(label='billing', prob=0.95, margin=0.89, escalate=False)The output is a distribution over the schema you passed in, so an off-schema label is not unlikely — it is unrepresentable.
Several questions about one input cost one encode and nothing else:
s = State(enc, "I was charged twice for last month")
s.ask(department=route, urgency=anger, churn_risk=churn)
# {'department': Decision(label='billing', prob=0.95, ...),
# 'urgency': {'score': 1.74, 'confidence': 0.72, ...},
# 'churn_risk': Decision(label=True, prob=0.81, ...)}| ms per ticket | |
|---|---|
| three questions, three separate calls | 3.46 |
three questions, one State |
1.13 |
| the three questions alone, vector already computed | 0.025 |
The encode is 1.10 ms and the three typed decisions on top of it are 25 µs together — a fourth question costs about 8 µs. Other engines engineer this: kev caches a state representation across questions, Laya batches ten to get the per-question cost down. In an embedding architecture it is not an optimisation, it is the shape of the thing.
| signal | question it answers | |
|---|---|---|
| ambiguous between labels | margin (post-softmax) |
which of these? |
| outside the schema | sim (raw max cosine) |
any of these? |
The softmax normalises magnitude away, so margin alone cannot see the second
case: "I'd like to speak to a manager" scores a higher margin than the
genuinely ambiguous "it broke". On CLINC, sim separates out-of-scope at
AUROC 0.994 against 0.945 for margin. Keeping them apart is also what lets the
engine refuse rather than escalate, which is the right action when no label
is correct.
This is where the embedding approach earns its keep. Most decision models
score options through the language-model head — at option letters, at a
[MASK] per option, at a single-token label. That is better at picking which
one. But it softmaxes over the option set, and a softmax has no way to say
none of these: whatever the model thought of the input in absolute terms is
normalised away before anything downstream sees it. Answering the second
question then needs a separately trained head.
Here it falls out of the geometry. sim is the raw cosine, before any
normalisation — a scalar that is simply low when nothing in the schema
resembles the input, at no cost and with nothing trained. Of the five engines
in RELATED_WORK.md, none publishes an out-of-scope number.
.venv/bin/python calibrate.py --k 16 --no-description # fit on validation
.venv/bin/python evaluate.py --k 16 # score on testAgainst the zero-shot baseline, which embeds one handwritten sentence per class:
| description only | 16 examples | |
|---|---|---|
| in-scope accuracy | 0.880 | 0.933 |
| errors | 36/300 | 20/300 |
| OOS AUROC | 0.989 | 0.994 |
| escalation needed for ≤5% error | 17.5% | 1.5% |
A 44% reduction in errors for 160 labels and no training — and it mostly dissolves the escalation problem, so tier 2 is called for a handful of inputs rather than a sixth of traffic.
.venv/bin/python serve.py --preload-tickets # 127.0.0.1:1977POST /v1/systemone is the System One API — the endpoint kev serves from a
0.8–9B model. One state, a dict of typed questions, all of them answered from a
single embedding:
curl -s localhost:1977/v1/systemone -H 'Content-Type: application/json' -d '{
"state": "I have waited three weeks and nobody has replied. I want to cancel.",
"questions": {
"department": {"type": "choice", "criteria": {
"billing": "payments, charges and refunds",
"support": "a problem with the product"}},
"urgency": {"type": "score", "criteria": [
"calm and polite", "mildly annoyed", "clearly frustrated",
"furious and threatening to leave"]},
"churn_risk": {"type": "noul", "criteria": {
"true": "this customer is about to cancel",
"false": "this customer is satisfied and staying"}}}}'| req/s | server-side p50 | p95 | |
|---|---|---|---|
| 1 client, one question | 225 | 2.25 ms | 4.07 ms |
| 16 clients, one question | 708 | 1.21 ms | 2.56 ms |
| 16 clients, three questions | 660 (1,980 decisions/s) | 1.28 ms | — |
Two things are specific to serving an embedding engine:
Schemas are compiled, and compiling costs encodes. An option is a vector here, not a string in a prompt, so a question whose options the server has not seen costs one encode per option first — 33 ms for the three questions above, 1.5 ms every time after. Compiled schemas are cached by a hash of their criteria.
Probabilities are uncalibrated unless you have fitted them. The engine's
most repeated finding is that thresholds do not transfer between schemas and
the failure is silent, so a server that accepted any schema and returned
confident-looking numbers would ship that bug to everyone. Answers carry
"calibrated": false and "escalate": null — never false — until a
calibration is registered for that exact question:
curl -s localhost:1977/v1/schemas -H 'Content-Type: application/json' \
-d '{"question": {...}, "calibration": {"temp":0.036,"min_sim":0.194,"min_margin":0.65}}'The registration binds to a hash of the criteria. Change one description and it is a different set of vectors, so the calibration stops applying and says so.
CLINC is the set this project argues with, but nobody else reports on it. The
same engine on MASSIVE — 60 intents, mteb/amazon_massive_intent, which other
decision models do publish on:
.venv/bin/python calibrate.py --dataset massive --k 16 --no-description \
--out calibration_massive.json
.venv/bin/python evaluate.py --dataset massive --k 16 \
--calibration calibration_massive.json0.710 accuracy on 2,974 test items, ECE 0.032, from the same 22.6M-parameter encoder with nothing trained. Laya reports 0.783 on MASSIVE intent from 421M parameters and an RL stage — 7.3 points ahead at 19× the size, on its own harness and split rather than head-to-head.
MASSIVE ships no out-of-scope class, so refusal cannot be measured there at all;
calibrate.py warns and records oos_validated: false rather than writing a
threshold that looks fitted. The 0.994 above is CLINC's.
Identical input, identical answer — bit-for-bit. No sampling anywhere in the
path, so a restarted server agrees with the one it replaced, and option order
cannot matter because a dot product has no order. TypeSafe report Jev agreeing
with itself 90.8% of the time on a repeated question, and recommend refusing
to act below p=0.60 to reach 99.2% — which costs 25.8% of traffic to human
review. stability.py checks this one here.
That is not a claim about being better at the task — see the Limits below, one of which is considerably worse than a prompted model's.
Irrelevant text is mixed into the answer, not ignored. Mean pooling averages
every token in the window, so appending one paragraph of unrelated boilerplate
to a short input changes 37% of decisions (stability.py). This is Jev's
"large irrelevant state" failure mode, and pooling suffers from it more directly
than attention does. Filter state in code before sending it.
Calibration does not transfer between schemas, and the failure is silent.
calibration.json is fitted for the CLINC banking routes; applied to the
support tickets its min_sim of 0.3085 refuses genuinely in-scope tickets at
0.27–0.28. Every schema needs its own fit. Choice() takes calibration= for
exactly this reason.
A confident wrong answer is invisible to a confidence threshold. On the
CLINC test set, 7 of 300 in-scope items (2.3%) are answered wrongly at p ≥ 0.90
— about a third of all errors — and min_sim and min_margin flag none of
them. That is not a threshold needing tuning. All seven are adjacent-intent
confusions (transactions vs. credit_limit, report_lost_card vs.
report_fraud), so the embedding sits close to the wrong anchor because it
genuinely should, and every signal the engine computes says the answer is fine.
This is the floor on what a cascade can do: escalation fixes uncertainty, and
these decisions are not uncertain. FINDINGS.md has the seven and the
temperature sweep behind them.
minilm.py MiniLM-L6 forward pass, written out so the graph is ours
build.py convert to encoder.mlpackage (verifies against HuggingFace)
build_encoder.py compile any BERT-architecture encoder (--model, --pooling)
where.py per-op device assignment from Core ML's compute planner
drift.py fp16 ANE vs fp32 torch, measured on decisions
warmup.py what an idle gap costs, per compute unit
serve.py the System One API over HTTP (stdlib only)
stability.py determinism, and what irrelevant text costs
system1.py Encoder + Choice / Boolean / Score primitives
fmserve.py minimal stdlib client for `fm serve` (tier 2)
cascade.py the two-tier demo
tickets.py demo schema: 5 routes, fixtures, and a validation set
calibrate_tickets.py fit the demo schema's thresholds
evalset.py CLINC150 loading and splits
massive.py MASSIVE loading (60 intents, no out-of-scope class)
anchors.py class representations: descriptions vs example centroids
calibrate.py fit temp / min_sim / min_margin on validation
evaluate.py test-set accuracy, ECE, OOS AUROC
sweep_anchors.py the k sweep behind the anchors table
demos/bluesky/ every link on Bluesky, sorted live on the ANE
FINDINGS.md what was measured, and which guesses were wrong
RELATED_WORK.md six other decision engines, and what they change here
NEXT_STEPS.md what is worth doing next, and what is deliberately not
Model packages and tokenizers are build artifacts and gitignored; run
build.py or build_encoder.py to regenerate. The fitted calibration*.json
files are results, not artifacts, and are tracked.