Skip to content

Latest commit

 

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

systemANE

Typed decisions on the Apple Neural Engine. 1.4 ms each, with 94.6% of the graph executing on the ANE, from a 22.6M-parameter model with no fine-tuning.

On CLINC150 — 1,300 human-written utterances collected independently of this project, 300 in-scope plus 1,000 genuinely out-of-scope — that gets 93.3% accuracy, ECE 0.041, and 0.994 AUROC at spotting out-of-scope input.

It is a System-1 decision engine: typed decisions over a sentence encoder, paired with a second tier that only runs when the first is not sure.

tier 1   MiniLM-L6 (22.6M) on the ANE     1.4 ms    answers most inputs
tier 2   Apple Foundation Model (~3B)    1212 ms    answers what tier 1 escalates

Tier 1 is not smart. It is cheap enough to run on everything, and — once calibrated — honest enough to say when it should not decide. That is the whole idea; the accuracy lives in tier 2.

Proof of concept. Tier 2 is served by fm serve, the Apple Foundation Models CLI built into macOS 27 at /usr/bin/fm — there is nothing external to install for it. fmserve.py talks to it over stdlib HTTP only.

It runs on the Neural Engine

The ANE is easy to claim and hard to verify — a Core ML model that silently falls back to CPU still returns the right answer, just slower. Four things, each measured by a script in this repo.

It lands there. where.py asks Core ML's own compute planner for the per-op device assignment rather than inferring it from timing: 157 of 166 ops (94.6%) on the MLNeuralEngineComputeDevice. The nine on CPU are five fp32↔fp16 casts, the embedding gather — a table lookup is not an ANE operation — and three mask ops (greater_equal, add, select). Every matmul, softmax and layernorm is on the ANE, including 21 of the 22 adds.

fp16 costs almost nothing. The usual objection to the ANE is precision, so drift.py runs the fp32 torch path against the fp16 Core ML path end to end, anchors and queries both, at two label counts:

identical decision broke a correct answer
CLINC, K=10, 1,300 rows 300/300 in-scope 0
MASSIVE, K=60, 2,974 rows 2,971/2,974 1

The one loss is "please turn lights off", iot_hue_lightoff → iot_hue_lighton, reported at p=0.489 on both paths — a coin flip in fp32 too. Every CLINC disagreement is an out-of-scope item, which has no correct label to lose.

What 1.4 ms buys. Running on everything, rather than on what you can afford. And because the cost is per input rather than per decision — one encode, then a dot product per question — three typed questions about one ticket cost 1.13 ms, of which the three decisions themselves are 25 µs:

three questions, three separate calls          3.46 ms
three questions, one State                     1.13 ms
the three questions alone, vector precomputed  0.025 ms

Under load, serve.py does 708 req/s at 16 clients on a laptop, p50 1.21 ms — against decider-2b's published 431 req/s at 64 clients on an NVIDIA GH200. Different work per decision and not a benchmark either side agreed to, but the shape of the difference is the point.

That 1.4 ms is the saturated figure. Like every Core ML backend, the encoder goes cold between calls: a 1 ms gap costs 44%, and an application deciding something every few hundred milliseconds should expect ~3 ms. The ANE pays the smallest penalty of the three, so its lead widens when calls are sporadic — 2.8× over CPU and 2.7× over GPU at a 200 ms gap. warmup.py has the table.

What is not measured: energy. laya-coreml reports 0.154 J per decision on the ANE against 0.429 J on a compiled GPU path — 2.78× the energy for only 1.39× the latency — which suggests the ANE's real argument is power, not speed, and that this README makes the weaker half of the case. Attributing joules to a 1.4 ms operation honestly is harder than it sounds, so it is recorded as open in NEXT_STEPS.md rather than guessed at.

Watching it work

▶ demos/bluesky/ANE-jetstreaming.mp4 — 24 seconds of every link posted to Bluesky being sorted as it arrives.

  systemANE · the Bluesky firehose, sorted on the Neural Engine
  ─────────────────────────────────────────────────────────────────────
  00:22   posts 1078  48.9/s   cards 178  8.1/s   shown 78  refused 100  56%
  encode 3.15 ms (warm 1.44)  capacity 695 posts/s warm  using 2.5% of one thread
  ─────────────────────────────────────────────────────────────────────
     news 0.53  new  [penncapital-star.com] Trump likes data centers. Congress...
    music 0.53 rule  [youtube.com        ] All Star - Smash Mouth (SKA PUNK Cover)
     news 0.34  new  [propublica.org     ] The FBI Anti-Corruption Squad Was Circling...

rule means the link's domain is in a lookup table. new means the encoder placed it from the card's title alone — which is 81.6% of the stream, since a 25-minute capture turned up 2,997 distinct domains and 66% of them appeared exactly once. Anchors are built from domains that do carry a rule, so the labels are free and nobody on this project wrote them.

Tested by splitting on domain, never on post — centroids from the Guardian, Globo, AP and Spiegel, scored on the NYT, BBC, Reuters and Le Monde: 0.821 accuracy, 0.704 macro-F1 over 823 held-out cards from sites never seen in training, against a 0.599 majority baseline.

The errors are on screen too, and they are the ones that number predicts — a book listing, "Continental Drift by Mai-Linh Hong", goes to music, because "title by person" is how a song is credited. Roughly one row in five is wrong.

It is also the case this engine was built for rather than a flattering one: in ~/Downloads a two-line host rule placed 99% of files and the model was decoration. A long tail is what gives an encoder a job, and this is one. demos/bluesky/README.md has the harvest, the rules and what went wrong.

Quick start

python3.12 -m venv .venv && .venv/bin/pip install -r requirements.txt
.venv/bin/python build.py       # convert MiniLM -> encoder.mlpackage (once)
.venv/bin/python where.py       # confirm it lands on the ANE

fm serve &                      # tier 2, in another shell (port 1976)
.venv/bin/python cascade.py     # the two-tier demo

cascade.py --no-tier2 runs tier 1 alone, and needs no Foundation Model at all. Python 3.12 — coremltools has no 3.14 wheels. macOS 27 on Apple silicon: the ANE for tier 1, the built-in Foundation Model for tier 2.

What it does

cascade.py, verbatim, with fm serve running:

ticket                                    tier1        p  marg   sim  why           tier2     ms
------------------------------------------------------------------------------------------------
I was charged twice for last month        billing   0.95  0.89  0.41
can't log in after the update, it says... auth      1.00  1.00  0.59
the app crashes every time I open the...  bug       1.00  1.00  0.54
would be great if you supported dark mode feature   1.00  0.99  0.32  structural    feature  2477
where is my package, it was supposed t... shipping  1.00  1.00  0.54
my subscription renewed but I cancelle... billing   0.99  0.98  0.31
2FA codes never arrive on my phone        auth      1.00  1.00  0.39
the checkout page throws a 500 when I ... billing   0.94  0.89  0.28
it broke                                  bug       0.75  0.52  0.27  ambiguous     bug       575
I need help with my account               auth      0.99  0.97  0.44
this isn't working                        auth      0.76  0.65  0.22  ambiguous     bug       585
something is wrong                        bug       0.96  0.94  0.31
can someone call me                       feature   0.65  0.47  0.14  out-of-scope  (none)
what are your office hours                bug       0.44  0.13  0.07  out-of-scope  (none)
I'd like to speak to a manager            feature   0.91  0.88  0.17  out-of-scope  (none)
happy holidays everyone                   shipping  0.46  0.21  0.08  out-of-scope  (none)

tier 1 answered 9/16, refused 4/16 (out of scope), escalated 3/16
tier 1 latency: mean 1.42 ms
tier 2 latency: mean 1212 ms (856x tier 1)
cascade total 3660 ms vs 19399 ms if every ticket went to the LLM (5.3x saving)

Three things in that table are the point:

  • The clear tickets are routed in ~1.4 ms each. No LLM.
  • The four out-of-scope inputs are refused, not guessed at. They are not escalated either: tier 2 is bound to the same five labels, so it would also be forced to pick a wrong one. Refusing is cheaper and more correct.
  • Escalation fixed a wrong answer. "this isn't working" was heading for auth at p=0.76; the margin was under threshold, tier 2 was asked, and it came back bug.

Thresholds come from calibration_tickets.json, fitted by calibrate_tickets.py on tickets.VALIDATION — 72 tickets disjoint from the sixteen above, so this is a held-out test and not a restatement of the fit. Those 72 are hand-written by the author, though, which is why the headline number comes from CLINC150 instead.

Typed decisions, not generated text

enc   = Encoder()
route = Choice(enc, ROUTES)           # option descriptions embedded once
d     = route("I was charged twice")  # ~1.4 ms
# Decision(label='billing', prob=0.95, margin=0.89, escalate=False)

The output is a distribution over the schema you passed in, so an off-schema label is not unlikely — it is unrepresentable.

Several questions about one input cost one encode and nothing else:

s = State(enc, "I was charged twice for last month")
s.ask(department=route, urgency=anger, churn_risk=churn)
# {'department': Decision(label='billing', prob=0.95, ...),
#  'urgency':    {'score': 1.74, 'confidence': 0.72, ...},
#  'churn_risk': Decision(label=True, prob=0.81, ...)}
ms per ticket
three questions, three separate calls 3.46
three questions, one State 1.13
the three questions alone, vector already computed 0.025

The encode is 1.10 ms and the three typed decisions on top of it are 25 µs together — a fourth question costs about 8 µs. Other engines engineer this: kev caches a state representation across questions, Laya batches ten to get the per-question cost down. In an embedding architecture it is not an optimisation, it is the shape of the thing.

Two signals, because there are two ways to not know

signal question it answers
ambiguous between labels margin (post-softmax) which of these?
outside the schema sim (raw max cosine) any of these?

The softmax normalises magnitude away, so margin alone cannot see the second case: "I'd like to speak to a manager" scores a higher margin than the genuinely ambiguous "it broke". On CLINC, sim separates out-of-scope at AUROC 0.994 against 0.945 for margin. Keeping them apart is also what lets the engine refuse rather than escalate, which is the right action when no label is correct.

This is where the embedding approach earns its keep. Most decision models score options through the language-model head — at option letters, at a [MASK] per option, at a single-token label. That is better at picking which one. But it softmaxes over the option set, and a softmax has no way to say none of these: whatever the model thought of the input in absolute terms is normalised away before anything downstream sees it. Answering the second question then needs a separately trained head.

Here it falls out of the geometry. sim is the raw cosine, before any normalisation — a scalar that is simply low when nothing in the schema resembles the input, at no cost and with nothing trained. Of the five engines in RELATED_WORK.md, none publishes an out-of-scope number.

Reproducing the headline

.venv/bin/python calibrate.py --k 16 --no-description   # fit on validation
.venv/bin/python evaluate.py  --k 16                    # score on test

Against the zero-shot baseline, which embeds one handwritten sentence per class:

description only 16 examples
in-scope accuracy 0.880 0.933
errors 36/300 20/300
OOS AUROC 0.989 0.994
escalation needed for ≤5% error 17.5% 1.5%

A 44% reduction in errors for 160 labels and no training — and it mostly dissolves the escalation problem, so tier 2 is called for a handful of inputs rather than a sixth of traffic.

As a server

.venv/bin/python serve.py --preload-tickets    # 127.0.0.1:1977

POST /v1/systemone is the System One API — the endpoint kev serves from a 0.8–9B model. One state, a dict of typed questions, all of them answered from a single embedding:

curl -s localhost:1977/v1/systemone -H 'Content-Type: application/json' -d '{
  "state": "I have waited three weeks and nobody has replied. I want to cancel.",
  "questions": {
    "department": {"type": "choice", "criteria": {
        "billing": "payments, charges and refunds",
        "support": "a problem with the product"}},
    "urgency":    {"type": "score", "criteria": [
        "calm and polite", "mildly annoyed", "clearly frustrated",
        "furious and threatening to leave"]},
    "churn_risk": {"type": "noul", "criteria": {
        "true": "this customer is about to cancel",
        "false": "this customer is satisfied and staying"}}}}'
req/s server-side p50 p95
1 client, one question 225 2.25 ms 4.07 ms
16 clients, one question 708 1.21 ms 2.56 ms
16 clients, three questions 660 (1,980 decisions/s) 1.28 ms —

Two things are specific to serving an embedding engine:

Schemas are compiled, and compiling costs encodes. An option is a vector here, not a string in a prompt, so a question whose options the server has not seen costs one encode per option first — 33 ms for the three questions above, 1.5 ms every time after. Compiled schemas are cached by a hash of their criteria.

Probabilities are uncalibrated unless you have fitted them. The engine's most repeated finding is that thresholds do not transfer between schemas and the failure is silent, so a server that accepted any schema and returned confident-looking numbers would ship that bug to everyone. Answers carry "calibrated": false and "escalate": null — never false — until a calibration is registered for that exact question:

curl -s localhost:1977/v1/schemas -H 'Content-Type: application/json' \
  -d '{"question": {...}, "calibration": {"temp":0.036,"min_sim":0.194,"min_margin":0.65}}'

The registration binds to a hash of the criteria. Change one description and it is a different set of vectors, so the calibration stops applying and says so.

A second dataset, for a number that compares

CLINC is the set this project argues with, but nobody else reports on it. The same engine on MASSIVE — 60 intents, mteb/amazon_massive_intent, which other decision models do publish on:

.venv/bin/python calibrate.py --dataset massive --k 16 --no-description \
    --out calibration_massive.json
.venv/bin/python evaluate.py  --dataset massive --k 16 \
    --calibration calibration_massive.json

0.710 accuracy on 2,974 test items, ECE 0.032, from the same 22.6M-parameter encoder with nothing trained. Laya reports 0.783 on MASSIVE intent from 421M parameters and an RL stage — 7.3 points ahead at 19× the size, on its own harness and split rather than head-to-head.

MASSIVE ships no out-of-scope class, so refusal cannot be measured there at all; calibrate.py warns and records oos_validated: false rather than writing a threshold that looks fitted. The 0.994 above is CLINC's.

One property that comes from the architecture

Identical input, identical answer — bit-for-bit. No sampling anywhere in the path, so a restarted server agrees with the one it replaced, and option order cannot matter because a dot product has no order. TypeSafe report Jev agreeing with itself 90.8% of the time on a repeated question, and recommend refusing to act below p=0.60 to reach 99.2% — which costs 25.8% of traffic to human review. stability.py checks this one here.

That is not a claim about being better at the task — see the Limits below, one of which is considerably worse than a prompted model's.

Limits

Irrelevant text is mixed into the answer, not ignored. Mean pooling averages every token in the window, so appending one paragraph of unrelated boilerplate to a short input changes 37% of decisions (stability.py). This is Jev's "large irrelevant state" failure mode, and pooling suffers from it more directly than attention does. Filter state in code before sending it.

Calibration does not transfer between schemas, and the failure is silent. calibration.json is fitted for the CLINC banking routes; applied to the support tickets its min_sim of 0.3085 refuses genuinely in-scope tickets at 0.27–0.28. Every schema needs its own fit. Choice() takes calibration= for exactly this reason.

A confident wrong answer is invisible to a confidence threshold. On the CLINC test set, 7 of 300 in-scope items (2.3%) are answered wrongly at p ≥ 0.90 — about a third of all errors — and min_sim and min_margin flag none of them. That is not a threshold needing tuning. All seven are adjacent-intent confusions (transactions vs. credit_limit, report_lost_card vs. report_fraud), so the embedding sits close to the wrong anchor because it genuinely should, and every signal the engine computes says the answer is fine.

This is the floor on what a cascade can do: escalation fixes uncertainty, and these decisions are not uncertain. FINDINGS.md has the seven and the temperature sweep behind them.

Layout

minilm.py             MiniLM-L6 forward pass, written out so the graph is ours
build.py              convert to encoder.mlpackage (verifies against HuggingFace)
build_encoder.py      compile any BERT-architecture encoder (--model, --pooling)
where.py              per-op device assignment from Core ML's compute planner
drift.py              fp16 ANE vs fp32 torch, measured on decisions
warmup.py             what an idle gap costs, per compute unit
serve.py              the System One API over HTTP (stdlib only)
stability.py          determinism, and what irrelevant text costs
system1.py            Encoder + Choice / Boolean / Score primitives
fmserve.py            minimal stdlib client for `fm serve` (tier 2)
cascade.py            the two-tier demo

tickets.py            demo schema: 5 routes, fixtures, and a validation set
calibrate_tickets.py  fit the demo schema's thresholds

evalset.py            CLINC150 loading and splits
massive.py            MASSIVE loading (60 intents, no out-of-scope class)
anchors.py            class representations: descriptions vs example centroids
calibrate.py          fit temp / min_sim / min_margin on validation
evaluate.py           test-set accuracy, ECE, OOS AUROC
sweep_anchors.py      the k sweep behind the anchors table

demos/bluesky/        every link on Bluesky, sorted live on the ANE

FINDINGS.md           what was measured, and which guesses were wrong
RELATED_WORK.md       six other decision engines, and what they change here
NEXT_STEPS.md         what is worth doing next, and what is deliberately not

Model packages and tokenizers are build artifacts and gitignored; run build.py or build_encoder.py to regenerate. The fitted calibration*.json files are results, not artifacts, and are tracked.

About

Using Apple's Neural Engine as a fast and free local "System One" Decision Engine (macOS 27) - "Honey, we have Jev at home"

Resources

Stars

14 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages