System 1 is "a machine for jumping to conclusions." — Daniel Kahneman. system-one is that machine, taught to jump carefully.
An open learning project inspired by TypeSafe's Jev. The goal is a small fine-tuned model, built on an existing backbone such as SmolLM, Qwen, or MiniCPM, with task-specific changes and efficient local inference. Apple MLX is the primary training and serving backend; a Rust runtime is an option if profiling justifies it.
Jev takes its name from the Jevons paradox:
make a resource cheap enough and consumption explodes. Its product category is
the "System One" model — Daniel Kahneman's term, from
Thinking, Fast and Slow,
for the fast, automatic, intuitive judgment faculty. This project takes the
category as its name: a System 1 built in the open — fine-tuned from open
models, calibrated on purpose, honest about what it knows. (It launched briefly
as system-one, Kahneman's nickname; the category name won.) The internal Python
package keeps the name jev for wire-compatibility clarity: the served API is
Jev's request/answer shape.
The model should answer narrow, request-defined questions about structured state using Choice / Score / Noul probabilities, without generating text. Code owns workflow logic; the model supplies semantic judgments and useful uncertainty.
Requires Apple Silicon (MLX) and uv.
Models on HuggingFace:
mpuig/system-one-qwen3-0.6b
(volume tier — the adapter also ships in this repo) ·
mpuig/system-one-minicpm5-2b-q8
(quality tier, 2.5 GB). Each carries its identity-bound temperature.json.
git clone https://github.com/mpuig/system-one && cd system-one
uv sync
# volume tier (0.6B). The adapter and temperature file ship in this repo;
# the Qwen base model (~1.2 GB) downloads from HuggingFace on first run.
uv run python scripts/serve.py --model Qwen/Qwen3-0.6B \
--adapter adapters/qwen3-0.6b-structured-v1-synthfiltered-rps \
--temperature release/temperature-qwen3-0.6b.json --port 8399
# ask three typed questions in one request:
curl -s -X POST http://127.0.0.1:8399/v1/systemone \
-H "Content-Type: application/json" -d '{
"state": "I have been charged twice for my flight to Berlin and nobody is answering the phone. I want my money back immediately.",
"questions": {
"refund_requested": {"type": "noul", "instructions": "Does the customer explicitly request a refund?"},
"request_type": {"type": "choice", "instructions": "What is the main request?",
"criteria": {"refund": "The customer wants money returned.",
"rebooking": "The customer wants a replacement flight.",
"information": "The customer wants information only."}},
"frustration": {"type": "score", "instructions": "How frustrated is the customer?",
"criteria": ["Calm; neutral tone.", "Annoyed; complains but cooperative.", "Angry; threats or ultimatums."]}}}'Actual output (rounded) — every question gets a typed answer with probabilities; nothing is generated or parsed:
{"model": "Qwen/Qwen3-0.6B@f9fabd98...",
"answers": {
"refund_requested": {"noul": 0.681},
"request_type": {"choice": "refund",
"probabilities": {"refund": 0.993, "rebooking": 0.007, "information": 0.0003},
"confidence": 0.989},
"frustration": {"score": 1.72,
"probabilities": {"0": 0.065, "1": 0.150, "2": 0.785},
"legend": {"0": "Calm; neutral tone.", "1": "Annoyed; complains but cooperative.", "2": "Angry; threats or ultimatums."},
"confidence": 0.579}},
"usage": {"input_tokens": 355, "output_tokens": 0}}Note the shape of calibrated honesty: decisive where the evidence is explicit (request_type at 0.99), hedged where it is thinner — these are probabilities to threshold and act on, not performative certainty. Criteria descriptions matter: the same questions with bare one-word options answer far less confidently.
The quality tier (MiniCPM5-2B fused to 8-bit — +6.8 accuracy points over the 0.6B on the fresh reserved test, CI [+4.5, +9.2]) is a 2.5 GB download:
hf download mpuig/system-one-minicpm5-2b-q8 --local-dir models/system-one-minicpm5-2b-q8
uv run python scripts/serve.py --model models/system-one-minicpm5-2b-q8 \
--temperature release/temperature-minicpm5-2b-q8.json --port 8399Details in release/MODEL_CARD_minicpm5-2b-q8.md. Per-workload calibration —
fitting temperatures to your traffic from ~100 labeled decisions — is one
call: POST /v1/calibrations (see docs/SERVING.md).
Eight ready-to-run use-case examples — support triage, LLM guardrails,
model routing, RAG reranking, citation checking, moderation, claims intake,
lead scoring — live in examples/, each with the actual output of
five systems — from the bare backbone to pinned Jev itself — including the
honest misses. The correctness ladder (24 authored questions; an illustration,
not a benchmark):
| Example | 2B base (no FT) | 0.6B FT | 2B FT | 2B FT q8 (served) | Jev 1.13.0 |
|---|---|---|---|---|---|
| citation-check | 0/2 | 0/2 | 2/2 | 2/2 | 1/2 |
| claims-intake | 3/4 | 3/4 | 4/4 | 4/4 | 4/4 |
| lead-scoring | 3/3 | 1/3 | 3/3 | 3/3 | 3/3 |
| llm-guardrail | 2/3 | 2/3 | 2/3 | 2/3 | 3/3 |
| model-routing | 2/3 | 3/3 | 2/3 | 2/3 | 3/3 |
| moderation | 2/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| rag-rerank | 2/2 | 2/2 | 2/2 | 2/2 | 2/2 |
| support-triage | 3/4 | 3/4 | 4/4 | 4/4 | 4/4 |
| Total | 17/24 | 17/24 | 22/24 | 22/24 | 23/24 |
Fine-tuning buys +5 questions on the same backbone; 8-bit quantization costs
nothing; Jev leads by one. Details, probabilities, and the misses of every
system: examples/.
Not affiliated with or endorsed by TypeSafe. "Jev-compatible" describes the request/answer shape (the official TypeScript SDK runs against this server unchanged in smoke tests), not certified behavioral equivalence. Training data targets include outputs from a pinned Jev API version; this is disclosed in the model cards. Code is MIT-licensed; base models keep their own licenses.
The repository contains a Python/Apple MLX research prototype, not a production-ready decision service. It implements versioned readouts, LoRA training, matched experiments, fitted temperatures, and a bounded local HTTP service.
- Actual structured-v1 SmolLM2-135M and Qwen3-0.6B training runs are complete; historical SmolLM3-3B results remain separately labeled.
- The API supports the basic Jev request/answer shapes and has a TypeScript SDK smoke test. Full behavioral and API compatibility is not established.
- Calibration and unfamiliar-rubric generalization are research goals, not guarantees.
structured-v1rendering preserves JSON and primitive identity. Historical adapters automatically uselegacy-v0; incompatible overrides are rejected.- Canonical data, split/provenance checks, grouped comparisons, independent Score levels, and up-to-255-option candidate readout are implemented and tested.
- The local service has bounded admission/caches, deadlines, model discovery, and a model-owning worker. See Serving for limits and caveats.
- Portable export and a Rust runtime are later, evidence-gated options.
See Architecture for current behavior and compatibility gaps, and Roadmap for the model and serving plan.
Beyond the served tiers, the repo is a full experimental workbench. Everything below also requires Apple Silicon; models download from Hugging Face on first use.
uv sync
# One request with all three primitives; no adapter required.
uv run python scripts/demo_request.py --model HuggingFaceTB/SmolLM2-135M
# Historical v0 baseline. New bare-model calls otherwise default to structured-v1.
uv run python scripts/eval_baseline.py --model mlx-community/SmolLM3-3B-Base-bf16 \
--task sst2 --n 200 --calibrate --renderer legacy-v0
# Development server; keep running and use a second terminal for the SDK test.
uv run python scripts/serve.py --model HuggingFaceTB/SmolLM2-135M --port 8399# Requires Node with TypeScript type-stripping support.
cd tests/ts && npm install && node test.tsAdd --adapter adapters/smollm3-3b to the 3B evaluation/server commands only after
training or obtaining that matching adapter. Datasets and full-weight models are
not bundled in git; the selected LoRA adapters (both tiers plus experiment
controls) are. See Training. For launch commands for every
local adapter, side-by-side servers, and request/output examples, use
Serving.
# Validates hashes and derives group-disjoint train/development/calibration/test.
# Choose a fresh output directory; existing datasets are never overwritten.
uv run python scripts/prepare_data.py --out-dir data/kev-v1
# A new external-data evaluator saves per-example predictions and a run report.
uv run python scripts/eval_dataset.py --model HuggingFaceTB/SmolLM2-135M \
--data data/kev-v1/development.jsonl --out-dir data/evals/v1-smoke --n 20
# Fast tests; GPU/model tests are opt-in (see Evaluation).
uv run python -m unittest discover -s tests/python -vThe audited local preparation produced 8,769 training, 1,128 development, 1,103 calibration, and 1,048 test questions. Do not mix these with the old corpora without cross-source leakage checks. See Training for the v1 LoRA command.
The headline numbers, each from a reserved test the model never influenced and spent exactly once: the frozen 0.6B volume tier reached 77.6% / ECE 0.049 on the untouched in-family test (development: 82.0% / 0.020 — dev numbers steer experiments and are optimistic by construction); the MiniCPM-q8 quality tier beat it by +6.8 accuracy points, CI [+4.5, +9.2], on a fresh out-of-family gold test, with better NLL and Brier. Quantization to 8-bit was quality-free (zero argmax flips, median 1.61x serving speedup). The most important negative result is also measured: in-family calibration does not survive distribution shift (confident errors 14-19% at t>=0.9 out-of-family vs ~2% in-family) — which is why the runtime ships per-workload temperature fitting, whose resampled evaluation cut that to 3-4%. An earlier 42-case rubric benchmark was retired after retraction — it swung on seed alone. Full protocols, confidence intervals, and every retraction: Experiments and Decisions.
Historical 3B experiments report SST-2 accuracy of 0.925 with gold-label LoRA and contextual correction (n=200). SST-2 shares its question and criteria with the IMDB training task: this is cross-dataset transfer, not unseen-question validation. A distilled adapter reports tweet-emotion ECE 0.065, versus 0.134 for gold-label LoRA, but the runs differ in data and lack uncertainty intervals. See Evaluation for all results and limitations.
The local data audit found valid Kev train/test downloads, validation overlap
across training sources, and two Nimble files containing only 404: Not Found.
The new preparer uses only the verified Kev files; it does not mix in those invalid
or overlapping legacy artifacts. See Data.
- Use-case examples — eight runnable request templates after Jev's use-case catalog, with real two-tier outputs
- Example benchmark — 24 local Qwen requests versus fresh Jev 1.13.0 and historical cached answers; a small diagnostic, not full parity evidence
- Overview — goals, scope, evidence, and sources
- Architecture — current implementation and compatibility gaps
- Serving — SmolLM/Qwen launch commands, requests, options, and bounded deployment
- Experiments — controlled model/readout/calibration experiments
- Data — external inventory, integrity checks, and split risks
- Training — v1 experiments, calibration, and historical v0 recipes
- Evaluation — metrics, commands, historical results, and required tests
- Roadmap — staged model and runtime implementation plan
- Decisions — experiment history and revised decisions
TYPESAFE_API_KEY is needed only for live Jev calls: distillation, uncached
agreement evaluation, and SDK tests with --with-jev. .env, datasets under
data/, and full-weight models are gitignored; the selected LoRA adapters,
corpus manifests, and release temperature artifacts are tracked. Check upstream
data/model licenses and TypeSafe's terms before distributing data or distilled
weights.