Reliability testing for probabilistic AI decisions: models that answer a typed question — pick an option, say yes or no, assign a level — with a probability for each answer.
You describe the decision once in a small Decision Contract, run a dataset through any backend, and get one reproducible reliability report with CI-friendly PASS / WARN / FAIL gates. DecGuard measures accuracy and calibration, fuzzes behavioral invariants, compares model versions, and checks production records offline.
Status: alpha (v0.1). The contract format and reports are versioned; the
systemonebackend is validated end to end against a real Kev-0.8B server and real Jev through OpenRouter, on Linux and macOS.
pip install decguard # or: uv tool install decguardPython 3.11+. Four runtime dependencies (pydantic, typer, httpx, PyYAML); no GPU, server or account.
git clone --depth 1 https://github.com/suffro/decguard && cd decguard
decguard validate examples/refund/decguard.yaml
decguard test examples/refund/decguard.yaml --allDecGuard 0.1.1 · refund_request (choice) · PASS
Cases 11 total · 11 decided · 0 errored · 10 labeled
Accuracy 1.000 macro-F1 1.000
Calibration ECE 0.139 Brier 0.049 NLL 0.159 (10 bins)
Selective confidence >= 0.8: coverage 0.727 abstention 0.273 accuracy 1.000
...
Checks
PASS min_accuracy accuracy 1 >= 0.9
PASS max_ece ece 0.139 <= 0.15
PASS option_order.max_violation_rate option_order.violation_rate 0 <= 0 (per case: max_tv_distance 0.03)
...
PASS: all gates hold
A contract (reference):
schema_version: "0.1"
decision:
name: refund_request
description: How should this customer refund request be handled?
type: choice # choice | noul | score
options: [refund, reject, review]
backend:
provider: systemone # or: mock, http, an installed plugin
model: typesafe/jev-1.13
url: https://openrouter.ai/api/alpha/decisions
bearer_token_env: OPENROUTER_API_KEY
dataset: cases.jsonl # {"input": ..., "expected": ...?, "metadata": ...?} per line
evaluation:
confidence_threshold: 0.8 # below this, a decision counts as an abstention
probability_tolerance: 0.02 # Jev rounds probabilities to two decimals
requirements: # hard gates -> exit code 1
min_accuracy: 0.95
max_ece: 0.05
warnings: # soft gates -> WARN
max_latency_p95_ms: 300
properties: # metamorphic checks (decguard fuzz / test --all)
option_order:
max_tv_distance: 0.03The quickstart guide walks through writing your first contract and dataset.
| Command | What it does |
|---|---|
decguard validate |
Check the contract, backend settings and dataset offline. |
decguard test |
Golden-dataset metrics (accuracy, calibration, coverage, latency, errors) and gates; --all adds properties. |
decguard fuzz |
Metamorphic properties: option order, label format, irrelevant/repeated context, whitespace, paraphrase, noul inversion, score monotonicity. Seeded and minimized. |
decguard replay |
Re-send stored property failures; exit 1 if they still fail. |
decguard diff |
Compare a baseline and a candidate run: flips, distribution and confidence shifts, calibration, errors, segments. |
decguard report |
Show a stored report; --contract re-applies edited gates without re-running the model. |
decguard check |
Analyze collected production records offline: calibration, drift, metadata segments. |
decguard run |
Make one decision and apply the contract's policy: accept, abstain, fallback or human_review. |
Exit codes: 0 pass (or warn), 1 a reliability gate failed, 2 configuration or runtime error.
Backends: an offline mock, a generic http protocol, systemone for Kev
(self-hosted), Jev through OpenRouter and TypeSafe endpoints — validated with real Kev
and Jev inference on Linux and macOS — plus Python plugins.
The full documentation is a VitePress site in docs/:
- Getting started: What is DecGuard? · Installation · Quickstart
- Concepts: Decisions and results · Golden datasets · Requirements and warnings
- Testing: Metrics · Metamorphic properties · Fuzzing and replay · Regression diffs
- Production: Production checks · Policies and Python SDK
- Backends: Overview · System One: Kev and Jev · HTTP · Custom backends
- Reference: CLI · Decision Contract · Reports and formats
- CI with GitHub Actions · Architecture
To browse it locally (Node.js 18+):
cd docs
npm ci
npm run devNot a model, a model server, a generic LLM evaluation framework or a hosted service. It owns the decision-specific reliability layer — contract, normalized results, metrics, gates, reports — and stays backend-neutral.
See CONTRIBUTING.md and the changelog. Licensed under Apache-2.0.