Deterministic, offline-first and local-first agentic API engineering.
Português (Brasil) · Portable distribution in English · Distribuição portátil em português
Documentation index · Índice de documentação em PT-BR
API Forge covers discovery, construction, evolution, migration, contracts, gRPC, testing, performance, TPS validation, observability, data access and governed agentic execution. Given an existing project or contract, it produces provenance-backed IR, evidence, plans, verifiable tasks and bounded decisions instead of guesses.
Resultado correto, verificável e reproduzível por token consumido.
python -m pip install -e '.[dev]' # Python >=3.12,<3.13
# Optional all-in-one local extras:
python -m pip install -e '.[all]'
# Optional visual terminal UX:
python -m pip install -e '.[tui]'The package can live in a user-selected virtual environment, prefix, mounted
volume or container. Add its executable directory to PATH, or invoke it by
absolute path. Set APIFORGE_HOME, APIFORGE_CONFIG and APIFORGE_CACHE when
the default project state location is not writable. See
the portable distribution guide.
The hostless first-run surface is:
apiforge inspect
apiforge init
apiforge status
apiforge doctor
apiforge context resolve --scope repo
The portable/workspace feature has paired English and Brazilian Portuguese guides. Start with the step-by-step guide for your language:
- Portable distribution and workspace — English
- Distribuição portátil e workspace — PT-BR
- Platform usage — English
- Uso da plataforma — PT-BR
- Experience and interoperability — English
- Experiência e interoperabilidade — PT-BR
- Economy guide — English
- Guia de economia — PT-BR
- Host parity — English
- Paridade entre hosts — PT-BR
- Evolution map — English
- Mapa de evolução — PT-BR
- Security policy — English
- Política de segurança — PT-BR
- Product closure specification — English
- Especificação de encerramento — PT-BR
- Architecture and MVP boundaries — English
- Arquitetura e fronteiras do MVP — PT-BR
- Capability matrix — English
- Matriz de capacidades — PT-BR
- Change-control host — English
- Host de change-control — PT-BR
- Devin integration — English
- Integração Devin — PT-BR
- Observability integrations — English
- Integrações de observabilidade — PT-BR
- Priority roadmap — English
- Roadmap prioritário — PT-BR
It works offline and without a configured agent host. Claude, Codex, Devin,
Copilot and MCP are optional adapters; unavailable capabilities remain explicit
in doctor and context payloads. workspace.yaml registers independent
repositories without requiring a monorepo.
The adaptive-routing program is complete and available in the portable, hostless core:
- Wave 1 adds
RoutingPlan/v1with explicit primary, fallback, parallel, reviewer, critic and referee roles. - Wave 2 adds multidimensional scorecards, freshness-aware signals, receipts for observations and an optional adversarial evaluation gate.
- Wave 3 adds versioned local expertise packs, family/implementation routing and explicit refusal when required local knowledge is unavailable.
Runs preserve routing.json for the compatible decision trace and
routing-plan.json for the bounded execution plan. The three waves do not
require a host, network, model SDK or external mutation. The evolution map
records the post-ship items that remain deferred, including high-level
ask/improve/migrate/fix orchestration, complete relationship inference,
remote Knowledge Pack updates, symlink installation, automatic host-file
overwrite, task-level precedence, distributed workspace debate and total host
parity.
The repository also carries the native Caveman/Cavekit assets under vendor/,
with pinned provenance and a SHA-256 manifest. RTK project filters live under
.rtk/; the optional RTK binary is not silently installed by API Forge.
The post-wave routing program is now closed as three additive, offline-first SDD shipments:
- A — Risk-Aware Routing and Task Complexity: explicit risk and task complexity govern verification depth, roles and bounded fallback behavior.
- B — Scorecard-Adaptive Routing: fresh local scorecards provide bounded champion/challenger/unresolved evidence without automatic promotion.
- C — Graph-Aware Impact: bounded explicit graph evidence adds conservative impact gates, candidate selection evidence and one explainable brief.
A, B and C remain additive: C cannot weaken A's risk/complexity gate or B's
scorecard evidence. The canonical C assessment is persisted as
graph-impact.json; stale, missing and unresolved graph evidence remains
visible and conservative. See the GraphImpactAssessment/v1 contract
and the C archive.
python scripts/vendor_caveman.py --check
apiforge agentops nativeThe deterministic supervisor turns intent into sealed TaskSpec work, runs
independent tasks in a sandbox, preserves full artifacts, invokes verification
and refuses DONE when evidence is missing. Debate is triggered by risk,
divergence, missing proof or explicit human request. The main building blocks
are:
task,runtime,sandbox,verification,briefandevals;- native provenance graph and measured context/token accounting;
- contract intelligence for OpenAPI and gRPC;
- offline Digital Twin scenarios for validation, auth, timeout, 5xx, rate limit, idempotent retry and contract mismatch;
- declarative performance plans for load, stress, spike, soak and capacity;
- OTel-first health correlation with SLO, error budget and performance signals;
- read-only plans for OTel, Datadog, Dynatrace and CloudWatch.
- relational access for PostgreSQL, MySQL/MariaDB and RDS/Aurora;
- streaming IR for Kafka/MSK, Kinesis, RabbitMQ, NATS and Pulsar;
- messaging IR for SQS, SNS, EventBridge and Kinesis;
- analytical access for OpenSearch/Elasticsearch and Redshift;
- advanced observed-risk profiles for Redis, DynamoDB, MongoDB/DocumentDB and Neptune.
The complete gap analysis and sequencing are maintained in docs/API_FORGE_EVOLUTION_MAP.md. It maps the remaining work across the agentic runtime, contracts, Java/Go/Python evolution, performance, observability, AWS, databases, security, evals and host parity. The next cycle is selected by dependencies and risk, not by an ad-hoc sequence of isolated “next steps”.
The execution sequence and acceptance gates are tracked in
docs/roadmaps/API_FORGE_PRIORITY_EXECUTION_ROADMAP.md.
Adapters now expose an explicit execution envelope (static, fixture,
live_read_only or live_mutation) with evidence level, input hashes,
limitations and unresolved diagnostics. Static analysis is never reported as
runtime proof. Local commands run through an allowlisted, shell-free sandbox;
durable control-plane steps support worker leases, heartbeats and recovery.
The first integrated evolution cycle is complete through Phase G. Runtime, migration, capacity, observability, data governance, security and agentic quality are now represented by local contracts, focused tests and SDD evidence.
The next interoperability program is also implemented as an additive local
surface: apiforge tui TASK opens the execution-first Textual UX, while
--fallback or a host without Textual uses the same Rich/JSON projection.
Knowledge freshness, four-host capability negotiation, observed Python matrix
receipts, modular experience CLI projections and bounded adaptive debate are
all evidence-driven; missing external receipts remain unresolved.
The Python core and CLI are shared across Claude Code, GPT/Codex, Devin and Copilot. Repository mirrors expose the same skills and agents where each host supports them. Host hooks, MCP lifecycle, slash commands and subagent APIs are host-specific and are not falsely reported as identical.
apiforge agentops parity
apiforge agentops activation-plan --host claude
apiforge agentops activation-plan --host gpt-codex
apiforge agentops activation-plan --host devin
apiforge agentops activation-plan --host copilotActivation plans are plan_only and approval-gated; they do not alter user
configuration. Devin has an additional payload-first adapter for Desktop, CLI
and Cloud:
apiforge devin probe
apiforge devin capabilities
apiforge devin payload "Review the current API evolution slice" --surface cli --task-kind reviewSee docs/HOST_PARITY.md and docs/integrations/API_FORGE_DEVIN.md.
The platform-completion foundation provides one evidence-backed workflow for API, database, messaging, CI/CD, cloud and front-end work. Agents understand the declared need and recommend practices or architectures with facts, assumptions, alternatives, trade-offs, risks, limitations and a verifier. It does not require a guided wizard and does not turn a parser, prompt or fixture into runtime proof.
Start with the public capability boundary:
apiforge capabilities list
apiforge capabilities verifyPara governar uma mudança de API ligada a Git e CI/CD:
apiforge change-control run \
--bundle tests/fixtures/api_git_cicd/change_bundle.json \
--out-dir .apiforge/change-control
apiforge change-control verify --run-dir .apiforge/change-control
apiforge change-control publish --run-dir .apiforge/change-control
apiforge change-control surface --run-dir .apiforge/change-control --surface ide \
--out .apiforge/change-control/ide.jsonEsse fluxo é read-only, reproduzível por replay e termina com status ok,
review, blocked ou failed, sem traceback governável na CLI. Merge, push,
dispatch de workflow, deploy e autofix continuam fora da fronteira.
collect escreve um bundle sanitizado e um
af-change-collection-receipt/1; publish emite JUnit, Markdown, SARIF e HTML
e serve expõe a mesma decisão através de um host UI/IDE read-only local ou
remoto, autenticado e TLS.
For a real case, use the complete chain:
analyze -> next-step -> graph -> evidence -> brief
The detailed commands, six vertical proof cells, agent output contract and Git/CI/CD/IDE/UI integration limits are documented in docs/guides/API_FORGE_PLATFORM_USAGE.md, docs/capabilities/API_FORGE_CAPABILITY_MATRIX.md and docs/architecture/API_FORGE_PLATFORM_COMPLETION.md.
The supported production boundary is explicit: local/static and authenticated
read-only remote paths are available; provider-specific freshness, live runtime
guarantees and production performance claims require an external receipt and
independent verifier. The dedicated CI host can request a PR auto-merge only
when the repository variable APIFORGE_AUTO_MERGE=true is explicitly set;
agents and the core never receive that authority.
apiforge contract-intel impact \
--protocol openapi \
--baseline tests/fixtures/openapi/orders-v1.yaml \
--candidate tests/fixtures/openapi/orders-v2-breaking.yaml
apiforge contract-intel twin \
--protocol openapi \
--contract tests/fixtures/openapi/orders-v1.yaml \
--scenario dependency-timeout
apiforge perf plan \
--subject orders --endpoint 'POST /orders' --target-tps 100
apiforge observability health \
--source telemetry.json --service orders
apiforge observability read-plan \
--provider dynatrace --service orders \
--start 2026-09-22T00:00:00Z --end 2026-09-22T01:00:00ZThe corresponding read-only gates can also be composed by integrations:
from apiforge.perf.capacity import assess_capacity
from apiforge.observability.readiness import assess_export_readiness
from apiforge.data_governance import assess_data_access
from apiforge.safety import assess_api_safety
from apiforge.quality import assess_agentic_qualityThey produce versioned ready/passed, review, blocked or
inconclusive outcomes. Missing evidence is named rather than defaulted.
These commands are offline by default. Real provider reads, credentials, load-generator execution and external mutations remain explicit adapters and policy-gated phases.
Specialized source models are offline and language-neutral:
apiforge model rds-access --path .
apiforge model kafka-access --path .
apiforge model msk-access --path .
apiforge model sqs-access --path .
apiforge model eventbridge-access --path .
apiforge model rabbitmq-access --path .
apiforge model nats-access --path .
apiforge model pulsar-access --path .
apiforge model opensearch-access --path .
apiforge model redshift-access --path .collect rds and collect msk produce hashed AWS posture dumps. No API Forge
command starts consumers, publishes messages, executes database queries or
changes broker/database topology by default.
Observability integrations now have optional OTel, Datadog and Dynatrace payload exporters plus host-owned authentication bindings. They remain disabled unless an allowlisted HTTPS endpoint, available credential, explicit approval and an authenticated host callback are all supplied.
apiforge analyze \
--contract tests/fixtures/openapi/orders-v1.yaml \
--project tests/fixtures/fastapi_orders \
--out-dir .apiforge/casePrints a compact JSON summary and writes a case directory:
| Artifact | Meaning |
|---|---|
api-ir.json |
Canonical API-IR: operations keyed by method+path, contract and code projections, provenance, input hashes, diagnostics |
facts.json |
Immutable evidence facts (code routes, contract operations) with source path/line/sha256 |
findings.json |
Rule verdicts — confirmed or unresolved, always evidence-backed |
changes.json |
Contract diff results — only when --baseline is given |
case.json |
Manifest written last; sha256 of every declared artifact |
Individual stages: apiforge discover --project ., apiforge model build --contract c.yaml --project ., apiforge diff contract --baseline a.yaml --candidate b.yaml, apiforge judge --contract c.yaml --project ..
Frameworks: analyze --framework fastapi|spring|go|auto — auto counts
.java/.py/.go files (none → AF-INPUT-FRAMEWORK-UNKNOWN). The Spring adapter
parses *.java with tree-sitter (@RestController/@RequestMapping/
@GetMapping…, RouterFunctions.route, JAX-RS); non-literal annotation args
emit AF-SPRING-UNRESOLVED-ROUTE, never inference. The Go adapter covers
chi (Route scopes, verb calls), net/http (Handle/HandleFunc with Go
1.22 "METHOD /path" patterns) and gin/echo verb calls; Mount, middleware
and non-literal paths emit AF-GO-UNRESOLVED-ROUTE. Every JSON command accepts
--detail-level summary|normal|full; summary drops verbose text fields
while refusal codes and fact_ids survive.
| Command | Purpose |
|---|---|
apiforge policy check --verb fs.delete --class destructive |
Evaluate an action against the policy catalog (allow/gate/deny) |
apiforge sdd check --root docs/sdd [--strict] |
Validate phase frontmatter, upstream hash cascade and evidence gates |
apiforge sdd status --root docs/sdd |
Per-feature phase status summary |
apiforge sdd stamp --artifact f/plan.md --upstream f/contract.md |
Write the upstream sha256 into frontmatter (line surgery only) |
apiforge sdd set-phase --root R --feature F --phase P --status S [--strict] |
Transition a phase; gates need evidence or a recorded override |
apiforge sandbox apply --root . --diff change.diff |
Apply a diff to copied before/after trees and report the finding delta |
apiforge sandbox clean --root . |
Remove .apiforge/sandbox only |
apiforge evidence emit --case .apiforge/case --out receipt.json |
Receipt binding artifact paths to sha256 (proves correspondence, not authorship) |
apiforge evidence verify --receipt receipt.json |
Re-hash every artifact the receipt lists |
apiforge next-step --findings findings.json --phase verify |
Route the dominant finding area to the specialist agent |
apiforge change-control run --bundle B --out-dir D |
Run the governed API/Git/CI/CD replay flow |
apiforge change-control collect --repository R --base-sha S --head-sha S |
Collect GitHub context through the GET-only adapter and write a collection receipt |
apiforge change-control verify --run-dir D |
Verify the local change-control artifact references |
apiforge change-control publish --run-dir D |
Emit JUnit, Markdown, SARIF and HTML projections |
apiforge change-control surface --run-dir D --surface ide|ui |
Export the canonical IDE/UI projection |
apiforge change-control serve --run-dir D |
Serve the local read-only IDE/UI host |
apiforge integration github-issues --repository owner/repo |
Read GitHub issues with a GET-only receipt |
apiforge integration health --url https://host/readyz --out receipt.json |
Record remote health and freshness |
apiforge integration json --url https://tracker/api/issues --out receipt.json |
Read Jira/Linear-like JSON without mutation |
apiforge integration verify-receipt --receipt receipt.json --now <ISO8601> |
Verify external receipt freshness |
apiforge platform verify-runtime --out runtime-receipt.json |
Execute allowlisted local probes for six verticals |
apiforge rules list [--area SECURITY] |
List catalog rules — the knowledge base every finding cites |
apiforge rules lookup AF-SEC-001 |
Print one rule's rationale/remediation/reference |
apiforge collect api-gateway --api-id X --out dump/ |
Fetch API Gateway config into an offline dump (needs pip install apiforge[aws]; the only family that touches AWS) |
apiforge model api-gateway --path dump/ |
Read the dump into facts — offline, no credentials |
apiforge build endpoint --contract c.yaml --operation-id X --project . |
Synthesize a Spring skeleton, prove it in the sandbox (main tree untouched) |
apiforge build endpoint ... --into-worktree NAME --approve |
Promote generated files into a policy-gated git worktree |
apiforge playbook api-governance-reviewer |
Render the coordinator's executor decomposition — works without dispatch |
apiforge economy report [--root .] |
Measured call sizes; detail_level_effect shows what summary saves |
apiforge context funnel --case .apiforge/case |
Measured bytes per case stage (api-ir → facts → findings → summary) |
apiforge context capsule --target "POST /orders" [--budget-bytes N] [--level L3|L4] |
ContextCapsule/v1: minimal evidence for one operation as hash-verified ctx://sha256/… refs under a byte budget |
apiforge context expand ctx://sha256/<hex> [--run-id R] |
One ctx object, re-hashed before it is returned |
apiforge context quality --capsule <json> --run-id R [--gate strict|evidence|permissive] |
ContextQualityReport/v1 + ContextSufficiencyResult/v1: 13 measured metrics (unresolved never carries a value) and the minimum-sufficient prune decision from the recorded ledger |
apiforge evals context-quality |
4-case corpus: metric catalog and sufficiency gates vs declared expectations |
apiforge memory rank [--term …] [--environment E] |
MemoryRankedResult/v1: §15 deterministic retrieval with per-signal score decomposition; semantic bonus is optional, never required |
apiforge memory quarantine-list, apiforge memory quarantine-resolve --candidate-id C --verdict persist|reject --by R |
§14 quarantine review: trust-insufficient candidates park in quarantine.jsonl; a human release re-runs the full gate pipeline |
Trust Plane (apiforge.trust) |
TrustUnit/v1 annotation for every context boundary, deterministic taint propagate(), allowlist-first authorize() over rules/tool_risk.yaml — DATA IS NOT INSTRUCTION enforced by contract |
apiforge economy stats [--run-id R] [--transcript T] |
Bytes attributed per run and source; tokens unresolved without a transcript |
apiforge economy record-usage --run-id R [--transcript T] [--estimate N --method M], economy ledger --run-id R |
§19–§20 token ledger: per-field provider usage rolled up per basis (observed/estimated/unresolved never mix) at run/task/agent granularity |
apiforge economy pricing [--pricing yaml], economy cost --provider P --model M --accounting <json> |
§21 declared ProviderPricing catalog resolved by effective_at; AF-ECONOMY-PRICING-MISSING refuses inferred prices |
apiforge economy reconcile --run-id R --estimate <json> |
§22 BudgetReconciliation/v1: estimated vs observed tokens/cost/tool_calls/elapsed_ms with calibration error; gaps stay unresolved |
apiforge governor decide --profile … --risk … [--budget …] |
§23 GovernorDecision/v1: risk floor raises the profile (never lowers), budget/security clamps named in clamped_by, absent inputs in unresolved |
| `apiforge governor gain | stop --action ` |
apiforge governor recover --failure-class C --attempt N, governor loop-check --fingerprints … |
§26–§27 RecoveryDecision/LoopDetection: closed failure-class ladder (unknown refuses AF-GOV-FAILURE-CLASS-UNKNOWN), repeated strategy fingerprints blocked AF-GOV-LOOP-DETECTED |
apiforge evals agent-governor |
4-case corpus: risk floor, security clamp, gain/stop, recovery ladder, loop block |
apiforge control routes, control shadow |
§28–§29 lifecycle state per route + recorded parallel-run observations (candidate vs legacy + sorted difference) |
apiforge control eval --route R [--trigger …], `control promote |
demote --route R` |
apiforge evals control-plane |
3-case corpus: shadow never governs, promotion gates, fallback + terminal refuse |
apiforge route model --inputs <json> [--scorecards F], route scorecard --evaluations F |
§33–§34 ModelRouteDecision: hard constraints then weighted rank over §34 scorecards segmented by task class |
apiforge route promote --route model_routing --evidence <json> [--evaluations F] |
§35 promotion through the control-plane lifecycle — refuses AF-ROUTE-PROMOTION-EVIDENCE without min_evaluations + quality_floor real traffic |
apiforge knowledge adaptive --query Q [--semantic], knowledge rewrite --query Q --deterministic-hits N |
§36–§39 L0→L4 ladder (stops at first sufficient level, L3 optional via local adapter) + gated query rewriting |
apiforge evals model-routing, evals retrieval |
3-case router corpus (constraints, floors, insufficient evals) + §38 strategy comparison on recall/precision/latency/tokens/cost |
apiforge runtime telemetry-export, telemetry-validate --otlp F, telemetry-ids, telemetry-collector-check --endpoint U --output-file F |
§50–§52 OTLP ExportTraceServiceRequest export with gen_ai.* semantics over all 18 operations, §51 CorrelationIds + W3C traceparent, deterministic structural acceptance, real-collector probe |
apiforge evals telemetry-otlp |
3-case corpus: full §50 operation coverage, §51 ids carried into OTLP attributes, malformed input stays unresolved |
apiforge agentops inspect <run>, agentops compare A B, agentops waste <run> |
§53–§57 AgentOps plane — RunInspection sectioned report, RunComparison over the 8-axis set, WasteReport from rules/agentops_waste.yaml with findings labeled observed/estimated/hypothesis; missing data stays unresolved, never zero-filled |
apiforge evals agentops |
4-case corpus: sections, waste detection, compare verdicts, missing-run honesty |
apiforge mcp audit, mcp disclose --task T, mcp benchmark |
§40–§43 tool-surface engineering — measured audit findings, task→active-tool-set advisory router, response-byte/token ranking; docs/mcp-compliance.md covers §44–§45 against spec 2025-11-25 |
apiforge evals tool-surface |
4-case corpus: audit baseline, disclosure routing, paging honesty, benchmark samples |
apiforge forge capabilities|health|submit|status|inspect|attach|result|evidence|handoff |
§46–§48 Forge Protocol — versioned public boundary (forge-protocol/v1) over the governed runtime; capability/engine/risk gates refuse honestly (AF-FORGE-*), tasks persist under .apiforge/forge/, handoff prepares ForgeHandoff records for declared peers; docs/architecture/forge-kernel-boundary.md covers §49 (analysis, no extraction) |
apiforge evals forge-protocol |
4-case corpus: lifecycle + completed-run + refusal gates + unresolved/unattached honesty |
apiforge evals trace-grading, evals security-adversarial, evals memory-evals |
§23–§25 eval plane — rubric-graded AgentSpan traces (rules/trace_rubric.yaml), 9 synthesized attacks against real defense surfaces, 9 memory cases over the §24 axes against the governed store |
apiforge evals live, evals frontier --report R [--latencies L --costs C] |
§23 declared layer: deterministic tier observed, provider tier deferred_external (never gates CI); quality×cost×latency Pareto — cost stays unresolved without provider data |
apiforge knowledge drift <domain> [--receipt …] --now …, knowledge impact |
§29 knowledge engine — drift rollup over read-only receipts (verified/stale/conflicted/deprecated/unresolved); declared source→pack→rule→skill→eval relation graph with named unresolved edges |
apiforge doctor --agentic |
Cross-plane agentic health: case/memory/trust/telemetry/evals/sdd/mcp/economy sections, findings with unlocks, unobservable planes named unresolved |
apiforge lab scenarios, evals knowledge-drift |
§28 opt-in scenario catalog (13 kinds, 8 covered, 5 declared gaps — never fabricated); §29 drift eval corpus (5 cases) |
apiforge economy explain <run_id> |
Why each ref was spent, from recorded provenance rules (no model call) |
apiforge evals economy [--record-baseline] |
12-case benchmark: evidence recall floor + median byte reduction vs recorded baseline |
apiforge runtime run|resume|debate <task> --profile economy|balanced|deep |
Economy Plane: profile preference with risk floor, trimmed optional roles, call cap + verification reserve, L0–L5 ladder; economy block in the result |
apiforge sdd classify --description … --path … [--baseline … --candidate …] [--write <feature>] |
Deterministic change risk (micro/low/medium/high) → minimum SDD profile; sdd check refuses profiles below it |
apiforge evals economy-routing |
15-case benchmark: role invariant, effective profile, critical ⇒ deep, economy cheaper for low risk |
apiforge context delta --base A [--head B] | --changed F [--invalidate] |
DeltaSlice/v1: changed files → graph nodes → impacted operations → capsule targets; read-only git |
apiforge cache stats|invalidate, apiforge context gc [--apply] |
Layered advisory cache: freshness per layer, dependency-aware invalidation, report-first gc |
apiforge evals cache |
10-case benchmark: warm hit rate 1.0, byte-identical vs --no-cache, zero stale reuse, invalidation precision/recall 1.0 |
apiforge knowledge select --intent "..." [--capability C] [--framework F] |
ExpertiseSelection/v1: only packs named by a trigger; no trigger loads nothing |
apiforge debate submit ... [--disagree point=reason] [--risk R] [--confidence X], apiforge debate packet |
PositionDelta/v1 submissions and a RefereePacket/v1 over one shared capsule |
apiforge agents audit |
AgentUniqueness/v1 per agent: unique capability/expertise/validator/tool/decision role, keep or merge-candidate (report only) |
apiforge evals selective-agentics |
13 cases: exact pack selection, per-role bytes ≤ 60% of naive, referee packet ≤ 50%, shadow share ±2 pp, deterministic audit |
apiforge --output compact <verb> |
Minified payload with null/empty pruned; lossless by rule (APIFORGE_OUTPUT) |
apiforge slice tests --input <pytest.log or junit.xml>, apiforge slice log --input <ci.log> |
TestSlice/v1 / ErrorSlice/v1: every failure or signature, full log behind ctx:// |
apiforge mcp surface [--surface full/compact], apiforge-mcp --surface compact or --host claude |
ToolSurface/v1: measured tool bytes; compact MCP = six gateways, all tools reachable |
apiforge agentops projection --host <host> |
HostProjection/v1: declared surface/output per host plus a verb-first map |
apiforge evals tool-economy |
10 cases: compact output lossless ≤ 65%, slicers recall 1.0 ≤ 20%, compact surface ≤ 25%, discover top-5 |
apiforge evals economy-matrix [--out report.json] |
EconomyMatrix/v1: 16 canonical contract changes × 3 profiles; quality, evidence, cost, context and latency kept apart; mutants and holdout gated |
apiforge evals gate --baseline A --candidate B |
EvaluationGate/v1: ship only with zero safety regression and quality regressions within the tolerance |
apiforge evals replay --corpus evals/corpus/economy-replay [--profile P] |
ReplayReport/v1: stored decisions re-planned under the current policy, no provider calls |
apiforge economy roi --root R |
RoleROI/v1: calls, facts added and outcome changes of each extra capability |
apiforge verify plan --changed F --risk R [--breaking] |
VerificationPlan/v1: ladder V0–V5 by risk and the impacted tests with reasons; never executes |
apiforge knowledge search --query Q [--tier 1/2/3] |
RetrievalResult/v1: declared query expansion, ranked passages with signals, top 3 then top 5 |
apiforge evidence resolve evidence://finding/<id> |
EvidenceNode/v1: one node and its one-hop neighbors as refs |
apiforge economy doctor (or apiforge doctor --economy) |
EconomyDoctor/v1: what makes runs pay more than needed, with unlocks |
apiforge economy providers, apiforge economy tier --capability C --risk R |
ProviderCapability/v1, TierDecision/v1: T0–T3, cheaper only with benchmark evidence |
apiforge agentops prompt --capability C |
PromptEnvelope/v1: stable hashed prefix + run suffix |
apiforge workspace locality --target <repo> [--transitive] |
LocalityPlan/v1: target, direct neighbors, transitive deferred |
apiforge workspace graph [--infer --run-id R] |
WorkspaceGraph/v1: declared relations; --infer adds static cross-repo calls/publishes_to_consumer edges (inferred, confidence, caller and callee file:line) and an audited workspace.infer ledger row |
apiforge field record|annotate|verify|report|export |
FieldRun/v2, FieldReport/v2: pre-registered field validation joined from existing run artifacts; the first record seals docs/field/cycle.lock.json (FieldCycleIdentity/v1), record --executor and verify --verifier must be independent actors, receipts go stale on any later annotation, and H1 plus follow-up SDDs are decided only when the cycle is ready (Wilson 95% CI, theme qualification, A/B deltas; docs/field/README.md) |
apiforge evals economy-extras |
15 cases across the seven wave-7 verbs |
apiforge knowledge watch --manifest upstream.json --now T |
FreshnessWatch/v1: packs whose upstream fingerprint, version, expiry or window says refresh_needed; never fetches |
apiforge evidence gate --question Q [--offline] |
LiveEvidenceDecision/v1: static questions stay local; live_read_only only for runtime questions; live_mutation refused |
apiforge verify escalate --static S --test T (or --test-slice) |
VerificationEscalation/v1: static → test → stop; read-only runtime only after an inconclusive test |
apiforge economy phase-budget --profile P [--usage U] |
PhaseBudgetPlan/v1: envelope split across SDD phases; contract/verify/secure protected |
apiforge runtime checkpoint <task> <run> |
EconomyCheckpoint/v1: spend a resume continues; resume never lowers the profile |
apiforge evals economy-freshness |
16 cases across the five wave-8 verbs |
apiforge evals economy-hardening [--repo-root .] |
Path containment, class-pool budget via plan_roles, contract envelope invariant, token coverage, phase status and build_delta on the analyzed fixture — the oracle runs production code |
apiforge evals agentic-quality [--responses-dir D] [--min-accuracy F] [--baseline R] |
Recorded specialist verdicts vs ground truth under each profile; absolute floor per profile (default 1.0) plus non-regression vs deep and an optional baseline report; claim_scope: recorded-agentic-outputs |
apiforge-mcp |
MCP server for the read/compose verbs (needs pip install apiforge[mcp]; every tool takes detail_level) |
apiforge collect lambda --function-name X --out dump/ |
Fetch Lambda config into a dump (Code.Location never persisted) |
apiforge collect sqs|sns|eventbridge|iam-role|cognito|waf ... --out dump/ |
Messaging/identity collectors — same offline-dump contract |
apiforge collect dynamodb|docdb|neptune|stepfunctions|cloudwatch|xray|kms|secrets|vpc-endpoints|s3 ... --out dump/ |
Datastore/ops collectors — secrets reads metadata only, S3 posture only (objects never listed), unset S3 config recorded as measured absence |
apiforge collect alb|ecs|eks|ec2|msk|elasticache ... --out dump/ |
Compute/LB collectors — describe-only posture (never user-data, console output or broker payloads) |
apiforge model sqs|sns|eventbridge|iam-role|cognito|waf --path dump/ |
Dump → aws.<svc>.* facts; boolean measures record declared absence |
apiforge model dynamodb|docdb|neptune|stepfunctions|cloudwatch|xray|kms|secrets|vpc-endpoints|s3 --path dump/ |
Same contract for the datastore/ops dumps; KMS rotation rule applies only to customer-managed keys |
apiforge model alb|ecs|eks|ec2|msk|elasticache --path dump/ |
Same contract for the compute/LB dumps; plain-HTTP only flags on internet-facing ALBs |
apiforge model lambda --path dump/ |
Lambda facts offline — env var names only, values never read |
apiforge model terraform --path infra/ |
API Gateway + Lambda resources from HCL; ${...} → AF-TF-UNRESOLVED |
apiforge model sam --path template.yaml |
Serverless resources; !Ref/!Sub → AF-SAM-UNRESOLVED |
apiforge model pact|schemathesis|k6|coverage --path r.json |
Test-tool reports → test.* facts (tools never run) |
apiforge model locust|jmeter|gatling|vegeta|wrk|hey|pytest-benchmark --path r |
Load-tool reports → test.<tool>.summary facts — RPS/TPS kept distinct, microbenchmarks never read as load runs |
apiforge model zap|semgrep|trivy|gitleaks --path r.json |
Security reports → sec.* facts (gitleaks never emits secrets) |
apiforge report build --case .apiforge/case --out report.json |
Compose the release evidence bundle |
apiforge report sign --report report.json |
Pin body/evidence/catalog hashes into the signature block |
apiforge report verify --report report.json |
Name the diverged part (body|evidence|catalog|signature_version); exit 4 |
apiforge report keygen --name k [--keys-dir D] |
Ed25519 keypair; report sign --key/verify --pubkey prove key possession, never identity |
apiforge dispatch run --coordinator C --case DIR |
Execute the playbook's dispatchable steps; missing inputs land in pending, never crash |
apiforge run tool semgrep|trivy|gitleaks|k6 --target T --out R |
Allowlisted scanner execution (fixed argv, no shell, --dry-run prints argv) |
apiforge run list |
Tool registry — declared metadata (license, capabilities, modes, evidence producer) plus measured install status; import-only tools refuse run with AF-RUN-IMPORT-ONLY |
| `apiforge debate open | submit |
apiforge tui TASK --root . |
Execution-first Textual TUI with governance, evidence and debate navigation (pip install -e '.[tui]') |
apiforge tui TASK --root . --fallback |
Headless Rich/JSON projection with AF-TUI-UNAVAILABLE unlock metadata |
| `apiforge experience status | doctor |
apiforge agentops negotiate --capability mcp [--host claude] |
Evidence-aware host capability intersection from local declarations |
apiforge knowledge freshness DOMAIN --receipt R --now ISO8601 |
Verify freshness without mutating Knowledge Packs |
apiforge migration matrix --ecosystem python --receipt R |
Publish observed compatibility cells; unexecuted versions stay unresolved |
apiforge plan strangler --project P --contract C |
Per-route cut plan; migrated routes name the parity evidence still due |
apiforge plan architecture --profile w.json |
Decision engine: ranks AWS primitives per role over a declared WorkloadProfile; every rejection names its cause, cost stays cost_to_validate |
apiforge contract list|show <name> |
Versioned canonical contracts (JSON schema per <Name>/v1) |
apiforge knowledge list|show|check |
Domain packs — source authority dates, runtime matrices, declared evals; check cross-validates rule ids |
| `apiforge task create | review |
apiforge brief show --task <id> |
OutcomeBrief — DONE is refused while gaps or missing acceptance remain |
apiforge graph build --case D --out G |
Canonical provenance graph (nodes.jsonl/edges.jsonl) — same inputs, same bytes |
| `apiforge graph query | impact |
apiforge graph export --graph G --out D [--format jsonl|neptune|rdf] |
Byte-identical copy, Neptune Gremlin-load CSV (vertices.csv/edges.csv, all :String) or RDF N-Triples (graph.nt); projections are grammar-validated and digested in export.json |
apiforge model redis --path P |
Static Redis/Valkey call-site scan (py/java/go) → data.redis.* facts + data_access_ir; binding: name is named, never proven |
apiforge model elasticache-access --path P |
Same Redis-protocol scan with the inventory declared as ElastiCache — provider named, never inferred |
apiforge model mongo|dynamodb-access --path P |
MongoDB/DocDB and DynamoDB call-site scans → data.<db>.* facts + data_access_ir; composite postures (full_scan, unfiltered_write, unbounded) come only from declared arguments |
apiforge model graph-access|neptune-access|neo4j-access --path P |
Gremlin/openCypher/SPARQL call sites (Python AST; Java/Go/TypeScript patterns) → data.graph.query facts + graph_access_ir with a DomainGraphSketch; judged by AF-DATA-013 and AF-GDB-001..010 |
apiforge model graph-explain --path DUMP [--format F] [--synthetic] |
Neptune Gremlin explain/profile, openCypher/SPARQL explain or Neo4j EXPLAIN/PROFILE dump → graph_plan_ir + data.graph.plan fact (AF-GDB-020..025); a call site without a plan keeps cardinality unresolved |
apiforge collect neptune-explain --endpoint E --language gremlin|opencypher --query Q --out D [--profile --reader-endpoint E] |
Read-only neptunedata allowlist; default plans never run the query; --profile runs it only for mutation-free literal text on a declared reader |
apiforge evals graph-quality |
Per rule × language precision/recall of graph rules over evals/corpus/graph-quality; exit 1 below thresholds.yaml |
apiforge model otel --path export.json |
OTLP/JSON trace export → perf.otel.* facts + performance_run; incomplete spans named unresolved |
apiforge perf compare --baseline A --candidate B --threshold-pct N [--repeat-baseline dir] |
compare_runs/detect_regression over two PerformanceRuns; added/removed/insufficient_data always named; --repeat-baseline measures the noise floor and suppresses deltas inside it |
apiforge perf verdict --run run.json [--repeat-baseline dir] |
passed/failed/inconclusive per run — validity conditions (baseline, generator saturation, TPS = completed transactions) name unevaluable evidence, never guess |
apiforge perf memory add --run r.json / perf memory search --subject S --tool T |
Append-only PerformanceRun store at .apiforge/perf/runs.jsonl (hash-backed); search filters declared fields, never infers |
apiforge perf suggest --case C |
suggest_fix: composes findings + catalog remediation into an ActionPlan with proposed_diff — never writes |
apiforge perf scenario --tool k6|jmeter|locust --scenario s.json |
Generate the tool's script for a declared scenario — deterministic template, never executes |
apiforge perf chaos |
List the declared controlled failure-injection scenarios (CHAOS-001..013) — injection is never executed |
apiforge model resilience --path <project> |
Static resilience scan (timeouts, retries, pools, breaker/shutdown/idempotency declarations) → resilience.* facts; heuristic, blind spots named |
| `apiforge autonomy status | set |
| `apiforge index build | status --project P [--findings f.json]` |
| — | analyze/discover extractors run through .apiforge/cache/ — hits are recorded in the ledger and named in the payload |
Twenty-five coordinator profiles in agents/*.md (one per specialty, each with a
nine-section contract, access, owned apiforge_tools, rule_areas and
executors) plus five executors in
agents/executors/*.md (af-inventory, af-extractor, af-judge,
af-verifier, af-synthesizer). AGENT_PROTOCOL.md is the operating
contract every profile points at; next-step routes by data, playbook
is the dispatch floor, and .apiforge/economy.jsonl measures every call.
Data and messaging are owned by api-data-access-architect (relational,
key-value, document, graph, search and analytical stores) and
api-event-driven-architect (queues, topics, streams and orchestration). The
roster has 25 agents rendered from agents/*.md to .claude/agents/,
.agents/agents/ and .codex/agents/ (apiforge agents sync|check|lint).
| Phase | Delivered local gate | Status |
|---|---|---|
| A | runtime review, bounded dynamic scheduler and checkpoints | complete |
| B | migration readiness for Java/Go/Python runtime plans | complete |
| C | capacity envelope from validated TPS/SLO evidence | complete |
| D | Datadog/Dynatrace/OTel export preflight | complete |
| E | Redis/MongoDB/DynamoDB/Neptune access governance | complete |
| F | API security and resilience control gate | complete |
| G | golden/holdout evaluation and host-parity aggregation | complete |
| H | product documentation and release consolidation | current |
The specialization sequence is also complete: relational/RDS, Kafka/MSK, AWS messaging, Redis/DynamoDB profiles, OpenSearch/Redshift, RabbitMQ/NATS/ Pulsar and advanced MongoDB/Neptune profiles.
The economy program (waves 0–8) is complete: measured ledger, context capsules, economy profiles with risk floors, layered cache and deltas, selective agentics, compact tool output, economy evals, targeted verification, live-evidence gating, knowledge freshness watch, per-SDD-phase budgets and resume checkpoints. See the economy guide; safety and evidence never enter the budget.
Three hardening rounds followed an external review (the program is now
closed; GitHub governance lives in
docs/guides/API_FORGE_GITHUB_GOVERNANCE.en.md): every case, fact, graph
and caller-supplied path is confined to the project and declared workspace
repositories; context, call and token budgets are contract invariants
(class pools, one ControlPlane counter including shadow challengers, token
coverage over model-facing rows); L0 early stop needs re-hashed proof
receipts; and the certification evals fail on the bugs they certify
against (evals economy-hardening runs production code, evals agentic-quality enforces an absolute floor plus baseline non-regression).
Each phase has an artifact chain under docs/sdd/, focused tests and a
rollback decision. External AWS, datastore, vendor and host mutations remain
adapter-owned and approval-gated.
| Code | Meaning |
|---|---|
| 0 | success |
| 2 | input/validation refusal (AF-INPUT-NOT-FOUND, AF-OPENAPI-*, AF-CASE-* storage) |
| 3 | integrity failure or denied/blocked governance result (AF-CASE-HASH-MISMATCH, AF-POLICY-DENY, sdd check not ok, ...) |
| 4 | confirmed finding at --fail-on severity or worse |
- No network. No SDK, no model call, no
$refresolution over the wire. - No code execution. FastAPI sources are parsed with
ast.parse, never imported. - Read-only inputs. Analyzed files are only read;
out_diroverlapping an input is refused. - Reproducible. Two runs on the same inputs produce byte-identical artifacts.
Supported: OpenAPI 3.1 (JSON or strict YAML — aliases, merge keys, duplicate
keys and custom tags rejected); FastAPI FastAPI/APIRouter decorators with
literal paths and include_router, including cross-file imports; local refs
#/components/schemas/... in the contract diff.
Unsupported (emits named unresolved diagnostics, never guesses): dynamic
route prefixes/paths, non-literal include targets, external or cyclic $ref,
arbitrary JSON Schema inference. See docs/architecture/mvp-boundaries.md.
python -c "from apiforge.case.service import load_case; print(load_case('.apiforge/case').case_id)"load_case re-hashes every declared artifact; tampering raises
CaseIntegrityError (AF-CASE-HASH-MISMATCH).
python -m pip install -e '.[dev]'
python scripts/vendor_caveman.py --check
python scripts/validate_skills.py
apiforge agents check --root .
apiforge capabilities verify
python -m ruff check src tests
python -m ruff format --check src tests
python -m mypy src/apiforge
python -m pytest -q
python scripts/check_release.pyThe same commands run in .github/workflows/ci.yml
on every push, every pull request targeting main and manual
workflow dispatch. The workflow is read-only, uses Python 3.12, cancels stale
runs for the same ref and uploads the JUnit test report when available. A push
to a non-main branch that finishes every validation green opens or reuses a PR
to main. A separate host-owned job may request auto-merge only when the
repository variable APIFORGE_AUTO_MERGE=true is explicitly enabled; the core
and agents cannot merge, push, deploy or dispatch.
The repository must authorize that operation with a least-privilege
APIFORGE_PR_TOKEN secret or the GitHub Actions setting that allows workflows
to create and approve pull requests.
Security reports belong in SECURITY.md, preferably through a
private GitHub Security Advisory. API Forge is released under the
MIT License.
