Foundry turns a ticket into a reviewed pull request, safely, by letting an approved AI agent do the work under supervision.
Raw tickets go in. Reviewed pull requests come out. Nothing unsafe makes it through the forge.
That's the whole pitch. Foundry is not another coding AI. It is the thing that sits above your coding AI and decides whether a piece of work is actually ready, what context it needs, whether it's safe to hand to an agent, who has to approve it, and what happened afterwards. The agent (Cursor, Claude, OpenAI, whatever) is the muscle. Foundry is the brain and the seatbelt.
The quick version: it's a bit like Terraform, but for shipping code with AI agents. You describe the intent, Foundry produces a plan, a human approves it, and only then does anything actually happen. Plan, approve, apply.
The formal product statement lives in VISION.md. This README is the practical, slightly more caffeinated version.
Anthropic's Claude Fable 5 and Mythos 5 can work autonomously for longer than any model before them - Stripe reported a codebase-wide migration that "would have taken a whole team over two months" done in a day. GitHub's early-access verdict points at the same future: developers handing increasingly ambitious work to agents and trusting the results across the software lifecycle.
Which means raw capability is no longer the bottleneck. The bottleneck is everything around the agent: was the ticket actually ready, which repo does this belong in, is this change safe to delegate, who signed off, what did it cost, and why did the agent do that? An agent that can run for days unsupervised is exactly the kind of agent you should not run unsupervised.
Foundry is designed for precisely this class of model. The more autonomous the agent, the more the gates matter:
- A Mythos-class model will happily take on the migration ticket; Foundry is what makes
migrations/**a hard policy block instead of a hope. - Long-horizon autonomy means more decisions made out of sight; Foundry writes down every one of them - content-hashed artifacts, policy decisions with reasons, a full audit timeline.
- Frontier models retry and self-correct; Foundry makes every retry a fresh, capped, budgeted policy decision instead of an unbounded loop.
Point the claude_code provider at Fable 5 (or cursor_* at Cursor's agents, or the signed webhook at anything else) and you get full agentic engineering with the seatbelt on: the model does the work, the humans keep the keys, and the audit trail keeps the receipts. True end-to-end agentic engineering is a governance problem, and that problem is the product.
git clone https://github.com/JeremySNR/Project-foundry && cd Project-foundry
python -m venv .venv
source .venv/bin/activate # macOS / Linux
# .venv\Scripts\activate # Windows (PowerShell or cmd)
pip install -e .
python scripts/demo.pyNo credentials, no network, no Docker. The demo drives the real production code path - orchestrator, policy engine, audit trail - through the whole story: a vague ticket gets bounced with drafted acceptance criteria, the improved ticket gets planned and gated, a human approves, the agent's PR fails CI and gets fixed by a governed retry, the PR merges, and a second run gets hard-blocked for touching migrations/. It ends with the receipts: the full audit trail and every policy decision. (--slow paces it for screen recording.)
Picture the journey of a ticket today: someone writes "add customer favourites", an engineer reads it, fills in the gaps in their head, maybe asks an AI to write it, eyeballs the result, opens a PR. Foundry makes every one of those invisible steps explicit and governed:
Linear ticket
-> Foundry reads it and asks: is this even clear enough to build?
-> gathers context (which repo, which files, what tests)
-> classifies risk (does this touch auth? payments? customer data?)
-> writes a delivery plan a human can actually sign off on
-> a human approves (this is a real step, not an afterthought)
-> Foundry hands the approved plan to a coding agent
-> the agent opens a PR
-> CI, CodeRabbit and humans review it
-> if CI fails or changes are requested, Foundry re-dispatches the agent
with the failure context (policy-gated, capped, audited)
-> Linear gets updated with status, summary and next action
And it deliberately stops at a reviewed PR. No auto-merge, no auto-deploy to production, no autonomous database migrations, no touching auth or payments without a human saying yes. The brakes are the product, not a missing feature.
The loop is also restartable where it should be: a ticket parked for clarification (or rejected, blocked, failed) can be re-triggered after it's improved - one active run per issue, not one run per issue forever.
The whole governed loop is built and tested, with swappable parts at every layer so no single vendor or piece of infra can hold you hostage. Nothing here needs the network or a paid API key to run the test suite, because every external thing hides behind a seam with a fake on the other side.
| Piece | What it does |
|---|---|
foundry.schemas |
The contracts for everything a run produces: the ticket snapshot, analysis, context, risk, plan, PR state, agent job. Pydantic, strict, validated. |
foundry.engines |
The intelligence. Deterministic heuristics by default; OpenAITicketAnalyzer (GPT-5.5) when you want a real brain at the gate. It judges readiness and missing info, it does not write code (that's the agent's job). |
foundry.policy |
The hard rules, as actual rules and not vibes. A pure-Python engine plus a matching OPA/Rego bundle. This is what blocks unsafe or unclear work. |
foundry.orchestrator |
The state machine that runs one ticket through the whole loop and writes down every decision. |
foundry.drivers |
One seam for how a run executes: inline in-process today, durable Temporal later, same interface. |
foundry.workflows |
The Temporal version of the loop: crash-proof, retries, and it'll happily wait days for an approval. |
foundry.agents |
The coding-agent abstraction. Foundry doesn't mind which blaster you bring: manual, a test fake, Cursor two ways, Claude Code via GitHub Actions, or any agent behind a signed webhook (see below). |
foundry.connectors |
Adapters for the tools Foundry talks to. Trackers: Linear, GitHub Issues, Jira. SCMs: GitHub and GitLab (watch the PR/MR, pull failing check summaries). |
foundry.api |
FastAPI app: signed Linear/GitHub/Jira/GitLab webhooks, Slack + Microsoft Teams interactivity approvals, approval commands, run status, the per-run decision timeline, the epic view for parent/child runs (GET /runs/{id}/epic, plus the epic board GET /epics), compliance evidence packs (GET /runs/{id}/evidence, plus the org-wide date-range archive GET /evidence and the cross-run epic export GET /runs/{id}/epic/evidence), and the dashboard. |
foundry.config |
The customisation story: a YAML file plus environment variables (see below). |
The cleanest way to hand work off is the Cursor Linear integration. Once a plan is approved, CursorViaLinearProvider drops an @Cursor comment with the governed instructions onto the Linear issue. Cursor's own integration runs the cloud agent, shows live status in Linear, and opens the PR. Foundry then watches that PR via the GitHub webhook and keeps Linear in sync. Delegated agents pick their own branch names, so PR-to-run correlation falls back to the Linear issue key embedded in the branch or PR title - the loop closes either way. Foundry stays the control plane and never tries to be the agent. For triggers that don't come through Linear, CursorCloudAgentProvider calls the Cursor API directly (POST /v0/agents).
Set agent.provider in the YAML and Foundry dispatches approved work elsewhere, no code changes:
claude_code- fires a GitHub Actionsworkflow_dispatchin the target repo; the repo runs Claude Code headless with the governed instructions (reference workflow inexamples/claude-code-runner.yml). The Anthropic key lives in the repo's secrets - Foundry never holds it.webhook- POSTs the HMAC-signed job input to your endpoint. Wire up Codex CLI, Aider, an internal tool, anything: do the work on the branch named in the payload, open a PR, and Foundry's GitHub webhook takes it from there.cursor_cloud/cursor_via_linear/manual- as above, or record the job for a human.auto- learned dispatch (issue #33): don't pin one agent, let the delivery-memory scorecards pick. Each run routes to the candidate with the best majority-merged history for its work-type/repo, falling back to a configured default until an agent has earned the pick (the same min-sample floor + >50%-merged gate the recommendation enforces). The choice is recorded as an explainableselectionreason on the run's audit trail, and a retry always reuses the agent that opened the PR. See Delivery memory below.
Every provider goes through the same create_job path, so the secret-leak guard and the policy gate apply no matter whose agent does the typing. (The gate is provider-agnostic - it decides whether and in which mode to dispatch, never which agent - so learned dispatch needed no policy/Rego change.)
The tracker and the SCM are seams too, not assumptions:
- Linear (default) - the original flow: signed webhook in, comments and state back.
- GitHub Issues (
tracker.provider: github_issues) - the issue is the ticket. Trigger with thefoundry:candidatelabel, approve with a/foundry approvecomment, and Foundry writes its analysis back as comments and tracks pipeline position withfoundry:status:labels. Approvers are keyed by GitHub login (the webhook signature plus GitHub's own identity authenticates the actor). Issue keys are synthesised from the repo name plus a short hash of the fullowner/repopath (CUSTOMEREB-42) so PR correlation works unchanged and similarly-named repos never collide. - Jira (
tracker.provider: jira) - same trigger/command semantics over Jira Cloud webhooks (/webhooks/jira). Jira keys (ACME-42) already match the correlation pattern.set_statefires the matching workflow transition when one exists and otherwise leaves your workflow alone. - GitLab - point a project webhook at
/webhooks/gitlab(merge request + pipeline events,X-Gitlab-Tokenauth) and merge requests close the loop exactly like GitHub PRs, including CI-failure remediation. SetFOUNDRY_GITLAB_API_TOKENso MR diffs are fetched and the same file-based gates (forbidden paths, oversize, sensitive areas) apply; without it GitLab MRs are diff-blind, just as GitHub PRs are withoutFOUNDRY_GITHUB_API_TOKEN.
Approvals don't have to happen in the tracker, either:
- Slack - approvers who live in chat can approve/reject/stop from an interactive message. Point Slack interactivity at
/webhooks/slackand setFOUNDRY_SLACK_SIGNING_SECRET; each button click is verified against Slack's v0 request signature (with replay-age protection) and then driven through the same policy gate, role checks, and audit writes as every other surface. The actor is the Slack-signeduser.id, so key approvers by Slack user id (as GitHub Issues keys them by login). Fail-closed: no signing secret, no endpoint. Outbound Slack has landed too (issue #32,connectors/slack.py): setFOUNDRY_SLACK_BOT_TOKEN+FOUNDRY_SLACK_CHANNELand Foundry posts the interactive approval message and run-status updates (parked / blocked / PR open / merged) back into the channel. - Microsoft Teams - the Teams-native twin of the Slack surface (issue #173,
api/teams.py/connectors/teams.py). Register a Teams Outgoing Webhook pointed at/webhooks/teamsand setFOUNDRY_TEAMS_SECURITY_TOKEN; each delivery is verified against Teams'Authorization: HMAC <signature>body signature, then driven through the same policy gate, role checks, and audit writes as every other surface. The approver @mentions the bot and typesapprove/reject/stop <issue-id>; the actor is the Teams-signedaadObjectId(key approvers by that id), roles from config. Fail-closed: no token, no endpoint. Outbound Teams: setFOUNDRY_TEAMS_WEBHOOK_URL(a Teams Incoming Webhook) and Foundry posts the approval Adaptive Card + the same run-status updates as an Adaptive Card. Slack and Teams can both be wired at once — they fan out together, best-effort per surface.
The webhook payload shapes are pinned by fixtures in tests/fixtures/ - spec-derived from the providers' webhook docs today, and meant to be replaced by redacted live captures over time. If a live integration ever disagrees with the mapping, the fix is a redacted capture plus a test, no credentials needed.
A PR that opens and then fails CI used to be where automation stalled. Now: a failing check suite or a changes-requested review re-dispatches the agent onto the same branch with the failure context (failing check names and summaries pulled from GitHub). Every retry passes the policy gate as retry_agent - approvals are re-checked, attempts are counted against remediation.max_agent_retries, and projected spend is checked against budget.max_cost_per_run. The cap binds at first dispatch too, not just on retries; providers that don't report cost_usd count budget.estimated_cost_per_dispatch per attempt as a proxy. Past the cap, the run parks at review required with a comment saying a human is needed. Forbidden-path blocks are never retried - blocked stays blocked.
GET /dashboard serves a zero-build, read-only page over the audit data: a live fleet strip at the top (runs in flight, the approval-queue depth and how long the oldest has waited — plus how many have breached the SLA when one is configured, agents running and how long the oldest has been running — plus how many have breached the execution SLA, and the costliest running agent's spend — plus how many have burned past the budget SLA, PRs open and how long the oldest has been awaiting review — plus how many have breached the review SLA, and spend committed by runs still in flight - the "what is every agent doing right now" view, backed by GET /metrics/fleet), an approval-queue panel (every run parked on a human, oldest first, each with its wait age and SLA breaches highlighted, backed by GET /metrics/approvals - issue #37), an execution-queue panel (every in-flight agent run — AGENT_RUNNING, agent working but no PR yet — oldest first, each with its run-time age and agent spend so far, time- and budget-SLA breaches highlighted, the hung/runaway-agent signal — which agent is the expensive one and whether it has burned past $N (dashboard.execution_cost_sla_usd) before the hard policy.max_cost_per_run cap blocks it — backed by GET /metrics/executions - issue #37), a review-queue panel (every open PR — PR_OPEN, agent shipped a PR and now awaiting review/CI — oldest first, each with its review-latency age and review-SLA breaches highlighted, the "PRs sitting unreviewed for N hours" signal, plus a "stale since last push" age with its own staleness-SLA breaches — the "which open PRs has nobody touched in N hours?" abandoned-vs-active signal — backed by GET /metrics/reviews - issue #37), a failure/triage panel (the failure-side complement to those three queues — every recently BLOCKED or EXECUTION_FAILED run, newest first, each with how long ago it failed and why, backed by GET /metrics/failures — the "what broke and needs a human" cut; bounded to a recent window since a failed run is terminal, read-only, a blocked run stays blocked - issue #37), a failures-by-category panel (the same recent failures rolled up by reason, most-frequent first — counts per category with a blocked/failed split, backed by GET /metrics/failures/by-category — the "what is the systemic blocker?" view beside the per-run feed, issue #37), a failures-by-repo panel (the same recent failures rolled up by routed repo, most-frequent first — counts per repo with a blocked/failed split, backed by GET /metrics/failures/by-repo — the "is one repo the systemic blocker?" view, the failure-side mirror of the delivery-by-repo table, issue #37), a failures-by-work-type panel (the same recent failures rolled up by work type, most-frequent first — counts per work type with a blocked/failed split, backed by GET /metrics/failures/by-work-type — the "do bugs fail while features ship?" view, the failure-side mirror of the delivery-by-work-type table, issue #37), a failures-by-category trend strip (each failure reason's failures-by-week series as a sparkline on one shared scale, backed by GET /metrics/failures/by-category/trends — the in-page render of the per-reason failure trend, "is a specific blocker — policy_denied, budget_exceeded, … — spiking or fading over time?", the failure-red sibling of the delivery per-repo/per-work-type trend strips, issue #37), a failures-by-repo trend strip (each routed repo's failures-by-week series as a sparkline on one shared scale, backed by GET /metrics/failures/by-repo/trends — "is this repo's failure rate climbing or fading over time?", the by-repo dimension of the failure trend, issue #37), a failures-by-work-type trend strip (each work type's failures-by-week series as a sparkline on one shared scale, backed by GET /metrics/failures/by-work-type/trends — "is this kind of work's failure rate climbing or fading over time, are bugs failing more while features ship?", the by-work-type dimension of the failure trend; completes the in-page failure-trend strip set — by-category/by-repo/by-work-type — mirroring how the delivery surface has both a by-repo and a by-work-type trend strip, issue #37), every run with status badges (filterable down to the approval queue - just what is waiting on a human right now), a delivery-metrics strip, a delivery-trend table (PRs shipped vs blocked, plus the per-week median and p90 tail sign-off and shipping latency — "is sign-off / shipping getting slower week over week?" — by week, backed by GET /metrics/delivery/trends), the agent scorecards and a per-agent merge-confidence trend sparkline strip (backed by GET /metrics/agents/trends - is each agent improving or sliding?), a delivery-by-repo table (PRs shipped, blocks, merge rate, spend, time-to-merge and time-to-approval per routed repo, backed by GET /metrics/delivery/by-repo - where work ships, stalls and spends) with a per-repo trend sparkline strip (PRs shipped by week per repo on one shared axis, backed by GET /metrics/delivery/by-repo/trends - is each repo speeding up or stalling?), a delivery-by-work-type table (PRs shipped, blocks, merge rate, spend, time-to-merge and time-to-approval per kind of work, backed by GET /metrics/delivery/by-work-type - do bugs ship while features stall, where does the spend go by category?) with a per-work-type trend sparkline strip (PRs shipped by week per work type on one shared axis, backed by GET /metrics/delivery/by-work-type/trends - are features speeding up or slipping vs bugs?) - both delivery trend strips' per-week bars also show that week's median and p90 sign-off and shipping latency on hover, the per-cell render of the same latency the delivery-trend table shows as columns - an epic board (multi-repo runs rolled up to one status with their child runs, backed by GET /epics - issue #35), a policy-gate panel (the effective gate this deployment enforces - the confidence threshold, protected paths, per-repo/per-path required-approval roles, the N-of-M approval counts, the risk-escalation surface (the diff-path globs and ticket-text keywords that flag a sensitive area and so demand its approval roles) and retry/budget caps, plus which backend enforces them, backed by GET /metrics/policy - the in-app twin of foundry-policy explain, issue #31), a compliance-check panel (when a baseline is configured via dashboard.policy_baseline: a control-by-control PASS/FAIL of whether this deployment's gate is at least as strict as that baseline preset — soc2/pci-dss/… — with the overall verdict, backed by GET /metrics/policy/check - the always-on, in-app twin of foundry-policy check, so "has our gate drifted below SOC 2?" is answerable on the dashboard, issue #31), an audit-integrity panel (the org-wide VERIFIED/TAMPERED verdict over the cross-row audit hash chain — every run in the window recomputed, with click-through to any failing run's timeline, backed by GET /metrics/integrity - the always-on, in-app twin of foundry-evidence verify, so "is our audit trail intact right now?" is answerable on the dashboard without a terminal; read-only, it reports a broken chain but blocks no run, issue #34), and per run the full decision timeline - artifacts, policy decisions with reasons, audit events, agent jobs and spend. It answers "why did the agent do that?" in one click. Token-gated by FOUNDRY_API_TOKEN and disabled when none is configured, same fail-closed posture as the API (the JSON equivalent is GET /runs/{id}/timeline); when OIDC browser login is configured an operator can sign in via SSO at /dashboard/login instead of pasting a token (issue #34), and the resulting read-only session cookie is rejected on the approval endpoint so it can't be CSRF'd into an approval. GET /metrics/fleet is a snapshot of the runs' current state (no time window), distinct from the historical delivery metrics below, which aggregate finished runs over a window.
One run targets one repo - but the migration ticket above spans several. orchestrator.intake_epic(ticket, ...) is the producer (#35): it splits an epic ticket into one child run per repo, each opened through the ordinary intake path so it is analysed, risk-classified, planned, independently policy-gated and approved, and rolled up under one parent run. The split is deterministic (engines/decomposition.py, the same no-model philosophy as the heuristic analyzer): it reads an explicit Repositories: section (- billing-api: migrate the ledger writes bullets, checkbox markers tolerated) or, failing that, the ticket's known_repositories (≥2). The epic's acceptance criteria are carried into every child, and each child is scoped to exactly one repo so it routes confidently. A ticket that names fewer than two repos is not an epic and runs as a single ordinary run. The parent/child runs surface in the epic board and the cross-run evidence export below. (Follow-up tracked on the issue: wiring epic detection into a webhook/trigger auto-path - intake_epic is an explicit entrypoint for now - and LLM-assisted decomposition.)
GET /runs/{id}/evidence (token-gated, ?format=html for a rendered page or ?format=pdf for a downloadable PDF — the optional [pdf] extra) exports a single run's full chain - ticket, plan, risk assessment, approvals with identities and granted roles, every policy decision, agent jobs and the PR - as a one-click procurement artifact. It bundles an integrity check with three parts: each artifact's content hash is recomputed from its stored payload; the append-only audit sequence is checked for gaps; and a cross-row linked hash chain over the audit events is recomputed, where each event commits to the previous event's hash, so a dropped, reordered, edited, or inserted row is detectable - not just a missing sequence number. (Events written before the chain existed read back as un-chained rather than failing the check, so it's safe to enable on an existing database. It's tamper-evidence: a wholesale rewrite that re-hashes every downstream event would still verify, since the chain has no external anchor - we don't oversell it.) The run's evidence is mapped onto named controls (SOC 2 CC8.1, ISO/IEC 27001:2022 A.8.32, EU AI Act Article 14 by default), each marked satisfied or showing exactly which evidence section is missing. The mappings are config, not code - override them under compliance.control_mappings in foundry.yaml.
GET /evidence (token-gated, ?format=html or ?format=pdf) is the org-wide version: every run created in a date range, each as the same per-run pack, plus a rollup over the whole window - aggregate integrity (which runs, if any, fail their hash/sequence check), a status breakdown, and per-control coverage (how many runs in the range satisfy each configured control). Bound the window with ISO 8601 from/to (from inclusive, to exclusive; a date-only to like 2026-06-14 covers the whole day) or with days (the last N days); with nothing supplied it defaults to the last 90 days. This is the "hand the auditor one file for the quarter" export.
All three evidence exports — per-run, org-wide, and the cross-run epic export below — also render to PDF via ?format=pdf (or --format pdf on the CLI), for the auditor who wants a fileable/signable document rather than a web page. PDF rendering is the optional, lazily-imported [pdf] extra (pure-Python fpdf2, no system binary or network — installed alongside the offline core only when you want it; a PDF request without it returns a clear 503/error rather than a crash). The PDF is built from the same evidence-pack data as the JSON and HTML, so the three formats agree by construction.
GET /runs/{id}/epic/evidence (token-gated, ?format=html or ?format=pdf) is the cross-run cut for an epic (#35): the parent run plus every child run it decomposed into - one per repo - bundled into a single document, so a codebase-wide migration's full chain is one auditable artifact rather than one file per repo. It resolves the epic root first (so it works called on a child too), carries the compute_epic_rollup status and explicit root/child linkage, and rolls the same aggregate summary - integrity, status breakdown, per-control coverage - across the whole epic. A run with no children exports cleanly as a one-run epic.
The foundry-evidence CLI is the offline twin of those three endpoints: it reads the same content-hashed trail straight from the database and produces the same packs from the same builders/renderers, so an auditor or security team can get a JSON or HTML evidence pack without standing up the API or holding a bearer token. foundry-evidence run <run_id> exports one run; foundry-evidence epic <run_id> exports an epic's whole cross-run chain (resolving the root first, so it works on a child id too); foundry-evidence archive [--from ISO] [--to ISO] [--days N] exports the org-wide date-range archive (same from-inclusive / to-exclusive / date-only-to-covers-the-day bounds as GET /evidence). Each takes --format json|html|pdf (default json; pdf needs the [pdf] extra) and --output PATH (default stdout); control mappings come from committed config (compliance.control_mappings), never from input - exactly like the API.
foundry-evidence verify turns that integrity check into a CI gate - the audit-side sibling of foundry-policy check. foundry-evidence verify <run_id> recomputes the tamper-evidence verdict for one run; foundry-evidence verify [--from ISO] [--to ISO] [--days N] verifies every run in a window (same date bounds as archive, default the last 90 days). Either way it prints a machine-readable JSON verdict and exits non-zero if any run's audit chain fails to verify, so a pipeline can break the build the moment a stored trail has been tampered with - turning tamper-evidence from a passive property into an active control. It runs the same verify_integrity that backs a pack's integrity block, so the gate verdict can't drift from the export. The same verdict is also surfaced continuously in-app: GET /metrics/integrity (and the dashboard's audit-integrity panel) is the always-on twin of verify, running the same build_integrity_archive over a window so an auditor sees "is the audit trail intact right now?" on the dashboard, the audit-trail sibling of how GET /metrics/policy/check mirrors foundry-policy check.
The same offline-twin pattern covers the operational fleet views (issue #37): foundry-memory fleet prints the live fleet_status snapshot (the offline twin of GET /metrics/fleet - runs in flight, approval/execution/review queue depths with their oldest ages and SLA breaches, committed spend, by-status breakdown, honouring the same dashboard.*_sla_seconds knobs as the dashboard) and foundry-memory failures [--days N] prints the triage feed (the offline twin of GET /metrics/failures - recently BLOCKED/EXECUTION_FAILED runs, newest first, each with how long ago it failed and why), with foundry-memory failures-by-category [--days N] the matching aggregate (the offline twin of GET /metrics/failures/by-category - the same recent failures rolled up by reason, most-frequent first, so the systemic blocker is visible at a glance), foundry-memory failures-by-repo [--days N] the repo-axis aggregate (the offline twin of GET /metrics/failures/by-repo - the same recent failures rolled up by routed repo, most-frequent first, so the "is one repo the systemic blocker?" question is answerable), foundry-memory failures-by-work-type [--days N] the work-type-axis aggregate (the offline twin of GET /metrics/failures/by-work-type - the same recent failures rolled up by work type, most-frequent first, so the "do bugs fail while features ship?" question is answerable) and foundry-memory failures-trends [--bucket day|week] [--days N] the over-time cut (the offline twin of GET /metrics/failures/trends - the same recent failures bucketed by when they failed onto one shared, zero-filled time axis with a blocked/failed split, the "are we failing more than usual, is something spiking?" direction-of-travel signal the recency feed and the point-in-time by-category roll-up can't show), and foundry-memory failures-by-category-trends [--bucket day|week] [--days N] the by-reason over-time cut (the offline twin of GET /metrics/failures/by-category/trends - the same recent failures grouped by reason and bucketed onto one shared zero-filled time axis, the by-category dimension of the org-wide failure trend, so "is a specific blocker spiking or fading over time?" - the question the org-wide trend and the point-in-time by-category roll-up each can't answer - is visible per reason), and foundry-memory failures-by-repo-trends [--bucket day|week] [--days N] the by-repo over-time cut (the offline twin of GET /metrics/failures/by-repo/trends - the same recent failures grouped by routed repo and bucketed onto one shared zero-filled time axis, the by-repo dimension of the org-wide failure trend, so "is this repo's failure rate climbing or fading over time?" - the question the org-wide trend and the point-in-time by-repo roll-up each can't answer - is visible per repo), and foundry-memory failures-by-work-type-trends [--bucket day|week] [--days N] the by-work-type over-time cut (the offline twin of GET /metrics/failures/by-work-type/trends - the same recent failures grouped by work type and bucketed onto one shared zero-filled time axis, the by-work-type dimension of the org-wide failure trend, so "is this kind of work's failure rate climbing or fading over time, are bugs failing more while features ship?" - the question the org-wide trend and the point-in-time by-work-type roll-up each can't answer - is visible per work type, completing the failure surface's per-dimension trend set). The three in-flight queue drill-downs behind the snapshot's counts have offline twins too: foundry-memory approvals (twin of GET /metrics/approvals - runs parked on a human, oldest first, with wait age + SLA), foundry-memory executions (twin of GET /metrics/executions - in-flight agent runs by run-time age and agent spend so far, the hung/runaway-agent signal, flagging any over the dashboard.execution_cost_sla_usd budget SLA) and foundry-memory reviews (twin of GET /metrics/reviews - open PRs by review-latency age, each also carrying its "stale since last push" age). The delivery cut has the same offline twins: foundry-memory delivery [--days N] (twin of GET /metrics/delivery - PRs shipped, blocks and why, retries, spend, time-to-merge, time-to-approval, routing precision by confidence band), foundry-memory delivery-trends [--bucket week|day] [--days N] (twin of GET /metrics/delivery/trends), foundry-memory delivery-by-repo [--days N] (twin of GET /metrics/delivery/by-repo), foundry-memory delivery-by-work-type [--days N] (twin of GET /metrics/delivery/by-work-type - the same outcomes grouped by kind of work) and foundry-memory delivery-by-repo-trends [--bucket week|day] [--days N] (twin of GET /metrics/delivery/by-repo/trends). All read the DB directly and call the same memory/metrics.py derivations the endpoints serve, with the same --days/--bucket defaults, so an on-call engineer or auditor with DB access but no running API or bearer token can still answer "is everything healthy right now?", "what just broke and needs a human?", "what is the oldest thing waiting, and is it overdue?" and "what did we ship in the window, where, by what kind of work, at what cost, and is throughput trending up or down?".
A run used to end at "PR merged" and the data died there. Now every finished run is distilled into an outcome row - time to merge, retries consumed, escalations, spend, and a block-reason taxonomy - derived entirely from the audit trail (so foundry-memory backfill rebuilds it for runs that finished before the table existed). That history feeds back in two ways:
- Routing priors. With the catalog enricher, "14 of 16 of this team's feature tickets merged in billing-service" becomes a routing signal with an audit-friendly reason string. It is bounded on every side: a minimum sample size before history speaks at all, a smoothed (never triumphalist) success rate that must clear 50%, and a confidence cap of 89 so an explicit repo association on the ticket (90) always wins.
memory.priors_enabled: falseswitches it off. - ROI evidence.
GET /metrics/delivery?days=90(token-gated) answers the question a buyer actually asks - PRs shipped, blocked and why, median/p90 time to merge, median/p90 time to approval (intake→sign-off latency for runs that reached approval - the governance-bottleneck signal in a human-gated system, the historical complement to the live approval queue), retries, escalations, total agent spend - plus observed routing precision by confidence band, which is the data you need before movingpolicy.repo_confidence_thresholdoff its default. The dashboard shows the same numbers in a strip above the run list, andfoundry-memory show-priorsprints the mined history.GET /metrics/delivery/trends?days=90&bucket=week(orbucket=day) returns the same throughput/blocks/spend bucketed over time - the "is delivery trending up or down?" view - which the dashboard renders as a trend table; each period also carries a time-to-approval distribution (median/p90), the over-time "is sign-off getting slower week over week?" cut the single-window aggregate can't show, and the per-repo/per-work-type trend series carry the same per-period approval latency.GET /metrics/delivery/by-repo?days=90(token-gated) returns the same outcomes grouped by routed repo - PRs shipped, blocks, merge rate, spend, time-to-merge and time-to-approval per repo - the "which repos are we shipping to, which stall, and where is the spend going?" cut the org-wide aggregate can't show (unrouted runs bucket under an explicit(unrouted)label); the dashboard renders it as a delivery-by-repo table.GET /metrics/delivery/by-work-type?days=90(token-gated) returns the same outcomes grouped by work type instead of repo - PRs shipped, blocks, merge rate, spend, time-to-merge and time-to-approval per kind of work (feature / bug / tech_debt / …) - the "do bugs ship reliably while features stall, and where does the spend go by category of work?" cut (runs whose work type was never classified bucket under an explicit(unclassified)label); the dashboard renders it as a delivery-by-work-type table.GET /metrics/delivery/by-repo/trends?days=90&bucket=week(orbucket=day) is the temporal cut of that - the same per-repo outcomes bucketed over time on one shared, zero-filled axis, so "is this repo shipping more or stalling?" is answerable per repo (the way/metrics/delivery/trendsis to/metrics/delivery); the dashboard renders it as a per-repo sparkline strip.GET /metrics/delivery/by-work-type/trends?days=90&bucket=week(orbucket=day) is the work-type dimension of the same temporal cut - the same per-work-type outcomes bucketed over time on one shared, zero-filled axis, so "are features shipping more or slipping vs bugs?" is answerable per kind of work; the dashboard renders it as a per-work-type sparkline strip beside the by-work-type table. - Agent scorecards.
GET /metrics/agents?days=90(token-gated) turns the same outcome rows into which agent to trust: per provider, broken down by work type and repo, the smoothed merge rate, retries consumed, and spend. GitHub will never tell you Cursor outperforms Copilot on your billing service; Foundry can, with receipts, and it compounds with every run.foundry-memory show-scorecardsprints it and the dashboard carries it alongside the metrics strip.GET /metrics/agents/trends?days=90&bucket=week(orbucket=day, orfoundry-memory show-scorecard-trends) buckets that same per-provider merge rate over time on one shared axis - the snapshot says "Cursor merges 70%", the trend says whether that is 50%→90% (route more to it) or 90%→50% (pull back); the dashboard renders it as a per-agent confidence sparkline strip.GET /metrics/agents/recommendation?work_type=feature&repo=billing-service(orfoundry-memory recommend-agent --work-type feature --repo billing-service) goes one step further and turns those scorecards into a single, explainable provider pick - same delivery-memory guard rails as the routing priors (min-sample floor, a >50%-merged history gate, an optional candidate allow-list so it only suggests agents you can actually dispatch). Settingagent.provider: auto(issue #33) turns that recommendation into the actual routing decision: the orchestrator callsrecommend_providerat first dispatch, routes to its pick over the configuredauto_candidates, and falls back toauto_fallbackuntil an agent earns the pick - the same guard rails, now acting, not just reporting. The pick is recorded as an explainableselectionblock on theAGENT_STARTEDaudit event, a retry never re-routes (it reuses the agent that opened the PR), and the kill switch is simply not settingauto. Because the policy gate is provider-agnostic, this needed no gate/Rego change.
Blocks are never auto-judged: a blocked run whose issue later merges in a fresh run is reported as superseded, which is the honest proxy for "the gate held and a human fixed the input".
Two kinds of config, kept deliberately separate:
- Behaviour goes in a YAML file. Which analyzer, the policy thresholds, the trigger label, who's allowed to approve. Commit this.
- Secrets go in the environment. Webhook signing secrets, API tokens, the database URL. Never commit these.
Copy foundry.example.yaml to foundry.yaml, edit it, and point FOUNDRY_CONFIG at it. The layering is: built-in defaults, then your YAML, then environment variables on top, so each deployment can override the sensitive and operational bits without editing the file.
analyzer:
provider: openai # or "heuristic" for the no-key default
model: gpt-5.5
risk:
provider: llm # or "heuristic" (default): keywords + globs, no key.
# "llm" adds cited evidence to the audit trail and can
# only ESCALATE over the deterministic floor, never lower it.
extra_sensitive_keywords: # teach the ticket-text heuristic your domain vocab (#31):
payments: ["pan", "cardholder"] # area -> extra keywords, merged ON TOP of the
customer_data: ["member record"] # built-in floor. Strictly additive (only ever
# flags MORE areas) - the ticket-text twin of
# sensitive_path_globs below.
custom_risk_categories: # NEW, dynamically-named risk categories beyond the fixed
crypto_keys: # seven areas (#155). Slug name (can't collide w/ a built-in)
keywords: ["hsm", "signing key"] # ticket-text triggers -> roles at the gate, and/or
path_globs: ["**/crypto/**"] # diff-path triggers -> REVIEW_REQUIRED escalation.
required_roles: ["security"] # Escalate-only: only ever ADDS a required approval,
gdpr_subject_data: # never drops a built-in's role. No new gate rule /
keywords: ["data subject"] # Rego change (roles ride the existing resolved-roles
required_roles: ["security"] # field both backends already read).
policy:
repo_confidence_threshold: 70 # block work we can't confidently place in a repo
max_files_changed: 12 # bigger PRs go to a human
forbidden_globs: ["infra/**", "**/infra/**", "migrations/**", "**/migrations/**", "**/.env*", "**/secrets/**"]
repo_forbidden_globs: # per-repo extras for monorepos; additive on top
payments-service: ["**/ledger/**"] # of forbidden_globs, scoped to the routed repo
repo_required_roles: # per-repo approval roles (#31), additive on top of
payments-service: ["security"] # risk-derived roles; can only make approval stricter
min_approvals: 1 # N-of-M "two-person rule" (#31): distinct human sign-offs
repo_min_approvals: # per-repo override; effective minimum is max(global, this)
payments-service: 2 # so a repo can only ever demand *more*, never fewer
path_required_roles: # per-PATH approval roles (#31/#35): diff-aware, monorepo
"**/billing/**": ["security"] # subtree -> role; a PR touching it that no approver
# signed for escalates to REVIEW_REQUIRED (additive).
# Like forbidden_globs, a bare glob ("billing/**") is
# depth-agnostic - it matches the dir at any nesting (#179).
change_freeze_windows: # time windows (#31): hold AUTONOMOUS retries during a freeze
- reason: "Weekend release blackout" # recurring weekly (weekdays + start/end local time,
weekdays: ["sat", "sun"] # IANA tz) or an absolute starts_at/ends_at range.
start: "00:00" # During an active window an auto-retry parks at
end: "23:59" # REVIEW_REQUIRED instead of re-dispatching (additive;
tz: "UTC" # the initial human-approved dispatch is not gated).
enforce_plan_scope: true # plan-vs-diff drift (#31): a PR changing files outside
# the LLM plan's expected_files_or_areas escalates to
# REVIEW_REQUIRED (additive; inert for the template planner,
# which declares no expected areas). Set false to disable.
enforce_plan_out_of_scope: true # plan out-of-scope (#169): a PR changing a path/area the
# plan listed under out_of_scope (promised not to touch)
# escalates to REVIEW_REQUIRED - a stronger off-plan signal
# (additive; inert when the plan declares none). Set false to disable.
enforce_plan_tests: false # test-plan satisfaction (#169): a plan promising tests but a diff
# touching no test file (per test_path_globs) escalates to
# REVIEW_REQUIRED (additive; inert when the plan promises no tests).
# Default OFF - heuristic; set true to enable.
# test_path_globs: ["**/test_*.py", "**/*_test.*", "tests/**", ...] # what counts as a test file
plan_satisfaction: # plan-satisfaction judge (#169): the headline plan-aware gate.
provider: none # "llm" runs an escalate-only LLM pass over the approved plan's
model: gpt-5.5 # intent (goal/scope/steps) + the PR diff summary, escalating to
# REVIEW_REQUIRED when the change doesn't plausibly satisfy the
# plan - beyond file containment. Degrade-to-noop on any LLM error.
# Default "none" = no judge (deterministic plan checks only).
sensitive_path_globs: # diff-aware risk: PRs touching these escalate (a bare glob
auth: ["**/auth/**", "**/login/**", "**/sso/**"] # like "auth/**" is depth-agnostic, matching
payments: ["**/billing/**", "**/stripe/**"] # the dir at any nesting - escalate-only (#179)
remediation:
max_agent_retries: 2 # CI-failure/review retries before a human takes over
retry_on: ["ci_failed", "changes_requested"]
budget:
max_cost_per_run: 25.0 # deny dispatch (first + retries) once projected spend hits this
estimated_cost_per_dispatch: 0.0 # proxy cost for providers that don't report spend (0 = off)
agent:
provider: cursor_via_linear # or cursor_cloud / claude_code / webhook / manual / auto
# provider: auto # learned dispatch (#33): pick per run by scorecard
# auto_candidates: [claude_code, cursor_cloud] # the agents auto may route between
# auto_fallback: manual # used until a candidate earns the pick
# auto_min_samples: 3 # min runs before a candidate is eligible
tracker:
provider: linear # or github_issues / jira
triggers:
label: "foundry:candidate" # runs only start on an explicit opt-in
decomposition:
provider: heuristic # or llm - infer prose-described epic splits (grounded, degrade-to-floor, #35)
epics:
auto_decompose: false # split a multi-repo ticket into per-repo child runs at intake (#35)
approval:
approvers: # roles are config, never request payload
- email: "lead@example.com"
roles: ["engineering"]
- email: "security@example.com"
roles: ["security"]Rather than start from a blank policy block, copy a vetted preset: Foundry ships a starter policy library (baseline, soc2, change-management, pci-dss) built only from the knobs above. foundry-policy presets lists them, foundry-policy show soc2 prints one to paste into your config, and foundry-policy explain soc2 shows the gate knobs it resolves to (threshold, protected paths, per-repo sign-offs, the two-person-rule approval count, per-path roles, caps). explain introspects your own deployment config too — foundry-policy explain --config foundry.yaml (or a path argument, or $FOUNDRY_CONFIG) prints the effective gate your config resolves to, not just a preset's, so you can answer "what does our gate actually enforce?" without standing up a run; add --format json for a machine-readable {source, kind, policy, active_freeze} object. explain (and the panel below) also list any configured change-freeze windows — the times autonomous retries are held for a human — and, as the offline twin of GET /metrics/policy, report whether one is in effect right now: the active window is marked [ACTIVE NOW] with a "CHANGE FREEZE ACTIVE" banner in text (and a top-level active_freeze field — {description, reason} or null — in --format json), so "why aren't autonomous retries firing right now?" is answerable from a terminal with no running API or bearer token. The dashboard surfaces the same effective gate in-app — a policy-gate panel backed by GET /metrics/policy — so an operator or auditor can see what the live deployment enforces without shell access; that endpoint reports the same active_freeze field, evaluated at request time, so a "CHANGE FREEZE ACTIVE" banner appears on the panel whenever a blackout window is in effect right now — answering "why aren't autonomous retries firing?" at a glance. The pci-dss preset exercises the full modern surface — an org-wide two-person rule (min_approvals, separation of duties), per-path security sign-off on cardholder/crypto subtrees, and protected key material. The presets are copy-to-adopt — browsing the library never changes a running deployment — and each is strict-or-stricter than the built-in defaults, so adopting one only ever tightens the gate.
Adopt-then-drift is the real risk: you copy a preset, edit it for your repos and approvers, and have no way to prove the result is still at least as strict as the control you started from. foundry-policy check --config foundry.yaml --against soc2 closes that loop — it loads your config and a baseline (a preset name or another config file) and reports, control by control, whether your config meets or exceeds it (repo_confidence_threshold/min_approvals must be ≥, file/retry/cost caps must be ≤, forbidden globs and per-repo/per-path required roles must be a superset, N-of-M counts compared on the gate's max(global, per-repo) effective value, and retry_on must be a subset — the autonomous-retry triggers, where a smaller set is stricter, so your config may auto-retry on no more failure reasons than the baseline). It exits non-zero when your config is weaker, so it drops straight into a compliance CI pipeline as a gate ("does our deployment still satisfy our SOC 2 / PCI-DSS baseline?"). Add --format json for a machine-readable result (the same per-control verdicts plus a weaknesses list on stdout) when a CI step wants to parse the outcome rather than scrape the text; the exit code is identical in both modes (0 meets/exceeds, 1 weaker, 2 usage/config error). Each finding in the json output carries typed fields alongside the human detail — a comparator (>=, <=, superset, or subset — the last for retry_on, whose missing/missing_items then carry the subject's disallowed extra triggers), the numeric subject/baseline values for the scalar knobs (null for an absent cap, or on collection knobs where a single number doesn't apply), a missing list naming exactly which baseline globs/roles/windows a collection knob fails to cover, and a structured missing_items breakdown of that same shortfall ({key, items} per shortfall repo/path for the per-repo/per-path role and glob knobs, {key, subject, baseline} for repo_min_approvals, {item} per missing global glob/window) so a consumer can read the gap by key without parsing the "<key>: <items>" prose — so a pipeline can compare the values directly (e.g. annotate the failing control with the numbers) instead of parsing the sentence. The check is read-only decision-support over the already-gated knobs — it changes no policy. The dashboard runs the same check continuously: set dashboard.policy_baseline (a preset name or a config path) and a compliance-check panel backed by GET /metrics/policy/check shows the control-by-control PASS/FAIL against that baseline in-app — the always-on twin of the CLI gate, so drift below your committed baseline is visible to an auditor on the dashboard, not just caught in a one-shot CI run.
The presets above are copy-to-adopt. When a separate team (say security) wants to own a set of rules and have them apply live without forking the app or hand-editing your main config, point policy.bundle_path (or FOUNDRY_POLICY_BUNDLE_PATH) at a user policy bundle — a separately-authored file with the same policy:/risk: knobs. At load, Foundry merges the bundle on top of your resolved config as a strictly-additive overlay: your config plus the built-in gate rules stay a non-overridable floor, so the bundle can only ever make the gate stricter — it can add protected paths, required-approval roles (per repo or per path), risk-escalation globs/keywords and change-freeze windows, raise the confidence and N-of-M approval thresholds, and tighten the file/retry/cost caps; it can narrow which failures auto-retry, but it can never drop a protected path, lower an approval count, loosen a cap, or add an autonomous-retry trigger. A bundle value that would weaken the floor is simply ignored (the stricter of base-and-bundle wins), and as defence-in-depth the merged result is re-checked against your base config with the same strictness comparator foundry-policy check uses — a bundle that would weaken any built-in denial or approval requirement fails closed at load. The bundle is loaded from a configured path, never request input, and may carry only policy knobs (anything else is rejected); it changes only knob values both the Python engine and the Rego bundle already read, so it needs no foundry.rego change and the two backends stay in lock-step. foundry-policy explain --config foundry.yaml shows the post-overlay gate, since the overlay is part of what your config resolves to.
Secrets via env:
| Env var | What it's for |
|---|---|
FOUNDRY_CONFIG |
Path to your YAML file. |
FOUNDRY_DATABASE_URL |
SQLAlchemy URL. SQLite by default, Postgres in prod. |
FOUNDRY_LINEAR_WEBHOOK_SECRET |
Verifies inbound Linear webhooks. |
FOUNDRY_GITHUB_WEBHOOK_SECRET |
Verifies inbound GitHub webhooks. |
FOUNDRY_LINEAR_API_TOKEN |
Turns on the live Linear connector (write-back). |
FOUNDRY_GITHUB_API_TOKEN |
Turns on the live GitHub connector (PR files; also the GitHub Issues tracker). |
FOUNDRY_JIRA_WEBHOOK_SECRET |
Enables /webhooks/jira (token-compared; endpoint disabled without it). Jira has no body signature, so this is an approver-level credential (the actor identity comes from the payload) — header-only by default; ?token= query delivery needs tracker.jira_allow_query_token: true. See SECURITY.md. |
FOUNDRY_JIRA_BASE_URL / ..._EMAIL / ..._API_TOKEN |
Jira Cloud credentials when the tracker is jira. |
FOUNDRY_GITLAB_WEBHOOK_SECRET |
Enables /webhooks/gitlab (X-Gitlab-Token; endpoint disabled without it). |
FOUNDRY_GITLAB_API_TOKEN / ..._BASE |
Fetches MR diffs so GitLab MRs run the same file-based gates as GitHub PRs (without it, MRs are diff-blind). ..._BASE overrides the API root for self-managed GitLab. |
FOUNDRY_SLACK_SIGNING_SECRET |
Enables /webhooks/slack (Slack v0 request-signing + replay-age; endpoint disabled without it). |
FOUNDRY_SLACK_BOT_TOKEN / FOUNDRY_SLACK_CHANNEL |
Enables outbound Slack: posts the interactive approval message + status updates (parked/blocked/PR open/merged). Fail-closed — both (token + channel, the latter also settable via notifications.slack_channel) required, else no notifier. |
FOUNDRY_TEAMS_SECURITY_TOKEN |
Enables /webhooks/teams (Microsoft Teams Outgoing-Webhook HMAC body-signing; endpoint disabled without it). The approver @mentions the bot and types approve/reject/stop <issue-id>; identity is the Teams-signed AAD/object id (configure approvers by that id), roles from config — the same policy-gated decision path as Slack/Linear/REST. |
FOUNDRY_TEAMS_WEBHOOK_URL |
Enables outbound Teams: posts the approval Adaptive Card + status updates to a Teams Incoming Webhook. Fail-closed (unset ⇒ no Teams notifier). When both Slack and Teams are wired they fan out together (best-effort per surface). |
FOUNDRY_API_TOKEN |
Bearer token for the REST approval endpoint, the timeline API, the delivery-metrics API, the compliance evidence-pack endpoint and the dashboard. Unset = those are disabled (fail closed) unless OIDC is configured; approvals still work via signed Linear comments. |
FOUNDRY_ARTIFACT_ENCRYPTION_KEY |
Fernet key to encrypt artifact payloads at rest (comma-separate to rotate; the first key encrypts). Needs the crypto extra; unset = plaintext (the default). Content hashes stay over plaintext, so audit integrity is unaffected. After turning a key on or rotating one, run foundry-db reencrypt-artifacts to re-wrap existing rows under the current key so the old key can be retired. |
FOUNDRY_OIDC_ISSUER / ..._AUDIENCE / ..._JWKS_URI |
Enables OIDC bearer auth on the token-gated API as an alternative/addition to FOUNDRY_API_TOKEN (issue #34): a valid JWT from your IdP is accepted alongside the static token. All three are required together (also settable under auth.oidc in YAML). Optionally FOUNDRY_OIDC_ALGORITHMS (comma-separated allow-list, default RS256) and FOUNDRY_OIDC_LEEWAY_SECONDS (clock-skew tolerance, default 60). Needs the oidc extra (pip install 'project-foundry[oidc]'; bundled in the Docker image). |
FOUNDRY_OIDC_SUBJECT_CLAIM / ..._GROUP_CLAIM |
For OIDC-authenticated REST approvals (issue #34): which verified claim names identify the approver (default email, falling back to sub) and carry IdP group membership (default groups). The approver identity is then bound to the verified token, not the request body. The IdP-group → approver-role map itself is YAML-only — auth.oidc.group_role_map ({group → roles}); a verified member of a mapped group is granted those roles on top of any committed approval.approvers grant. |
FOUNDRY_OIDC_ORG_CLAIM |
Enables multi-tenancy / row-level isolation (issue #156): the verified OIDC claim that names the caller's org (e.g. org or tenant; also auth.oidc.org_claim in YAML). When set, every read and write in that request is isolated to that org — a unit of work scoped to one tenant can neither read nor write another's rows. The org is taken only from a cryptographically-verified principal — the bearer token's claim, or (for the browser dashboard) the org stamped into the signed SSO session cookie at login — never the request payload. Unset (the default) ⇒ single-tenant: every row lives under the default org and behaviour is unchanged. |
FOUNDRY_WEBHOOK_ORG_SECRETS |
Routes webhook deliveries to a non-default org under multi-tenancy (issue #34). Webhooks carry no OIDC token, so without this every webhook-created run / PR observation lands in the default org. Give each tenant a dedicated webhook secret as comma-separated org=secret pairs (e.g. acme=whsec_a,globex=whsec_b); a delivery signed (HMAC: Linear/GitHub) or tokened (GitLab/Jira) with org X's secret binds tenant X for the whole handler, while the global webhook secrets keep mapping to the default org. The org is which committed secret matched, never the payload (invariant #5). Each secret must be distinct from the others and from the global webhook secrets (rejected at startup otherwise). Unset (the default) ⇒ single-tenant: every delivery runs in the default org, unchanged. |
FOUNDRY_OIDC_CLIENT_SECRET / FOUNDRY_SESSION_SECRET |
Enable browser SSO login for the dashboard (issue #34): an operator signs in via your IdP (OAuth2 authorization-code + PKCE) at /dashboard/login instead of pasting a token. The non-secret parts go under auth.oidc (client_id / authorization_endpoint / token_endpoint / redirect_uri, all-or-nothing; optional scopes / session_ttl_seconds / cookie_secure, also env-settable as FOUNDRY_OIDC_CLIENT_ID etc.); these two secrets are env-only and required to wire login (the OAuth client secret, and the HMAC key signing the session cookie). The session cookie authenticates the dashboard's read calls only — it is rejected on the approval endpoint (CSRF-safe). When FOUNDRY_OIDC_ORG_CLAIM is set (multi-tenancy), the cookie also carries the verified org so the dashboard's reads are scoped to the operator's own tenant without a bearer token (preserved across sliding-session re-mints). Needs the oidc + http extras. Optionally set auth.oidc.session_max_lifetime_seconds (also FOUNDRY_OIDC_SESSION_MAX_LIFETIME_SECONDS) to enable sliding sessions: session_ttl_seconds then becomes the idle timeout and this the absolute cap, so an active operator's session is kept alive on each request (never past login + cap, forcing a periodic re-login) while an abandoned one still expires; unset = the prior fixed-TTL cookie. Optionally set auth.oidc.end_session_endpoint (+ post_logout_redirect_uri, also FOUNDRY_OIDC_END_SESSION_ENDPOINT / ..._POST_LOGOUT_REDIRECT_URI) for RP-Initiated (federated) logout — /dashboard/logout then ends the IdP SSO session too, not just the local cookie, so a shared workstation can't silently re-authenticate. |
FOUNDRY_SCIM_BEARER_TOKEN |
Enables the SCIM 2.0 provisioning surface (issue #157): /scim/v2/Users + /scim/v2/Groups (create / read / list / replace / PATCH / delete) so your IdP can provision and de-provision the users and groups Foundry recognises. A machine-to-machine credential, distinct from FOUNDRY_API_TOKEN; unset = the surface is disabled (403) and approver-role resolution never consults the directory, so a non-SCIM deployment is unchanged. SCIM provisions membership, never roles — an active provisioned user's group memberships map to approver roles through the committed auth.oidc.group_role_map (the group displayName is the key), and deactivating a user (active:false) or removing it from a group revokes that authority. |
FOUNDRY_AGENT_PROVIDER |
Overrides agent.provider from the YAML (manual / cursor_* / claude_code / webhook / auto). |
FOUNDRY_AGENT_AUTO_CANDIDATES / ..._FALLBACK / ..._MIN_SAMPLES |
Learned dispatch (issue #33), used when the provider is auto: a comma-separated candidate allow-list the scorecard routes between, the fallback agent used until a candidate earns the pick, and the min-sample floor. Also settable under agent.auto_* in YAML. |
FOUNDRY_CURSOR_API_TOKEN |
Needed when the provider is cursor_cloud. |
FOUNDRY_AGENT_WEBHOOK_URL / ..._SECRET |
Needed when the provider is webhook; the secret HMAC-signs the job payload. |
OPENAI_API_KEY |
Needed when the analyzer provider is openai. |
TEMPORAL_ADDRESS |
The Temporal server, for durable runs. |
FOUNDRY_CONTEXT_PROVIDER |
Overrides context.provider (static, catalog or code). |
FOUNDRY_CONTEXT_ORG |
GitHub org for foundry-catalog sync; overrides context.org. |
The same code runs on a laptop (SQLite, heuristics, no keys) and in production (Postgres, GPT-5.5, live Linear and GitHub) with nothing changing but config. That's the point.
The fastest path to a deployed instance is docs/quickstart.md - zero to governed PR in ~30 minutes with docker compose up (API + Postgres, optional Temporal profile, dashboard included). A bare docker compose up boots on a fresh clone with no copy step (it mounts the committed foundry.example.yaml); the API container applies Alembic migrations on startup, so Postgres gets its schema — and its alembic_version stamp — without a manual step.
Tagged releases (vX.Y.Z) publish a container image to GHCR (ghcr.io/jeremysnr/project-foundry) automatically, gated on the full test suite.
Foundry dogfoods itself. This repository is its own dogfood target: foundry.dogfood.yaml points Foundry at this repo with GitHub Issues as the tracker, and scripts/smoke_e2e_github.py drives one live governed run — intake, plan, approval, policy gate, dispatch, audit timeline — against a real issue using only a GitHub token. The full runbook (including deploying it webhook-driven and flipping the agent from manual to claude_code) is docs/dogfooding.md.
For development:
python -m venv .venv
source .venv/bin/activate # macOS / Linux; on Windows: .venv\Scripts\activate
pip install -e ".[test]"
pytestServe the API:
pip install -e ".[server,http]"
export FOUNDRY_CONFIG=foundry.yaml
export FOUNDRY_LINEAR_WEBHOOK_SECRET=... # and friends
uvicorn foundry.api.app:app_from_env --factoryTo use the catalog-backed context enricher (context.provider: catalog), populate the repo
catalog first and then keep it fresh with a periodic sweep:
export FOUNDRY_GITHUB_API_TOKEN=...
foundry-catalog sync --org <your-github-org> --bootstrapRun this on a schedule (e.g. daily cron or a Temporal workflow) so the catalog stays current. The sync is stateful and budget-aware: interrupted sweeps resume automatically on the next run.
The code-aware enricher (context.provider: code) goes further: the sync also records each
repo's file tree (one Git Trees API call), test layout, CODEOWNERS rules and root dependency
manifests — foundry-catalog sync --code-facts, implied when the provider is code. Routing
then matches tickets against actual code paths, reason strings cite concrete files and owners
("Code evidence: src/billing/invoice.py; owners: @org/payments"), and the context bundle carries
candidate files, the test layout and inferred test commands for the plan. Worst case the sync
spends 9 API calls per repo instead of 3; the same budget and resume semantics apply.
With planner.provider: llm the planner consumes that code-aware context and produces a
file-level plan: named files to touch, where the tests live, the commands to verify, and a
populated expected_files_or_areas. It's only consulted for a buildable, confidently-routed run;
the goal/scope/branch and the guardrail block (forbidden paths, no migrations, stop conditions)
stay deterministic — the model enriches the plan but can't relax a constraint — and an LLM failure
degrades to the deterministic template plan. The template planner remains the no-key default.
Epic decomposition has the same shape. The deterministic producer splits a multi-repo ticket via
an explicit Repositories: section or ≥2 associated repos; with decomposition.provider: llm,
an epic described only in prose — "migrate the ledger in billing-api and the checkout in
customer-web" — is recovered by inference. The deterministic decomposer stays a hard floor: the
model is consulted only when the floor declines, every repo it proposes must already appear in the
ticket text (no invented repos), fewer than two grounded repos degrades to the floor, and each
child still runs the full policy gate and its own approval — so the LLM can only add a split,
never weaken one.
Optional extras, install what you need:
.[llm]GPT-5.5 analyzer.[http]live Linear and GitHub transports.[workflow]Temporal durable execution.[postgres]Postgres driver + Alembic migrations. On Postgres, Alembic is the single schema owner: runalembic upgrade head(the Docker image does this automatically on startup;make migratedoes it by hand). SQLite dev/test databases have no migration step and are bootstrapped in-process..[pdf]PDF rendering for the compliance evidence exports (pure-Pythonfpdf2; JSON/HTML never need it).[otel]OpenTelemetry tracing (without it, the spans are free no-ops)
There is also a live end-to-end smoke test (scripts/smoke_e2e.py) that drives a real Linear issue through approval, agent dispatch and PR observation. It is gated on FOUNDRY_E2E=1 plus real credentials and never runs in CI.
There's also a real OPA bundle in src/foundry/policy/foundry.rego; run opa test src/foundry/policy if you have the OPA CLI. It's kept in lock-step with the Python engine and the two are tested against the same cases.
These are enforced, tested, and not negotiable by a prompt:
- No acceptance criteria, no build. Even if the model swears it's ready. (And when Foundry bounces a ticket, it drafts the acceptance criteria for you - clarification is a 30-second edit, not a rejection.)
- If we can't confidently say which repo this belongs in, we stop and ask.
- Production deploys and database migrations cannot run autonomously. Full stop, for now.
auto_mergeandproduction_deployare modelled as policy actions that are denied unconditionally - "never" is an enforced, audited decision, not an absence of code. - The policy gate is default-deny: an action it doesn't recognise is refused.
- No agent runs without a human's sign-off, full stop. The gate itself requires at least one recorded approval before any autonomous action - the human-in-the-loop promise is a policy rule, not just an orchestration detail, so a code path that reached the gate without an approval would still be refused. (Sensitive areas need specific roles on top of that; see the next point.) For a stricter two-person rule, set
policy.min_approvals(or a per-repopolicy.repo_min_approvals): the run accumulates distinct sign-offs — each its own audited approval, a duplicate from the same person refused — and only proceeds once enough humans have approved. The minimum is a one-way ratchet (a per-repo value can only raise it), so it can only ever make approval stricter. When more than one sign-off is required, the approval prompt (the tracker comment and the Slack message) says so up front ("Approvers required: N distinct sign-offs"), so the first approver isn't surprised by a run that stays parked after they approve. And once a partial sign-off lands, the next approver is re-pinged on the same surfaces ("N of M distinct sign-offs collected — K more required"), so a two-person-rule run doesn't go silent between approvals. - Auth, payments, PII and customer data need a human approval before an agent goes near them - and the approval has to come from someone whose configured role covers it. Roles live in committed YAML; an API caller cannot claim "security" for themselves. A sign-off from someone whose role doesn't cover the work is refused up front - never recorded and then quietly blocked at dispatch, so the audit trail never shows an approval for work that was actually denied. Roles can be scoped to a repo (
policy.repo_required_roles, resolved at intake) or, for monorepos, to a path (policy.path_required_roles): a PR whose diff touches a configured subtree (e.g.**/billing/**) that no approver signed for escalates to human review when the PR opens. Both are strictly additive — they only ever add a required sign-off, never drop one. - Approval surfaces are authenticated, full stop. Linear comments arrive over a signed webhook with the actor identity from Linear; the REST endpoint needs a bearer token (the static
FOUNDRY_API_TOKENor, when configured, an OIDC JWT from your IdP - signature, issuer, audience, expiry and an RS256-only algorithm allow-list all verified) and is disabled outright when no credential is configured. On the OIDC path the approver identity is bound to the verified token, not the request body, and IdP groups map to approver roles via committed config (auth.oidc.group_role_map) - so SSO group membership, not a hand-maintained email list, can grant approval authority, while the role a run requires is unchanged. When SCIM provisioning is enabled (FOUNDRY_SCIM_BEARER_TOKEN, issue #157) those identities and groups are managed by your IdP rather than hand-maintained: SCIM provisions membership, the committed group→role map still owns authority, and de-provisioning a user (or removing it from a group) revokes that authority — a deactivated approver is refused outright. - A captured webhook can't be replayed into action. Every delivery is deduped against a durable, bounded table (
(provider, delivery_id), shared across workers, pruned on a TTL), so a redelivered approval or CI-failure event is dropped instead of re-driving state. Setwebhook.replay_max_age_secondsto additionally reject deliveries older than a window for providers that carry a timestamp (Linear). - The network surfaces are rate limited. Signatures stop unauthorised callers; a coarse per-client cap (on by default, configurable under
rate_limit:) stops a flood of authorised-looking ones - a replayed webhook in a loop, a runaway integration, a token brute-force - from exhausting the process. Webhooks and the API get independent budgets so a burst on one can't starve the other. - Run records are not public.
GET /runsandGET /runs/{id}need the same bearer token (or SSO session) as every other read, because a run record names who requested it, who approved it and what the agent spent. With noFOUNDRY_API_TOKENand no OIDC configured they are disabled outright (403) rather than falling open; only/healthzand the signed webhook intake answer unauthenticated. The bundled dashboard already sends the token, so nothing changes there. - Risk is checked twice: once from the ticket (before dispatch) and again from the diff (after the PR opens). A ticket that said "fix the button" whose PR touches
auth/escalates to human review - and the guardrails re-run on every push, so an agent can't open a clean PR and sneak files in later. Withrisk.provider: llm, a model pass writes its cited reasoning into the audit trail ("touches session issuance inauth/tokens.py") - and it may only escalate over the deterministic keyword/glob floor, never downgrade it. - Bigger-than-expected PRs and anything touching forbidden paths get bounced to a human.
- When a code-aware planner (
planner.provider: llm) scoped the run to specific files/areas, a PR that strays outside that approved scope is bounced to a human too (policy.enforce_plan_scope, on by default) - the "agent went off-plan" signal. It's inert for the default template planner, which declares no expected scope. A PR that reaches into a path/area the plan explicitly marked out of scope is bounced too (policy.enforce_plan_out_of_scope, on by default) - a stronger off-plan signal, likewise inert until a planner declares out-of-scope entries. Those checks reason about file containment; turning onpolicy.plan_satisfaction.provider: llmadds the headline plan-aware gate that judges whether the diff actually does what the plan said (goal/scope/steps vs. the PR summary) and bounces it to a human when it doesn't - escalate-only, and an LLM outage is a no-op (it never blocks or releases a run). - The agent may retry its own failing PR, but every retry is a fresh policy decision: approvals re-checked, attempts capped, budget capped, all audited. Past the cap, a human takes over. A forbidden-path block is never retried.
- No auto-merge. Ever, in this version.
- Secrets never end up in an agent prompt; job inputs are scanned before dispatch.
- Every decision, every artifact, every approval is content-hashed and written down, so you can always answer "why did the agent do that?".
- Artifact payloads can be encrypted at rest (
FOUNDRY_ARTIFACT_ENCRYPTION_KEY, opt-in) without touching the audit guarantees: hashes are computed over plaintext, so integrity verification holds with or without the key. Thefoundry-db reencrypt-artifactsmaintenance command re-wraps rows written before a key was enabled (or under a rotated-away key) so encryption covers historical data and a retired key can be dropped — plaintext-preserving, all-orgs, idempotent, with--dry-run.
These aren't suggestions, they're the creed. This is the Way.
Tracker --webhook--> Foundry API (Linear, GitHub Issues, Jira)
|
v
RunDriver (inline now, Temporal-backed later)
|
v
FoundryOrchestrator
|
analyse -> enrich -> classify risk -> plan
|
Policy gate (OPA-style hard rules)
|
Human approval (in the tracker)
|
CodingAgentProvider (Cursor x2, Claude Code, webhook, manual)
|
PR / MR opens
|
SCM --webhook--> Foundry watches the PR, updates the tracker
(GitHub, GitLab)
src/foundry/
config.py YAML + env settings
observability.py OpenTelemetry spans (no-op without the extra)
schemas/ the run artifact contracts (+ enums in common.py)
engines/ analyzer / enrichment / risk / planner, plus the GPT-5.5 analyzer,
the escalate-only LLM risk classifier (llm_risk.py), and the
file-level LLM planner (llm_planner.py)
orchestrator.py the state machine that runs a ticket end to end
drivers.py the RunDriver seam (inline today, Temporal attaches here)
workflows/ decisions.py (pure) + the Temporal workflow, activities, worker
policy/ the Python engine + foundry.rego (kept in sync)
agents/ provider abstraction: manual, fake, Cursor (two ways), Claude Code, webhook
connectors/ Linear, GitHub, GitHub Issues, Jira, GitLab, live HTTP transports
db/ SQLAlchemy models (runs, artifacts, audit, policy, jobs) + at-rest artifact encryption (encryption.py) + foundry-db re-encrypt maintenance (maintenance.py, cli.py)
audit/ content hashing + the verifiable trail
api/ the FastAPI app, webhook security, payload mapping, dashboard
tests/ one module per package, plus the gated Temporal/Postgres/E2E tests
tests/fixtures/ spec-derived webhook payloads pinning every payload mapping
migrations/ Alembic migrations — the sole schema owner on Postgres; SQLite dev uses init_schema/create_all
examples/ reference Claude Code runner workflow
scripts/ demo.py (offline narrated demo), the live E2E smoke test, wait_for_temporal.py (real-server CI readiness)
docs/ quickstart
Apache-2.0. See LICENSE, CONTRIBUTING.md and SECURITY.md. The short version of the contribution rules: the safety gates are the product - PRs that weaken a gate, an approval requirement or the audit trail don't merge, and any policy change lands in the Python engine and the Rego bundle together, with tests on both.
Foundry takes its name from the Mandalorian forge, where the Armorer works raw beskar into something built to last and keeps to a strict creed the whole time. It fit a little too well. This thing takes raw tickets, forges them into solid reviewed work, and won't break its own rules to get there. The policy gate is the Armorer, the safety rules are the creed, and the coding agents are the ones swinging the hammer. Foundry just makes sure nobody melts something important. (If none of that means anything to you, no harm done, it still ships PRs.)
The loop is complete, closed (the agent now fixes its own failing CI under governance), multi-vendor on every side (three trackers, two SCMs, five agent providers), visible (the dashboard), deployable (docker compose up, Alembic migrations, Postgres in CI) and released (GHCR image on tags). The durable Temporal driver is now CI-proven end-to-end against a real temporalio/auto-setup server (the docker-compose Temporal profile), not just the in-memory time-skipping harness. What's left is hardening against live traffic: battle-testing the webhook payload mappings with the E2E smoke script. The long game, per the vision, is to grow this from ticket-to-PR into a full Engineering OS: planning, build, test, deploy, observability and incidents, all under the same control plane. One honest loop first, though.
Forged in the covert. Raw ore in, beskar out.