Model-agnostic platform for verifiable agentic software delivery.
Workflows, traceability, evidence-native testing, agent evals, and verify-before-done loops.
English | 日本語
Takumi is an open-source orchestration layer for AI software delivery. It is not a single coding agent, a Pi wrapper, or a DeepSeek fork. Takumi coordinates pluggable agent runtimes through explicit workflows, quality gates, artifacts, traceability, and independent verification.
Agents propose. Takumi verifies.
Takumi is designed for teams that need AI agents to produce software changes that can be reviewed, tested, resumed, audited, and delivered with evidence. Japanese SI / V-model delivery is included as a first-party workflow, not a boundary of the platform.
- Run structured software delivery workflows from requirements to evidence.
- Swap agent runtimes without changing the platform core.
- Keep traceability from requirements to design, tests, evidence, and review.
- Gate agent output with real tests and independently collected evidence.
- Evaluate agent reliability with held-out verifier SWE tasks.
- Run a Manage-Execute-Audit loop for long-horizon agent work.
Install from source. npm package publishing is planned, but not shipped yet.
git clone https://github.com/sscodeai/takumi.git
cd takumi
pnpm install
pnpm buildRun a deterministic local task with no API key:
pnpm exec takumi init
pnpm exec takumi run "Implement user login API"Run a workflow with a real runtime:
export COMMANDCODE_API_KEY=...
pnpm exec takumi run requirements.md --workflow jp-si-standard --runtime deepseekRun the MEA loop:
pnpm exec takumi loop "Fix the sumEven bug and add tests" --runtime deepseek --max-rounds 5| Problem | Takumi approach |
|---|---|
| AI tools generate code but lose the delivery trail | Traceability matrix across requirements, design, tests, evidence, and review |
| Products lock users into one model or harness | Runtime adapter API for fake, Pi, DeepSeek, CLI bridges, and future runtimes |
| Tests and evidence are treated as afterthoughts | Evidence-native pipeline for unit and integration test deliverables |
| Agents can claim "done" too early | Quality gates, held-out verifier tests, and independent verification |
| Long tasks lose state across context windows | Manage-Execute-Audit loop with persistent task state |
| Domain delivery processes are hard to encode | Extensible workflow and skill system, with Japanese SI as a built-in example |
Takumi
|
Orchestration Core
|
+-------------------+-------------------+
| | |
Skills Tools Workflows
| | |
+-------------------+-------------------+
|
Runtime API
|
+-----------------+-----------------+
| | |
Pi DeepSeek CLI
| |
More runtimes Any harness
The same principle applies to where work comes from (ADR-006):
Work request
|
TaskBoardProvider
|
+-------------+---------+---------+-------------+
| | | |
GitHub GitLab Jira Notion Redmine ...one MCP/REST
| | | | | adapter
Issues + PRs Issues + MRs Status select prop Status workflow
|
└──> DeliveryProvider (ADR-007)
|
GitHub / GitLab — plain push, one pull request, checks, merge of the
reviewed commit only; no capability, no silent no-op
A board keeps its own delivery state (a label, a status, a column); Takumi keeps the execution evidence. A provider declares what it can do, and a capability it does not have is an explicit error, never a silent no-op. Delivery is a separate port because the two are genuinely different systems: a Jira board can pair with a GitHub delivery, and a Notion database has no pull request at all.
takumi/
├── apps/cli/ # takumi CLI: init, run, loop, runtime list, extension list
├── apps/console/ # lightweight web console with SSE live logs
├── packages/core/ # workflow engine, runtime API, artifacts, traceability, MEA loop, sandbox
├── runtimes/ # fake, pi, deepseek, cli adapters
├── boards/ # task-board providers: fake, github, gitlab, jira, notion, redmine
├── deliveries/ # delivery providers: fake, github, gitlab
├── extensions/ # skills, tools, workflows
├── eval/ # agent reliability evaluation tasks
├── bench/ # system benchmark baselines
├── examples/ # end-to-end delivery examples, including Japanese SI
└── docs/ # ADRs (006-013: boards, delivery, rails, pilot, metrics, filing, scope, review), notes
- Small Core: orchestration, task/event/artifact models, runtime abstraction, extension loading, approval, and audit.
- Four extension kinds: Skill, Tool Plugin, Workflow Plugin, Runtime Adapter.
- Board agnostic: work comes from a
TaskBoardProvider(GitHub, GitLab, Jira, Notion, or one MCP/REST adapter for the rest). The board owns the delivery state; Takumi owns the execution evidence. - Harness agnostic:
runTask,cancel,getStatus,getUsage,getArtifactsover a unified event stream. - Traceability by default:
REQ-001 -> DESIGN-001 -> UT-001 -> EVIDENCE-001. - Human in the loop: workflows can declare approval gates.
- Independent verification: agent claims are not trusted as completion evidence.
| Capability | Status |
|---|---|
| Pluggable runtimes: fake / Pi / DeepSeek / CLI bridge | Done |
| Declarative workflows, including Japanese SI / V-model | Done |
| Skills for requirements, basic design, detailed design, tests, evidence, review | Done |
| Traceability matrix generation | Done |
| Approval gates | Done |
| Durable resume from audit records | Done |
| Parallel workflow step execution | Done |
| Sandbox abstraction with unshare support | Done |
| Web console with live logs | Done |
| Task-board providers: fake / GitHub / GitLab / Jira / Notion / Redmine, one shared contract suite | Done |
Read-only board view: takumi board --provider <id> |
Done |
| Delivery providers: fake / GitHub / GitLab — plain push, one PR, merge of the reviewed commit only | Done |
Runnable board -> delivery -> merge demo with in-memory providers: node scripts/board-delivery-demo.mjs |
Done |
Pilot safety rails (ADR-008): exclusive slot lock, closed event registry, state bootstrap (takumi board --check/--bootstrap) |
Done |
Pilot tick (ADR-009): takumi pilot --once — select, lock, worktree, agent, deliver, review; systemd/cron shape |
Done |
| Pilot metrics (ADR-010): JSON counters + a Prometheus textfile, and progress writes throttled | Done |
Filing work (ADR-011): createWork on every board, idempotent by marker; a red pipeline files its own item |
Done |
Scoping a tick (ADR-012): query scope through each board's own search, fail-closed where it cannot search |
Done |
| OpenHands as an agent (spike report): proven through the existing pilot seam, zero new abstraction | Done |
Deterministic reviewer (ADR-013): reviewMode: rules — test-weakening and protected paths blocked, fail-closed review |
Done |
| Committed-credential rule (ADR-019): a secret in the diff blocks the delivery and asks for ROTATION, not deletion | Done |
Sidecar reviewer (ADR-021): reviewer: semgrep — a rule set pinned in the repository, a pinnable tool version, findings with file:line, and a refusal (never a clean review) when it cannot run |
Done |
Board mirrors (ADR-017): pilot.boardMirrors — one authority, N write-only projections (state, text, labels, comments) onto e.g. a Notion database, with takumi pilot --resync to rebuild them; verified against a live Notion database |
Done |
A page's body as CONTENT (ADR-023): Notion writes spec.body as paragraph blocks, reads it back on every read (readContent), and repairs an incomplete page append-only, so a rebuild really rebuilds |
Done |
Report surface (ADR-022): reporter: reviewdog — findings posted on the host's review surface, diff-filtered, with a NOT_RUN that is recorded instead of a delivery that fails |
Done |
Workflow review (ADR-024): a rule_review step runs the deterministic rules on the delivery path (a block OR a human finding stops it), the quality gate fails closed on output it cannot parse, and a reviewer's verdict is a structured REVIEW_VERDICT: marker — the prose decides nothing, a missing verdict is not a pass |
Done |
| Agent Eval with held-out verifier tests and repair loop | Done |
| MEA loop: Manage, Execute, Audit | Done |
| Golden Path E2E: Spring Boot + Vue inventory system, 53 Java files, 59 tests green | Done |
Takumi ships an Agent Eval framework under eval/: 23 real SWE tasks across
easy, hard, trap, no-self-test, complex, and implicit constraint groups. Each
task has machine-verifiable verifier tests that are excluded from the agent
prompt context. Agent-written tests are never treated as the ground truth.
TAKUMI_EVAL_HARNESS=deepseek TAKUMI_EVAL_MODEL=deepseek/deepseek-v4-flash node eval/scripts/run-eval.mjs
TAKUMI_EVAL_HARNESS=deepseek TAKUMI_EVAL_MODEL=deepseek/deepseek-v4-pro node eval/scripts/run-eval.mjs --tasks=ts-complex
TAKUMI_EVAL_HARNESS=pi node eval/scripts/run-eval.mjs| Model | First-pass | False completion | Repair triggered |
|---|---|---|---|
| deepseek v4-flash | 23/23 | 0% | 0 |
| deepseek v4-pro, complex | 4/5 | 0% | 1 natural repair |
| Pi, opencode-zen | 6/6 | 0% | 0 |
See docs/bugs-fixed.md for every defect found in the board and delivery layers — symptom, root cause, the commit that fixed it and the test that pins it — docs/evaluation.md for methodology, raw caveats, and docs/acceptance-report.md for the acceptance gates (what each guard claims, the test that pins it, and which layers have run against a real host).
Takumi's Manager Loop follows the Manage-Execute-Audit pattern:
- Manager owns persistent task state and decides the next subtask.
- Executor works in a fresh context and may modify the environment.
- Auditor independently verifies the resulting state.
- Only clean audit evidence can mark a task record complete.
This is implemented in packages/core/src/manager-loop.ts
and exposed through takumi loop.
Inspired by LongHorizon-Harness:
- GitHub: https://github.com/AMAP-ML/LongHorizon-Harness
- Paper: https://arxiv.org/abs/2608.01964
- Website: https://lh-harness.pages.dev/
- examples/pi10: Golden Path enterprise delivery project based on a Japanese SI / V-model workflow, including requirements, design documents, Spring Boot backend, Vue frontend, unit/integration tests, and evidence summaries.
- examples/minimal-vmodel: smaller V-model example for workflow and traceability experiments.
- examples/openhands-agent.sh: run the OpenHands CLI as a pilot agent command (see
docs/openhands-spike-report.md).
Takumi is currently a Developer Preview.
- Agent Eval results are preliminary: 23 tasks is useful evidence, not a statistically large benchmark.
takumi loopis v1: manager decisions are still LLM-driven and non-deterministic.- Pi runtime is opt-in and depends on an external Pi SDK.
- Traceability currently relies on ID naming conventions rather than structural foreign keys.
- Tool plugins run with user privileges unless a sandbox is explicitly selected.
- Real-repo-scale evaluation is still on the roadmap.
- The board layer (ADR-006) covers the work source; the delivery layer (ADR-007) covers branch, pull request, checks and merge. Boards differ in what they can express, on purpose: a Notion database has no pull requests and cannot EDIT a comment, and it records labels only when
labelsPropertynames the column — so each of those operations fails closed instead of quietly doing nothing. - Providers are exercised live where it matters most and offline everywhere else: the GitLab board, the GitLab delivery and the Notion mirror have run against real instances (a private project and a Notion database, including a live item delivered end to end and an 8-item projection), while GitHub, Jira and Redmine are covered by the offline suites only. Expect to adjust API details on first use — pagination beyond the first page of ITEMS (
listWorkstill reads one page of 100), site-specific status/property names, self-hosted base URLs. - Claiming is not atomic on any of these boards (
atomicClaim: falseeverywhere): two runs sharing one account can both believe they claimed an item, which is why a local slot lock — one runner per item — remains the caller's job. - A conflicting base merge is aborted and handed to the review session; takumi never resolves a conflict by rewriting history.
- Board mirrors (ADR-017) are wired, tested through the CLI against the fake board, and verified LIVE against a Notion database (a real delivery projected, and an 8-item resync that is idempotent). A mirror projects the item's STATE, its TEXT, its LABELS and its COMMENTS; the state record is deliberately not mirrored, since it is the control-flow surface. Notion cannot create its own column options, so the six delivery states must exist in the database before the first projection.
- A Notion read costs one extra request per page (
readContent, on by default) because a page's body is content, not a property:listWorkpays it once per row, andreadContent: falseis the documented way to trade the body for cheaper reads. - The report surface (ADR-022) is verified with reviewdog's
localreporter and through the CLI; posting to a LIVE host needs reviewdog's own-conf(project + merge-request ids) and a token in the environment, which is not wired yet. - The semgrep sidecar (ADR-021) costs ~333MB per host and is a hard dependency for
reviewer: semgrep: a host without it stops rather than merging unverified. Its baseline discrimination on a repository with PRE-EXISTING findings is not covered by a test yet, its timing is measured only on small repositories, and an agent can still silence it with an inline# nosemgrep(the rules that catch a self-written suppression are not written yet). - Redmine needs its
statusMap(CLI:--status-map "ready=New,pr_open=In Progress") and, for the run record, a text custom field (--state-field). An issue whose status maps to nothing is REPORTED with the fix, never silently dropped from the board. - Workflow review (ADR-024): the deterministic rules now gate a workflow (
rule_review), but the workflow path does not yet write ADR-018's review digest into a board state record — that binding belongs to the delivery loop, so a workflow delivery that merges through the pilot gets it and one that delivers directly does not. The rules themselves stay path-based heuristics: a missed test file means a missed finding (the same posture as CI), which is why they are the gate and the prose reviewer only reads.
- Core types, runtime abstraction, and extension discovery
- Workflow engine with DAG, approval gates, retry, and capability validation
- Fake, Pi, DeepSeek, and CLI runtime adapters
- Japanese SI skills and V-model workflow
- Agent Eval and system benchmarks
- Durable resume, parallel execution, web console, sandbox, and MEA loop
- Task-board providers (fake / GitHub / GitLab / Jira / Notion / Redmine) with one shared contract suite
- DeliveryProvider port: branch, plain push, one pull request, checks, merge of the reviewed head
- Delivery adapters for GitHub and GitLab
- One MCP/REST board adapter for the long tail (Backlog, Plane, in-house systems)
- A
deliveries/gitadapter for a bare remote with no review surface (ADR-015: the push is read back from the remote, the base moves forward to exactly the reviewed head, and nothing claims a pull request or a pipeline that does not exist) - A
runtimes/openhandsadapter (ADR-016): events from the--jsonstream, per-task tokens and cost from OpenHands' own conversation accounting, the verbatim transcript as an artifact - Claude runtime adapter
- Excel, Word, and Playwright tool plugins
- Jira and GitHub tool plugins (the board layer already covers issues)
- npm package publishing — prepared but blocked on a name (
takumiis taken on npm) and a token; see docs/RELEASING.md - Real-repo-scale evals
Takumi is released under the MIT License.