Problem
Original text:
「我發現接下來一個重要的東西是看一個功能的 speed and accuracy,要能夠讓一個功能最快最準是目標,我在想是要在idd還是要另外創立一個profiling 的plugin呢」
— Source: maintainer, /idd-issue invocation, 2026-10-05
The next goal is to measure a feature's speed and accuracy, and to drive it toward being as fast and as accurate as possible. Before building anything, one placement decision has to be made: does this live inside IDD (as part of the issue lifecycle), or in a separate profiling plugin?
Type
feature — with an open placement decision (to be settled in diagnose / discuss before any implementation)
Expected
- A repeatable way to measure, for one feature, how fast it is and how accurate it is, so that "fastest and most accurate" becomes a checkable target instead of an impression.
- A recorded decision on where this capability lives: an IDD lifecycle step or skill, a standalone plugin that IDD consumes, or a split between the two.
Actual
Context relevant to the placement decision
- Name collision:
idd-verify --profile (code / prose / academic) already exists and means a reviewer-lens configuration, not performance profiling.
idd-route record writes routing-stats.jsonl — process metrics (round trips, findings per verify), not metrics of the feature itself.
- Measurement-first tools already exist outside IDD, each for one domain:
bestocr, bestasr, llm-context-benchmark, parallel-ai-agents:ensemble-eval.
- The dev rule
deep-integration-over-hardcode applies if the answer is "separate plugin": IDD would declare a dependency on it rather than carry its own copy.
Impact
Every feature whose output can be right or wrong — gate verdicts, classifiers, OCR / ASR, parsers — currently re-invents its own measurement, and speed is not measured at all. Settling the placement first avoids building the same harness twice.
Priority
P2 — direction-setting; assumed scheduled after the open #316 / PR #318 lands (change on request).
Clarity Surface(idd-clarify run 2026-10-05T08:45:43Z)
| Type |
Source |
Question for you |
Status |
| ambiguity |
「看一個功能的 speed and accuracy」 |
這裡的「一個功能」是哪個層級:一支 helper 函式、一個 skill、一個 MCP tool,還是整個 plugin? |
resolved @ 2026-10-05T09:52:03Z (reason: 使用者未定 → 移交 diagnose 判定層級;首個實例是 che-apple-mail-mcp 的 create_draft(一個 MCP tool 操作,有多種實作方式)) |
| ambiguity |
「speed」 |
speed 要量的是執行時間、token 與費用,還是 API 呼叫次數? |
resolved @ 2026-10-05T09:52:03Z (reason: 分開看;先量執行時間(wall-clock),token/費用與 API 次數之後另看) |
| ambiguity |
「accuracy」 |
accuracy 要用單一個準確率,還是像 #316 那樣把 false positive 和 false negative 分開看(兩個方向的代價不同)? |
resolved @ 2026-10-05T09:52:03Z (reason: FP 與 FN 分開看;要跑很多次(replicate),兩個方向各自估錯誤率) |
| missing-context |
「accuracy」 |
accuracy 的真值從哪裡來:人工判讀的凍結語料、事先寫好的標準答案,還是跨模型一致? |
resolved @ 2026-10-05T09:52:03Z (reason: 大量 replicate;每次結果仍需一個判對錯的 oracle,依功能而定(create_draft 可機械檢查)) |
| ambiguity |
「讓一個功能最快最準」 |
快和準衝突的時候,哪一個優先?還是兩個都列出來讓人自己取捨? |
resolved @ 2026-10-05T09:52:03Z (reason: 看情況;像 create_draft 這類 process 要求完全正確:正確是硬限制,在零錯誤的前提下求最快(不是兩者取捨)) |
| ambiguity |
「profiling 的plugin」 |
IDD 已有 idd-verify --profile(指審查視角的組合),新的效能量測要避開 profiling 這個名字嗎? |
resolved @ 2026-10-05T09:52:03Z (reason: 避開 profiling 這個名字;使用者另提可能有耦合問題(記於 comment)) |
Current Status
Phase: created
Last updated: 2026-10-05 by /idd-comment → /idd-update
Key Decisions
Scope Changes
Blocking
- (none) — next:
/idd-diagnose #369 (placement decision: inside IDD vs standalone tool IDD consumes)
Commits
Problem
The next goal is to measure a feature's speed and accuracy, and to drive it toward being as fast and as accurate as possible. Before building anything, one placement decision has to be made: does this live inside IDD (as part of the issue lifecycle), or in a separate profiling plugin?
Type
feature — with an open placement decision (to be settled in diagnose / discuss before any implementation)
Expected
Actual
idd-verifychecks fidelity — does the diff satisfy the issue — not performance.### Blockingsections) with hand-reviewed truth, false-positive / false-negative counts, and a rule that a reader-rule change must flip a measured corpus row. None of that machinery is reusable outside bug: #298 的修正只落在 idd-list — 另三個 Complexity consumer 未動,且 idd-list 自身 Step 5 與 Step 3.7 互相矛盾 #316.Context relevant to the placement decision
idd-verify --profile(code / prose / academic) already exists and means a reviewer-lens configuration, not performance profiling.idd-route recordwritesrouting-stats.jsonl— process metrics (round trips, findings per verify), not metrics of the feature itself.bestocr,bestasr,llm-context-benchmark,parallel-ai-agents:ensemble-eval.deep-integration-over-hardcodeapplies if the answer is "separate plugin": IDD would declare a dependency on it rather than carry its own copy.Impact
Every feature whose output can be right or wrong — gate verdicts, classifiers, OCR / ASR, parsers — currently re-invents its own measurement, and speed is not measured at all. Settling the placement first avoids building the same harness twice.
Priority
P2 — direction-setting; assumed scheduled after the open #316 / PR #318 lands (change on request).
Clarity Surface(idd-clarify run 2026-10-05T08:45:43Z)
create_draft(一個 MCP tool 操作,有多種實作方式))create_draft可機械檢查))create_draft這類 process 要求完全正確:正確是硬限制,在零錯誤的前提下求最快(不是兩者取捨))idd-verify --profile(指審查視角的組合),新的效能量測要避開 profiling 這個名字嗎?Current Status
Phase: created
Last updated: 2026-10-05 by /idd-comment → /idd-update
Key Decisions
create_draft, correctness is a hard constraint — fastest among fully correct candidates (lexicographic, not a trade-off)idd-verify --profile); coupling concern recorded as a placement inputche-apple-mail-mcpcreate_draft(GUI path vs Experimental create_draft path: write the draft directly into Mail's store, trigger upload by toggling its read status che-apple-mail-mcp#472 direct write)Scope Changes
Blocking
/idd-diagnose #369(placement decision: inside IDD vs standalone tool IDD consumes)Commits