DMI-configurator: structured configuration, YAML, and web UI - #122
Conversation
There was a problem hiding this comment.
🟡 Changes recommended
The implementation/documentation contract has a few concrete mismatches (e.g., config version enforcement and outdated plan snippets) plus a dependency-surface cleanup to address before safe approval.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Pull request overview
Adds DMI-configurator, a local, no-build-step web UI and supporting Python configuration pipeline to generate a typed, validated, canonical YAML DMI capture configuration from a model descriptor plus user selections—while keeping the existing runtime configuration interface intact (YAML is a front-end serialization format + compiler path, not a runtime rewrite).
Changes:
- Introduces a structured configuration stack (
DMIConfigschema, YAML parse/dump, validation, compilation, compatibility bridge). - Adds a FastAPI backend + static frontend UI (packaged assets) and a
dmiCLI forui+describe-model. - Adds a new runtime-adjacent selection filter (
filter_by_layers) to support layer-range authored configs.
File summaries
| File | Description |
|---|---|
| tests/test_configurator_api.py | Contract tests for configurator HTTP API + static asset serving. |
| tests/test_configuration_yaml.py | YAML canonicalization + round-trip + golden-file coverage. |
| tests/test_configuration_validation.py | Model-aware validation tests (availability, layer bounds, schedule checks). |
| tests/test_configuration_schema.py | Unit tests for schema dataclasses (LayerSelection, ModelTopology, policies). |
| tests/test_configuration_manifest.py | Descriptor parsing/IO tests and ModelShape translation checks. |
| tests/test_configuration_introspect.py | Tests for deriving descriptors from HF-style configs and rejection cases. |
| tests/test_configuration_compatibility.py | Tests for legacy selection bridge + compilation + new layer filtering. |
| tests/golden/qwen3-moe.yaml | Golden config artifact for MoE model case. |
| tests/golden/qwen3-basic.yaml | Golden config artifact for minimal/basic config case. |
| tests/golden/qwen3-attention.yaml | Golden config artifact for attention-over-range config case. |
| tests/data/moe-decoder.model.yaml | Synthetic MoE descriptor fixture for availability testing. |
| src/dmi/ui/static/styles.css | New UI styling (light/dark, layout, state chips, issue rendering). |
| src/dmi/ui/static/index.html | New single-page UI shell and controls layout. |
| src/dmi/ui/static/architecture.js | SVG architecture renderer driven by server-provided metadata. |
| src/dmi/ui/static/app.js | UI state management + server round-trips for validate/parse/serialize/save. |
| src/dmi/ui/server.py | Uvicorn entrypoint wrapper for serving configurator. |
| src/dmi/ui/app.py | FastAPI app + endpoints delegating to configuration layer; static mounting. |
| src/dmi/ui/init.py | Lazy imports so dmi.ui doesn’t require FastAPI at import time. |
| src/dmi/hooks/selection.py | Adds hook_belongs_to_layers + filter_by_layers (additive runtime filter). |
| src/dmi/configuration/yaml.py | YAML parse/dump + canonicalization + disk IO helpers. |
| src/dmi/configuration/validation.py | Validation producing field-addressed issues for UI display. |
| src/dmi/configuration/schema.py | Canonical typed configuration + descriptor schema definitions. |
| src/dmi/configuration/manifest.py | Descriptor parsing/IO and translation to DMI ModelShapeConfig. |
| src/dmi/configuration/introspect.py | Descriptor derivation from HF configs/model ids with lazy transformers import. |
| src/dmi/configuration/errors.py | Configuration-layer exception types. |
| src/dmi/configuration/compiler.py | Compiles DMIConfig to CompiledDMIConfig via existing selector + layer filter. |
| src/dmi/configuration/compatibility.py | Bridge between structured observations and legacy selection string. |
| src/dmi/configuration/catalog_adapter.py | UI-facing hook metadata projection + availability/reason computation. |
| src/dmi/configuration/architecture.py | Defines diagram nodes and produces model/layout payloads for the UI. |
| src/dmi/configuration/init.py | Consolidated public API exports for configuration subsystem. |
| src/dmi/cli.py | Adds dmi ui and dmi describe-model CLI commands. |
| pyproject.toml | Adds [ui] optional deps, dmi console script, and UI static package data. |
| examples/model_descriptors/README.md | Documentation for descriptors + how to generate/use them. |
| examples/model_descriptors/llama3-8b.yaml | Shipped example descriptor used by tests/UI. |
| docs/dmi-configurator-plan.md | Design/implementation plan and rationale for configurator architecture. |
Review details
Suppressed comments (2)
docs/dmi-configurator-plan.md:244
- The CompiledDMIConfig snippet documents a
runtime: RuntimeConfigfield, but the implementedCompiledDMIConfig(src/dmi/configuration/compiler.py) only carrieshook_specs,schedule, and optionalpolicy. Keeping the snippet aligned avoids implying runtime transport settings are part of config compilation.
@dataclass
class CompiledDMIConfig:
hook_specs: list[HookSpec]
schedule: CaptureSchedule
runtime: RuntimeConfig
policy: RuntimePolicy | None
docs/dmi-configurator-plan.md:255
- The
compile_configexample callsfilter_by_layers(specs, config.observations.layers), but the implemented function takes explicitstart/endints (and compilation also no longer passes aruntimefield). Update the snippet so readers can copy/paste it without hitting a TypeError or referencing non-existent fields.
def compile_config(config, model_context):
specs = select_hook_specs(
model_context.specs,
",".join(config.observations.hooks),
cfg=model_context.shape,
)
specs = filter_by_layers(specs, config.observations.layers)
return CompiledDMIConfig(
- Files reviewed: 35/35 changed files
- Comments generated: 3
- Review effort level: Lite
💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
There was a problem hiding this comment.
🟡 Changes recommended
The backend save endpoint and YAML parser currently mishandle error/type cases (uncaught ConfigurationError in /api/config/save and or {} / or [] defaulting that can silently bypass structural validation).
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Review details
- Files reviewed: 48/48 changed files
- Comments generated: 3
- Review effort level: Lite
There was a problem hiding this comment.
🟢 Approval recommended
The changes are additive, well-covered by CPU-marked tests across schema/YAML/validation/UI API surfaces, and the runtime layer-range wiring is exercised end-to-end.
Review details
Suppressed comments (1)
Previously missed (1) — in code that hasn't changed since the last review.
pyproject.toml:20
- Optional:
tests/test_configurator_api.pyimportsfastapi.testclient.TestClient, which depends onhttpxin many FastAPI/Starlette setups. If someone installs only the[ui]extra, the UI will work but running the API contract tests may fail with anImportErrorifhttpxisn’t pulled transitively; consider addinghttpxto theuiextra (or adding an explicitimportorskip("httpx")guard in the tests) to make the optional stack more self-contained for contributors.
- Files reviewed: 48/48 changed files
- Comments generated: 0 new
- Review effort level: Lite
|
All 14 inline findings are addressed at dd3bb32 (responses on each thread). Two notes: the vLLM layer-range fix (3919326699) needs a matching revision of the pinned third_party/vllm-integration submodule — in this repo the failure is now a diagnosable error rather than a silent range drop. The 6 Copilot findings were already fixed at c8875a8/de37754 and are pinned as regression tests in tests/test_pr122_review_findings.py. CPU suite at head: 1373 passed. |
Samfisheryu
left a comment
There was a problem hiding this comment.
Requesting changes for confirmed runtime and estimator correctness issues. I reproduced each finding inline. The existing CPU suite is healthy (1,387 passed), and the supported Hugging Face torch.compile schedule guard itself behaves correctly; the blocking gaps are configuration application, adapter ownership, and packed/vLLM parity.
|
Independent review round: 8 adversarial reviewers, all confirmed findings fixed at 3dc68ca. Review lenses: config pipeline, HTTP security/correctness, schedule-gate runtime, estimator fidelity, frontend JS, descriptor/schema integrity, test honesty (mutation-tested), PR-body coherence. Every finding was verified by execution before being flagged; every confirmed one is fixed with a test in this commit. Fixed:
Known ceilings (disclosed, deferred): the vLLM integration's direct commit path is not schedule-gated (separate repo, DMI-vLLM-Integration#21 adds the layers keyword; gating is next); torch.compile/CUDA-graph replay bakes the gate at first trace (compiled decode should stay on default schedules until enforcement moves native). CPU suite: 1414 passed. |
Samfisheryu
left a comment
There was a problem hiding this comment.
Thanks — I independently re-ran the original repros on 27c40b3. Several fixes are substantive (live layer rejection, design-time labelling, save serialization, ordinary duplicate-key and hook-type validation, port handling, and clean CLI errors). Keeping changes requested for the still-incomplete runtime capability boundary and the confirmed residual implementation bugs below.
|
All 8 follow-up findings are fixed at 450b340 (replies on each thread), with regression tests in tests/test_pr122_round4_review_findings.py. CPU suite: 1679 passed; both CI checks green. Two known ceilings remain disclosed in code rather than fixed here: the vLLM direct-commit path still bypasses the schedule gate (needs the graph-safe integration change, separate repo), and the token-range index records attempted rather than captured traffic. |
|
All review threads on this PR are now resolved (43/43: 6 Copilot, 14 XbzOnGit, 23 Samfisheryu), each with its fix commit, mechanism, and pinning test recorded on the thread. Since the last review round: packed estimator honesty, the verbatim vLLM partition rule, one-owner attach enforcement, fail-closed encoder subtypes, ring/workload/YAML boundary hardening, offset/warmup disclosure, plus the CUDA-graph schedule disclosure, CLI error narrowing, and selection taxonomy. Suite: 1679 passed; both CI checks green on 450b340. Requesting re-review. |
The review round surfaced three rules that deserve a written rationale rather than a thread of replies: the configuration artifact is strict end-to-end; the runtime surface is enforced at the smallest possible layer (the adapter driver's before_forward) rather than scattered across every emit point; and the HTTP surface trusts the loopback bind but refuses the network without a per-launch X-DMI-Token. The one-owner attach invariant is the same rule applied to a model. The alternatives weighed (stricter loopback defaults, TLS/OAuth on the local tool, rejecting packed configs that name a schedule, hooking the gate at every emit point, an inline JSON or frontmatter idiom) all get a sentence here so the next round does not relitigate them.
* Accept a layer range in VLLMAdaptor.attach_model DMI-configurator (ProjectDMX/DMI#122) forwards a selected layer range as attach_model(model, hook_selection, layers=LayerSelection(...)). The adaptor refused any ranged configuration with TypeError, and Packed/vLLM is the configurator's default backend. The keyword is applied in _VLLMHookSelection.from_model so the filter hits BOTH sides: the installed local specs and the model-wide candidate-rank formula sets. A signature-only fix would let the per-rank byte plans disagree with what the filtered local specs actually capture. Global hooks (layer_no < 0) are never restricted, matching the HF adapter. A range that matches nothing raises a DMI ConfigValidationError naming the range instead of silently compiling to an empty selection. Closes #20. Covered by tests/test_attach_layer_range.py, which runs the selection logic on CPU behind a vLLM import stub, so CI needs no vLLM install; the six existing suites that need the real release were already failing on this machine before this change (no vllm distribution) and are unaffected. * Keep the adapter importable and attachable on DMI without the layer-range API The official-vLLM CI matrix runs this repo against released DMI (v1.1.0) and DMI main, neither of which exports hook_belongs_to_layers -- the module-level import broke collection of every test module there. The layer-range API imports now live inside the layers branch (with a clear upgrade-DMI error when a range is passed without them), and attach_model forwards the keyword to super() only when a range is configured, so un-ranged attaches behave exactly as before on old DMI. The new tests skip cleanly where dmi.configuration is absent and pin the old-DMI contract by simulating the missing facade name. * Stub vLLM in tests only when vLLM is not installed The old guard (sys.modules + hasattr distributed) skipped the stub only when vLLM was already imported; an installed-but-unimported vLLM got shadowed for the whole session, so other tests' view of vllm depended on collection order. importlib.util.find_spec decides on installation, not import state: real vLLM always wins, the stub only exists for CPU CI where vLLM is absent. The suite's assertions are behavior-based on the DMI side and hold under either vLLM. Review thread: copilot on PR #21 (test_attach_layer_range.py:23). --------- Co-authored-by: Alan Liu <zaoxing@users.noreply.github.com>
There was a problem hiding this comment.
🟡 Changes recommended
BackendAdapter.before_forward() does not advance the step counter when build_step_context() returns None, which can break warmup/stride accounting and capture steps more often than the configured schedule.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Review details
- Files reviewed: 63/63 changed files
- Comments generated: 2
- Review effort level: Lite
There was a problem hiding this comment.
🟡 Changes recommended
The public integration API docs currently mis-specify the filter_by_layers signature/return type, which can mislead downstream integrators even if runtime behavior is correct.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Review details
- Files reviewed: 63/63 changed files
- Comments generated: 3
- Review effort level: Lite
…rrowing, selection taxonomy - A non-default capture schedule now disables CUDA-graph compilation for the generate() call, mirroring the capacity-overflow path: the gate is a Python branch in HookPoint.forward and replay bakes it at the first decode trace, so a refusal (or capture) at step one would apply to every later step. Default schedules are unaffected. Pinned by driving generate_with_monitoring with a stubbed adapter and inspecting the kwargs the model receives. - cli.main no longer catches bare RuntimeError -- that class carries genuine bugs (RecursionError is one) and swallowed them into one stderr line. The optional-dependency failure gets its own UIDependencyError; internal raise sites updated. - compile_config re-raises an unknown hook selection as ConfigurationError instead of bare ValueError, so callers funneling on the taxonomy (the UI's 400 mapping) see a 400, not a 500. CPU suite: 1588 passed (after the main merge).
…hip, boundaries Samfisheryu's follow-up review kept CHANGES_REQUESTED on eight residuals; all are fixed here with tests, RED first. Ownership: - attach_config refuses a second adapter on an already-owned model BEFORE any mutation (no schedule install, no attach call), naming the current owner; same-adapter reconfiguration is allowed; detach by a non-owner leaves the marker (HF adapter path pinned). Packed honesty (the vLLM runtime executes no schedule): - capture_prefill/capture_decode no longer shrink packed volume, peak, or sustained figures -- the runtime captures both phases regardless. The authored intent stays visible as an explicit warning; batched figures are unchanged. - vLLM PP partition rule ported from 0.27.1's get_pp_indices (remainder backward from the second-to-last stage, never the last; cited cases [1,2,1], [2,3,3,2], 32/40 verified) plus an explicit, validated pp_layer_counts override for VLLM_PP_LAYER_PARTITION deployments. - Offset/warmup disclosure no longer depends on a stride being set. Boundaries: - Ring: only absent/None defaults pinned_bytes; every explicit value is validated (exact non-bool int, >= 0), and check_ring_fit refuses negatives itself. - Workload.cache_max_len requires an exact integer. - YAML: unhashable mapping keys raise ConstructorError (a 400), and the descriptor file path uses the same strict loader as configs. - Introspect: any model_type containing "encoder" fails closed instead of riding a decoder family prefix (qwen2_audio_encoder); decoder+head configs still accepted. CPU suite: 1679 passed.
- The per-request token-range index records every driven step including schedule-refused ones; it is the attempted-traffic index, not the captured-traffic index. - The _accepts_layers probe cannot distinguish forwarding **kwargs from tolerating them; the backstop is the v1 contract plus compile-time live-hook rejection, stated where the probe lives.
Attention-weight hooks (pattern, attn_scores) attend at most window KV positions on sliding-window models, so they are shaped against min(context, window) instead of the full context. Only kv_dim is affected, and only those shapes consume it. Exact-int >= 1 validation, assumption text when set, no-op when absent or above context. FP8 KV-cache quantization was investigated and deliberately not modeled: hooks capture forward activations in model dtype, so cache-store quantization changes zero hook payload bytes.
…mates" The knob's core semantic was incorrect in the unsafe direction: compute_hook_shape returns full [heads, q_len, kv_dim] for pattern/attn_scores with no window input, plan_step reserves exactly that, and these hooks fire meaningfully only in eager mode -- where SDPA materializes the full scores matrix and masks it rather than shrinking it. Capping kv_dim by the window made the estimate report less than the runtime reserves (up to context/window under-count), and check_ring_fit would bless a ring the runtime overflows. The FP8 non-modeling analysis in the original message stands; the window half is withdrawn until a backend materializes band-shaped scores.
…y banner - Save disables itself for the request round trip (re-armed in finally). - Layer rail ticks are keyboard-operable checkboxes sharing one picking function with the click path; architecture nodes already had this. - Estimate panel marks itself loading with aria-busy, cleared on either guarded outcome. - Status chip is a polite live region; ring meter exposes value semantics. - Token-gated binds get a persistent curl-only banner revealed on 401 instead of per-panel curl-hint errors.
- _reject_unknown gains an error= kwarg so the config half (ConfigurationError) and the descriptor half (DescriptorError) share one body. As a bonus this stops the int-YAML-key AttributeError the descriptor half had -- the manifest copy already used repr() but the config half still did not. - save_config and save_descriptor share _write_text_atomic(path, text, what, error): the only delta is which exception class wraps an OSError. CPU suite: 1685 passed.
The review round surfaced three rules that deserve a written rationale rather than a thread of replies: the configuration artifact is strict end-to-end; the runtime surface is enforced at the smallest possible layer (the adapter driver's before_forward) rather than scattered across every emit point; and the HTTP surface trusts the loopback bind but refuses the network without a per-launch X-DMI-Token. The one-owner attach invariant is the same rule applied to a model. The alternatives weighed (stricter loopback defaults, TLS/OAuth on the local tool, rejecting packed configs that name a schedule, hooking the gate at every emit point, an inline JSON or frontmatter idiom) all get a sentence here so the next round does not relitigate them.
The yaml.py/manifest.py simplification shipped first without its test updates, which broke CI: the old atomic-save test raised ENOSPC from dump_config, and the new save_config evaluates dump_config before _write_text_atomic's try-block, so the OSError escaped unwrapped. The test now patches Path.write_text on the .tmp sibling instead, which is the real I/O boundary the atomic write protects. Rest of the round: - UI_DEPENDENCY_MESSAGE moves to errors.py beside UIDependencyError, so app.py and server.py quote one string. - The auth banner reuses .notice plus one margin-top override instead of duplicating seven properties. - Copy calls a shared serializeCurrentState() rather than reissuing the /api/config/serialize POST shape refreshOutput already builds. - The two test-side brace matchers (_function_body / _plain_function_body) collapse into one with an optional async prefix. CPU suite: 1685 passed.
DMI-vLLM-Integration PR #21 merged (squash 23717cb): attach_model accepts layers=LayerSelection(...), applies it to the local specs and the model-wide candidate-rank sets, and stays importable on DMI builds without the layer-range facade. The pin moves old-main 29f26c3 -> 23717cb, picking up both the V2-runner promotion (#19) and the layer range (#21).
The warning told the UI that attach_config refuses a ranged configuration on Packed (vLLM) until the integration accepted ranges -- that is PR #21, now merged and pinned. The estimator test flips to pin the absence of the warning (RED against the stale block), and attach_config's fail-closed guard stays for genuinely older integration checkouts, with its comment updated to name them. CPU suite: 1689 passed.
…vLLM The recipe for proving a saved .dmi.yaml artifact drives a real vLLM 0.27.1 model through dmi.configuration.attach_config existed only in the verification session: build the CI-style env (vllm 0.27.1 + DMI --no-deps + integration --no-deps), subclass the V1 DMXGPUWorker to suppress the self-attach (V2 refuses subclasses), run the engine core in-process, and assert ownership, ranged spec layers, disabled out-of-range hook points, and a completed generate. Also correct the install step for shared checkouts: the native build emits ABI-suffixed extensions per interpreter, so 'make clean' is destructive there.
a2ed118 to
4def7cc
Compare
dmi ui/describe-model needed transformers installed by hand to resolve a model directory or bare HF model id; pip install -e ".[hf]" makes that reproducible, and the missing-transformers error now names it.
Keep the top-level entry point consistent with the configurator plan doc.
Samfisheryu
left a comment
There was a problem hiding this comment.
Re-review at 393f4f3: 354 tests passed across the eight relevant configuration/runtime/estimator/API regression suites. The pinned vLLM adapter now accepts and forwards layers.
One confirmed regression remains beyond the already disclosed feature limitations: the previously fixed None-context step accounting and its tests are absent from this head. I added the old/new execution results to the existing inline thread: the author's original four tests pass on 74d7be7, but three fail on 393f4f3. Please restore that fix; the prior "fixed" reply does not describe the current tree.
The vLLM direct-commit schedule gap and attempted-vs-captured token-range behavior remain acknowledged limitations, not newly reported implementation bugs. This recheck does not claim a full vLLM/HF graph or multi-GPU run.
Two conflicts, both where main's #127 (native capture write path) and this branch added to the same block; both resolved as the union, since the two sides configure unrelated things: - .github/workflows/python-checks.yml: main added the botocore/boto3 pip lines the native capture suites need at collection time; this branch added pyyaml and the [ui] extra (fastapi/uvicorn/httpx). Both jobs need both sets. Keeping only one side would trip main's own new gate, which fails the build when a cpu test skips for anything but absent hardware -- the configurator's API contract suite skips itself without fastapi/httpx, and the native suites error at collection without botocore. - Makefile: main added PYTEST_ARGS (so CI can pass --junit-xml to the gate); this branch added MODEL for the `make ui` target. Independent variables in the same header block. No src/ file was touched by both sides, so nothing else conflicted.
Samfisheryu
left a comment
There was a problem hiding this comment.
Full re-review at badbfca: no new reproduced implementation blocker beyond the previously acknowledged runtime schedule limitations. The restored None-context step-accounting fix is retained and its regressions pass; the pinned vLLM adapter accepts and forwards layers.
Validation: 2,038 CPU tests passed; 16 real single-GPU native/reference/ClickHouse Ring tests passed; GPT-2 HF eager and CUDA Graph correctness E2Es passed (2 tests, 56 subtests). Both CI checks are green. These E2Es validate their tested/default capture paths, not unsupported schedule enforcement.
Keeping the existing requested changes for the core capability boundary: vLLM's direct commit_step path still bypasses schedule gating, and non-default HF schedules still disable compilation instead of implementing graph-safe enforcement. The attempted-vs-captured token-range behavior also remains acknowledged. I am not treating documentation nits as blockers; this run does not claim full vLLM or multi-GPU E2E coverage.
The /polish loop's working files are per-run agent bookkeeping, not project artifacts: nothing in the tree, the build, or CI reads them, and they go stale once the branch they describe has landed. #133 removes them from main; dropping them here too so a merge of this branch cannot carry them back in.
_FakeEagerTransport was a hand-copied RingTransport, and it drifted: it lacked capture_step until the previous commit, so five GPU tests raised AttributeError after #122 added the capture-schedule gate. It also reimplemented production logic by precomputing effective_cap instead of going through the transport's cached property. The tests now build a real RingTransport(engine) over the fake engine, set force_eager, and replace only submit_cpu_direct with a list append. tests/test_hook_point_eager_cap_cache.py already does this. The fake engine gains the payload_tensor() the constructor needs, and each hook writes into that payload. Any attribute RingTransport gains in future is therefore present in these tests, and effective_cap is the production property. Every test's assertions are unchanged, and the now-unused effective_ring_bytes import is gone. GPU (RTX 4090, svc worktree, point.py and ring.py identical to origin/main): 20 passed. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
Nothing covered #122's capture-schedule gate on the eager path. When the schedule refuses a step, the driver skips plan/commit, so no ring space is reserved and no meta is pushed. If a producer still ran, it would write unreserved bytes and desync the task/meta FIFO. HookPoint.forward checks transport.capture_step after the `if not self.enabled` early return and before the CUDA probe, so a refused step never reaches the force_eager block. The new test arms a real RingTransport with force_eager=True and capture_step=False. It then calls the hook with one tensor over staging (flush + cpu_direct) and one that fits staging but not the current slack (flush + reserve + ring). It asserts that nothing was flushed, reserved, submitted to cpu_direct or dispatched to a producer. GPU (RTX 4090, svc worktree, point.py and ring.py identical to origin/main): 21 passed. On a scratch copy of the package with the gate moved below the eager block, the new test fails with flushes == 2, and the four counters read flushes=2 reserved=[32] direct=1 dispatched=1. The other 20 tests still pass there. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
* Give the eager-transport test fake the capture-schedule gate #122 added a capture-schedule gate to HookPoint.forward: it reads transport.capture_step on the active transport before anything else. The real RingTransport sets it to True. This file's _FakeEagerTransport never gained the attribute, so all five eager-path tests have raised AttributeError at hooks/point.py since #122, before reaching the capacity and staging behaviour they exist to pin. They are marked gpu, and CI has no GPU runner, so nothing flagged it. The fake now arms the gate the way the real transport does. GPU (RTX 4090): 5 failed, 15 passed before; 20 passed after. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y * Build the eager-path tests on a real RingTransport _FakeEagerTransport was a hand-copied RingTransport, and it drifted: it lacked capture_step until the previous commit, so five GPU tests raised AttributeError after #122 added the capture-schedule gate. It also reimplemented production logic by precomputing effective_cap instead of going through the transport's cached property. The tests now build a real RingTransport(engine) over the fake engine, set force_eager, and replace only submit_cpu_direct with a list append. tests/test_hook_point_eager_cap_cache.py already does this. The fake engine gains the payload_tensor() the constructor needs, and each hook writes into that payload. Any attribute RingTransport gains in future is therefore present in these tests, and effective_cap is the production property. Every test's assertions are unchanged, and the now-unused effective_ring_bytes import is gone. GPU (RTX 4090, svc worktree, point.py and ring.py identical to origin/main): 20 passed. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y * Test that a refused capture step skips the eager safety net Nothing covered #122's capture-schedule gate on the eager path. When the schedule refuses a step, the driver skips plan/commit, so no ring space is reserved and no meta is pushed. If a producer still ran, it would write unreserved bytes and desync the task/meta FIFO. HookPoint.forward checks transport.capture_step after the `if not self.enabled` early return and before the CUDA probe, so a refused step never reaches the force_eager block. The new test arms a real RingTransport with force_eager=True and capture_step=False. It then calls the hook with one tensor over staging (flush + cpu_direct) and one that fits staging but not the current slack (flush + reserve + ring). It asserts that nothing was flushed, reserved, submitted to cpu_direct or dispatched to a producer. GPU (RTX 4090, svc worktree, point.py and ring.py identical to origin/main): 21 passed. On a scratch copy of the package with the gate moved below the eager block, the new test fails with flushes == 2, and the four counters read flushes=2 reserved=[32] direct=1 dispatched=1. The other 20 tests still pass there. Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y --------- Co-authored-by: Alan Liu <zaoxing@users.noreply.github.com>
What this is
DMI-configurator: a local web tool that turns a model descriptor plus a few visual selections into a validated DMI capture configuration.
One job:
The architecture visualization is the interaction surface; the generated YAML is the product. This is a configuration tool, not an observability dashboard.
See
docs/dmi-configurator-plan.mdfor the full design, its verification against the codebase, and the decisions taken.The central architectural decision
DMI does not switch from Python configuration to YAML. YAML is added as a serialization format in front of the existing mechanisms, via a loader/compiler pair:
The separation is the point. What is explicitly rejected:
yaml.safe_load()producing a dict that every module then indexes into. One parser, one typed object, everything else downstream.The frontend holds no DMI semantics — it derives its choices from the hook catalog through a backend adapter, and every validity answer is a server round-trip through the same Python the runtime uses. What the UI calls valid is what DMI calls valid.
What landed
dmi/configuration/schema.pydmi/configuration/manifest.pydmi/configuration/catalog_adapter.pydmi/configuration/architecture.pydmi/configuration/validation.pydmi/configuration/yaml.pydmi/configuration/compatibility.pydmi/configuration/compiler.pydmi/configuration/estimate.py(+POST /api/estimate)dmi/hooks/selection.py(filter_by_layers)dmi/ui/app.py,dmi/ui/server.pydmi/ui/static/(no build step)dmi/configuration/introspect.pydmi/cli.pydmi/adapters/base.py,dmi/adapters/types.py,dmi/hooks/point.py,dmi/transport/ring.pyPhases 1–7 of the plan, plus #121's runtime wiring (layer ranges applied at attach, estimator, launch ergonomics) and capture-schedule enforcement in the adapter driver. Phase 8 (runtime policy) is deferred by design.
Compatibility
The public API-v1 facade grows three exports (
filter_by_layers,hook_belongs_to_layers,LayerSelection);filter_by_layers/hook_belongs_to_layersfollow the same convention asfilter_by_pp_rank/filter_by_tp_rank.attach_modelgained a keyword-onlylayers=...parameter on the base class and the HF adapter — call sites that never pass it are unaffected, andattach_configdiagnoses an adapter that cannot accept it rather than crashing.The capture hot path does change behavior, beyond the additive filter:
StepContextgained aphasefield the HF adapter reports,before_forwardnow consults the capture schedule before every step (phase flags, strides, warmup — previously authored and ignored), the transport disarms producers on schedule-refused steps (a driver-level skip alone left HookPoints writing unreserved bytes), andHookPoint.forwardhonors that flag. Steps driven outsidebefore_forward(e.g. the pinned vLLM integration's direct commit path) are not schedule-gated yet; the estimator discloses what its reductions assume.The flat
dmx_hook_selectionform stays supported through a compatibility adapter rather than being made canonical.pyyamlandpydanticwere already core dependencies;fastapianduvicornare new behind a small[ui]extra. A normal DMI install pulls in no web framework.Deliberate constraints
runtimeblock in the YAML. The existing runtime surface isRingConfig, a native struct of transport parameters. Exposing a subset would imply a support contract that does not exist.layers.endis inclusive.LayerSelection(8, 15)selects eight layers; the UI labels it "Layers 8–15" and the label must not lie.decoder_transformer. Encoder-decoder configs and known non-decoder families (BERT/ViT/T5-style) are refused rather than mis-rendered.Verification
pytest -m cpu: 1414 passed (descriptor, configuration, serialization round-trip, integration, golden-file suites, plus the schedule-gate and API-contract suites).parse(serialize(config)) == configafter canonical normalization — what makes the YAML a configuration artifact rather than a one-way export.🤖 Generated with Claude Code (review round included 8 independent adversarial reviewers; all confirmed findings fixed on this branch)