Skip to content

DMI-configurator: structured configuration, YAML, and web UI - #122

Merged
zaoxing merged 32 commits into
mainfrom
feat/dmi-configurator
Sep 22, 2026
Merged

zaoxing merged 32 commits into
mainfrom
feat/dmi-configurator

Conversation

@zaoxing

@zaoxing zaoxing commented Sep 2, 2026 •

Copy link
Copy Markdown
Collaborator

What this is

DMI-configurator: a local web tool that turns a model descriptor plus a few visual selections into a validated DMI capture configuration.

One job:

Turn a model descriptor plus a user's visual selections into a validated DMI runtime YAML configuration.

The architecture visualization is the interaction surface; the generated YAML is the product. This is a configuration tool, not an observability dashboard.

See docs/dmi-configurator-plan.md for the full design, its verification against the codebase, and the decisions taken.

The central architectural decision

DMI does not switch from Python configuration to YAML. YAML is added as a serialization format in front of the existing mechanisms, via a loader/compiler pair:

config.yaml
  -> load_config()      knows YAML, not the runtime
  -> DMIConfig          canonical in-process representation
  -> compile_config()   knows the runtime, not YAML
  -> CompiledDMIConfig
  -> existing DMI runtime

The separation is the point. What is explicitly rejected: yaml.safe_load() producing a dict that every module then indexes into. One parser, one typed object, everything else downstream.

The frontend holds no DMI semantics — it derives its choices from the hook catalog through a backend adapter, and every validity answer is a server round-trip through the same Python the runtime uses. What the UI calls valid is what DMI calls valid.

What landed

Area Module
Typed schema dmi/configuration/schema.py
Descriptors dmi/configuration/manifest.py
Catalog projection dmi/configuration/catalog_adapter.py
Diagram metadata dmi/configuration/architecture.py
Validation dmi/configuration/validation.py
YAML dmi/configuration/yaml.py
Legacy bridge dmi/configuration/compatibility.py
Compilation dmi/configuration/compiler.py
Payload estimator dmi/configuration/estimate.py (+ POST /api/estimate)
Layer filter dmi/hooks/selection.py (filter_by_layers)
Backend dmi/ui/app.py, dmi/ui/server.py
Front end dmi/ui/static/ (no build step)
Descriptor derivation dmi/configuration/introspect.py
CLI dmi/cli.py
Capture-schedule enforcement dmi/adapters/base.py, dmi/adapters/types.py, dmi/hooks/point.py, dmi/transport/ring.py

Phases 1–7 of the plan, plus #121's runtime wiring (layer ranges applied at attach, estimator, launch ergonomics) and capture-schedule enforcement in the adapter driver. Phase 8 (runtime policy) is deferred by design.

Compatibility

The public API-v1 facade grows three exports (filter_by_layers, hook_belongs_to_layers, LayerSelection); filter_by_layers/hook_belongs_to_layers follow the same convention as filter_by_pp_rank/filter_by_tp_rank. attach_model gained a keyword-only layers=... parameter on the base class and the HF adapter — call sites that never pass it are unaffected, and attach_config diagnoses an adapter that cannot accept it rather than crashing.

The capture hot path does change behavior, beyond the additive filter: StepContext gained a phase field the HF adapter reports, before_forward now consults the capture schedule before every step (phase flags, strides, warmup — previously authored and ignored), the transport disarms producers on schedule-refused steps (a driver-level skip alone left HookPoints writing unreserved bytes), and HookPoint.forward honors that flag. Steps driven outside before_forward (e.g. the pinned vLLM integration's direct commit path) are not schedule-gated yet; the estimator discloses what its reductions assume.

The flat dmx_hook_selection form stays supported through a compatibility adapter rather than being made canonical.

pyyaml and pydantic were already core dependencies; fastapi and uvicorn are new behind a small [ui] extra. A normal DMI install pulls in no web framework.

Deliberate constraints

  • No runtime block in the YAML. The existing runtime surface is RingConfig, a native struct of transport parameters. Exposing a subset would imply a support contract that does not exist.
  • layers.end is inclusive. LayerSelection(8, 15) selects eight layers; the UI labels it "Layers 8–15" and the label must not lie.
  • Runtime policy does nothing yet, and says so. The rule from the plan: do not ship a UI where selecting "Performance" has no effect.
  • Only decoder_transformer. Encoder-decoder configs and known non-decoder families (BERT/ViT/T5-style) are refused rather than mis-rendered.

Verification

  • pytest -m cpu: 1414 passed (descriptor, configuration, serialization round-trip, integration, golden-file suites, plus the schedule-gate and API-contract suites).
  • The serialized round-trip is pinned as parse(serialize(config)) == config after canonical normalization — what makes the YAML a configuration artifact rather than a one-way export.
  • Live-verified during review: ClickHouse-backed e2e suites, Garage/S3, native backend e2e, and the HF correctness suite (gpt2) all green on this branch.

🤖 Generated with Claude Code (review round included 8 independent adversarial reviewers; all confirmed findings fixed on this branch)

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The implementation/documentation contract has a few concrete mismatches (e.g., config version enforcement and outdated plan snippets) plus a dependency-surface cleanup to address before safe approval.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Pull request overview

Adds DMI-configurator, a local, no-build-step web UI and supporting Python configuration pipeline to generate a typed, validated, canonical YAML DMI capture configuration from a model descriptor plus user selections—while keeping the existing runtime configuration interface intact (YAML is a front-end serialization format + compiler path, not a runtime rewrite).

Changes:

  • Introduces a structured configuration stack (DMIConfig schema, YAML parse/dump, validation, compilation, compatibility bridge).
  • Adds a FastAPI backend + static frontend UI (packaged assets) and a dmi CLI for ui + describe-model.
  • Adds a new runtime-adjacent selection filter (filter_by_layers) to support layer-range authored configs.
File summaries
File Description
tests/test_configurator_api.py Contract tests for configurator HTTP API + static asset serving.
tests/test_configuration_yaml.py YAML canonicalization + round-trip + golden-file coverage.
tests/test_configuration_validation.py Model-aware validation tests (availability, layer bounds, schedule checks).
tests/test_configuration_schema.py Unit tests for schema dataclasses (LayerSelection, ModelTopology, policies).
tests/test_configuration_manifest.py Descriptor parsing/IO tests and ModelShape translation checks.
tests/test_configuration_introspect.py Tests for deriving descriptors from HF-style configs and rejection cases.
tests/test_configuration_compatibility.py Tests for legacy selection bridge + compilation + new layer filtering.
tests/golden/qwen3-moe.yaml Golden config artifact for MoE model case.
tests/golden/qwen3-basic.yaml Golden config artifact for minimal/basic config case.
tests/golden/qwen3-attention.yaml Golden config artifact for attention-over-range config case.
tests/data/moe-decoder.model.yaml Synthetic MoE descriptor fixture for availability testing.
src/dmi/ui/static/styles.css New UI styling (light/dark, layout, state chips, issue rendering).
src/dmi/ui/static/index.html New single-page UI shell and controls layout.
src/dmi/ui/static/architecture.js SVG architecture renderer driven by server-provided metadata.
src/dmi/ui/static/app.js UI state management + server round-trips for validate/parse/serialize/save.
src/dmi/ui/server.py Uvicorn entrypoint wrapper for serving configurator.
src/dmi/ui/app.py FastAPI app + endpoints delegating to configuration layer; static mounting.
src/dmi/ui/init.py Lazy imports so dmi.ui doesn’t require FastAPI at import time.
src/dmi/hooks/selection.py Adds hook_belongs_to_layers + filter_by_layers (additive runtime filter).
src/dmi/configuration/yaml.py YAML parse/dump + canonicalization + disk IO helpers.
src/dmi/configuration/validation.py Validation producing field-addressed issues for UI display.
src/dmi/configuration/schema.py Canonical typed configuration + descriptor schema definitions.
src/dmi/configuration/manifest.py Descriptor parsing/IO and translation to DMI ModelShapeConfig.
src/dmi/configuration/introspect.py Descriptor derivation from HF configs/model ids with lazy transformers import.
src/dmi/configuration/errors.py Configuration-layer exception types.
src/dmi/configuration/compiler.py Compiles DMIConfig to CompiledDMIConfig via existing selector + layer filter.
src/dmi/configuration/compatibility.py Bridge between structured observations and legacy selection string.
src/dmi/configuration/catalog_adapter.py UI-facing hook metadata projection + availability/reason computation.
src/dmi/configuration/architecture.py Defines diagram nodes and produces model/layout payloads for the UI.
src/dmi/configuration/init.py Consolidated public API exports for configuration subsystem.
src/dmi/cli.py Adds dmi ui and dmi describe-model CLI commands.
pyproject.toml Adds [ui] optional deps, dmi console script, and UI static package data.
examples/model_descriptors/README.md Documentation for descriptors + how to generate/use them.
examples/model_descriptors/llama3-8b.yaml Shipped example descriptor used by tests/UI.
docs/dmi-configurator-plan.md Design/implementation plan and rationale for configurator architecture.
Review details

Suppressed comments (2)

docs/dmi-configurator-plan.md:244

  • The CompiledDMIConfig snippet documents a runtime: RuntimeConfig field, but the implemented CompiledDMIConfig (src/dmi/configuration/compiler.py) only carries hook_specs, schedule, and optional policy. Keeping the snippet aligned avoids implying runtime transport settings are part of config compilation.
@dataclass
class CompiledDMIConfig:
    hook_specs: list[HookSpec]
    schedule: CaptureSchedule
    runtime: RuntimeConfig
    policy: RuntimePolicy | None

docs/dmi-configurator-plan.md:255

  • The compile_config example calls filter_by_layers(specs, config.observations.layers), but the implemented function takes explicit start/end ints (and compilation also no longer passes a runtime field). Update the snippet so readers can copy/paste it without hitting a TypeError or referencing non-existent fields.
def compile_config(config, model_context):
    specs = select_hook_specs(
        model_context.specs,
        ",".join(config.observations.hooks),
        cfg=model_context.shape,
    )
    specs = filter_by_layers(specs, config.observations.layers)
    return CompiledDMIConfig(
  • Files reviewed: 35/35 changed files
  • Comments generated: 3
  • Review effort level: Lite

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/dmi/configuration/yaml.py Outdated
Comment thread docs/dmi-configurator-plan.md
Comment thread pyproject.toml

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The backend save endpoint and YAML parser currently mishandle error/type cases (uncaught ConfigurationError in /api/config/save and or {} / or [] defaulting that can silently bypass structural validation).

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Review details
  • Files reviewed: 48/48 changed files
  • Comments generated: 3
  • Review effort level: Lite

Comment thread src/dmi/configuration/yaml.py Outdated
Comment thread src/dmi/configuration/yaml.py Outdated
Comment thread src/dmi/ui/app.py Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

The changes are additive, well-covered by CPU-marked tests across schema/YAML/validation/UI API surfaces, and the runtime layer-range wiring is exercised end-to-end.

Review details

Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

pyproject.toml:20

  • Optional: tests/test_configurator_api.py imports fastapi.testclient.TestClient, which depends on httpx in many FastAPI/Starlette setups. If someone installs only the [ui] extra, the UI will work but running the API contract tests may fail with an ImportError if httpx isn’t pulled transitively; consider adding httpx to the ui extra (or adding an explicit importorskip("httpx") guard in the tests) to make the optional stack more self-contained for contributors.
  • Files reviewed: 48/48 changed files
  • Comments generated: 0 new
  • Review effort level: Lite

@zaoxing
zaoxing requested a review from XbzOnGit September 2, 2026 19:59

@XbzOnGit XbzOnGit left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Inline review of PR 122 at de37754. Labels used here: Blocking = should be fixed before merge; Major = important correctness or robustness issue.

Comment thread src/dmi/configuration/compiler.py
Comment thread src/dmi/configuration/validation.py
Comment thread src/dmi/configuration/estimate.py Outdated
Comment thread src/dmi/configuration/compiler.py
Comment thread src/dmi/ui/app.py Outdated
Comment thread src/dmi/configuration/schema.py
Comment thread src/dmi/configuration/introspect.py
Comment thread src/dmi/ui/static/app.js Outdated
Comment thread src/dmi/configuration/estimate.py Outdated
Comment thread src/dmi/configuration/estimate.py Outdated
@zaoxing

zaoxing commented Sep 3, 2026

Copy link
Copy Markdown
Collaborator Author

All 14 inline findings are addressed at dd3bb32 (responses on each thread). Two notes: the vLLM layer-range fix (3919326699) needs a matching revision of the pinned third_party/vllm-integration submodule — in this repo the failure is now a diagnosable error rather than a silent range drop. The 6 Copilot findings were already fixed at c8875a8/de37754 and are pinned as regression tests in tests/test_pr122_review_findings.py. CPU suite at head: 1373 passed.

@Samfisheryu Samfisheryu left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes for confirmed runtime and estimator correctness issues. I reproduced each finding inline. The existing CPU suite is healthy (1,387 passed), and the supported Hugging Face torch.compile schedule guard itself behaves correctly; the blocking gaps are configuration application, adapter ownership, and packed/vLLM parity.

Comment thread src/dmi/configuration/compiler.py
Comment thread src/dmi/configuration/compiler.py
Comment thread src/dmi/configuration/estimate.py Outdated
Comment thread src/dmi/configuration/estimate.py Outdated
Comment thread src/dmi/configuration/estimate.py
Comment thread src/dmi/configuration/yaml.py Outdated
Comment thread src/dmi/configuration/estimate.py
Comment thread src/dmi/ui/app.py Outdated
Comment thread src/dmi/ui/server.py
Comment thread src/dmi/cli.py
@zaoxing

zaoxing commented Sep 5, 2026

Copy link
Copy Markdown
Collaborator Author

Independent review round: 8 adversarial reviewers, all confirmed findings fixed at 3dc68ca.

Review lenses: config pipeline, HTTP security/correctness, schedule-gate runtime, estimator fidelity, frontend JS, descriptor/schema integrity, test honesty (mutation-tested), PR-body coherence. Every finding was verified by execution before being flagged; every confirmed one is fixed with a test in this commit.

Fixed:

  • HTTP: malformed bodies (non-string yaml, float/bool/Inf workload scalars, huge ring ints) were raw 500s — now 400s; non-ASCII token headers crashed the compare; /api/config is token-gated on network binds; CSRF on the default loopback bind (sendBeacon could blind-fire a config write) closed by refusing foreign-Origin writes.
  • Data safety: save was truncate-then-write — a mid-write ENOSPC destroyed the previous config; now temp+os.replace (config and descriptor both).
  • Estimator: ring fit ignored the task-entry cap (a wide full-preset model fit on bytes while every real step was OVERSIZED); PP split disagreed with vLLM's get_pp_indices on non-divisible layer counts; batched final_logits over-counted ~1000x (logits_to_keep=1 is what generate() materializes).
  • Gate: the no-phase warmup/stride tail is now test-pinned (mutation-tested gap); foreign phase values can't silently kill capture; final_logits with vocab_size=0 no longer validates/compiles as available.
  • Descriptor boundary: exact-int schema_version, unknown-key refusal in model/root, DescriptorError for bad ids, num_experts-without-top_k refused.
  • Frontend: boot-race guard (listeners fired before state existed), Copy stale-response stamp, tab a11y.
  • Docs: capture-policy-review's 'schedule enforced by no adapter' marked resolved; PR body no longer claims additive-only.

Known ceilings (disclosed, deferred): the vLLM integration's direct commit path is not schedule-gated (separate repo, DMI-vLLM-Integration#21 adds the layers keyword; gating is next); torch.compile/CUDA-graph replay bakes the gate at first trace (compiled decode should stay on default schedules until enforcement moves native).

CPU suite: 1414 passed.

@Samfisheryu Samfisheryu left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks — I independently re-ran the original repros on 27c40b3. Several fixes are substantive (live layer rejection, design-time labelling, save serialization, ordinary duplicate-key and hook-type validation, port handling, and clean CLI errors). Keeping changes requested for the still-incomplete runtime capability boundary and the confirmed residual implementation bugs below.

Comment thread src/dmi/configuration/estimate.py Outdated
Comment thread src/dmi/configuration/estimate.py
Comment thread src/dmi/adapters/base.py
Comment thread src/dmi/configuration/introspect.py
Comment thread src/dmi/configuration/estimate.py
Comment thread src/dmi/ui/app.py Outdated
Comment thread src/dmi/configuration/estimate.py
Comment thread src/dmi/configuration/yaml.py Outdated
@zaoxing

zaoxing commented Sep 6, 2026

Copy link
Copy Markdown
Collaborator Author

All 8 follow-up findings are fixed at 450b340 (replies on each thread), with regression tests in tests/test_pr122_round4_review_findings.py. CPU suite: 1679 passed; both CI checks green. Two known ceilings remain disclosed in code rather than fixed here: the vLLM direct-commit path still bypasses the schedule gate (needs the graph-safe integration change, separate repo), and the token-range index records attempted rather than captured traffic.

@zaoxing

zaoxing commented Sep 6, 2026

Copy link
Copy Markdown
Collaborator Author

All review threads on this PR are now resolved (43/43: 6 Copilot, 14 XbzOnGit, 23 Samfisheryu), each with its fix commit, mechanism, and pinning test recorded on the thread. Since the last review round: packed estimator honesty, the verbatim vLLM partition rule, one-owner attach enforcement, fail-closed encoder subtypes, ring/workload/YAML boundary hardening, offset/warmup disclosure, plus the CUDA-graph schedule disclosure, CLI error narrowing, and selection taxonomy. Suite: 1679 passed; both CI checks green on 450b340. Requesting re-review.

@zaoxing
zaoxing requested a review from Samfisheryu September 6, 2026 02:02
zaoxing added a commit that referenced this pull request Sep 6, 2026
The review round surfaced three rules that deserve a written rationale
rather than a thread of replies: the configuration artifact is strict
end-to-end; the runtime surface is enforced at the smallest possible
layer (the adapter driver's before_forward) rather than scattered
across every emit point; and the HTTP surface trusts the loopback bind
but refuses the network without a per-launch X-DMI-Token. The
one-owner attach invariant is the same rule applied to a model.

The alternatives weighed (stricter loopback defaults, TLS/OAuth on the
local tool, rejecting packed configs that name a schedule, hooking the
gate at every emit point, an inline JSON or frontmatter idiom) all
get a sentence here so the next round does not relitigate them.
zaoxing added a commit to ProjectDMX/DMI-vLLM-Integration that referenced this pull request Sep 6, 2026
* Accept a layer range in VLLMAdaptor.attach_model

DMI-configurator (ProjectDMX/DMI#122) forwards a selected layer range as
attach_model(model, hook_selection, layers=LayerSelection(...)). The
adaptor refused any ranged configuration with TypeError, and Packed/vLLM
is the configurator's default backend.

The keyword is applied in _VLLMHookSelection.from_model so the filter hits
BOTH sides: the installed local specs and the model-wide candidate-rank
formula sets. A signature-only fix would let the per-rank byte plans
disagree with what the filtered local specs actually capture. Global hooks
(layer_no < 0) are never restricted, matching the HF adapter. A range that
matches nothing raises a DMI ConfigValidationError naming the range instead
of silently compiling to an empty selection.

Closes #20. Covered by tests/test_attach_layer_range.py, which runs the
selection logic on CPU behind a vLLM import stub, so CI needs no vLLM
install; the six existing suites that need the real release were already
failing on this machine before this change (no vllm distribution) and are
unaffected.

* Keep the adapter importable and attachable on DMI without the layer-range API

The official-vLLM CI matrix runs this repo against released DMI (v1.1.0)
and DMI main, neither of which exports hook_belongs_to_layers -- the
module-level import broke collection of every test module there.

The layer-range API imports now live inside the layers branch (with a
clear upgrade-DMI error when a range is passed without them), and
attach_model forwards the keyword to super() only when a range is
configured, so un-ranged attaches behave exactly as before on old DMI.
The new tests skip cleanly where dmi.configuration is absent and pin the
old-DMI contract by simulating the missing facade name.

* Stub vLLM in tests only when vLLM is not installed

The old guard (sys.modules + hasattr distributed) skipped the stub only
when vLLM was already imported; an installed-but-unimported vLLM got
shadowed for the whole session, so other tests' view of vllm depended on
collection order. importlib.util.find_spec decides on installation, not
import state: real vLLM always wins, the stub only exists for CPU CI
where vLLM is absent. The suite's assertions are behavior-based on the
DMI side and hold under either vLLM.

Review thread: copilot on PR #21 (test_attach_layer_range.py:23).

---------

Co-authored-by: Alan Liu <zaoxing@users.noreply.github.com>
@zaoxing
zaoxing requested a lite review from Copilot September 6, 2026 15:05

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

BackendAdapter.before_forward() does not advance the step counter when build_step_context() returns None, which can break warmup/stride accounting and capture steps more often than the configured schedule.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Review details
  • Files reviewed: 63/63 changed files
  • Comments generated: 2
  • Review effort level: Lite

Comment thread src/dmi/adapters/base.py
Comment thread src/dmi/hooks/point.py

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The public integration API docs currently mis-specify the filter_by_layers signature/return type, which can mislead downstream integrators even if runtime behavior is correct.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Review details
  • Files reviewed: 63/63 changed files
  • Comments generated: 3
  • Review effort level: Lite

Comment thread docs/integration-api-v1.md
Comment thread src/dmi/transport/ring.py
Comment thread tests/test_ui_client_invariants.py
…rrowing, selection taxonomy

- A non-default capture schedule now disables CUDA-graph compilation for
  the generate() call, mirroring the capacity-overflow path: the gate is a
  Python branch in HookPoint.forward and replay bakes it at the first
  decode trace, so a refusal (or capture) at step one would apply to every
  later step. Default schedules are unaffected. Pinned by driving
  generate_with_monitoring with a stubbed adapter and inspecting the
  kwargs the model receives.
- cli.main no longer catches bare RuntimeError -- that class carries
  genuine bugs (RecursionError is one) and swallowed them into one stderr
  line. The optional-dependency failure gets its own UIDependencyError;
  internal raise sites updated.
- compile_config re-raises an unknown hook selection as
  ConfigurationError instead of bare ValueError, so callers funneling on
  the taxonomy (the UI's 400 mapping) see a 400, not a 500.

CPU suite: 1588 passed (after the main merge).
…hip, boundaries

Samfisheryu's follow-up review kept CHANGES_REQUESTED on eight residuals;
all are fixed here with tests, RED first.

Ownership:
- attach_config refuses a second adapter on an already-owned model BEFORE
  any mutation (no schedule install, no attach call), naming the current
  owner; same-adapter reconfiguration is allowed; detach by a non-owner
  leaves the marker (HF adapter path pinned).

Packed honesty (the vLLM runtime executes no schedule):
- capture_prefill/capture_decode no longer shrink packed volume, peak, or
  sustained figures -- the runtime captures both phases regardless. The
  authored intent stays visible as an explicit warning; batched figures
  are unchanged.
- vLLM PP partition rule ported from 0.27.1's get_pp_indices (remainder
  backward from the second-to-last stage, never the last; cited cases
  [1,2,1], [2,3,3,2], 32/40 verified) plus an explicit, validated
  pp_layer_counts override for VLLM_PP_LAYER_PARTITION deployments.
- Offset/warmup disclosure no longer depends on a stride being set.

Boundaries:
- Ring: only absent/None defaults pinned_bytes; every explicit value is
  validated (exact non-bool int, >= 0), and check_ring_fit refuses
  negatives itself.
- Workload.cache_max_len requires an exact integer.
- YAML: unhashable mapping keys raise ConstructorError (a 400), and the
  descriptor file path uses the same strict loader as configs.
- Introspect: any model_type containing "encoder" fails closed instead of
  riding a decoder family prefix (qwen2_audio_encoder); decoder+head
  configs still accepted.

CPU suite: 1679 passed.
- The per-request token-range index records every driven step including
  schedule-refused ones; it is the attempted-traffic index, not the
  captured-traffic index.
- The _accepts_layers probe cannot distinguish forwarding **kwargs from
  tolerating them; the backstop is the v1 contract plus compile-time
  live-hook rejection, stated where the probe lives.
Attention-weight hooks (pattern, attn_scores) attend at most window KV
positions on sliding-window models, so they are shaped against
min(context, window) instead of the full context. Only kv_dim is
affected, and only those shapes consume it. Exact-int >= 1 validation,
assumption text when set, no-op when absent or above context.

FP8 KV-cache quantization was investigated and deliberately not modeled:
hooks capture forward activations in model dtype, so cache-store
quantization changes zero hook payload bytes.
…mates"

The knob's core semantic was incorrect in the unsafe direction:
compute_hook_shape returns full [heads, q_len, kv_dim] for
pattern/attn_scores with no window input, plan_step reserves exactly
that, and these hooks fire meaningfully only in eager mode -- where SDPA
materializes the full scores matrix and masks it rather than shrinking
it. Capping kv_dim by the window made the estimate report less than the
runtime reserves (up to context/window under-count), and check_ring_fit
would bless a ring the runtime overflows. The FP8 non-modeling analysis
in the original message stands; the window half is withdrawn until a
backend materializes band-shaped scores.
…y banner

- Save disables itself for the request round trip (re-armed in finally).
- Layer rail ticks are keyboard-operable checkboxes sharing one picking
  function with the click path; architecture nodes already had this.
- Estimate panel marks itself loading with aria-busy, cleared on either
  guarded outcome.
- Status chip is a polite live region; ring meter exposes value semantics.
- Token-gated binds get a persistent curl-only banner revealed on 401
  instead of per-panel curl-hint errors.
- _reject_unknown gains an error= kwarg so the config half (ConfigurationError)
  and the descriptor half (DescriptorError) share one body. As a bonus
  this stops the int-YAML-key AttributeError the descriptor half had -- the
  manifest copy already used repr() but the config half still did not.
- save_config and save_descriptor share _write_text_atomic(path, text, what,
  error): the only delta is which exception class wraps an OSError.

CPU suite: 1685 passed.
The review round surfaced three rules that deserve a written rationale
rather than a thread of replies: the configuration artifact is strict
end-to-end; the runtime surface is enforced at the smallest possible
layer (the adapter driver's before_forward) rather than scattered
across every emit point; and the HTTP surface trusts the loopback bind
but refuses the network without a per-launch X-DMI-Token. The
one-owner attach invariant is the same rule applied to a model.

The alternatives weighed (stricter loopback defaults, TLS/OAuth on the
local tool, rejecting packed configs that name a schedule, hooking the
gate at every emit point, an inline JSON or frontmatter idiom) all
get a sentence here so the next round does not relitigate them.
The yaml.py/manifest.py simplification shipped first without its test
updates, which broke CI: the old atomic-save test raised ENOSPC from
dump_config, and the new save_config evaluates dump_config before
_write_text_atomic's try-block, so the OSError escaped unwrapped. The
test now patches Path.write_text on the .tmp sibling instead, which is
the real I/O boundary the atomic write protects.

Rest of the round:
- UI_DEPENDENCY_MESSAGE moves to errors.py beside UIDependencyError, so
  app.py and server.py quote one string.
- The auth banner reuses .notice plus one margin-top override instead
  of duplicating seven properties.
- Copy calls a shared serializeCurrentState() rather than reissuing the
  /api/config/serialize POST shape refreshOutput already builds.
- The two test-side brace matchers (_function_body /
  _plain_function_body) collapse into one with an optional async prefix.

CPU suite: 1685 passed.
DMI-vLLM-Integration PR #21 merged (squash 23717cb): attach_model
accepts layers=LayerSelection(...), applies it to the local specs and
the model-wide candidate-rank sets, and stays importable on DMI builds
without the layer-range facade. The pin moves old-main 29f26c3 ->
23717cb, picking up both the V2-runner promotion (#19) and the layer
range (#21).
The warning told the UI that attach_config refuses a ranged
configuration on Packed (vLLM) until the integration accepted ranges --
that is PR #21, now merged and pinned. The estimator test flips to pin
the absence of the warning (RED against the stale block), and
attach_config's fail-closed guard stays for genuinely older
integration checkouts, with its comment updated to name them.

CPU suite: 1689 passed.
…vLLM

The recipe for proving a saved .dmi.yaml artifact drives a real vLLM
0.27.1 model through dmi.configuration.attach_config existed only in
the verification session: build the CI-style env (vllm 0.27.1 + DMI
--no-deps + integration --no-deps), subclass the V1 DMXGPUWorker to
suppress the self-attach (V2 refuses subclasses), run the engine core
in-process, and assert ownership, ranged spec layers, disabled
out-of-range hook points, and a completed generate. Also correct the
install step for shared checkouts: the native build emits ABI-suffixed
extensions per interpreter, so 'make clean' is destructive there.
@zaoxing
zaoxing force-pushed the feat/dmi-configurator branch from a2ed118 to 4def7cc Compare September 7, 2026 20:04
dmi ui/describe-model needed transformers installed by hand to resolve a
model directory or bare HF model id; pip install -e ".[hf]" makes that
reproducible, and the missing-transformers error now names it.
Keep the top-level entry point consistent with the configurator plan doc.

@Samfisheryu Samfisheryu left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review at 393f4f3: 354 tests passed across the eight relevant configuration/runtime/estimator/API regression suites. The pinned vLLM adapter now accepts and forwards layers.

One confirmed regression remains beyond the already disclosed feature limitations: the previously fixed None-context step accounting and its tests are absent from this head. I added the old/new execution results to the existing inline thread: the author's original four tests pass on 74d7be7, but three fail on 393f4f3. Please restore that fix; the prior "fixed" reply does not describe the current tree.

The vLLM direct-commit schedule gap and attempted-vs-captured token-range behavior remain acknowledged limitations, not newly reported implementation bugs. This recheck does not claim a full vLLM/HF graph or multi-GPU run.

@zaoxing
zaoxing requested a review from Samfisheryu September 9, 2026 19:43
Samfisheryu and others added 2 commits September 10, 2026 11:55
Two conflicts, both where main's #127 (native capture write path) and this
branch added to the same block; both resolved as the union, since the two
sides configure unrelated things:

- .github/workflows/python-checks.yml: main added the botocore/boto3 pip
  lines the native capture suites need at collection time; this branch
  added pyyaml and the [ui] extra (fastapi/uvicorn/httpx). Both jobs need
  both sets. Keeping only one side would trip main's own new gate, which
  fails the build when a cpu test skips for anything but absent hardware --
  the configurator's API contract suite skips itself without fastapi/httpx,
  and the native suites error at collection without botocore.

- Makefile: main added PYTEST_ARGS (so CI can pass --junit-xml to the
  gate); this branch added MODEL for the `make ui` target. Independent
  variables in the same header block.

No src/ file was touched by both sides, so nothing else conflicted.

@Samfisheryu Samfisheryu left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Full re-review at badbfca: no new reproduced implementation blocker beyond the previously acknowledged runtime schedule limitations. The restored None-context step-accounting fix is retained and its regressions pass; the pinned vLLM adapter accepts and forwards layers.

Validation: 2,038 CPU tests passed; 16 real single-GPU native/reference/ClickHouse Ring tests passed; GPT-2 HF eager and CUDA Graph correctness E2Es passed (2 tests, 56 subtests). Both CI checks are green. These E2Es validate their tested/default capture paths, not unsupported schedule enforcement.

Keeping the existing requested changes for the core capability boundary: vLLM's direct commit_step path still bypasses schedule gating, and non-default HF schedules still disable compilation instead of implementing graph-safe enforcement. The attempted-vs-captured token-range behavior also remains acknowledged. I am not treating documentation nits as blockers; this run does not claim full vLLM or multi-GPU E2E coverage.

The /polish loop's working files are per-run agent bookkeeping, not project
artifacts: nothing in the tree, the build, or CI reads them, and they go
stale once the branch they describe has landed. #133 removes them from
main; dropping them here too so a merge of this branch cannot carry them
back in.
@zaoxing
zaoxing requested review from Samfisheryu and a lite review from Copilot September 16, 2026 02:15

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@zaoxing
zaoxing merged commit e03b3bb into main Sep 22, 2026
2 checks passed
zaoxing added a commit that referenced this pull request Sep 24, 2026
_FakeEagerTransport was a hand-copied RingTransport, and it drifted: it
lacked capture_step until the previous commit, so five GPU tests raised
AttributeError after #122 added the capture-schedule gate. It also
reimplemented production logic by precomputing effective_cap instead of
going through the transport's cached property.

The tests now build a real RingTransport(engine) over the fake engine,
set force_eager, and replace only submit_cpu_direct with a list append.
tests/test_hook_point_eager_cap_cache.py already does this. The fake
engine gains the payload_tensor() the constructor needs, and each hook
writes into that payload. Any attribute RingTransport gains in future is
therefore present in these tests, and effective_cap is the production
property. Every test's assertions are unchanged, and the now-unused
effective_ring_bytes import is gone.

GPU (RTX 4090, svc worktree, point.py and ring.py identical to
origin/main): 20 passed.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
zaoxing added a commit that referenced this pull request Sep 24, 2026
Nothing covered #122's capture-schedule gate on the eager path. When the
schedule refuses a step, the driver skips plan/commit, so no ring space
is reserved and no meta is pushed. If a producer still ran, it would
write unreserved bytes and desync the task/meta FIFO. HookPoint.forward
checks transport.capture_step after the `if not self.enabled` early
return and before the CUDA probe, so a refused step never reaches the
force_eager block.

The new test arms a real RingTransport with force_eager=True and
capture_step=False. It then calls the hook with one tensor over staging
(flush + cpu_direct) and one that fits staging but not the current slack
(flush + reserve + ring). It asserts that nothing was flushed, reserved,
submitted to cpu_direct or dispatched to a producer.

GPU (RTX 4090, svc worktree, point.py and ring.py identical to
origin/main): 21 passed. On a scratch copy of the package with the gate
moved below the eager block, the new test fails with flushes == 2, and
the four counters read flushes=2 reserved=[32] direct=1 dispatched=1.
The other 20 tests still pass there.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y
zaoxing added a commit that referenced this pull request Sep 24, 2026
* Give the eager-transport test fake the capture-schedule gate

#122 added a capture-schedule gate to HookPoint.forward: it reads
transport.capture_step on the active transport before anything else. The
real RingTransport sets it to True. This file's _FakeEagerTransport never
gained the attribute, so all five eager-path tests have raised
AttributeError at hooks/point.py since #122, before reaching the capacity
and staging behaviour they exist to pin. They are marked gpu, and CI has no
GPU runner, so nothing flagged it.

The fake now arms the gate the way the real transport does.

GPU (RTX 4090): 5 failed, 15 passed before; 20 passed after.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Build the eager-path tests on a real RingTransport

_FakeEagerTransport was a hand-copied RingTransport, and it drifted: it
lacked capture_step until the previous commit, so five GPU tests raised
AttributeError after #122 added the capture-schedule gate. It also
reimplemented production logic by precomputing effective_cap instead of
going through the transport's cached property.

The tests now build a real RingTransport(engine) over the fake engine,
set force_eager, and replace only submit_cpu_direct with a list append.
tests/test_hook_point_eager_cap_cache.py already does this. The fake
engine gains the payload_tensor() the constructor needs, and each hook
writes into that payload. Any attribute RingTransport gains in future is
therefore present in these tests, and effective_cap is the production
property. Every test's assertions are unchanged, and the now-unused
effective_ring_bytes import is gone.

GPU (RTX 4090, svc worktree, point.py and ring.py identical to
origin/main): 20 passed.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

* Test that a refused capture step skips the eager safety net

Nothing covered #122's capture-schedule gate on the eager path. When the
schedule refuses a step, the driver skips plan/commit, so no ring space
is reserved and no meta is pushed. If a producer still ran, it would
write unreserved bytes and desync the task/meta FIFO. HookPoint.forward
checks transport.capture_step after the `if not self.enabled` early
return and before the CUDA probe, so a refused step never reaches the
force_eager block.

The new test arms a real RingTransport with force_eager=True and
capture_step=False. It then calls the hook with one tensor over staging
(flush + cpu_direct) and one that fits staging but not the current slack
(flush + reserve + ring). It asserts that nothing was flushed, reserved,
submitted to cpu_direct or dispatched to a producer.

GPU (RTX 4090, svc worktree, point.py and ring.py identical to
origin/main): 21 passed. On a scratch copy of the package with the gate
moved below the eager block, the new test fails with flushes == 2, and
the four counters read flushes=2 reserved=[32] direct=1 dispatched=1.
The other 20 tests still pass there.

Claude-Session: https://claude.ai/code/session_01PcY9QS1FAehHTzkjdpvN6Y

---------

Co-authored-by: Alan Liu <zaoxing@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants