Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitmodules
Original file line number Diff line number Diff line change
Expand Up @@ -10,3 +10,6 @@
[submodule "third_party/DMI-Megatron-Integration"]
path = third_party/DMI-Megatron-Integration
url = https://github.com/ProjectDMX/DMI-Megatron-Integration.git
[submodule "third_party/sglang-integration"]
path = third_party/sglang-integration
url = https://github.com/ProjectDMX/DMI-SGLang-Integration.git
3 changes: 2 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@
> collaborators to explore downstream applications built on DMI such as **interpretability**, **speculative decoding**,
> **hallucination analysis**, **distillation**, **activation steering**, and beyond. If you're interested, please [contact us](mailto:ynn1999@umd.edu,sixianx@umd.edu,zaoxing@umd.edu).

> **Project Status — research preview.** DMI supports HuggingFace and official vLLM integrations across 20+ dense and MoE model families, including Qwen3, Llama, Gemma, Mistral, GPT-OSS, Phi, and Granite, with early support for emerging architectures such as Qwen3.6, Llama 4, DeepSeek V4 Flash, GLM-5.2, Kimi K3, and MiniMax-M2.7, plus Megatron-LM training. SGLang support is on the way. APIs may change. Contributions, bug reports, and feature requests are welcome.
> **Project Status — research preview.** DMI supports HuggingFace and official vLLM integrations across 20+ dense and MoE model families, including Qwen3, Llama, Gemma, Mistral, GPT-OSS, Phi, and Granite, with early support for emerging architectures such as Qwen3.6, Llama 4, DeepSeek V4 Flash, GLM-5.2, Kimi K3, and MiniMax-M2.7, plus Megatron-LM training, and an early SGLang integration (offline `Engine`, Qwen3 dense). APIs may change. Contributions, bug reports, and feature requests are welcome.

> **👀Technical Report Available:** https://arxiv.org/abs/2605.11093

Expand Down Expand Up @@ -153,6 +153,7 @@ for o in llm.generate(["The answer is"], SamplingParams(max_tokens=16)):
| **[Core installation](docs/install.md)** | Install DMI from source and build the native backend |
| **[HuggingFace](docs/huggingface.md)** | Run HF generation, monitored generation, and offline benchmark scripts |
| **[vLLM](docs/vllm.md)** | Run DMI through the vLLM offline API or `vllm serve` |
| **[SGLang](docs/sglang.md)** | Run DMI through the SGLang offline `Engine` API |
| **[Megatron-LM](docs/megatron.md)** | Run DMI during Megatron-LM training |

## Contribute
Expand Down
12 changes: 7 additions & 5 deletions docs/install.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,10 +31,11 @@ nvidia-smi

## 1. Clone the repository

The repo uses four git submodules: the DMI HuggingFace integration, the
version-matched DMI-vLLM integration, the version-matched DMI-Megatron
integration, and the `clickhouse-cpp` C++ client. The commands below fetch all
four repositories; they do not install any Python integration.
The repo uses five git submodules: the DMI HuggingFace integration, the
version-matched DMI-vLLM integration, the version-matched DMI-SGLang
integration, the version-matched DMI-Megatron integration, and the
`clickhouse-cpp` C++ client. The commands below fetch all five repositories;
they do not install any Python integration.

The command below creates one backend checkout. If you plan to use multiple
backends, repeat it with distinct target directories such as `DMI-hf`,
Expand All @@ -53,6 +54,7 @@ Expected submodule paths:

- `third_party/transformers/` — modified HF Transformers (`gpt2_p`, `qwen3_p`, `llama_p`)
- `third_party/vllm-integration/` — DMI integration for an unmodified official vLLM installation
- `third_party/sglang-integration/` — DMI integration for an unmodified official SGLang installation
- `third_party/DMI-Megatron-Integration/` — DMI integration with its pinned Megatron-LM fork at `third_party/DMI-Megatron-Integration/third_party/megatron-lm/`
- `third_party/clickhouse-cpp/` — ClickHouse C++ client linked into the native backend

Expand Down Expand Up @@ -215,7 +217,7 @@ ClickHouse; use the host benchmark separately with a running server.
## 6. Choose one backend

Continue with the [HuggingFace guide](huggingface.md), [vLLM guide](vllm.md),
or [Megatron-LM guide](megatron.md). Use a separate environment and checkout
[SGLang guide](sglang.md), or [Megatron-LM guide](megatron.md). Use a separate environment and checkout
for each backend. The HuggingFace path installs a modified Transformers
checkout, the vLLM path installs its own official dependency set, and the
Megatron-LM path installs its version-matched integration and pinned fork. Do
Expand Down
117 changes: 117 additions & 0 deletions docs/sglang.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,117 @@
# SGLang usage

DMI provides experimental support for official SGLang 0.5.19 through the version-matched integration
checkout pinned at `third_party/sglang-integration/`. It does not contain an
SGLang fork: the integration is an out-of-tree package that registers through
SGLang's general-plugin entry point (`sglang.srt.plugins`) and its external
model package mechanism.

## Install the SGLang backend

Use a dedicated environment and DMI checkout for SGLang. Do not install the
modified HuggingFace integration from `third_party/transformers/` or the vLLM
integration in this environment. Complete the [core installation](install.md)
first.

SGLang 0.5.19 is not published on PyPI (wheels stop at 0.5.10, which predates
the plugin framework), so install it from the pinned source tag, then install
the integration editable and rebuild the native backend against that
environment's PyTorch (2.13.0 / CUDA 13.0 for this release):

```bash
git clone --branch v0.5.19 --depth 1 https://github.com/sgl-project/sglang.git
pip install -e sglang/python
pip install -e third_party/sglang-integration/
make -C native clean
make -C native -j
python -c "from dmi.transport.native import RingConfig; print(RingConfig())"
```

The integration checkout's
[`docs/env-setup.md`](https://github.com/ProjectDMX/DMI-SGLang-Integration/blob/main/docs/env-setup.md)
records the exact dependency resolution used for qualification (uv, pinned
torch constraints, no Rust extensions).

The integration rejects unsupported versions and checks architecture/execution
modes before DMI allocation and weight loading. Upstream device/distributed
initialization may already have happened. Rejected modes include TP/PP/DP,
speculative decoding, LoRA, torch.compile, PD disaggregation, and others. See
the integration's
[port audit](https://github.com/ProjectDMX/DMI-SGLang-Integration/blob/main/docs/v0519-port.md)
for the qualified compatibility cells.

## Offline API

For the qualified offline, non-streaming generation API, use `DMIEngine`
instead of `sglang.Engine`. It forwards every
`ServerArgs` keyword unchanged, carries DMI settings to the spawned worker
processes, and tags each output with a lazy `dmi_internal` readback handle:

```python
from dmi_sglang_integration import DMIEngine

engine = DMIEngine(
model_path="Qwen/Qwen3-0.6B",
mem_fraction_static=0.5,
disable_prefill_cuda_graph=True, # DMI runs prefill eagerly anyway
dmx_model_id="my-run",
dmx_db_host="localhost",
dmx_hook_selection="resid_pre,final_ln,token_ids,final_logits",
dmx_ring_payload_mb=1024,
dmx_ring_pinned_mb=1024,
)
try:
outputs = engine.generate(
prompt=["The capital of France is"],
sampling_params={"temperature": 0, "max_new_tokens": 8},
rid=["request-0"],
)
print(outputs[0]["text"])
engine.dmi_stop_monitoring() # authoritative flush to ClickHouse
hidden = outputs[0]["dmi_internal"].hidden_states
finally:
engine.shutdown()
```

Call `dmi_stop_monitoring()` before `shutdown()`: SGLang terminates its worker
processes without a graceful path, so the explicit RPC is the only durable
flush point. Monitoring is terminal for the engine; create a new `DMIEngine`
for another capture.
If flush fails, repeated stop calls continue to report failure; shutdown is not
proof of successful persistence. Construct engines sequentially in one process.

Plain `sglang.Engine` also works when the integration package is installed:
set the `DMX_*` environment variables before constructing it. Set
`SGLANG_PLUGINS=none` (or `DMX_SGLANG_ENABLE=0`) to run an unmonitored
baseline in the same environment.

## Common configuration

| Setting | Environment variable | Default |
| --- | --- | --- |
| Capture identifier | `DMX_MODEL_ID` | model path |
| Hook selection | `DMX_HOOK_SELECTION` | `sglang-full` (every hook except attention matrices) |
| ClickHouse host / port / database / table | `DMX_DB_HOST`, `DMX_DB_PORT`, `DMX_DB_DATABASE`, `DMX_DB_TABLE` | unset (no persistence), `9000`, `default`, `offload` |
| GPU payload ring / pinned staging | `DMX_RING_PAYLOAD_MB`, `DMX_RING_PINNED_MB` | `1024`, `1024` |
| Device-side padding strip | `DMX_GPU_PADDING_STRIP` | `1` |
| System `libstdc++` preload for conda interpreters | `DMX_SGLANG_PRELOAD_LIBSTDCXX` | `auto` |

The table describes plain-plugin environment defaults. `DMIEngine` instead
defaults to a unique capture ID and ClickHouse on localhost, exports its
overrides only during worker startup, and restores the parent's environment.

## Troubleshooting

- **`GLIBCXX_3.4.30 not found` in the scheduler process** -- a conda-provided
interpreter binds conda's older `libstdc++` before DMI's native backend
loads. `DMIEngine` preloads the system library in worker processes
automatically; for plain `sglang.Engine` or `sglang serve`, export
`LD_PRELOAD=/usr/lib/x86_64-linux-gnu/libstdc++.so.6`.
- **`DMI monitoring is not active in this scheduler`** -- the plugin did not
register. Check `pip show DMI-SGLang-Integration`, `SGLANG_PLUGINS`, and
`DMX_SGLANG_ENABLE`.
- **Overlap scheduler rows** -- with the default overlap scheduler SGLang runs
one extra decode forward per request after its final decision; DMI captures
that forward too (one extra `token_ids` row holding the last generated token
and one unused `final_logits` decision). Pass `disable_overlap_schedule=True`
for a capture that matches the public output exactly.
1 change: 1 addition & 0 deletions third_party/sglang-integration
Submodule sglang-integration added at 97c8fb
Loading