diff --git a/.gitmodules b/.gitmodules index 7cdf51793..562d5ecf0 100644 --- a/.gitmodules +++ b/.gitmodules @@ -10,3 +10,6 @@ [submodule "third_party/DMI-Megatron-Integration"] path = third_party/DMI-Megatron-Integration url = https://github.com/ProjectDMX/DMI-Megatron-Integration.git +[submodule "third_party/sglang-integration"] + path = third_party/sglang-integration + url = https://github.com/ProjectDMX/DMI-SGLang-Integration.git diff --git a/README.md b/README.md index e24d07645..c8e883434 100644 --- a/README.md +++ b/README.md @@ -13,7 +13,7 @@ > collaborators to explore downstream applications built on DMI such as **interpretability**, **speculative decoding**, > **hallucination analysis**, **distillation**, **activation steering**, and beyond. If you're interested, please [contact us](mailto:ynn1999@umd.edu,sixianx@umd.edu,zaoxing@umd.edu). -> **Project Status โ€” research preview.** DMI supports HuggingFace and official vLLM integrations across 20+ dense and MoE model families, including Qwen3, Llama, Gemma, Mistral, GPT-OSS, Phi, and Granite, with early support for emerging architectures such as Qwen3.6, Llama 4, DeepSeek V4 Flash, GLM-5.2, Kimi K3, and MiniMax-M2.7, plus Megatron-LM training. SGLang support is on the way. APIs may change. Contributions, bug reports, and feature requests are welcome. +> **Project Status โ€” research preview.** DMI supports HuggingFace and official vLLM integrations across 20+ dense and MoE model families, including Qwen3, Llama, Gemma, Mistral, GPT-OSS, Phi, and Granite, with early support for emerging architectures such as Qwen3.6, Llama 4, DeepSeek V4 Flash, GLM-5.2, Kimi K3, and MiniMax-M2.7, plus Megatron-LM training, and an early SGLang integration (offline `Engine`, Qwen3 dense). APIs may change. Contributions, bug reports, and feature requests are welcome. > **๐Ÿ‘€Technical Report Available:** https://arxiv.org/abs/2605.11093 @@ -153,6 +153,7 @@ for o in llm.generate(["The answer is"], SamplingParams(max_tokens=16)): | **[Core installation](docs/install.md)** | Install DMI from source and build the native backend | | **[HuggingFace](docs/huggingface.md)** | Run HF generation, monitored generation, and offline benchmark scripts | | **[vLLM](docs/vllm.md)** | Run DMI through the vLLM offline API or `vllm serve` | +| **[SGLang](docs/sglang.md)** | Run DMI through the SGLang offline `Engine` API | | **[Megatron-LM](docs/megatron.md)** | Run DMI during Megatron-LM training | ## Contribute diff --git a/docs/install.md b/docs/install.md index 5d2dfd741..41f30e88a 100644 --- a/docs/install.md +++ b/docs/install.md @@ -31,10 +31,11 @@ nvidia-smi ## 1. Clone the repository -The repo uses four git submodules: the DMI HuggingFace integration, the -version-matched DMI-vLLM integration, the version-matched DMI-Megatron -integration, and the `clickhouse-cpp` C++ client. The commands below fetch all -four repositories; they do not install any Python integration. +The repo uses five git submodules: the DMI HuggingFace integration, the +version-matched DMI-vLLM integration, the version-matched DMI-SGLang +integration, the version-matched DMI-Megatron integration, and the +`clickhouse-cpp` C++ client. The commands below fetch all five repositories; +they do not install any Python integration. The command below creates one backend checkout. If you plan to use multiple backends, repeat it with distinct target directories such as `DMI-hf`, @@ -53,6 +54,7 @@ Expected submodule paths: - `third_party/transformers/` โ€” modified HF Transformers (`gpt2_p`, `qwen3_p`, `llama_p`) - `third_party/vllm-integration/` โ€” DMI integration for an unmodified official vLLM installation +- `third_party/sglang-integration/` โ€” DMI integration for an unmodified official SGLang installation - `third_party/DMI-Megatron-Integration/` โ€” DMI integration with its pinned Megatron-LM fork at `third_party/DMI-Megatron-Integration/third_party/megatron-lm/` - `third_party/clickhouse-cpp/` โ€” ClickHouse C++ client linked into the native backend @@ -215,7 +217,7 @@ ClickHouse; use the host benchmark separately with a running server. ## 6. Choose one backend Continue with the [HuggingFace guide](huggingface.md), [vLLM guide](vllm.md), -or [Megatron-LM guide](megatron.md). Use a separate environment and checkout +[SGLang guide](sglang.md), or [Megatron-LM guide](megatron.md). Use a separate environment and checkout for each backend. The HuggingFace path installs a modified Transformers checkout, the vLLM path installs its own official dependency set, and the Megatron-LM path installs its version-matched integration and pinned fork. Do diff --git a/docs/sglang.md b/docs/sglang.md new file mode 100644 index 000000000..8cb679ed8 --- /dev/null +++ b/docs/sglang.md @@ -0,0 +1,117 @@ +# SGLang usage + +DMI provides experimental support for official SGLang 0.5.19 through the version-matched integration +checkout pinned at `third_party/sglang-integration/`. It does not contain an +SGLang fork: the integration is an out-of-tree package that registers through +SGLang's general-plugin entry point (`sglang.srt.plugins`) and its external +model package mechanism. + +## Install the SGLang backend + +Use a dedicated environment and DMI checkout for SGLang. Do not install the +modified HuggingFace integration from `third_party/transformers/` or the vLLM +integration in this environment. Complete the [core installation](install.md) +first. + +SGLang 0.5.19 is not published on PyPI (wheels stop at 0.5.10, which predates +the plugin framework), so install it from the pinned source tag, then install +the integration editable and rebuild the native backend against that +environment's PyTorch (2.13.0 / CUDA 13.0 for this release): + +```bash +git clone --branch v0.5.19 --depth 1 https://github.com/sgl-project/sglang.git +pip install -e sglang/python +pip install -e third_party/sglang-integration/ +make -C native clean +make -C native -j +python -c "from dmi.transport.native import RingConfig; print(RingConfig())" +``` + +The integration checkout's +[`docs/env-setup.md`](https://github.com/ProjectDMX/DMI-SGLang-Integration/blob/main/docs/env-setup.md) +records the exact dependency resolution used for qualification (uv, pinned +torch constraints, no Rust extensions). + +The integration rejects unsupported versions and checks architecture/execution +modes before DMI allocation and weight loading. Upstream device/distributed +initialization may already have happened. Rejected modes include TP/PP/DP, +speculative decoding, LoRA, torch.compile, PD disaggregation, and others. See +the integration's +[port audit](https://github.com/ProjectDMX/DMI-SGLang-Integration/blob/main/docs/v0519-port.md) +for the qualified compatibility cells. + +## Offline API + +For the qualified offline, non-streaming generation API, use `DMIEngine` +instead of `sglang.Engine`. It forwards every +`ServerArgs` keyword unchanged, carries DMI settings to the spawned worker +processes, and tags each output with a lazy `dmi_internal` readback handle: + +```python +from dmi_sglang_integration import DMIEngine + +engine = DMIEngine( + model_path="Qwen/Qwen3-0.6B", + mem_fraction_static=0.5, + disable_prefill_cuda_graph=True, # DMI runs prefill eagerly anyway + dmx_model_id="my-run", + dmx_db_host="localhost", + dmx_hook_selection="resid_pre,final_ln,token_ids,final_logits", + dmx_ring_payload_mb=1024, + dmx_ring_pinned_mb=1024, +) +try: + outputs = engine.generate( + prompt=["The capital of France is"], + sampling_params={"temperature": 0, "max_new_tokens": 8}, + rid=["request-0"], + ) + print(outputs[0]["text"]) + engine.dmi_stop_monitoring() # authoritative flush to ClickHouse + hidden = outputs[0]["dmi_internal"].hidden_states +finally: + engine.shutdown() +``` + +Call `dmi_stop_monitoring()` before `shutdown()`: SGLang terminates its worker +processes without a graceful path, so the explicit RPC is the only durable +flush point. Monitoring is terminal for the engine; create a new `DMIEngine` +for another capture. +If flush fails, repeated stop calls continue to report failure; shutdown is not +proof of successful persistence. Construct engines sequentially in one process. + +Plain `sglang.Engine` also works when the integration package is installed: +set the `DMX_*` environment variables before constructing it. Set +`SGLANG_PLUGINS=none` (or `DMX_SGLANG_ENABLE=0`) to run an unmonitored +baseline in the same environment. + +## Common configuration + +| Setting | Environment variable | Default | +| --- | --- | --- | +| Capture identifier | `DMX_MODEL_ID` | model path | +| Hook selection | `DMX_HOOK_SELECTION` | `sglang-full` (every hook except attention matrices) | +| ClickHouse host / port / database / table | `DMX_DB_HOST`, `DMX_DB_PORT`, `DMX_DB_DATABASE`, `DMX_DB_TABLE` | unset (no persistence), `9000`, `default`, `offload` | +| GPU payload ring / pinned staging | `DMX_RING_PAYLOAD_MB`, `DMX_RING_PINNED_MB` | `1024`, `1024` | +| Device-side padding strip | `DMX_GPU_PADDING_STRIP` | `1` | +| System `libstdc++` preload for conda interpreters | `DMX_SGLANG_PRELOAD_LIBSTDCXX` | `auto` | + +The table describes plain-plugin environment defaults. `DMIEngine` instead +defaults to a unique capture ID and ClickHouse on localhost, exports its +overrides only during worker startup, and restores the parent's environment. + +## Troubleshooting + +- **`GLIBCXX_3.4.30 not found` in the scheduler process** -- a conda-provided + interpreter binds conda's older `libstdc++` before DMI's native backend + loads. `DMIEngine` preloads the system library in worker processes + automatically; for plain `sglang.Engine` or `sglang serve`, export + `LD_PRELOAD=/usr/lib/x86_64-linux-gnu/libstdc++.so.6`. +- **`DMI monitoring is not active in this scheduler`** -- the plugin did not + register. Check `pip show DMI-SGLang-Integration`, `SGLANG_PLUGINS`, and + `DMX_SGLANG_ENABLE`. +- **Overlap scheduler rows** -- with the default overlap scheduler SGLang runs + one extra decode forward per request after its final decision; DMI captures + that forward too (one extra `token_ids` row holding the last generated token + and one unused `final_logits` decision). Pass `disable_overlap_schedule=True` + for a capture that matches the public output exactly. diff --git a/third_party/sglang-integration b/third_party/sglang-integration new file mode 160000 index 000000000..97c8fb4fe --- /dev/null +++ b/third_party/sglang-integration @@ -0,0 +1 @@ +Subproject commit 97c8fb4fea1184fb06e1f9df4e73e4cc93899fa9