Conversation
Pins third_party/sglang-integration to the DMI-SGLang-Integration checkout for official SGLang 0.5.19 (out-of-tree plugin, no SGLang fork), mirroring third_party/vllm-integration, and documents installation, the offline DMIEngine API, configuration, and troubleshooting in docs/sglang.md. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
third_party/sglang-integration, pointing to the newly published ProjectDMX/DMI-SGLang-Integration at 97c8fb4fea1184fb06e1f9df4e73e4cc93899fa9.docs/sglang.mdand update README/install guidance for a dedicated SGLang environment, the offlineDMIEngineAPI, and explicit flush before shutdown.Scope and compatibility identity
Experimental offline, non-streaming Qwen3 dense, TP1/PP1 integration for official SGLang v0.5.19, source commit
0bcd822377da7b5718e674eaf9c870d349424dd1.Runtime used for qualification: DMI 1.2.0 / API v1, Python 3.12, torch 2.13.0+cu130, sgl-kernel 0.4.6.post1, FlashInfer 0.6.18, RTX 4090, BF16, ClickHouse 25.12.2.54. A matching native build and a version-isolated environment are required.
The original implementation report includes Qwen3-0.6B eager/graph/padded-graph/overlap and Qwen3-8B graph cells. The follow-up review independently reran only the bounded Qwen3-0.6B padded-graph cell; it did not independently reproduce all five cells.
The tested core baseline was
c222f18c4db55f6f2d7136b1d6608c0b540b273a(the SGLang branch adds documentation/submodule changes on top). Current main has since advanced; this PR does not claim a full runtime requalification of those later core changes.Review fixes included in the pinned integration
stream=Truebefore generation.Validation
[1, 2, 4], 3 requests, 8 greedy tokens each, overlap off, prefill graph off in both modes.resid_pre,final_ln,token_ids,final_logits.git diff --check; integration commit fetched successfully from the public remote before pinning.The GPU rerun covered environment/flush/storage fixes. The later positional-stream guard was covered by CPU rejection and nonstreaming-forwarding tests, not another GPU run. Full stock logits bitwise equality and all-hook capture are not claimed.
Evidence at the pinned integration commit:
Limits
TP/PP/DP beyond the supported single-device topology, MoE/other architectures, speculative decoding, LoRA, PD disaggregation, TBO, torch.compile/tc_piecewise decode, and non-generation models are rejected. Prefill graph is forced eager. Serving, streaming, cache-hit/chunked-prefill interaction matrices, quantization and full-hook stress are not qualified. Overlap can capture one executed lookahead forward beyond public completion; disable overlap for exact public-step alignment.
No release tag is created and no automatic merge is requested.