Skip to content

Add experimental SGLang 0.5.19 integration - #136

Open
zhaoyboo wants to merge 2 commits into
mainfrom
sglang-support
Open

zhaoyboo wants to merge 2 commits into
mainfrom
sglang-support

Conversation

@zhaoyboo

Copy link
Copy Markdown
Collaborator

Summary

  • Add third_party/sglang-integration, pointing to the newly published ProjectDMX/DMI-SGLang-Integration at 97c8fb4fea1184fb06e1f9df4e73e4cc93899fa9.
  • Add docs/sglang.md and update README/install guidance for a dedicated SGLang environment, the offline DMIEngine API, and explicit flush before shutdown.
  • Keep backend-specific code in the integration repository: this core PR changes only the submodule and documentation. It does not fork SGLang or change DMI native/API implementation.

Scope and compatibility identity

Experimental offline, non-streaming Qwen3 dense, TP1/PP1 integration for official SGLang v0.5.19, source commit 0bcd822377da7b5718e674eaf9c870d349424dd1.

Runtime used for qualification: DMI 1.2.0 / API v1, Python 3.12, torch 2.13.0+cu130, sgl-kernel 0.4.6.post1, FlashInfer 0.6.18, RTX 4090, BF16, ClickHouse 25.12.2.54. A matching native build and a version-isolated environment are required.

The original implementation report includes Qwen3-0.6B eager/graph/padded-graph/overlap and Qwen3-8B graph cells. The follow-up review independently reran only the bounded Qwen3-0.6B padded-graph cell; it did not independently reproduce all five cells.

The tested core baseline was c222f18c4db55f6f2d7136b1d6608c0b540b273a (the SGLang branch adds documentation/submodule changes on top). Current main has since advanced; this PR does not claim a full runtime requalification of those later core changes.

Review fixes included in the pinned integration

  • Restore parent DMX/LD_PRELOAD environment after worker startup; prevent sequential engines from inheriting auto-generated capture IDs or per-engine overrides.
  • Preserve failed flush status on both engine and worker; a repeated stop cannot turn a failed capture into an apparent success.
  • Reject positional as well as keyword stream=True before generation.
  • Check every storage row for the unique capture, rather than filtering away unexpected request IDs before validation.
  • Correct initialization-order documentation and explicitly exclude concurrent engine construction in one process.

Validation

  • Final CPU gate: 101 passed (16 new regressions).
  • Runtime audit: 23 resolved targets, 0 errors / 0 audit warnings. Known optional bailing MoE imports lack vLLM; those architectures are out of scope.
  • Independent stock/monitored GPU processes: Qwen3-0.6B, captured decode buckets [1, 2, 4], 3 requests, 8 greedy tokens each, overlap off, prefill graph off in both modes.
  • Public output equality; 234 logprob values equal at zero tolerance, max difference 0.0.
  • 744 exact-once stored rows, with token/range/type/argmax checks and a capture-wide storage query. Selected hooks: resid_pre,final_ln,token_ids,final_logits.
  • git diff --check; integration commit fetched successfully from the public remote before pinning.

The GPU rerun covered environment/flush/storage fixes. The later positional-stream guard was covered by CPU rejection and nonstreaming-forwarding tests, not another GPU run. Full stock logits bitwise equality and all-hook capture are not claimed.

Evidence at the pinned integration commit:

Limits

TP/PP/DP beyond the supported single-device topology, MoE/other architectures, speculative decoding, LoRA, PD disaggregation, TBO, torch.compile/tc_piecewise decode, and non-generation models are rejected. Prefill graph is forced eager. Serving, streaming, cache-hit/chunked-prefill interaction matrices, quantization and full-hook stress are not qualified. Overlap can capture one executed lookahead forward beyond public completion; disable overlap for exact public-step alignment.

No release tag is created and no automatic merge is requested.

zhaoyboo and others added 2 commits September 16, 2026 14:17
Pins third_party/sglang-integration to the DMI-SGLang-Integration checkout
for official SGLang 0.5.19 (out-of-tree plugin, no SGLang fork), mirroring
third_party/vllm-integration, and documents installation, the offline
DMIEngine API, configuration, and troubleshooting in docs/sglang.md.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant