R11: Native Vision
Native vision remains a separate roadmap lane after text support.
It requires:
- exact image and video preprocessing
- the 27-layer vision encoder
- spatial and temporal patch handling
- vision-to-text projection into the 5,120-wide decoder
- attachment identity and preprocessing receipts
- image and video reference fixtures
- native output parity against an upstream-supported runtime
- bounded image/video size, frame count, timeout, and memory admission
Until R11 lands, prompt marker projection must not be described as image or
video understanding.
Dependencies
- Blocked on stable R2 prompt marker semantics, R3 multimodal tensor inventory, R8 serving, and a separately qualified mmproj/vision artifact plan.
Scope rules
- This phase owns native image/video preprocessing, the 27-layer vision encoder, spatial/temporal patches, and projection into the 5,120-wide text decoder.
- Download and admit vision projector artifacts separately with immutable source, size, digest, dtype, memory, and processor provenance.
- Prompt marker rendering from R2 is not image/video understanding and must remain a refusal on text-only lanes.
- Keep media attachment identity, preprocessing configuration, runtime/backend identity, fallback truth, and bounded limits machine-legible.
Required evidence
- Synthetic processor/operator tests plus real image/video reference fixtures and native parity against a pinned upstream-supported runtime.
- Admission/refusal rows cover dimensions, frame counts, formats, timeouts, memory, malformed inputs, and missing projectors.
- Serving tests cover chat/responses attachment identity, streaming, tools with media, and text-only refusal behavior.
Validation and closure
- Run focused processor, vision-model, backend, and OpenAI serving tests on retained reference fixtures.
- Merge to main and comment with commit SHA, artifact/source digests, exact commands/results, parity cases, supported bounds, residency, fallback truth, and refusals.
- Close only after origin/main supports native bounded image and video understanding; marker-only behavior cannot close it.
Canonical references
- docs/qwen38/IMPLEMENTATION_ROADMAP.md, R11
- docs/qwen38/MODEL_FACTS.md
- docs/qwen38/UPSTREAM_ARTIFACT_INDEX.md
- docs/INFERENCE_ENGINE.md
R11: Native Vision
Native vision remains a separate roadmap lane after text support.
It requires:
Until R11 lands, prompt marker projection must not be described as image or
video understanding.
Dependencies
Scope rules
Required evidence
Validation and closure
Canonical references