Skip to content

[Qwen3.8 R11] Native image and video understanding #1154

Description

@AtlantisPleb

R11: Native Vision

Native vision remains a separate roadmap lane after text support.

It requires:

  • exact image and video preprocessing
  • the 27-layer vision encoder
  • spatial and temporal patch handling
  • vision-to-text projection into the 5,120-wide decoder
  • attachment identity and preprocessing receipts
  • image and video reference fixtures
  • native output parity against an upstream-supported runtime
  • bounded image/video size, frame count, timeout, and memory admission

Until R11 lands, prompt marker projection must not be described as image or
video understanding.

Dependencies

  • Blocked on stable R2 prompt marker semantics, R3 multimodal tensor inventory, R8 serving, and a separately qualified mmproj/vision artifact plan.

Scope rules

  • This phase owns native image/video preprocessing, the 27-layer vision encoder, spatial/temporal patches, and projection into the 5,120-wide text decoder.
  • Download and admit vision projector artifacts separately with immutable source, size, digest, dtype, memory, and processor provenance.
  • Prompt marker rendering from R2 is not image/video understanding and must remain a refusal on text-only lanes.
  • Keep media attachment identity, preprocessing configuration, runtime/backend identity, fallback truth, and bounded limits machine-legible.

Required evidence

  • Synthetic processor/operator tests plus real image/video reference fixtures and native parity against a pinned upstream-supported runtime.
  • Admission/refusal rows cover dimensions, frame counts, formats, timeouts, memory, malformed inputs, and missing projectors.
  • Serving tests cover chat/responses attachment identity, streaming, tools with media, and text-only refusal behavior.

Validation and closure

  • Run focused processor, vision-model, backend, and OpenAI serving tests on retained reference fixtures.
  • Merge to main and comment with commit SHA, artifact/source digests, exact commands/results, parity cases, supported bounds, residency, fallback truth, and refusals.
  • Close only after origin/main supports native bounded image and video understanding; marker-only behavior cannot close it.

Canonical references

  • docs/qwen38/IMPLEMENTATION_ROADMAP.md, R11
  • docs/qwen38/MODEL_FACTS.md
  • docs/qwen38/UPSTREAM_ARTIFACT_INDEX.md
  • docs/INFERENCE_ENGINE.md

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestpriority:P1Important follow-up after the P0 path is movingqwen38Qwen3.8 implementation roadmaproadmapRoadmap worktype:compatibilityCompatibility with an external reference format

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions