-
Notifications
You must be signed in to change notification settings - Fork 44
All issues
Issue creation is restricted in this repository
Issues
is:issue state:open
is:issue state:open
Search results
fix(test): mlxcel-core CUDA test binary crashes at the default thread count
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layersplatform:linuxLinux (CUDA / packaging) specificLinux (CUDA / packaging) specificpriority:mediumMedium priorityMedium prioritystatus:investigationFeasibility spike / under investigationFeasibility spike / under investigationtype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#1566 In lablup/mlxcel;perf: first-token latency is dominated by lazy weight materialization
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layerspriority:mediumMedium priorityMedium prioritystatus:readyReady to be worked onReady to be worked ontype:performancePerformance improvementsPerformance improvementsStatus: Open.#1564 In lablup/mlxcel;fix(inference): fused_sample_probs differs by 1 ULP at temperature 1.0 on sm_70
area:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)platform:linuxLinux (CUDA / packaging) specificLinux (CUDA / packaging) specificpriority:mediumMedium priorityMedium prioritystatus:investigationFeasibility spike / under investigationFeasibility spike / under investigationtype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#1563 In lablup/mlxcel;fix(cuda/quant): quantized prefill is not bitwise reproducible on sm_70, and its determinism test never runs
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layersplatform:linuxLinux (CUDA / packaging) specificLinux (CUDA / packaging) specificpriority:mediumMedium priorityMedium prioritystatus:investigationFeasibility spike / under investigationFeasibility spike / under investigationtype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#1558 In lablup/mlxcel;test(models): validate Inkling omni on real checkpoints
arch:moeSparse mixture-of-experts decoderSparse mixture-of-experts decoderarea:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)area:modelsModel architectures, weights, loading, metadataModel architectures, weights, loading, metadatamodelsize:xlarge4-bit checkpoint > 100GB; exceeds a 128GB dev box, needs bigger hardware4-bit checkpoint > 100GB; exceeds a 128GB dev box, needs bigger hardwaremodeltype:omniMulti-modal omni model (text + vision + audio)Multi-modal omni model (text + vision + audio)platform:macosmacOS (Apple Silicon) specificmacOS (Apple Silicon) specificpriority:mediumMedium priorityMedium prioritystatus:backlogIn the backlog, not yet readyIn the backlog, not yet readytype:testTest related changesTest related changesStatus: Open.#1549 In lablup/mlxcel;perf(cuda/quant): qmm_sm70 — Volta tensor-core MMA path for quantized GEMM
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layersplatform:linuxLinux (CUDA / packaging) specificLinux (CUDA / packaging) specificpriority:highHigh priorityHigh prioritystatus:investigationFeasibility spike / under investigationFeasibility spike / under investigationtype:performancePerformance improvementsPerformance improvementsStatus: Open.#1543 In lablup/mlxcel;perf(cuda): f16 activation policy below Ampere (bf16 has no Volta hardware path)
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layersplatform:linuxLinux (CUDA / packaging) specificLinux (CUDA / packaging) specificpriority:highHigh priorityHigh prioritystatus:investigationFeasibility spike / under investigationFeasibility spike / under investigationtype:performancePerformance improvementsPerformance improvementsStatus: Open.#1542 In lablup/mlxcel;epic: NVIDIA Volta (sm_70) inference acceleration program
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layersplatform:linuxLinux (CUDA / packaging) specificLinux (CUDA / packaging) specificpriority:mediumMedium priorityMedium prioritystatus:in-progressCurrently being worked onCurrently being worked ontype:performancePerformance improvementsPerformance improvementsStatus: Open.#1536 In lablup/mlxcel;feat(server): server-side MCP tool loop through the /v1/responses built-in tool contract
area:architectureArchitecture and code structure changesArchitecture and code structure changespriority:lowLow priorityLow prioritystatus:backlogIn the backlog, not yet readyIn the backlog, not yet readytype:enhancementNew features, capabilities, or significant additionsNew features, capabilities, or significant additionstype:securitySecurity vulnerability or fixSecurity vulnerability or fixStatus: Open.#1457 In lablup/mlxcel;feat(server): late-interaction scoring endpoint for multi-vector embedders
area:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)priority:lowLow priorityLow prioritystatus:readyReady to be worked onReady to be worked ontype:enhancementNew features, capabilities, or significant additionsNew features, capabilities, or significant additionsStatus: Open.#1426 In lablup/mlxcel;refactor(embeddings): hoist family-local weight-key handling into the shared embedding sanitizer
area:modelsModel architectures, weights, loading, metadataModel architectures, weights, loading, metadatapriority:lowLow priorityLow prioritystatus:readyReady to be worked onReady to be worked ontype:refactorCode restructuring without changing functionalityCode restructuring without changing functionalityStatus: Open.#1424 In lablup/mlxcel;fix(qwen2_5_vl): the generation loader does not normalize the raw HuggingFace Conv3d patch-embed layout
area:modelsModel architectures, weights, loading, metadataModel architectures, weights, loading, metadatamodeltype:vlmVision-language modelVision-language modelpriority:mediumMedium priorityMedium prioritystatus:readyReady to be worked onReady to be worked ontype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#1423 In lablup/mlxcel;