-
Notifications
You must be signed in to change notification settings - Fork 51
All issues
Issue creation is restricted in this repository
Issues
is:issue state:open
is:issue state:open
Search results
perf(qwen_vl): text decode falls to half of mlx-lm as context grows
area:modelsModel architectures, weights, loading, metadataModel architectures, weights, loading, metadatapriority:mediumMedium priorityMedium prioritystatus:readyReady to be worked onReady to be worked ontype:performancePerformance improvementsPerformance improvementsStatus: Open.#1686 In lablup/mlxcel;perf(m5): FalconH1 and GraniteMoeHybrid decode at a third of mlx-lm
arch:hybridHybrid attention stack (linear/sliding/full or attention+SSM interleave)Hybrid attention stack (linear/sliding/full or attention+SSM interleave)arch:ssmState-space / recurrent model (Mamba, RWKV)State-space / recurrent model (Mamba, RWKV)area:benchmarkBenchmark harness and performance measurement (bench_*.sh, /update-benchmarks)Benchmark harness and performance measurement (bench_*.sh, /update-benchmarks)area:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)area:modelsModel architectures, weights, loading, metadataModel architectures, weights, loading, metadataplatform:macosmacOS (Apple Silicon) specificmacOS (Apple Silicon) specificpriority:highHigh priorityHigh prioritystatus:readyReady to be worked onReady to be worked ontype:performancePerformance improvementsPerformance improvementsStatus: Open.#1685 In lablup/mlxcel;fix(minicpm_v4_6): image gets 16 vision tokens against the reference 64
area:modelsModel architectures, weights, loading, metadataModel architectures, weights, loading, metadatamodelsize:small4-bit checkpoint <= 10GB; fast iteration, safe for smoke tests4-bit checkpoint <= 10GB; fast iteration, safe for smoke testsmodeltype:vlmVision-language modelVision-language modelpriority:highHigh priorityHigh prioritystatus:readyReady to be worked onReady to be worked ontype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#1684 In lablup/mlxcel;fix(granite_vision): descriptive prompts draw a refusal, colour questions work
area:modelsModel architectures, weights, loading, metadataModel architectures, weights, loading, metadatamodelsize:small4-bit checkpoint <= 10GB; fast iteration, safe for smoke tests4-bit checkpoint <= 10GB; fast iteration, safe for smoke testsmodeltype:vlmVision-language modelVision-language modelpriority:highHigh priorityHigh prioritystatus:readyReady to be worked onReady to be worked ontype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#1683 In lablup/mlxcel;chore(ci): extend Dependabot to the python/ and webpage/site dependency roots
priority:lowLow priorityLow prioritystatus:readyReady to be worked onReady to be worked ontype:choreMaintenance tasks (build, CI, etc.)Maintenance tasks (build, CI, etc.)Status: Open.#1670 In lablup/mlxcel;test(python): cover the _common.py URL normalization helpers
priority:lowLow priorityLow prioritystatus:readyReady to be worked onReady to be worked ontype:testTest related changesTest related changesStatus: Open.#1674 In lablup/mlxcel;docs(recipes): add a recipes/README explaining the registry snapshot lifecycle
priority:lowLow priorityLow prioritystatus:readyReady to be worked onReady to be worked ontype:docsDocumentation improvements or additionsDocumentation improvements or additionsStatus: Open.#1677 In lablup/mlxcel;docs: update cascade-attention prose from "benchmark missing" to "measured and rejected"
priority:lowLow priorityLow prioritystatus:readyReady to be worked onReady to be worked ontype:docsDocumentation improvements or additionsDocumentation improvements or additionsStatus: Open.#1676 In lablup/mlxcel;docs(cli): sync the serve --tp-size family list with generate and mlxcel-server
priority:lowLow priorityLow prioritystatus:readyReady to be worked onReady to be worked ontype:docsDocumentation improvements or additionsDocumentation improvements or additionsStatus: Open.#1672 In lablup/mlxcel;docs(server): document that --api-prefix removes the health endpoints from the public set
priority:lowLow priorityLow prioritystatus:readyReady to be worked onReady to be worked ontype:docsDocumentation improvements or additionsDocumentation improvements or additionsStatus: Open.#1671 In lablup/mlxcel;chore(deps): add the shipped x86_64 Linux CUDA target to deny.toml graph targets
priority:lowLow priorityLow prioritystatus:readyReady to be worked onReady to be worked ontype:choreMaintenance tasks (build, CI, etc.)Maintenance tasks (build, CI, etc.)Status: Open.#1675 In lablup/mlxcel;docs(reports): Korean report 1059 is missing section 3.2 from the English original
priority:lowLow priorityLow prioritystatus:readyReady to be worked onReady to be worked ontype:docsDocumentation improvements or additionsDocumentation improvements or additionsStatus: Open.#1673 In lablup/mlxcel;