-
Notifications
You must be signed in to change notification settings - Fork 101
Pull requests: flagos-ai/vllm-plugin-FL
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
fix(worker): account for CUDA graph memory in KV cache sizing
core
#381
opened Aug 14, 2026 by
CherryLemon
Collaborator
Loading…
2 of 3 tasks
feat(hygon): support INT4 inference for Qwen3.6 and GLM-5.2
core
#379
opened Aug 13, 2026 by
lisa2342
Loading…
fix(metax): consolidate vLLM 0.24.0 compat shims into the plugin
core
#377
opened Aug 13, 2026 by
tengqm
Contributor
Loading…
feat(kunlunxin): add GDN fused post-conv Triton bypass for Kunlunxin XPU
core
#374
opened Aug 13, 2026 by
dongjibin1996
Loading…
[CICD]: Add Enflame S60 CI support
ci
docs
tests
#373
opened Aug 13, 2026 by
HermiaHuan
Collaborator
Loading…
feat(metax): support DeepSeek V4 W8A8 INT8 inference
core
#372
opened Aug 12, 2026 by
liyu030
Loading…
3 tasks
feat(ascend): add NPU device context, head_dim>192 attention fallback, and FLAGCX_PATH-free backend resolution
core
#369
opened Aug 12, 2026 by
Joiin0392
Loading…
1 of 3 tasks
feat(ascend): add build and smoke-test scripts for Ascend NPU deployment
#370
opened Aug 12, 2026 by
Joiin0392
Loading…
3 tasks done
[Kunlunxin] Fix garbled output with CUDA Graph
core
#368
opened Aug 11, 2026 by
0songHan
Loading…
3 tasks
[Sunrise] Add cuda-graph, int8 and more model support on Sunrise
core
#366
opened Aug 11, 2026 by
cyx11111
Loading…
3 tasks
ascend: blacklist lift_fresh/_to_copy (coreDim=0 on scalar tensor init)
core
#361
opened Aug 10, 2026 by
tengqm
Contributor
Loading…
Cherry-pick PR #297 to v0.3.0-dev: Add thead vendor attention backend
core
#360
opened Aug 10, 2026 by
physics31415926
Collaborator
•
Draft
support mla atten in flaggems
core
#359
opened Aug 10, 2026 by
ZenithJoyH
Contributor
Loading…
3 tasks
support Qwen3.5 FP8 Model on GCU (Enflame)
core
#358
opened Aug 9, 2026 by
liuyancong-enflame-tech
Contributor
Loading…
1 of 3 tasks
gcu: enable enflame GCU300 via vLLM native FLASH_ATTN backend
core
#357
opened Aug 9, 2026 by
tengqm
Contributor
Loading…
feat(cambricon): enable MLU590 (device map + graph class)
core
#355
opened Aug 8, 2026 by
tengqm
Contributor
Loading…
[hot-fix] fix GroupedTopKRouterFL routing and MLA decode method name
core
#352
opened Aug 7, 2026 by
tonyh168
Loading…
Qwen36_dense_moe performance surpasses vllm_Ascend
core
docs
tests
#349
opened Aug 7, 2026 by
appleinsky
Loading…
3 tasks
perf(metax): optimize single-request prefill attention
core
#345
opened Aug 5, 2026 by
liyu030
Loading…
2 of 3 tasks
feat(compilation): add BreakableCUDAGraphWrapper support for OOT vendor backends (vLLM 0.24.0+)
core
tests
#342
opened Aug 3, 2026 by
physics31415926
Collaborator
Loading…
2 of 3 tasks
feat(quantization): adapt W8A8 inference to vLLM 0.24
core
docs
tests
#336
opened Aug 3, 2026 by
rdzhu225
Contributor
Loading…
Previous Next
ProTip!
Filter pull requests by the default branch with base:main.