A curated list of ML systems papers posted on arXiv.
- Paper List for Machine Learning Systems (arXiv preprints)
- Data Processing
- Training System
- Inference System
- Attention Optimization
- Mixture of Experts (MoE)
- Communication Optimization & Network Infrastructure for Distributed ML
- Fault tolerance & Straggler mitigation
- GPU Memory Management & Optimization
- GPU Sharing
- Compiler
- GPU Kernel Optimization
- LLM Long Context
- Model Compression
- Federated Learning
- Privacy-Preserving ML
- ML APIs & Application-Side Optimization
- ML for Systems
- Energy Efficiency
- Retrieval-Augmented Generation (RAG)
- Simulation
- Systems for Agentic AI
- Multimodal
- Hybrid LLMs
- Others
Data pipeline optimization
- [arxiv'26] Lakestream: A Consistent and Brokerless Data Plane for Large Foundation Model Training
- [arxiv'25] Scalable and Performant Data Loading
- [arxiv'25] OVERLORD: Ultimate Scaling of DataLoader for Multi-Source Large Foundation Model Training
- [arxiv'25] The Streaming Batch Model for Efficient and Fault-Tolerant Heterogeneous Execution
- [arxiv'25] In-Network Preprocessing of Recommender Systems on Multi-Tenant SmartNICs
- [arxiv'24] TensorSocket: Shared Data Loading for Deep Learning Training
- [arxiv'24] Efficient Tabular Data Preprocessing of ML Pipelines
Preprocessing stalls
- [arxiv'24] PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
- [arxiv'23] Rinas: Training with Dataset Shuffling Can Be General and Fast
Specific workloads (GNN, DLRM)
- [arxiv'23] Towards Data-centric Graph Machine Learning: Review and Outlook
- [arxiv'23] FlexShard: Flexible Sharding for Industry-Scale Sequence Recommendation Models
- [arxiv'23] MTrainS: Improving DLRM training efficiency using heterogeneous memories
LLM data plane
- [arxiv'25] DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
- [arxiv'25] Mixtera: A Data Plane for Foundation Model Training
- [arxiv'26] CvxCluster: Solving Large, Complex, Granular Resource Allocation Problems 100-1000x Faster
- [arxiv'26] PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
- [arxiv'26] BOA Constrictor: Squeezing Performance out of GPUs in the Cloud via Budget-Optimal Allocation
- [arxiv'26] SkyNomad: On Using Multi-Region Spot Instances to Minimize AI Batch Job Cost
- [arxiv'25] Semantic-Aware Scheduling for GPU Clusters with Large Language Models
- [arxiv'25] Holistic Heterogeneous Scheduling for Autonomous Applications using Fine-grained, Multi-XPU Abstraction
- [arxiv'25] Tesserae: Scalable Placement Policies for Deep Learning Workloads
- [arxiv'25] LeMix: Unified Scheduling for LLM Training and Inference on Multi-GPU Systems
- [arxiv'25] TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
- [arxiv'24] Zeal: Rethinking Large-Scale Resource Allocation with "Decouple and Decompose"
- [arxiv'24] Rubick: Exploiting Job Reconfigurability for Deep Learning Cluster Scheduling
- [arxiv'23] Energy-Efficient GPU Clusters Scheduling for Deep Learning
- [arxiv'22] Singularity: Planet-Scale, Preemptive and Elastic Scheduling of AI Workloads
-
[arxiv'26] Tangram: Hiding GPU Heterogeneity for Efficient LLM Parallelization
-
[arxiv'26] Breaking the Bubble: Asynchronous Pipeline Parallel Training with Bounded Weight Inconsistency
-
[arxiv'26] Piper: A Programmable Distributed Training System
-
[arxiv'26] DynaTrain: Fast Online Parallelism Switching for Elastic LLM Training
-
[arxiv'26] JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials
-
[arxiv'26] A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability
-
[arxiv'26] Accelerating Compound LLM Training Workloads with Maestro
-
[arxiv'26] AutoScout: Structured Optimization for Automating ML System Configuration
-
[arxiv'26] NEST: Network- and Memory-Aware Device Placement For Distributed Deep Learning
-
[arxiv'26] veScale-FSDP: Flexible and High-Performance FSDP at Scale
-
[arxiv'26] DHP: Efficient Scaling of MLLM Training with Dynamic Hybrid Parallelism
-
[arxiv'26] Joint Training on AMD and NVIDIA GPUs
-
[Survey 🔍] [arxiv'26] Distributed Hybrid Parallelism for Large Language Models: Comparative Study and System Design Guide
-
[arxiv'26] tLoRA: Efficient Multi-LoRA Training with Elastic Shared Super-Models
-
[arxiv'26] TimelyFreeze: Adaptive Parameter Freezing Mechanism for Pipeline Parallelism
-
[arxiv'26] Mitigating Staleness in Asynchronous Pipeline Parallelism via Basis Rotation
-
[arxiv'25] Diving into 3D Parallelism with Heterogeneous Spot Instance GPUs: Design and Implications
-
[arxiv'25] SIGMA: An AI-Empowered Training Stack on Early-Life Hardware
-
[arxiv'25] BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models
-
[arxiv'25] AsyncHZP: Hierarchical ZeRO Parallelism with Asynchronous Scheduling for Scalable LLM Training
-
[arxiv'25] A Flexible Programmable Pipeline Parallelism Framework for Efficient DNN Training
-
[arxiv'25] SlimPack: Fine-Grained Asymmetric Packing for Balanced and Efficient Variable-Length LLM Training
-
[arxiv'25] AdaPtis: Reducing Pipeline Bubbles with Adaptive Pipeline Parallelism on Heterogeneous Models
-
[arxiv'25] HAPT: Heterogeneity-Aware Automated Parallel Training on Heterogeneous Clusters
-
[arxiv'25] Scaling Up Data Parallelism in Decentralized Deep Learning
-
[arxiv'25] Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
-
[arxiv'25] TrainVerify: Equivalence-Based Verification for Distributed LLM Training
-
[arxiv'25] Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage
-
[arxiv'25] ZenFlow: Enabling Stall-Free Offloading Training via Asynchronous Updates
-
[arxiv'25] Rethinking Dynamic Networks and Heterogeneous Computing with Automatic Parallelization
-
[arxiv'25] H2:Towards Efficient Large-Scale LLM Training on Hyper-Heterogeneous Cluster over 1,000 Chips
-
[arxiv'25] Balanced and Elastic End-to-end Training of Dynamic LLMs
-
[arxiv'25] ZenFlow: Enabling Stall-Free Offloading Training via Asynchronous Updates
-
[arxiv'25] Parallel Scaling Law for Language Models
-
[arxiv'25] Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed Clusters
-
[arxiv'25] You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models
-
[arxiv'25] WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training
-
[arxiv'25] Nonuniform-Tensor-Parallelism: Mitigating GPU failure impact for Scaled-up LLM Training
-
[arxiv'25] Cornstarch: Distributed Multimodal Training Must Be Multimodality-Aware
-
[arxiv'25] PipeOffload: Improving Scalability of Pipeline Parallelism with Memory Optimization
-
[arxiv'25] AutoHete: An Automatic and Efficient Heterogeneous Training System for LLMs
-
[arxiv'25] Astra: Efficient and Money-saving Automatic Parallel Strategies Search on Heterogeneous GPUs
-
[arxiv'25] Scaling Inference-Efficient Language Models
-
[arxiv'25] MiniMax-01: Scaling Foundation Models with Lightning Attention
-
[arxiv'24] Automatically Planning Optimal Parallel Strategy for Large Language Models
-
[arxiv'24] Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters
-
[arxiv'24] Scaling Deep Learning Training with MPMD Pipeline Parallelism
-
[arxiv'24] Demystifying Workload Imbalances in Large Transformer Model Training over Variable-length Sequences
-
[arxiv'24] HETHUB: A Distributed Training System with Heterogeneous Cluster for Large-Scale Models
-
[arxiv'24] Data-Centric and Heterogeneity-Adaptive Sequence Parallelism for Efficient LLM Training
-
[arxiv'24] Accelerating Large Language Model Training with 4D Parallelism and Memory Consumption Estimator
-
[arxiv'24] BitPipe: Bidirectional Interleaved Pipeline Parallelism for Accelerating Large Models Training
-
[arxiv'24] Cephalo: Harnessing Heterogeneous GPU Clusters for Training Transformer Models
-
[arxiv'24] SimpleFSDP: Simpler Fully Sharded Data Parallel with torch.compile
-
[arxiv'24] FusionLLM: A Decentralized LLM Training System on Geo-distributed GPUs with Adaptive Compression
-
[arxiv'24] PipeFill: Using GPUs During Bubbles in Pipeline-parallel LLM Training
-
[arxiv'24] Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
-
[arxiv'24] Efficient Multi-Task Large Model Training via Data Heterogeneity-aware Model Management
-
[arxiv'24] FlashFlex: Accommodating Large Language Model Training over Heterogeneous Environment
-
[arxiv'24] PARALLELGPUOS: A Concurrent OS-level GPU Checkpoint and Restore System using Validated Speculation
-
[arxiv'24] Unicron: Economizing Self-Healing LLM Training at Scale
-
[arxiv'24] TBA: Faster Large Language Model Training Using SSD-Based Activation Offloading
-
[arxiv'24] Optimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble Exploitation
-
[Survey 🔍] [arxiv'24] Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
-
[arxiv'24] LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
-
[arxiv'24] PAFT: A Parallel Training Paradigm for Effective LLM Fine-Tuning
-
[arxiv'24] BurstAttention: An Efficient Distributed Attention Framework for Extremely Long Sequences
-
[arxiv'24] Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
-
[arxiv'24] Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
-
[arxiv'24] GRAWA: Gradient-based Weighted Averaging for Distributed Training of Deep Learning Models
-
[arxiv'24] BitDelta: Your Fine-Tune May Only Be Worth One Bit
-
[arxiv'24] NutePrune: Efficient Progressive Pruning with Numerous Teachers for Large Language Models
-
[arxiv'24] Accelerating Parallel Sampling of Diffusion Models
-
[arxiv'24] Training DNN Models over Heterogeneous Clusters with Optimal Performance
-
[arxiv'24] Breaking MLPerf Training: A Case Study on Optimizing BERT
-
[arxiv'24] LocMoE: A Low-overhead MoE for Large Language Model Training
-
[arxiv'24] Re-evaluating the Memory-balanced Pipeline Parallelism: BPipe
-
[arxiv'23] SuperScaler: Supporting Flexible DNN Parallelization via a Unified Abstraction
-
[arxiv'23] ASPEN: High-Throughput LoRA Fine-Tuning of Large Language Models with a Single GPU
-
[arxiv'23] FlexModel: A Framework for Interpretability of Distributed Large Language Models
-
[arxiv'23] Holmes: Towards Distributed Training Across Clusters with Heterogeneous NIC Environment
-
[arxiv'23] RTP: Rethinking Tensor Parallelism with Memory Deduplication
-
[arxiv'23] FP8-LM: Training FP8 Large Language Models
-
[arxiv'23] Redco: A Lightweight Tool to Automate Distributed Training of LLMs on Any GPU/TPUs
-
[arxiv'23] FLM-101B: An Open LLM and How to Train It with $100K Budget
-
[arxiv'23] UniAP: Unifying Inter- and Intra-Layer Automatic Parallelism by Mixed Integer Quadratic Programming
-
[arxiv'23] Modeling Parallel Programs using Large Language Models
-
[arxiv'23] Proteus: Simulating the Performance of Distributed DNN Training
-
[arxiv'23] Automated Tensor Model Parallelism with Overlapped Communication for Efficient Foundation Model Training
-
[arxiv'23] Decoupled Model Schedule for Deep Learning Training
-
[arxiv'23] RAF: Holistic Compilation for Deep Learning Model Training
-
[arxiv'23] Ada-Grouper: Accelerating Pipeline Parallelism in Preempted Network by Adaptive Group-Scheduling for Micro-Batches
-
[arxiv'23] Does compressing activations help model parallel training?
-
[arxiv'23] Colossal-Auto: Unified Automation of Parallelization and Activation Checkpoint for Large-scale Models
-
[arxiv'23] Scaling Vision Transformers to 22 Billion Parameters
-
[arxiv'23] Auto-Parallelizing Large Models with Rhino: A Systematic Approach on Production AI Platform
-
[arxiv'23] TAP: Accelerating Large-Scale DNN Training Through Tensor Automatic Parallelisation
-
[arxiv'23] SuperScaler: Supporting Flexible DNN Parallelization via a Unified Abstraction
-
[arxiv'23] ATP: Adaptive Tensor Parallelism for Foundation Models
-
[arxiv'22] Colossal-AI: A Unified Deep Learning System For Large-Scale Parallel Training
-
[arxiv'22] Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
-
[arxiv'21] Amazon SageMaker Model Parallelism: A General and Flexible Framework for Large Model Training
-
[arxiv'21] GSPMD: General and Scalable Parallelization for ML Computation Graphs
-
[arxiv'20] Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
-
[arxiv'20] Scaling Laws for Neural Language Models
- [arxiv'26] AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning
- [arxiv'26] JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models
- [arxiv'26] LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
- [arxiv'26] Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents
- [arxiv'26] Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted Generation
- [arxiv'26] Spotlight: Synergizing Seed Exploration and Spot GPUs for DiT RL Post-Training
- [arxiv'26] Sparrow: Sparse Rollout for Stable and Efficient Long-context RL of Large Language Models
- [arxiv'26] Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training
- [arxiv'26] Libra: Efficient Resource Management for Agentic RL Post-Training
- [arxiv'26] Schedule-Level Shared-Prefix Reuse for LLM RL Training
- [arxiv'26] Polar: Agentic RL on Any Harness at Scale
- [arxiv'26] AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs
- [arxiv'26] ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning
- [arxiv'26] ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
- [arxiv'26] Heddle: A Distributed Orchestration System for Agentic RL Rollout
- [arxiv'26] Role-Based Fault Tolerance System for LLM RL Post-Training
- [arxiv'26] ARL-Tangram: Unleash the Resource Efficiency in Agentic Reinforcement Learning
- [arxiv'26] RLHFless: Serverless Computing for Efficient RLHF
- [arxiv'26] TVCACHE: A Stateful Tool-Value Cache for Post-Training LLM Agents
- [arxiv'26] Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs
- [arxiv'26] WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL
- [arxiv'26] RL over Commodity Networks: Overcoming the Bandwidth Barrier with Lossless Sparse Deltas
- [arxiv'26] Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
- [arxiv'26] Jet-RL: Enabling On-Policy FP8 Reinforcement Learning with Unified Training and Rollout Precision Flow
- [arxiv'26] Unleashing Efficient Asynchronous RL Post-Training via Staleness-Constrained Rollout Coordination
- [arxiv'26] OrchestrRL: Dynamic Compute and Network Orchestration for Disaggregated RL
- [arxiv'25] HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments
- [arxiv'25] ThreadWeaver: Adaptive Threading for Efficient Parallel Reasoning in Language Models
- [arxiv'25] RLHFSpec: Breaking the Efficiency Bottleneck in RLHF Training via Adaptive Drafting
- [arxiv'25] Fast LLM Post-training via Decoupled and Best-of-N Speculation
- [arxiv'25] Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter
- [arxiv'25] Beat the long tail: Distribution-Aware Speculative Decoding for RL Training
- [arxiv'25] WeChat-YATT: A Scalable, Simple, Efficient, and Production Ready Training Library
- [arxiv'25] The Path Not Taken: RLVR Provably Learns Off the Principals
- [arxiv'25] AReaL-Hex: Accommodating Asynchronous RL Training over Heterogeneous GPUs
- [arxiv'25] Ask a Strong LLM Judge when Your Reward Model is Uncertain
- [arxiv'25] RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs
- [arxiv'25] Laminar: A Scalable Asynchronous RL Post-Training Framework
- [arxiv'25] The Art of Scaling Reinforcement Learning Compute for LLMs
- [arxiv'25] xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning
- [arxiv'25] Hybrid Reinforcement: When Reward Is Sparse, It's Better to Be Dense
- [arxiv'25] Learning from Failures: Understanding LLM Alignment through Failure-Aware Inverse RL
- [arxiv'25] Spurious Rewards: Rethinking Training Signals in RLVR
- [arxiv'25] Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
- [arxiv'25] RL in the Wild: Characterizing RLVR Training in LLM Deployment
- [arxiv'25] APRIL: Active Partial Rollouts in Reinforcement Learning to tame long-tail generation
- [arxiv'25] ToRL: Scaling Tool-Integrated RL
- [arxiv'25] VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- [arxiv'25] Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
- [Survey 🔍] [arxiv'25] A Survey of Reinforcement Learning for Large Reasoning Models
- [arxiv'25] RewardDance: Reward Scaling in Visual Generation
- [arxiv'25] floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RL
- [arxiv'25] ParaThinker: Native Parallel Thinking as a New Paradigm to Scale LLM Test-time Compute
- [arxiv'25] History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
- [arxiv'25] SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling
- [arxiv'25] SPECS: Faster Test-Time Scaling through Speculative Drafts
- [arxiv'25] Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models
- [arxiv'25] ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
- [arxiv'25] GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- [arxiv'25] Multi-module GRPO: Composing Policy Gradients and Prompt Optimization for Language Model Programs
- [arxiv'25] Scaling RL to Long Videos
- [arxiv'25] Test-Time Training Done Right
- [arxiv'25] LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
- [arxiv'25] On-Policy RL with Optimal Reward Baseline
- [arxiv'25] StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
- [arxiv'25] DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- [arxiv'25] Reward Reasoning Model
- [arxiv'24] Optimizing RLHF Training for Large Language Models with Stage Fusion
For comprehensive list of GNN systems papers, refer to https://github.com/chwan1016/awesome-gnn-systems.
- [arxiv'26] Understanding and Reducing Metadata-Driven Host Overheads in Sampling-Based GNN Training
- [arxiv'25] Plexus: Taming Billion-edge Graphs with 3D Parallel GNN Training
- [arxiv'25] Armada: Memory-Efficient Distributed Training of Large-Scale Graph Neural Networks
- [arxiv'24] FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large Scale
- [arxiv'23] ReFresh: Reducing Memory Access from Exploiting Stable Historical Embeddings for Graph Neural Network Training
- [arxiv'23] Helios: An Efficient Out-of-core GNN Training System on Terabyte-scale Graphs with In-memory Performance
- [arxiv'23] GNNPipe: Accelerating Distributed Full-Graph GNN Training with Pipelined Model Parallelism
-
[arxiv'26] PLoRA: An NDP-Enhanced Pooled-Memory System for Cost-Efficient Multi-LoRA Serving
-
[arxiv'26] DeltaServe: Host-Agnostic Co-Serving of Inference and Fine-Tuning for LLMs
-
[arxiv'26] SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving
-
[arxiv'26] SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer
-
[arxiv'26] Denial of Deadline: Network-Driven Accuracy Collapse in Distributed Inference Pipelines
-
[arxiv'26] InstantInfer: Enabling Fast LLM Cold Start with Communicating Finite Automata
-
[arxiv'26] Roomie: Interference-Aware Colocation for Efficient Model Serving
-
[arxiv'26] HyMCache: A KV Cache Framework for Multi-Turn LLM Serving with CXL-Hybrid Memory
-
[arxiv'26] Xema: Efficient Diffusion Serving through Fine-Grained Memory Management and Auto-Configuration
-
[arxiv'26] FlashDiff: Efficient Regional Execution and Scheduling for Diffusion Model Serving
-
[arxiv'26] SiFAR: Synchronization-Free All-Reduce for Low-Latency LLM Inference
-
[arxiv'26] Akashic: A Low-Overhead LLM Inference Service with MemAttention
-
[arxiv'26] SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling
-
[arxiv'26] Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving
-
[arxiv'26] Tile-Level Activation Overlap for Efficient LLM Inference
-
[arxiv'26] Sangam: Efficiently Serving Diffusion LLMs with the AR Stack
-
[arxiv'26] CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving
-
[arxiv'26] DiLaServe: High SLO Attainment Serving for Diffusion Language Models
-
[arxiv'26] TraceLab: Characterizing Coding Agent Workloads for LLM Serving
-
[arxiv'26] Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving
-
[arxiv'26] SmoothAgent: Efficient Long-Horizon LLM-Based Agent Serving with Lookahead Context Engineering
-
[arxiv'26] Cache-Resident LLM Inference in GB-Scale Last-Level Caches
-
[arxiv'26] HERALD: High-Throughput Block Diffusion LLM Serving via CPU-GPU Cooperative KV Cache Retrieval
-
[arxiv'26] JetFlow: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting
-
[arxiv'26] AoiZora: Topology-Aware Auto-Parallel Optimization for Inference of Diffusion Transformers
-
[arxiv'26] SAC: Disaggregated KV Cache System for Sparse Attention LLMs with CXL
-
[arxiv'26] ShuntServe: Cost-Efficient LLM Serving on Heterogeneous Spot GPU Clusters
-
[arxiv'26] ReMP: Low-Downtime Runtime Model-Parallelism Reconfiguration for LLM Serving
-
[arxiv'26] TurboServe: Serving Streaming Video Generation Efficiently and Economically
-
[arxiv'26] RouteBalance: Fused Model Routing and Load Balancing for Heterogeneous LLM Serving
-
[arxiv'26] SwiftCache: Efficient LLM Serving for Multi-turn Conversations with Heterogeneous KV Cache Sharing
-
[arxiv'26] Fairness-Aware and Latency-Controllable Scheduling for Chunked-Prefill LLM Serving
-
[arxiv'26] FMplex: Model Virtualization for Serving Extensible Foundation Models
-
[arxiv'26] Tangram: Unlocking Non-Uniform KV Cache for Efficient Multi-turn LLM Serving
-
[arxiv'26] Lodestar: An Online-Learning LLM Inference Router
-
[arxiv'26] Beyond Greedy Chunking: SLO-Aware Sliding-Window Scheduling for LLM Inference
-
[arxiv'26] RTP-LLM: High-Performance Alibaba LLM Inference Engine
-
[arxiv'26] Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving
-
[arxiv'26] Scaling LLM Inference Beyond Amdahl`s Limits via Eliminating Non-Scalable Overheads
-
[arxiv'26] ObjectCache: Layerwise Object-Storage Retrieval for KV Cache Reuse
-
[arxiv'26] Optimus: Elastic Decoding for Efficient Diffusion LLM Serving
-
[arxiv'26] ObjectCache: Layerwise Object-Storage Retrieval for KV Cache Reuse
-
[arxiv'26] VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
-
[arxiv'26] Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches
-
[arxiv'26] An Interpretable Latency Model for Speculative Decoding in LLM Serving
-
[arxiv'26] Kairos: A Scalable Serving System for Physical AI
-
[arxiv'26] SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
-
[arxiv'26] Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
-
[arxiv'26] ZeRO-Prefill: Zero Redundancy Overheads in MoE Prefill Serving
-
[arxiv'26] PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers
-
[arxiv'26] Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs
-
[arxiv'26] Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving
-
[arxiv'26] DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference
-
[arxiv'26] Pythia: Toward Predictability-Driven Agent-Native LLM Serving
-
[arxiv'26] Continuous Semantic Caching for Low-Cost LLM Serving
-
[arxiv'26] FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
-
[arxiv'26] GENSERVE: Efficient Co-Serving of Heterogeneous Diffusion Model Workloads
-
[arxiv'26] Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics
-
[arxiv'26] Understand and Accelerate Memory Processing Pipeline for Disaggregated LLM Inference
-
[arxiv'26] Multi-stage Flow Scheduling for LLM Serving
-
[arxiv'26] TIDE: Token-Informed Depth Execution for Per-Token Early Exit in LLM Inference
-
[arxiv'26] WWW.Serve: Interconnecting Global LLM Services through Decentralization
-
[arxiv'26] Chimera: Latency- and Performance-Aware Multi-agent Serving for Heterogeneous LLMs
-
[arxiv'26] Swarm: Co-Activation Aware KVCache Offloading Across Multiple SSDs
-
[arxiv'26] FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
-
[arxiv'26] Serving Compound Inference Systems on Datacenter GPUs
-
[arxiv'26] StreamWise: Serving Multi-Modal Generation in Real-Time at Scale
-
[arxiv'26] O^3-LSM: Maximizing Disaggregated LSM Write Performance via Three-Layer Offloading
-
[arxiv'26] vLLM Semantic Router: Signal Driven Decision Routing for Mixture-of-Modality Models
-
[arxiv'26] Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator
-
[arxiv'26] SUN: Shared Use of Next-token Prediction for Efficient Multi-LLM Disaggregated Serving
-
[arxiv'26] LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding
-
[arxiv'26] FLYING SERVING: On-the-Fly Parallelism Switching for Large Language Model Serving
-
[arxiv'26] How Fast Can I Run My VLA? Demystifying VLA Inference Performance with VLA-Perf
-
[arxiv'26] BiScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
-
[arxiv'26] Efficient Multi-round LLM Inference over Disaggregated Serving
-
[arixv'26] SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining
-
[arxiv'26] PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
-
[arxiv'26] OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
-
[arxiv'26] PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
-
[arxiv'26] Plan, Verify and Fill: A Structured Parallel Decoding Approach for Diffusion Language Models
-
[arxiv'26] DFlash: Block Diffusion for Flash Speculative Decoding
-
[arxiv'26] VoxServe: Streaming-Centric Serving System for Speech Language Models
-
[arxiv'26] Laser: Unlocking Layer-Level Scheduling for Efficient Multi-SLO LLM Serving
-
[arxiv'26] PLA-Serve: A Prefill-Length-Aware LLM Serving System
-
[arxiv'26] AIConfigurator: Lightning-Fast Configuration Optimization for Multi-Framework LLM Serving
-
[arxiv'26] FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems
-
[arxiv'25] LIMINAL: Exploring The Frontiers of LLM Decode Performance
-
[arxiv'25] TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
-
[arxiv'25] L4: Low-Latency and Load-Balanced LLM Serving via Length-Aware Scheduling
-
[arxiv'25] Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
-
[arxiv'25] EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
-
[arxiv'25] MultiPath Transfer Engine: Breaking GPU and Host-Memory Bandwidth Bottlenecks in LLM Services
-
[arxiv'25] PROSERVE: Unified Multi-Priority Request Scheduling for LLM Serving
-
[arxiv'25] xGR: Efficient Generative Recommendation Serving at Scale
-
[arxiv'25] ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding
-
[arxiv'25] TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
-
[arxiv'25] AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving
-
[arxiv'25] Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding
-
[arxiv'25] OmniInfer: System-Wide Acceleration Techniques for Optimizing LLM Serving Throughput and Latency
-
[arxiv'25] OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
-
[arxiv'25] Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
-
[arxiv'25] CLO: Efficient LLM Inference System with CPU-Light KVCache Offloading via Algorithm-System Co-Design
-
[arxiv'25] FengHuang: Next-Generation Memory Orchestration for AI Inferencing
-
[arxiv'25] Synera: Synergistic LLM Serving across Device and Cloud at Scale
-
[arxiv'25] DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing
-
[arxiv'25] From Models to Operators: Rethinking Autoscaling Granularity for Large Generative Models
-
[arxiv'25] TapOut: A Bandit-Based Approach to Dynamic Speculative Decoding
-
[arxiv'25] FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
-
[arxiv'25] Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
-
[arxiv'25] SPAD: Specialized Prefill and Decode Hardware for Disaggregated LLM Inference
-
[arxiv'25] From Tokens to Layers: Redefining Stall-Free Scheduling for LLM Serving with Layered Prefill
-
[arxiv'25] TridentServe: A Stage-level Serving System for Diffusion Pipelines
-
[arxiv'25] MACE: A Hybrid LLM Serving System with Colocated SLO-aware Continuous Retraining Alignment
-
[arxiv'25] TetriServe: Efficient DiT Serving for Heterogeneous Image Generation
-
[arxiv'25] Parallax: Efficient LLM Inference Service over Decentralized Environment
-
[arxiv'25] RServe: Overlapping Encoding and Prefill for Efficient LMM Inference
-
[arxiv'25] Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
-
[arxiv'25] Shift Parallelism: Low-Latency, High-Throughput LLM Inference for Dynamic Workloads
-
[arxiv'25] FineServe: Precision-Aware KV Slab and Two-Level Scheduling for Heterogeneous Precision LLM Serving
-
[arxiv'25] AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
-
[arxiv'25] Predictable LLM Serving on GPU Clusters
-
[arxiv'25] Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
-
[arxiv'25] Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics
-
[arxiv'25] HyperFlexis: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling
-
[arxiv'25] Equinox: Holistic Fair Scheduling in Serving Large Language Models
-
[arxiv'25] Efficient Mixed-Precision Large Language Model Inference with TurboMind
-
[arxiv'25] Kairos: Low-latency Multi-Agent Serving with Shared LLMs and Excessive Loads in the Public Cloud
-
[arxiv'25] Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
-
[arxiv'25] Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
-
[arxiv'25] Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
-
[arxiv'25] Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding
-
[arxiv'25] Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
-
[arxiv'25] MIRAGE: KV Cache Optimization through Parameter Remapping for Multi-tenant LLM Serving
-
[arxiv'25] On Evaluating Performance of LLM Inference Serving Systems
-
[arxiv'25] PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
-
[arxiv'25] SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
-
[arxiv'25] Utility-Driven Speculative Decoding for Mixture-of-Experts
-
[arxiv'25] Cascadia: A Cascade Serving System for Large Language Models
-
[arxiv'25] Efficient and Workload-Aware LLM Serving via Runtime Layer Swapping and KV Cache Resizing
-
[arxiv'25] SkyLB: A Locality-Aware Cross-Region Load Balancer for LLM Inference
-
[arxiv'25] EmbAdvisor: Adaptive Cache Management for Sustainable LLM Serving
-
[arxiv'25] SCORPIO: Serving the Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference
-
[arxiv'25] HybridServe: Efficient Serving of Large AI Models with Confidence-Based Cascade Routing
-
[arxiv'25] ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
-
[arxiv'25] TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
-
[arxiv'25] Tilus: A Virtual Machine for Arbitrary Low-Precision GPGPU Computation in LLM Serving
-
[arxiv'25] ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
-
[arxiv'25] ELIS: Efficient LLM Iterative Scheduling System with Response Length Predictor
-
[arxiv'25] Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
-
[arxiv'25] Tempo: Application-aware LLM Serving with Mixed SLO Requirements
-
[arxiv'25] Ascendra: Dynamic Request Prioritization for Efficient LLM Serving
-
[arxiv'25] Streaming, Fast and Slow: Cognitive Load-Aware Streaming for Efficient LLM Serving
-
[arxiv'25] Bullet: Boosting GPU Utilization for LLM Serving via Dynamic Spatial-Temporal Orchestration
-
[Survey 🔍] [arxiv'25] Taming the Titans: A Survey of Efficient LLM Inference Serving
-
[arxiv'25] PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
-
[arxiv'25] Circinus: Efficient Query Planner for Compound ML Serving
-
[arxiv'25] HPU: High-Bandwidth Processing Unit for Scalable, Cost-effective LLM Inference via GPU Co-processing
-
[arxiv'25] SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference
-
[arxiv'25] gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
-
[arxiv'25] Optimizing SLO-oriented LLM Serving with PD-Multiplexing
-
[arxiv'25] SLO-Aware Scheduling for Large Language Model Inferences
-
[arxiv'25] Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
-
[arxiv'25] HELIOS: Adaptive Model And Early-Exit Selection for Efficient LLM Inference Serving
-
[arxiv'25] DynaServe: Unified and Elastic Tandem-Style Execution for Dynamic Disaggregated LLM Serving
-
[arxiv'25] Efficient LLM Serving on Hybrid Real-time and Best-effort Requests
-
[arxiv'25] Understanding and Optimizing Multi-Stage AI Inference Pipelines
-
[arxiv'24] Fast and Live Model Auto Scaling with O(1) Host Caching
-
[arxiv'25] WaferLLM: A Wafer-Scale LLM Inference System
-
[arxiv'25] Niyama : Breaking the Silos of LLM Inference Serving
-
[arxiv'25] PipeBoost: Resilient Pipelined Architecture for Fast Serverless LLM Scaling
-
[arxiv'25] Jenga: Effective Memory Management for Serving LLM with Heterogeneity
-
[arxiv'25] Collaborative Speculative Inference for Efficient LLM Inference Serving
-
[arxiv'25] Seesaw: High-throughput LLM Inference via Model Re-sharding
-
[arxiv'25] SpecServe: Efficient and SLO-Aware Large Language Model Serving with Adaptive Speculative Decoding
-
[arxiv'25] ADOR: A Design Exploration Framework for LLM Serving with Enhanced Latency and Throughput
-
[arxiv'25] Long-Context Inference with Retrieval-Augmented Speculative Decoding
-
[arxiv'25] Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM
-
[arxiv'25] KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse
-
[arxiv'25] Serving Models, Fast and Slow:Optimizing Heterogeneous LLM Inferencing Workloads at Scale
-
[arxiv'25] LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
-
[arxiv'25] HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading
-
[arxiv'25] Autellix: An Efficient Serving Engine for LLM Agents as General Programs
-
[arxiv'25] Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
-
[arxiv'25] MPIC: Position-Independent Multimodal Context Caching System for Efficient MLLM Serving
-
[arxiv'25] Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
-
[arxiv'25] HyGen: Efficient LLM Serving via Elastic Online-Offline Request Co-location
-
[arxiv'25] DeltaZip: Efficient Serving of Multiple Full-Model-Tuned LLMs
-
[arxiv'25] DeepFlow: Serverless Large Language Model Serving at Scale
-
[arxiv'25] AdaServe: SLO-Customized LLM Serving with Fine-Grained Speculative Decoding
-
[arxiv'25] EchoLM: Accelerating LLM Serving with Real-time Knowledge Distillation
-
[arxiv'25] OMEGA: A Low-Latency GNN Serving System for Large Graphs
-
[arxiv'25] PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
-
[arxiv'25] Hierarchical Autoscaling for Large Language Model Serving with Chiron
-
[arxiv'25] Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
-
[arxiv'25] Accelerated Diffusion Models via Speculative Sampling
-
[arxiv'24] LLM Inference Unveiled: Survey and Roofline Model Insights
-
[arxiv'24] Efficiently Serving LLM Reasoning Programs with Certaindex
-
[arxiv'24] LoL-PIM: Long-Context LLM Decoding with Scalable DRAM-PIM System
-
[arxiv'24] TimelyLLM: Segmented LLM Serving System for Time-sensitive Robotic Applications
-
[arxiv'24] Dovetail: A CPU/GPU Heterogeneous Speculative Decoding for LLM inference
-
[arxiv'24] SYMPHONY: Improving Memory Management for LLM Inference Workloads
-
[arxiv'24] A System for Microserving of LLMs
-
[arxiv'24] HashAttention: Semantic Sparsity for Faster Inference
-
[arxiv'24] SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices
-
[arxiv'24] Unifying KV Cache Compression for Large Language Models with LeanKV
-
[arxiv'24] PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
-
[arxiv'24] Optimizing Speculative Decoding for Serving Large Language Models Using Goodput
-
[arxiv'24] EcoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
-
[arxiv'24] EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
-
[arxiv'24] Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection
-
[arxiv'24] SuffixDecoding: A Model-Free Approach to Speeding Up Large Language Model Inference
-
[arxiv'24] V-LoRA: An Efficient and Flexible System Boosts Vision Applications with LoRA LMM
-
[arxiv'24] HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference
-
[arxiv'24] NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
-
[arxiv'24] ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
-
[arxiv'24] Is the GPU Half-Empty or Half-Full? Practical Scheduling Techniques for LLMs
-
[arxiv'24] POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
-
[arxiv'24] MagicPIG: LSH Sampling for Efficient LLM Generation
-
[arxiv'24] Revisiting SLO and Goodput Metrics in LLM Serving
-
[arxiv'24] EPIC: Efficient Position-Independent Context Caching for Serving Large Language Models
-
[arxiv'24] ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
-
[arxiv'24] SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
-
[arxiv'24] vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
-
[arxiv'24] DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
-
[arxiv'24] Missile: Fine-Grained, Hardware-Level GPU Resource Isolation for Multi-Tenant DNN Inference
-
[arxiv'24] P/D-Serve: Serving Disaggregated Large Language Model at Scale
-
[arxiv'24] MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models
-
[arxiv'24] LLMServingSim: A HW/SW Co-Simulation Infrastructure for LLM Inference Serving at Scale
-
[Survey 🔍] [arxiv'24] LLM Inference Serving: Survey of Recent Advances and Opportunities
-
[arxiv'24] Metron: Holistic Performance Evaluation Framework for LLM Inference Systems
-
[arxiv'24] Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
-
[arxiv'24] One Queue Is All You Need: Resolving Head-of-Line Blocking in Large Language Model Serving
-
[arxiv'24] MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
-
[arxiv'24] Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inference
-
[arxiv'24] The CAP Principle for LLM Serving
-
[arxiv'24] BlockLLM: Multi-tenant Finer-grained Serving for Large Language Models
-
[arxiv'24] Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
-
[arxiv'24] Learn To be Efficient: Build Structured Sparsity in Large Language Models
-
[arxiv'24] Sponge: Inference Serving with Dynamic SLOs Using In-Place Vertical Scaling
-
[arxiv'24] ALTO: An Efficient Network Orchestrator for Compound AI Systems
-
[arxiv'24] ATP: Enabling Fast LLM Serving via Attention on Top Principal Keys
-
[arxiv'24] Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
-
[arxiv'24] FlexLLM: A System for Co-Serving Large Language Model Inference and Parameter-Efficient Finetuning
-
[arxiv'24] Wisdom of Committee: Distilling from Foundation Model to SpecializedApplication Model
-
[arxiv'24] RelayAttention for Efficient Large Language Model Serving with Long System Prompts
-
[arxiv'24] LLM-PQ: Serving LLM on Heterogeneous Clusters with Phase-Aware Partition and Adaptive Quantization
-
[arxiv'24] APIServe: Efficient API Support for Large-Language Model Inferencing
-
[arxiv'24] ServerlessLLM: Locality-Enhanced Serverless Inference for Large Language Models
-
[arxiv'24] MoE-Infinity: Activation-Aware Expert Offloading for Efficient MoE Serving
-
[arxiv'24] FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
-
[arxiv'24] Accelerating Retrieval-Augmented Language Model Serving with Speculation
-
[arxiv'24] CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
-
[arxiv'24] Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
-
[arxiv'24] DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
-
[Survey 🔍] [arxiv'24] Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding
-
[arxiv'24] Learned Best-Effort LLM Serving
-
[arxiv'24] Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
-
[arxiv'23] DeltaZip: Multi-Tenant Language Model Serving via Delta Compression
-
[arxiv'23] Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding
-
[arxiv'23] Fairness in Serving Large Language Models
-
[arxiv'23] Moirai: Towards Optimal Placement for Distributed Inference on Heterogeneous Devices
-
[arxiv'23] Punica: Multi-Tenant LoRA Serving
-
[arxiv'23] Pipeline Parallelism for DNN Inference with Practical Performance Guarantees
-
[arxiv'23] SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
-
[arxiv'23] High-throughput Generative Inference of Large Language Models with a Single GPU
-
[arxiv'21] Supporting Massive DLRM Inference through Software Defined Memory
- [arxiv'26] AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
- [arxiv'26] FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling
- [arxiv'26] SageBwd: A Trainable Low-bit Attention
- [arxiv'26] Test-Time Training with KV Binding Is Secretly Linear Attention
- [arxiv'26] MonarchRT: Efficient Attention for Real-Time Video Generation
- [arxiv'26] PackInfer: Compute- and I/O-Efficient Attention for Batched LLM Inference
- [arxiv'25] BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding
- [arxiv'25] SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention [Code]
-
[arxiv'26] Incast-Free MoE Rate-Based Scheduling
-
[arxiv'26] MoX: Efficient MoE Routing on Direct-Connect Topologies
-
[arxiv'26] PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization
-
[arxiv'26] Less Experts, Faster Decoding: Cost-Aware Speculative Decoding for Mixture-of-Experts
-
[arxiv'26] Communication-Aware Placement and Pruning for Efficient Mixture-of-Experts Inference
-
[arxiv'26] ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving
-
[arxiv'26] Moebius: Serving Mixture-of-Expert Models with Seamless Runtime Parallelism Switch
-
[arxiv'26] ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill
-
[arxiv'26] FoMoE: Breaking the Full-Replica Barrier with a Federation of MoEs
-
[arxiv'26] Coordinated Scheduling for MoE LLM Serving
-
[arxiv'26] Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training
-
[arxiv'26] UltraEP: Unleash MoE Training and Inference on Rack-Scale Nodes with Near-Optimal Load Balancing
-
[arxiv'26] ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving
-
[arxiv'26] PithTrain: A Compact and Agent-Native MoE Training System
-
[arxiv'26] HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs
-
[arxiv'26] HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs
-
[arxiv'26] NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding
-
[arxiv'26] GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems
-
[arxiv'26] EMO: Frustratingly Easy Progressive Training of Extendable MoE
-
[arxiv'26] Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models
-
[arxiv'26] ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning
-
[arxiv'26] DisagMoE: Computation-Communication overlapped MoE Training via Disaggregated AF-Pipe Parallelism
-
[arxiv'26] Federation of Experts: Communication Efficient Distributed Inference for Large Language Models
-
[arxiv'26] Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference
-
[arxiv'26] Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism
-
[arxiv'26] ZeRO-Prefill: Zero Redundancy Overheads in MoE Prefill Serving
-
[arxiv'26] Rethinking Network Topologies for Cost-Effective Mixture-of-Experts LLM Serving
-
[arxiv'26] UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training
-
[arxiv'26] Holistic Scaling Laws for Optimal Mixture-of-Experts Architecture Optimization
-
[arxiv'26] Speculating Experts Accelerates Inference for Mixture-of-Experts
-
[arxiv'26] NCCL EP: Towards a Unified Expert Parallel Communication API for NCCL
-
[arxiv'26] The qs Inequality: Quantifying the Double Penalty of Mixture-of-Experts at Inference
-
[arxiv'26] Scalable Training of Mixture-of-Experts Models with Megatron Core
-
[arxiv'26] MoEless: Efficient MoE LLM Serving via Serverless Computing
-
[arxiv'26] Revealing the Challenges of Attention-FFN Disaggregation for Modern MoE Models and Hardware Systems
-
[arxiv'26] Practical FP4 Training for Large-Scale MoE Models on Hopper GPUs
-
[arxiv'26] MoE-Spec: Expert Budgeting for Efficient Speculative Decoding
-
[arxiv'26] MELINOE: Fine-Tuning Enables Memory-Efficient Inference for Mixture-of-Experts Models
-
[arxiv'26] XShare: Collaborative in-Batch Expert Sharing for Faster MoE Inference
-
[arxiv'26] OmniMoE: An Efficient MoE by Orchestrating Atomic Experts at Scale
-
[arxiv'26] Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE Parallelism
-
[arxiv'26] Dynamic Expert Sharing: Decoupling Memory from Parallelism in Mixture-of-Experts Diffusion LLMs
-
[arxiv'26] LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts
-
[arxiv'26] Least-Loaded Expert Parallelism: Load Balancing An Imbalanced Mixture-of-Experts
-
[arxiv'26] MoE-DisCo:Low Economy Cost Training Mixture-of-Experts Models
-
[arxiv'26] MoEBlaze: Breaking the Memory Wall for Efficient MoE Training on Modern GPUs
-
[arxiv'26] Making MoE-based LLM Inference Resilient with Tarragon
-
[arxiv'25] FUSCO: High-Performance Distributed Data Shuffling via Transformation-Communication Fusion
-
[arxiv'25] Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
-
[arxiv'25] Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
-
[arxiv'25] SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
-
[arxiv'25] Janus: Disaggregating Attention and Experts for Scalable MoE Inference
-
[arxiv'25] Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
-
[arxiv'25] Context-Aware Mixture-of-Experts Inference on CXL-Enabled GPU-NDP Systems
-
[arxiv'25] MicroMoE: Fine-Grained Load Balancing for Mixture-of-Experts with Token Scheduling
-
[arxiv'25] MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
-
[arxiv'25] Pre-Attention Expert Prediction and Prefetching for Mixture-of-Experts Large Language Models
-
[arxiv'25] FarSkip-Collective: Unhobbling Blocking Communication in Mixture of Experts Models
-
[arxiv'25] DualSparse-MoE: Coordinating Tensor/Neuron-Level Sparsity with Expert Partition and Reconstruction
-
[arxiv'25] BuddyMoE: Exploiting Expert Redundancy to Accelerate Memory-Constrained Mixture-of-Experts Inference
-
[arxiv'25] Opportunistic Expert Activation: Batch-Aware Expert Routing for Faster Decode Without Retraining
-
[arxiv'25] MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs
-
[arxiv'25] ReXMoE: Reusing Experts with Minimal Overhead in Mixture-of-Experts
-
[arxiv'25] MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
-
[arxiv'25] GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
-
[arxiv'25] Orders in Chaos: Enhancing Large-Scale MoE LLM Serving with Data Movement Forecasting
-
[arxiv'25] ElasticMoE: An Efficient Auto Scaling Method for Mixture-of-Experts Models
-
[arxiv'25] MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
-
[arxiv'25] DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert Pruning
-
[arxiv'25] Symphony-MoE: Harmonizing Disparate Pre-trained Models into a Coherent Mixture-of-Experts
-
[arxiv'25] Steering MoE LLMs via Expert (De)Activation
-
[arxiv'25] HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing
-
[arxiv'25] LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference
-
[arxiv'25] LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference
-
[arxiv'25] LongCat-Flash Technical Report
-
[arxiv'25] Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
-
[arxiv'25] HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
-
[arxiv'25] MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
-
[arxiv'25] Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
-
[arxiv'25] Chain-of-Experts: Unlocking the Communication Power of Mixture-of-Experts Models
-
[arxiv'25] HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap
-
[arxiv'25] PiKV: KV Cache Management System for Mixture of Experts
-
[arxiv'25] BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
-
[arxiv'25] Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
-
[arxiv'25] The New LLM Bottleneck: A Systems Perspective on Latent Attention and Mixture-of-Experts
-
[arxiv'25] Muon is Scalable for LLM Training
-
[arxiv'25] Long-Tailed Distribution-Aware Router For Mixture-of-Experts in Large Vision-Language Model
-
[arxiv'25] Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
-
[arxiv'25] HarMoEny: Efficient Multi-GPU Inference of MoE Models
-
[arxiv'25] Load Balancing Mixture of Experts with Similarity Preserving Routers
-
[arxiv'25] MoE-GPS: Guidlines for Prediction Strategy for Dynamic Expert Duplication in MoE Load Balancing
-
[arxiv'25] EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models
-
[arxiv'25] CoMoE: Contrastive Representation for Mixture-of-Experts in Parameter-Efficient Fine-tuning
-
[arxiv'25] PreMoe: Lightening MoEs on Constrained Memory by Expert Pruning and Retrieval
-
[arxiv'25] Not All Models Suit Expert Offloading: On Local Routing Consistency of Mixture-of-Expert Models
-
[arxiv'25] Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
-
[arxiv'25] PT-MoE: An Efficient Finetuning Framework for Integrating Mixture-of-Experts into Prompt Tuning
-
[arxiv'25] Faster MoE LLM Inference for Extremely Large Models
-
[arxiv'25] Accelerating Mixture-of-Experts Training with Adaptive Expert Replication
-
[arxiv'25] MoE-Gen: High-Throughput MoE Inference on a Single GPU with Module-Based Batching
-
[arxiv'25] Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
-
[arxiv'25] Unveiling Hidden Collaboration within Mixture-of-Experts in Large Language Models
-
[arxiv'25] Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
-
[arxiv'25] MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource Constraints
-
[arxiv'25] C3PO: Critical-Layer, Core-Expert, Collaborative Pathway Optimization for Test-Time Expert Re-Mixing
-
[arxiv'25] Cluster-Driven Expert Pruning for Mixture-of-Experts Large Language Models
-
[arxiv'25] S'MoRE: Structural Mixture of Residual Experts for LLM Fine-tuning
-
[arxiv'25] Finding Fantastic Experts in MoEs: A Unified Study for Expert Dropping Strategies and Observations
-
[arxiv'25] HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
-
[arxiv'25] ProMoE: Fast MoE-based LLM Serving using Proactive Caching
-
[arxiv'25] Mixture of Lookup Experts
-
[arxiv'25] eMoE: Task-aware Memory Efficient Mixture-of-Experts-Based (MoE) Model Inference
-
[arxiv'25] Continual Pre-training of MoEs: How robust is your router?
-
[arxiv'25] Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs
-
[arxiv'25] Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts
-
[arxiv'25] CoSMoEs: Compact Sparse Mixture of Experts
-
[arxiv'25] Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
-
[arxiv'25] BigMac: A Communication-Efficient Mixture-of-Experts Model Structure for Fast Training and Inference
-
[arxiv'25] DSMoE: Matrix-Partitioned Experts with Dynamic Routing for Computation-Efficient Dense LLMs
-
[arxiv'25] MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing
-
[arxiv'25] Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
-
[arxiv'25] Fair-MoE: Fairness-Oriented Mixture of Experts in Vision-Language Models
-
[arxiv'25] fMoE: Fine-Grained Expert Offloading for Large Mixture-of-Experts Serving
-
[arxiv'25] Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism
-
[arxiv'25] BTS: Harmonizing Specialized Experts into a Generalist LLM
-
[arxiv'25] Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
-
[arxiv'25] Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
-
[arxiv'24] DeepSeek-V3 Technical Report
-
[arxiv'24] HEXA-MoE: Efficient and Heterogeneous-aware MoE Acceleration with ZERO Computation Redundancy
-
[arxiv'24] Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts
-
[arxiv'24] ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing
-
[Survey 🔍] [arxiv'24] A Survey on Inference Optimization Techniques for Mixture of Experts Models
-
[arxiv'24] DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
-
[arxiv'24] Llama 3 Meets MoE: Efficient Upcycling
-
[arxiv'24] Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
-
[arxiv'24] Mixture of A Million Experts
-
[arxiv'24] MoE-CAP: Cost-Accuracy-Performance Benchmarking for Mixture-of-Experts Systems
-
[arxiv'24] Toward Inference-optimal Mixture-of-Expert Large Language Models
-
[arxiv'24] Expert-Token Resonance: Redefining MoE Routing through Affinity-Driven Active Selection
-
[arxiv'24] MoE Jetpack: From Dense Checkpoints to Adaptive Mixture of Experts for Vision Tasks
-
[arxiv'24] Prediction Is All MoE Needs: Expert Load Distribution Goes from Fluctuating to Stabilizing
-
[arxiv'24] DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
-
[arxiv'24] UOE: Unlearning One Expert Is Enough For Mixture-of-experts LLMS
-
[arxiv'24] Dense Backpropagation Improves Routing for Sparsely-Gated Mixture-of-Experts
-
[arxiv'24] MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
-
[arxiv'24] Pipeline MoE: A Flexible MoE Implementation with Pipeline Parallelism
-
[arxiv'24] EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
-
[arxiv'24] Shortcut-connected Expert Parallelism for Accelerating Mixture of Experts
-
[arxiv'24] Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection
-
[arxiv'24] HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference
-
[arxiv'24] Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent
-
[arxiv'24] Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent
-
[arxiv'24] Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
-
[arxiv'24] Demystifying the Compression of Mixture-of-Experts Through a Unified Framework
-
[arxiv'24] Optimizing Mixture-of-Experts Inference Time Combining Model Deployment and Communication Scheduling
-
[arxiv'24] MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
-
[arxiv'24] Duo-LLM: A Framework for Studying Adaptive Computation in Large Language Models
-
[arxiv'24] MoH: Multi-Head Attention as Mixture-of-Head Attention
-
[arxiv'24] AT-MoE: Adaptive Task-planning Mixture of Experts via LoRA Approach
-
[arxiv'24] Aria: An Open Multimodal Native Mixture-of-Experts Model
-
[arxiv'24] MC-MoE: Mixture Compressor for Mixture-of-Experts LLMs Gains More
-
[arxiv'24] MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
-
[arxiv'24] Upcycling Large Language Models into Mixture of Experts
-
[arxiv'24] No Need to Talk: Asynchronous Mixture of Language Models
-
[arxiv'24] Lazarus: Resilient and Elastic Training of Mixture-of-Experts Models with Adaptive Expert Placement
-
[arxiv'24] HMoE: Heterogeneous Mixture of Experts for Language Modeling
-
[arxiv'24] FedMoE: Personalized Federated Learning via Heterogeneous Mixture of Experts
-
[arxiv'24] AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies
-
[arxiv'24] Layerwise Recurrent Router for Mixture-of-Experts
-
[arxiv'24] Partial Experts Checkpoint: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training
-
[arxiv'24] MoDE: Effective Multi-task Parameter Efficient Fine-Tuning with a Mixture of Dyadic Experts
-
[arxiv'24] Self-MoE: Towards Compositional Large Language Models with Self-Specialized Experts
-
[arxiv'24] Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models
-
[arxiv'24] CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
-
[arxiv'24] AdaMoLE: Fine-Tuning Large Language Models with Adaptive Mixture of Low-Rank Adaptation Experts
-
[arxiv'24] MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA based Mixture of Experts
-
[arxiv'24] Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
-
[arxiv'24] MoE-Infinity: Activation-Aware Expert Offloading for Efficient MoE Serving
-
[arxiv'23] Towards MoE Deployment: Mitigating Inefficiencies in Mixture-of-Expert (MoE) Inference
-
[arxiv'23] Fast Inference of Mixture-of-Experts Language Models with Offloading
-
[arxiv'22] ST-MoE: Designing Stable and Transferable Sparse Expert Models
- [arxiv'26] NUNA: Characterizing and Mitigating Non-Uniform Network Access in Multi-Die GPU Scale-Up Systems
- [arxiv'26] StrataCL: Fabric-Native Communication Library for Production Supernodes
- [arxiv'26] X-Stage: An Overlooked Pipeline Stage for Communication-Computation Overlap in DiT Inference
- [arxiv'26] Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives
- [arxiv'26] Scalable Optimal Transport Algorithm for Network Alignment
- [arxiv'26] SiFAR: Synchronization-Free All-Reduce for Low-Latency LLM Inference
- [arxiv'26] CTA-Pipelining: A Latency-Oriented Spatial Scaling Method for Multi-GPU Systems
- [arxiv'26] Toward a Unified GPU-Aware OpenSHMEM Specification
- [arxiv'26] The Multipath Reliable Connection (MRC) Transport
- [arxiv'26] Eidola: Modeling Multi-GPU Network Communication Traffic in Distributed AI Workloads
- [arxiv'26] PCCL: Process Group-Aware Scalable and Generic Collective Algorithm Synthesizer
- [arxiv'26] Demystifying NVSHMEM: A System-Level Analysis on Symmetric Memory and Device-Initiated Operations in GPU Communication
- [arxiv'26] HetCCL: Enabling Collective Communication For Mixed-Vendor Heterogeneous Clusters
- [arxiv'26] ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training
- [arxiv'26] Extreme-Scale Interconnection Networks
- [arxiv'26] Revisiting Bruck: Phase-Efficient All-to-All Communication in Reconfigurable Networks
- [arxiv'26] Exploiting Multicast for Accelerating Collective Communication
- [arxiv'26] Rethinking Network Topologies for Cost-Effective Mixture-of-Experts LLM Serving
- [arxiv'26] Harvest: Adaptive Photonic Switching Schedules for Collective Communication in Scale-up Domains
- [arxiv'26] UCCL-Zip: Lossless Compression Supercharged GPU Communication
- [arxiv'26] From Skew to Symmetry: Node-Interconnect Multi-Path Balancing with Execution-time Planning for Modern GPU Clusters
- [arxiv'26] To Reconfigure or Not to Reconfigure: Optimizing All-to-All Collectives in Circuit-Switched Photonic Interconnects
- [arxiv'26] NCCL EP: Towards a Unified Expert Parallel Communication API for NCCL
- [arxiv'26] Network Design for Wafer-Scale Systems with Wafer-on-Wafer Hybrid Bonding
- [arxiv'26] Spritz: Path-Aware Load Balancing in Low-Diameter Networks
- [arxiv'26] ACOS: Arrays of Cheap Optical Switches
- [arxiv'26] Bring Your Own Objective: Inter-operability of Network Objectives in Datacenters
- [arxiv'26] RL over Commodity Networks: Overcoming the Bandwidth Barrier with Lossless Sparse Deltas
- [arxiv'26] DynamiQ: Accelerating Gradient Synchronization using Compressed Multi-hop All-reduce
- [arxiv'26] MonkeyTree: Near-Minimal Congestion for Multi-tenant Training via Migration
- [arxiv'26] HetCCL: Accelerating LLM Training with Heterogeneous GPUs
- [arxiv'26] AutoOverlap: Enabling Fine-Grained Overlap of Computation and Communication with Chunk-Based Scheduling
- [arxiv'26] Heterogeneous Low-Bandwidth Pre-Training of LLMs
- [arxiv'25] Analyzing Communication Predictability in LLM Training
- [arxiv'25] UCCL-EP: Portable Expert-Parallel Communication
- [arxiv'25] Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
- [arxiv'25] Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design
- [arxiv'25] FarSkip-Collective: Unhobbling Blocking Communication in Mixture of Experts Models
- [arxiv'25] GPU-Initiated Networking for NCCL
- [arxiv'25] DMA Collectives for Efficient ML Communication Offloads
- [arxiv'25] Collective Communication for 100k+ GPUs
- [arxiv'25] Uno: A One-Stop Solution for Inter- and Intra-Datacenter Congestion Control and Reliable Connectivity
- [arxiv'25] MSCCL++: Rethinking GPU Communication Abstractions for Cutting-edge AI Applications
- [arxiv'25] Toward Co-adapting Machine Learning Job Shape and Cluster Topology
- [arxiv'25] Efficient AllReduce with Stragglers
- [arxiv'25] TASP: Topology-aware Sequence Parallelism
- [arxiv'25] Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality
- [arxiv'25] RoCE BALBOA: Service-enhanced Data Center RDMA for SmartNICs
- [arxiv'25] RailX: A Flexible, Scalable, and Low-Cost Network Architecture for Hyper-Scale LLM Training Systems
- [arxiv'25] Demystifying NCCL: An In-depth Analysis of GPU Communication Protocols and Algorithms
- [arxiv'25] NoLoCo: No-all-reduce Low Communication Training Method for Large Models
- [arxiv'25] TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
- [arxiv'25] FLASH: Fast All-to-All Communication in GPU Clusters
- [arxiv'25] MCMComm: Hardware-Software Co-Optimization for End-to-End Communication in Multi-Chip-Modules
- [arxiv'25] GenTorrent: Scaling Large Language Model Serving with An Overley Network
- [arxiv'25] Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
- [arxiv'25] FlashOverlap: A Lightweight Design for Efficiently Overlapping Communication and Computation
- [arxiv'25] An Extensible Software Transport Layer for GPU Networking (
UCCL) [Code] - [arxiv'25] HeteroPod: XPU-Accelerated Infrastructure Offloading for Commodity Cloud-Native Applications
- [Survey 🔍] [arxiv'25] GPU-centric Communication Schemes for HPC and ML Applications
- [arxiv'25] UB-Mesh: a Hierarchically Localized nD-FullMesh Datacenter Network Architecture
- [arxiv'25] Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo
- [arxiv'25] InfinitePOD: Building Datacenter-Scale High-Bandwidth Domain for LLM with Optical Circuit Switching Transceivers
- [arxiv'25] In-Network Preprocessing of Recommender Systems on Multi-Tenant SmartNICs
- [arxiv'25] Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning
- [arxiv'25] The Power of Negative Zero: Datatype Customization for Quantized Large Language Models
- [arxiv'25] mFabric: An Efficient and Scalable Fabric for Mixture-of-Experts Training
- [arxiv'24] TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication
- [arxiv'24] The Landscape of GPU-Centric Communication
- [arxiv'24] Revisiting the Time Cost Model of AllReduce
- [arxiv'24] LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs
- [arxiv'24] LumosCore: Highly Scalable LLM Clusters with Optical Interconnect
- [arxiv'24] HiCCL: A Hierarchical Collective Communication Library
- [arxiv'24] CSPS: A Communication-Efficient Sequence-Parallelism based Serving System for Transformer based Models with Long Prompts
- [arxiv'24] Domino: Eliminating Communication in LLM Training via Generic Tensor Slicing and Overlapping
- [arxiv'24] Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
- [arxiv'24] Demystifying the Communication Characteristics for Distributed Transformer Models
- [arxiv'24] MLTCP: Congestion Control for DNN Training
- [arxiv'24] ForestColl: Efficient Collective Communications on Heterogeneous Network Fabrics
- [Survey 🔍] [arxiv'23] Communication-Efficient Distributed Deep Learning: A Comprehensive Survey
- [arxiv'23] Optimized Network Architectures for Large Language Model Training with Billions of Parameters
- [arxiv'23] FlexShard: Flexible Sharding for Industry-Scale Sequence Recommendation Models
- [arxiv'23] Rethinking Memory and Communication Cost for Efficient Large Language Model Training
- [arxiv'23] Zen: Near-Optimal Sparse Tensor Synchronization for Distributed DNN Training
- [arxiv'23] Optimized Network Architectures for Large Language Model Training with Billions of Parameters
- [arxiv'23] TACOS: Topology-Aware Collective Algorithm Synthesizer for Distributed Training
- [arxiv'26] Concordia: JIT-Compiled Persistent-Kernel Checkpointing for Fault-Tolerant LLM Inference
- [arxiv'26] ARGUS: Production-Scale Tracing and Performance Diagnosis for over 10,000-GPU Clusters
- [arxiv'26] LUMEN: Coordinated Failure Recovery for Distributed LLM Serving
- [arxiv'26] Don't Let a Few Network Failures Slow the Entire AllReduce
- [arxiv'26] Characterization-Guided GPU Fault Resilience in NVIDIA MPS
- [arxiv'26] Orbax: Distributed Checkpointing with JAX
- [arxiv'26] LiveR: Fine-Grained Elasticity via Live Reconfiguration for Model Training
- [arxiv'26] TierCheck: Tiered Checkpointing for Fault Tolerance in Large Language Model Training
- [arxiv'26] From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs
- [arxiv'26] ResiHP: Taming LLM Training Failures with Dynamic Hybrid
- [arxiv'26] Resilient AI Supercomputer Networking using MRC and SRv6
- [arxiv'26] Role-Based Fault Tolerance System for LLM RL Post-Training
- [arxiv'26] ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
- [arxiv'26] Towards Resiliency in Large Language Model Serving with KevlarFlow
- [arxiv'26] Training LLMs with Fault Tolerant HSDP on 100,000 GPUs
- [arxiv'26] Making MoE-based LLM Inference Resilient with Tarragon
- [arxiv'25] TTrace: Lightweight Error Checking and Diagnosis for Distributed Training
- [arxiv'25] Reliable and Resilient Collective Communication Library for LLM Training and Serving
- [arxiv'25] SHIFT: An RDMA Failure-Resilient Layer for Distributed Training
- [arxiv'25] FFTrainer: Fast Failover in Large-Language Model Training with Almost-Free State Management
- [arxiv'25] FailSafe: High-performance Resilient Serving
- [arxiv'25] GoCkpt: Gradient-Assisted Multi-Step overlapped Checkpointing for Efficient LLM Training
- [arxiv'25] MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
- [arxiv'25] ElasWave: An Elastic-Native System for Scalable Hybrid-Parallel Training
- [arxiv'25] Efficient AllReduce with Stragglers
- [arxiv'25] Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
- [arxiv'25] Nonuniform-Tensor-Parallelism: Mitigating GPU failure impact for Scaled-up LLM Training
- [arxiv'25] Enhancing reliability in AI inference services: An empirical study on real production incidents
- [arxiv'25] Characterizing GPU Resilience and Impact on AI/HPC Systems
- [arxiv'24] FALCON: Pinpointing and Mitigating Stragglers for Large-Scale Hybrid-Parallel Training
- [arxiv'24] TrainMover: Efficient ML Training Live Migration with No Memory Overhead
- [arxiv'24] Cloud Atlas: Efficient Fault Localization for Cloud Systems using Language Models and Causal Insight
- [arxiv'24] ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development
- [arxiv'24] Universal Checkpointing: Efficient and Flexible Checkpointing for Large Scale Distributed Training
- [arxiv'24] Lazarus: Resilient and Elastic Training of Mixture-of-Experts Models with Adaptive Expert Placement
- [arxiv'24] PARALLELGPUOS: A Concurrent OS-level GPU Checkpoint and Restore System using Validated Speculation
- [arxiv'23] Unicron: Economizing Self-Healing LLM Training at Scale
- [arxiv'26] Xema: Efficient Diffusion Serving through Fine-Grained Memory Management and Auto-Configuration
- [arxiv'26] FORGE: Fused On-Register Gradient Elimination for Memory-Efficient LLM Training
- [arxiv'25] CARMA: Collocation-Aware Resource Manager with GPU Memory Estimator
- [arxiv'25] Reducing GPU Memory Fragmentation via Spatio-Temporal Planning for Efficient Large-Scale Model Training
- [arxiv'25] Memory Analysis on the Training Course of DeepSeek Models
- [arxiv'24] Accelerating Large Language Model Training with 4D Parallelism and Memory Consumption Estimator
- [arxiv'23] Rethinking Memory and Communication Cost for Efficient Large Language Model Training
- [arxiv'23] Quantized Distributed Training of Large Models with Convergence Guarantees (
QSDP) - [arxiv'23] Does compressing activations help model parallel training?
- [arxiv'16] Training Deep Nets with Sublinear Memory Cost
- [arxiv'26] Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs
- [arxiv'26] MuxTune: Efficient Multi-Task LLM Fine-Tuning in Multi-Tenant Datacenters via Spatial-Temporal Backbone Multiplexing
- [arxiv'25] MSched: GPU Multitasking via Proactive Memory Scheduling
- [arxiv'25] Towards Efficient and Practical GPU Multitasking in the Era of LLM
- [arxiv'25] Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
- [arxiv'24] PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
- [arxiv'24] Tally: Non-Intrusive Performance Isolation for Concurrent Deep Learning Workloads
- [arxiv'23] GACER: Granularity-Aware ConcurrEncy Regulation for Multi-Tenant Deep Learning
- [arxiv'23] MuxFlow: Efficient and Safe GPU Sharing in Large-Scale Production Deep Learning Clusters
- [arxiv'26] VDCores: Resource Decoupled Programming and Execution for Asynchronous GPU
- [arxiv'26] DITRON: Distributed Multi-level Tiling Compiler for Parallel Tensor Programs
- [arxiv'26] Prism: Symbolic Superoptimization of Tensor Programs
- [arxiv'26] Nautilus: An Auto-Scheduling Tensor Compiler for Efficient Tiled GPU Kernels
- [arxiv'26] Compiler-First State Space Duality and Portable O(1) Autoregressive Caching for Inference
- [arxiv'26] PolyBlocks: A Compiler Infrastructure for AI Chips and Programming Frameworks
- [arxiv'26] ZipFlow: a Compiler-based Framework to Unleash Compressed Data Movement for Modern GPUs
- [arxiv'26] Axe: A Simple Unified Layout Abstraction for Machine Learning Compilers
- [arxiv'25] Tawa: Automatic Warp Specialization for Modern GPUs with Asynchronous References
- [arxiv'25] Dato: A Task-Based Programming Model for Dataflow Accelerators
- [arxiv'25] Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
- [arxiv'25] TileLang: A Composable Tiled Programming Model for AI Systems
- [arxiv'25] Hexcute: A Tile-based Programming Language with Automatic Layout and Task-Mapping Synthesis
- [arxiv'25] DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
- [arxiv'25] Hercules: A Compiler for Productive Programming of Heterogeneous Systems
- [arxiv'26] SparseDitto: Customizing GPU Kernels for Different Sparsity Patterns with LLM-Based Agentic System
- [arxiv'26] KernelBrain: Coarse-to-Fine, Budget-Aware Search for Agentic GPU Kernel Optimization
- [arxiv'26] TileSight: A First-Principles Tile-Centric Analytical GPU Performance Model from Cores to Clusters
- [arxiv'26] KernelBench-Verified: Do LLM-Generated Kernels Actually Beat PyTorch?
- [arxiv'26] ExpertPlex: A High-Goodput Disaggregated Serving System for MoE LLMs with Adaptive Persistent Kernels
- [arxiv'26] Are LLM-Generated GPU Kernels Production-Ready? A Trace-Driven Benchmark and Optimization Agent
- [arxiv'26] SpecGen: Accelerating Agentic Kernel Optimization with Speculative Generation
- [arxiv'26] AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel Synthesis
- [arxiv'26] FastKernels: Benchmarking GPU Kernel Generation in Production
- [arxiv'26] CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs
- [arxiv'26] ARGUS: Agentic GPU Optimization Guided by Data-Flow Invariants
- [arxiv'26] Ragged Paged Attention: A High-Performance and Flexible LLM Inference Kernel for TPU
- [arxiv'26] Fleet: Hierarchical Task-based Abstraction for Megakernels on Multi-Die GPUs
- [arxiv'26] WaveTune: Wave-aware Bilinear Modeling for Efficient GPU Kernel Auto-tuning
- [arxiv'26] Improving Efficiency of GPU Kernel Optimization Agents using a Domain-Specific Language and Speed-of-Light Guidance
- [arxiv'26] Kernel-Smith: A Unified Recipe for Evolutionary Kernel Optimization
- [arxiv'26] AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search
- [arxiv'26] SOL-ExecBench: Speed-of-Light Benchmarking for Real-World GPU Kernels Against Hardware Limits
- [arxiv'26] CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
- [arxiv'26] K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model
- [arxiv'26] Deep Kernel Fusion for Transformers
- [arxiv'26] Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
- [arxiv'25] Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
- [arxiv'25] KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta
- [arxiv'25] Memory-Efficient Acceleration of Block Low-Rank Foundation Models on Resource Constrained GPUs
- [arxiv'25] FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection
- [arxiv'25] Flash Multi-Head Feed-Forward Network
- [arxiv'25] Iris: First-Class Multi-GPU Programming Experience in Triton
- [arxiv'25] AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization
- [arxiv'25] ParallelKittens: Systematic and Practical Simplification of Multi-GPU AI Kernels
- [arxiv'25] HipKittens: Fast and Furious AMD Kernels
- [arxiv'25] LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving
- [arxiv'25] TileLang: A Composable Tiled Programming Model for AI Systems
- [arxiv'24] ACS: Concurrent Kernel Execution on Irregular, Input-Dependent Computational Graphs
- [arxiv'24] Flex Attention: A Programming Model for Generating Optimized Attention Kernels
- [arxiv'23] Stream-K: Work-centric Parallel Decomposition for Dense Matrix-Matrix Multiplication on the GPU
- [arxiv'21] Characterizing Concurrency Mechanisms for NVIDIA GPUs under Deep Learning Workloads
- [arxiv'26] Libra: Taming Attention Workload Skew in Long-Context LLM Training with Bounded Sequence Pool
- [arxiv'26] HCMS: Head-Chunked Multi-Stream Pipeline for Communication-Computation Overlap in Long-Sequence Parallel Attention
- [arxiv'26] Lynx: Progressive Speculative Quantization for accelerating KV Transfer in Long-Context Inference
- [arxiv'26] KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding
- [arxiv'26] MosaicKV: Serving Long-Context LLM with Dynamic Two-D KV Cache Compression
- [arxiv'26] FlashCP: Load-Balanced Communication-Efficient Context Parallelism for LLM Training
- [arxiv'26] RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention
- [arxiv'26] SparDA: Sparse Decoupled Attention for Efficient Long-Context LLM Inference
- [arxiv'26] Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
- [arxiv'26] Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving
- [arxiv'25] Optimizing Long-context LLM Serving via Fine-grained Sequence Parallelism
- [arxiv'25] Long-Context Attention Benchmark: From Kernel Efficiency to Distributed Context Parallelism
- [arxiv'25] Efficient Long-context Language Model Training by Core Attention Disaggregation
- [arxiv'25] Data-Centric Elastic Pipeline Parallelism for Efficient Long-Context LLM Training
- [arxiv'25] Strata: Hierarchical Context Caching for Long Context Language Model Serving
- [arxiv'25] TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving
- [arxiv'25] HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
- [arxiv'25] SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling
- [arxiv'25] Training Long-Context LLMs Efficiently via Chunk-wise Optimization
- [arxiv'25] SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM Training
- [arxiv'25] XAttention: Block Sparse Attention with Antidiagonal Scoring
- [arxiv'25] SPPO:Efficient Long-sequence LLM Training via Adaptive Sequence Pipeline Parallel Offloading
- [arxiv'25] ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs
- [arxiv'25] Long-Context Inference with Retrieval-Augmented Speculative Decoding
- [arxiv'25] ParallelComp: Parallel Long-Context Compressor for Length Extrapolation
- [arxiv'25] LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
- [arxiv'25] MoBA: Mixture of Block Attention for Long-Context LLMs
- [arxiv'25] Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs
- [arxiv'25] APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
- [arxiv'25] Twilight: Adaptive Attention Sparsity with Hierarchical Top-p Pruning
- [arxiv'25] Adjoint sharding for very long context training of state space models
- [arxiv'24] LoL-PIM: Long-Context LLM Decoding with Scalable DRAM-PIM System
- [arxiv'24] Data-Centric and Heterogeneity-Adaptive Sequence Parallelism for Efficient LLM Training
- [arxiv'24] USP: A Unified Sequence Parallelism Approach for Long Context Generative AI
- [arxiv'24] Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer
- [arxiv'24] ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
- [arxiv'24] Mnemosyne: Parallelization Strategies for Efficiently Serving Multi-Million Context Length LLM Inference Requests Without Approximations
- [arxiv'24] CSPS: A Communication-Efficient Sequence-Parallelism based Serving System for Transformer based Models with Long Prompts
- [arxiv'24] FocusLLM: Scaling LLM's Context by Parallel Decoding
For comprehensive list of quantization papers, refer to https://github.com/Efficient-ML/Awesome-Model-Quantization.
- [arxiv'26] PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization
- [arxiv'26] TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs
- [arxiv'26] TriAxialKV: Toward Extreme Low-Precision KV-Cache Quantization for Agentic Inference Tasks
- [arxiv'26] Grid Games: The Power of Multiple Grids for Quantizing Large Language Models
- [arxiv'26] SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
- [arxiv'26] Sparse-BitNet: 1.58-bit LLMs are Naturally Friendly to Semi-Structured Sparsity
- [arxiv'26] Practical FP4 Training for Large-Scale MoE Models on Hopper GPUs
- [arxiv'26] Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery
- [arxiv'25] Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
- [arxiv'25] MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
- [arxiv'25] TAH-QUANT: Effective Activation Quantization in Pipeline Parallelism over Slow Network
- [arxiv'25] DECA: A Near-Core LLM Decompression Accelerator Supporting Out-of-Order Invocation
- [arxiv'25] ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition
- [arxiv'24] Accelerating Distributed Deep Learning using Lossless Homomorphic Compression
- [arxiv'24] FedMoE: Personalized Federated Learning via Heterogeneous Mixture of Experts
- [arxiv'24] FedEx: Expediting Federated Learning over Heterogeneous Mobile Devices by Overlapping and Participant Selection
- [arxiv'24] Decoupled Vertical Federated Learning for Practical Training on Vertically Partitioned Data
- [arxiv'23] CAFE: Carbon-Aware Federated Learning in Geographically Distributed Data Centers
- [arxiv'23] Federated Learning of Large Language Models with Parameter-Efficient Prompt Tuning and Adaptive Optimization
- [arxiv'26] SimpleTool: Parallel Decoding for Real-Time LLM Function Calling
- [arxiv'24] APIServe: Efficient API Support for Large-Language Model Inferencing
- [arxiv'26] CommBench: Can LLMs Write Correct and Efficient GPU Communication Code?
- [arxiv'26] SparseDitto: Customizing GPU Kernels for Different Sparsity Patterns with LLM-Based Agentic System
- [arxiv'26] KernelBrain: Coarse-to-Fine, Budget-Aware Search for Agentic GPU Kernel Optimization
- [arxiv'26] Harness Engineering for LLM-Driven GPU Kernel Generation
- [arxiv'26] KernelBench-Verified: Do LLM-Generated Kernels Actually Beat PyTorch?
- [arxiv'26] Are LLM-Generated GPU Kernels Production-Ready? A Trace-Driven Benchmark and Optimization Agent
- [arxiv'26] SOLAR: AI-Powered Speed-of-Light Performance Analysis
- [arxiv'26] SpecGen: Accelerating Agentic Kernel Optimization with Speculative Generation
- [arxiv'26] Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems
- [arxiv'26] PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
- [arxiv'26] VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
- [arxiv'26] KEET: Explaining Performance of GPU Kernels Using LLM Agents
- [arxiv'26] Assistants, Not Architects: The Role of LLMs in Networked Systems Design
- [arxiv'26] Improving Efficiency of GPU Kernel Optimization Agents using a Domain-Specific Language and Speed-of-Light Guidance
- [arxiv'26] Kernel-Smith: A Unified Recipe for Evolutionary Kernel Optimization
- [arxiv'26] AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search
- [arxiv'26] Improving Coherence and Persistence in Agentic AI for System Optimization
- [arxiv'26] CUCo: An Agentic Framework for Compute and Communication Co-design
- [arxiv'26] StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning
- [arxiv'26] K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model
- [arxiv'26] AdaEvolve: Adaptive LLM Driven Zeroth-Order Optimization
- [arxiv'25] AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization
- [arxiv'25] ASAP: an Agentic Solution to Auto-optimize Performance of Large-Scale LLM Training
- [arxiv'25] Barbarians at the Gate: How AI is Upending Systems Research [Code]
- [arxiv'25] SuperCoder: Assembly Program Superoptimization with Large Language Models
- [arxiv'24] Cloud Atlas: Efficient Fault Localization for Cloud Systems using Language Models and Causal Insight
- [arxiv'24] LLMTune: Accelerate Database Knob Tuning with Large Language Models
- [arxiv'24] LLM-Enhanced Data Management
- [arxiv'24] MPIrigen: MPI Code Generation through Domain-Specific Language Models
- [arxiv'24] Can Large Language Models Write Parallel Code?
- [arxiv'23] LLM-Assisted Code Cleaning For Training Accurate Code Generators
- [arxiv'23] Large Language Models for Compiler Optimization
- [arxiv'26] PowerScale: Energy-Efficient Geo-Distributed Model Training with Federated Datacenter Power
- [arxiv'26] Energy Calculus: A Compositional Algebra of Energy in Computational Systems
- [arxiv'26] Energy-Aware Scheduling for Serverless LLM Serving on Shared GPUs
- [arxiv'26] The Language-Energy Divide: Measuring Energy Costs of Multilingual LLM Inference
- [arxiv'26] Power-Flexible AI Data Centers: A New Paradigm for Grid-Responsive Compute
- [arxiv'26] PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
- [arxiv'26] EnergyLens: Predictive Energy-Aware Exploration for Multi-GPU LLM Inference Optimization
- [arxiv'26] OpenG2G: A Simulation Platform for AI Datacenter-Grid Runtime Coordination
- [arxiv'26] Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips
- [arxiv'26] From Barrier to Bridge: The Case for AI Data Center/Power Grid Co-Design
- [arxiv'26] The Energy Cost of Execution-Idle in GPU Clusters
- [arxiv'26] Where Do the Joules Go? Diagnosing Inference Energy Consumption
- [arxiv'26] Kareus: Joint Reduction of Dynamic and Static Energy in Large Model Training
- [arxiv'26] GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
- [arxiv'25] VoltanaLLM: Feedback-Driven Frequency Control and State-Space Routing for Energy-Efficient LLM Serving
- [arxiv'25] GreenLLM: SLO-Aware Dynamic Frequency Scaling for Energy-Efficient LLM Serving
- [arxiv'25] Power Stabilization for AI Training Datacenters
- [arxiv'25] The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization
- [arxiv'25] EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
- [arxiv'25] EcoServe: Designing Carbon-Aware AI Inference Systems
- [arxiv'25] Life-Cycle Emissions of AI Hardware: A Cradle-To-Grave Approach and Generational Trends
- [arxiv'24] GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
- [arxiv'24] EaCO: Resource Sharing Dynamics and Its Impact on Energy Efficiency for DNN Training
- [arxiv'24] DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
- [arxiv'23] CAFE: Carbon-Aware Federated Learning in Geographically Distributed Data Centers
- [arxiv'26] RAG-Stack: Co-Optimizing RAG Serving Performance and Quality
- [arxiv'26] QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving
- [arxiv'26] PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
- [arxiv'25] Patchwork: A Unified Framework for RAG Serving
- [arxiv'25] Accelerating Retrieval-Augmented Language Model Serving with Speculation
- [arxiv'25] RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation Serving
- [arxiv'25] Long-Context Inference with Retrieval-Augmented Speculative Decoding
- [arxiv'24] Towards Understanding Systems Trade-offs in Retrieval-Augmented Generation Model Inference
- [arxiv'24] RAGServe: Fast Quality-Aware RAG Systems with Configuration Adaptation
- [arxiv'24] Dehallucinating Parallel Context Extension for Retrieval-Augmented Generation
- [arxiv'24] Accelerating Retrieval-Augmented Language Model Serving with Speculation
- [arxiv'26] KernelSight-LM: A Kernel-Level LLM Inference Simulator
- [arxiv'26] Simulating Unified Tensor Resharding in heterogeneous AI systems
- [arxiv'26] AGENTSERVESIM: A Hardware-aware Simulator for Multi-Turn LLM Agent Serving
- [arxiv'26] ASTRA-sim 3.0: Next-Level Distributed Machine Learning Simulations via High-Fidelity GPU and Infrastructure Modeling
- [arxiv'26] A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM
- [arxiv'26] Frontier: Towards Comprehensive and Accurate LLM Inference Simulation
- [arxiv'26] LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure
- [arxiv'26] SynPerf: A Hybrid Analytical-ML Framework for GPU Performance Prediction
- [arxiv'26] Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving
- [arxiv'25] Scalable Synthesis of distributed LLM workloads through Symbolic Tensor Graphs
- [arxiv'25] Frontier: Simulating the Next Generation of LLM Inference Systems
- [arxiv'25] Frontier: Simulating the Next Generation of LLM Inference Systems
- [arxiv'25] Maya: Optimizing Deep Learning Training Workloads using Emulated Virtual Accelerators
- [arxiv'26] When Does Disaggregation Pay? Simulating Prefill--Decode--Attention--FFN Specialization for Agentic LLM Inference
- [arxiv'26] AtumAI: A Principled Framework for Agentic Generation of Datacenter Control-Plane Policies
- [arxiv'26] Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework
- [arxiv'26] TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI
- [arxiv'26] SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving
- [arxiv'26] Scalable LLM Agent Tool Access in the Cloud
- [arxiv'26] A Workflow-Aware Serving Layer for Agentic Applications
- [arxiv'26] SPORK: Self-Speculative Forking to Accelerate Agentic LLM Inference
- [arxiv'26] SwarmX: Agentic Scheduling for Low-Latency Agentic Systems
- [arxiv'26] CoAgent: Concurrency Control for Multi-Agent Systems
- [arxiv'26] Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents
- [arxiv'26] Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloads
- [arxiv'26] Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI
- [arxiv'26] Agentic AI Workload Characteristics
- [arxiv'26] VineLM: Trie-Based Fine-Grained Control for Agentic Workflows
- [arxiv'26] HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
- [arxiv'26] xGoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
- [arxiv'26] MARS: Efficient, Adaptive Co-Scheduling for Heterogeneous Agentic Systems
- [arxiv'26] Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
- [arxiv'26] KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving
- [arxiv'26] Heddle: A Distributed Orchestration System for Agentic RL Rollout
- [arxiv'26] Efficient LLM Serving for Agentic Workflows: A Data Systems Perspective
- [arxiv'26] Orla: A Library for Serving LLM-Based Multi-Agent Systems
- [arxiv'26] SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
- [arxiv'26] CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
- [arxiv'26] ArchAgent: Agentic AI-driven Computer Architecture Discovery
- [arxiv'26] Pancake: Hierarchical Memory System for Multi-Agent LLM Serving
- [arxiv'26] DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference
- [arxiv'26] ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System
- [arxiv'26] REDSearcher: A Scalable and Cost-Efficient Framework for Long-Horizon Search Agents
- [arxiv'26] LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM Agents
- [arxiv'26] VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents
- [arxiv'26] ToolCaching: Towards Efficient Caching for LLM Tool-calling
- [arxiv'26] Toward Efficient Agents: Memory, Tool learning, and Planning
- [arxiv'26] Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference
- [arxiv'26] Beyond Max Tokens: Stealthy Resource Amplification via Tool Calling Chains in LLM Agents
- [arxiv'26] XGrammar 2: Dynamic and Efficient Structured Generation Engine for Agentic LLMs
- [arxiv'26] Nalar: An agent serving framework
- [arxiv'26] Software-Defined Agentic Serving
- [arxiv'25] Speculative Actions: A Lossless Framework for Faster Agentic Systems
- [arxiv'25] ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
- [arxiv'25] Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- [arxiv'25] Towards Efficient Agents: A Co-Design of Inference Architecture and System
- [arxiv'25] Beyond Training: Enabling Self-Evolution of Agents with MOBIMEM
- [arxiv'25] Optimizing Agentic Language Model Inference via Speculative Tool Calls
- [arxiv'25] Astraea: A State-Aware Scheduling Engine for LLM-Powered Agents
- [arxiv'25] Measuring Agents in Production
- [arxiv'25] Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework
- [arxiv'25] Aragog: Just-in-Time Model Routing for Scalable Serving of Agentic Workflows
- [arxiv'25] AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization
- [arxiv'25] Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
- [arxiv'25] Sherlock: Reliable and Efficient Agentic Workflow Execution
- [arxiv'25] A CPU-Centric Perspective on Agentic AI
- [arxiv'25] Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution
- [arxiv'25] MobiAgent: A Systematic Framework for Customizable Mobile Agents
- [arxiv'25] rStar2-Agent: Agentic Reasoning Technical Report
- [arxiv'25] Efficient and Scalable Agentic AI with Heterogeneous Systems
- [arxiv'25] Agent.xpu: Efficient Scheduling of Agentic LLM Workloads on Heterogeneous SoC
- [arxiv'25] GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents
- [arxiv'25] The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks
- [arxiv'24] AI Metropolis: Scaling Large Language Model-based Multi-Agent Simulation with Out-of-order Execution
https://github.com/friedrichor/Awesome-Multimodal-Papers
- [arxiv'26] Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference
- [arxiv'26] HorizonServe: Coordinating Request Scheduling with GPU Sharing for Omni-Model Serving
- [arxiv'26] ACID: Adaptive Caching for vIDeo generation
- [arxiv'26] Kaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space Correlations
- [arxiv'26] Sangam: Efficiently Serving Diffusion LLMs with the AR Stack
- [arxiv'26] LiveServe: Interaction-Aware Serving for Real-Time Omni-Modal LLMs
- [arxiv'26] Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse
- [arxiv'26] AoiZora: Topology-Aware Auto-Parallel Optimization for Inference of Diffusion Transformers
- [arxiv'26] M*: A Modular, Extensible, Serving System for Multimodal Models
- [arxiv'26] Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain
- [arxiv'26] Heterogeneous Parallelism for Multimodal Large Language Model Training
- [arxiv'26] DisagFusion: Asynchronous Pipeline Parallelism and Elastic Scheduling for Disaggregated Diffusion Serving
- [arxiv'26] Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing
- [arxiv'26] Cornserve: A Distributed Serving System for Any-to-Any Multimodal Models
- [arxiv'26] BitDance: Scaling Autoregressive Generative Models with Binary Tokens
- [arxiv'26] vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models
- [arxiv'26] VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents
- [arxiv'26] EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
- [arxiv'25] Cornserve: Efficiently Serving Any-to-Any Multimodal Models
- [arxiv'25] FoundationMotion: Auto-Labeling and Reasoning about Spatial Movement in Videos
- [arxiv'25] MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
- [arxiv'25] FlowMM: Cross-Modal Information Flow Guided KV Cache Merging for Efficient Multimodal Context Inference
- [arxiv'25] OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
- [arxiv'25] Fast-dLLM v2: Efficient Block-Diffusion LLM
- [arxiv'25] Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
- [arxiv'25] Mordal: Automated Pretrained Model Selection for Vision Language Models
- [arxiv'25] Dimple: Discrete Diffusion Multimodal Large Language Model with Parallel Decoding
- [arxiv'24] LlamaFusion: Adapting Pretrained Language Models for Multimodal Generation
- [Survey 🔍] [arxiv'24] A Survey of Resource-efficient LLM and Multimodal Foundation Models
- [arxiv'26] TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure
- [arxiv'26] Application-Driven Architecture Exploration for Cross-Layer Heterogeneous Systems
- [arxiv'26] Enabling Multi-Dimensional Distributed Trace Comparison with Contrast
- [arxiv'26] LLM-as-a-Verifier: A General-Purpose Verification Framework
- [arxiv'26] Eiger: An Efficient Library for GPU-based Data Analytics
- [arxiv'26] ROSA: A Robotics Foundation Model Serving System for Robot Factories
- [arxiv'26] Instant GPU Efficiency Visibility at Fleet Scale
- [arxiv'26] The Price Reversal Phenomenon: When Cheaper Reasoning Models End Up Costing More
- [arxiv'26] Challenges and Design Considerations for Finding CUDA Bugs Through GPU-Native Fuzzing
- [arxiv'26] AI+HW 2035: Shaping the Next Decade
- [arxiv'26] SKYLIGHT: A Scalable Hundred-Channel 3D Photonic In-Memory Tensor Core Architecture for Real-time AI Inference
- [arxiv'26] Semantic Search At LinkedIn
- [arxiv'26] Long-term Monitoring of Kernel and Hardware Events to Understand Latency Variance
- [arxiv'25] Cyclotron: Compilation of Recurrences to Distributed and Systolic Architectures
- [arxiv'25] Streaming Tensor Program: A streaming abstraction for dynamic parallelism
- [arxiv'25] OckBench: Measuring the Efficiency of LLM Reasoning
- [arxiv'25] vAttention: Verified Sparse Attention
- [arxiv'25] Training Large Language Models To Reason In Parallel With Global Forking Tokens
- [arxiv'25] How to Train Your Advisor: Steering Black-Box LLMs with Advisor Models
- [arxiv'25] Slm-mux: Orchestrating small language models for reasoning
- [arxiv'25] Hybrid Architectures for Language Models: Systematic Analysis and Design Insights
- [arxiv'25] Less is More: Recursive Reasoning with Tiny Networks
- [arxiv'25] ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
- [arxiv'25] Rethinking Thinking Tokens: LLMs as Improvement Operators
- [arxiv'25] Generalized Parallel Scaling with Interdependent Generations
- [arxiv'25] Composer: A Search Framework for Hybrid Neural Architecture Design
- [arxiv'25] dParallel: Learnable Parallel Decoding for dLLMs
- [arxiv'25] AI Factories: It's time to rethink the Cloud-HPC divide
- [arxiv'25] Efficient Training-Free Online Routing for High-Volume Multi-LLM Serving
- [arxiv'25] SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences
- [arxiv'25] Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
- [arxiv'25] LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
- [arxiv'25] DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis
- [arxiv'25] Less Is More: Training-Free Sparse Attention with Global Locality for Efficient Reasoning
- [arxiv'25] Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
- [arxiv'25] LobRA: Multi-tenant Fine-tuning over Heterogeneous Data
- [arxiv'25] Copilot Arena: A Platform for Code LLM Evaluation in the Wild
- [arxiv'25] ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism
- [arxiv'25] Libra: Synergizing CUDA and Tensor Cores for High-Performance Sparse Matrix Multiplication
- [arxiv'25] Prompt-to-Leaderboard: Prompt-Adaptive LLM Evaluations [Code]
- [arxiv'25] SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
- [arxiv'25] Reinforcement Pre-Training
- [arxiv'25] MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models
- [arxiv'25] Thinking Short and Right Over Thinking Long: Serving LLM Reasoning Efficiently and Accurately
- [arxiv'25] Faster Video Diffusion with Trainable Sparse Attention
- [arxiv'25] SSR: Speculative Parallel Scaling Reasoning in Test-time
- [arxiv'25] Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought
- [arxiv'25] Think Only When You Need with Large Hybrid-Reasoning Models
- [arxiv'25] Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers
- [arxiv'25] Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
- [arxiv'25] Process Reward Models That Think
- [arxiv'25] Seed-Thinking-v1.5: Advancing Superb Reasoning Models with Reinforcement Learning
- [arxiv'25] Sleep-time Compute: Beyond Inference Scaling at Test-time
- [arxiv'25] SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning
- [arxiv'25] Scaling Laws for Native Multimodal Models Scaling Laws for Native Multimodal Models
- [arxiv'25] OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens
- [arxiv'25] NotebookOS: A Notebook Operating System for Interactive Training with On-Demand GPUs
- [arxiv'25] Alchemist: Towards the Design of Efficient Online Continual Learning System
- [arxiv'25] Linear Attention for Efficient Bidirectional Sequence Modeling
- [arxiv'25] S*: Test Time Scaling for Code Generation
- [arxiv'25] Optimizing Model Selection for Compound AI Systems
- [arxiv'25] Copilot Arena: A Platform for Code LLM Evaluation in the Wild
- [arxiv'25] Efficient-vDiT: Efficient Video Diffusion Transformers with Attention Tile
- [arxiv'25] BARE: Combining Base and Instruction-Tuned Language Models for Better Synthetic Data Generation
- [arxiv'25] Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
- [arxiv'25] Adaptive Semantic Prompt Caching with VectorQ
- [arxiv'25] Measuring GPU utilization one level deeper
- [arxiv'24] Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
- [arxiv'24] Debunking the CUDA Myth Towards GPU-based AI Systems
- [arxiv'24] XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models
- [arxiv'24] Scorch: A Library for Sparse Deep Learning
- [arxiv'24] Drowning in Documents: Consequences of Scaling Reranker Inference
- [arxiv'24] Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs Questions
- [arxiv'24] Computational Bottlenecks of Training Small-scale Large Language Models
- [Survey 🔍] [arxiv'24] A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness
- [arxiv'24] Stochastic Monkeys at Play: Random Augmentations Cheaply Break LLM Safety Alignment
- [arxiv'24] DroidSpeak: Enhancing Cross-LLM Communication
- [arxiv'24] Disaggregating Embedding Recommendation Systems with FlexEMR
- [arxiv'24] JudgeBench: A Benchmark for Evaluating LLM-based Judges
- [arxiv'24] You Only Need One Step: Fast Super-Resolution with Stable Diffusion via Scale Distillation
- [arxiv'24] Computing in the Era of Large Generative Models: From Cloud-Native to AI-Native
- [arxiv'23] Efficiently Programming Large Language Models using SGLang
- [arxiv'23] Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- [arxiv'22] Training language models to follow instructions with human feedback