Skip to content
 
 

Latest commit

 

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

[ICML'26] TWLA: Achieving Ternary Weights and Low-Bit Activations for LLMs via Post-Training Quantization

Zhixiong Zhao*, Zukang Xu*, Zhixuan Chen, Xing Hu, Zhe Jiang and Dawei Yang


Paper arXiv Code GitHub

🔥🔥🔥 News

  • 2026-05-18: Code is released. ⭐️⭐️⭐️
  • 2026-05-01: TWLA is accepted at ICML 2026. 🎉🎉🎉
  • 2025-01-29: This repo is released.

Abstract: Large language models (LLMs) exhibit exceptional general language processing capabilities, but their memory and compute costs hinder deployment. Ternarization has emerged as a promising compression technique, offering significant reductions in model size and inference complexity. However, existing methods struggle with heavy-tailed activation distributions and therefore keep activations in high precision, fundamentally limiting end-to-end inference acceleration. To overcome this limitation, we propose TWLA, a post-training quantization (PTQ) framework that achieves 1.58-bit weight compression and 4-bit activation quantization while maintaining high accuracy. TWLA comprises three components: (1) Euclidean-to-Manifold Asymmetric Ternary Quantizer (E2M-ATQ) minimizes layer-output error under weight ternarization via a two-stage optimization from Euclidean initialization to manifold relocation; (2) Kronecker Orthogonal Tri-Modal Shaping (KOTMS) applies a Kronecker-structured orthogonal rotation to reshape weights into ternary-friendly tri-modal distributions, while the shared rotation statistically suppresses activation outliers; and (3) Inter-Layer Aware Activation Mixed Precision (ILA-AMP) explicitly introduces adjacent-layer second-order interaction costs in bit allocation and jointly optimizes for the layer-wise disparity of activation quantization gains induced by the shared orthogonal transform, preventing cascades triggered by a few weak layers. Extensive experiments demonstrate that TWLA is a PTQ method that maintains high accuracy under the W1.58A4 configuration, while delivering significant inference acceleration. The code is available at TWLA.


Figure 1 in the main paper demonstrates that our proposed TWLA remains robust under both weight-only and weight–activation quantization (at equal memory cost), while other methods degrade substantially with 4-bit activation quantization.

Dependencies

# Clone the github repo and go to the default directory 'TWLA'.
conda create -n twla python=3.9
conda activate twla
pip install torch torchvision torchaudio
pip install -r requirements.txt

🔗 Contents

  1. Post-training quantization and evaluation
  2. Results
  3. Citation
  4. Acknowledgements

Post-training quantization with PPL evaluation (Example: Qwen3-8B)

KOTMS

    python scripts/KOTMS.py \
    --model Qwen/Qwen3-8B \
    --export_rotated checkpoints/qwen3_8b_rotated.pt \
    --ngpus 4 \
    --use_gmm \
    --gmm_iters 100 \
    --gmm_lr_r 1e-2 \
    --gmm_lr_l 1e-2

ILA-AMP

    python scripts/ILA_AMP.py \
    --model Qwen/Qwen3-8B \
    --import_rotated checkpoints/qwen3_8b_rotated.pt \
    --dp_cache dp_cache/qwen3_8b \
    --dp_ngpus 4

Run-TWLA

Weight-only (W1.58A16)

    python run_twla.py \
    --model Qwen/Qwen3-8B \
    --import_rotated checkpoints/qwen3_8b_rotated.pt \
    --dp_cache dp_cache/qwen3_8b \
    --save_quant_model save_models/Qwen3-8B \
    --eval_qa \
    --abits 16

Weight-Activation (W1.58A4)

    python run_twla.py \
    --model Qwen/Qwen3-8B \
    --import_rotated checkpoints/qwen3_8b_rotated.pt \
    --dp_cache dp_cache/qwen3_8b \
    --load_quant_model save_models/Qwen3-8B \
    --eval_qa \
    --dp_avg_abits 4

Evaluation on zero-shot QA datasets

We use lm-evaluation-harness kit to evaluate performance on QA datasets. Please refer to their framework for evaluating quantized models.

🔎 Results

TWLA achieves superior perplexity performance on WikiText2 datasets and superior average accuracy on 7 zero-shot QA datasets. (click to expand)

Citation

If you find the code helpful in your research or work, please cite the following paper.

@misc{zhao2026twlaachievingternaryweights,
      title={TWLA: Achieving Ternary Weights and Low-Bit Activations for LLMs via Post-Training Quantization}, 
      author={Zhixiong Zhao and Zukang Xu and Zhixuan Chen and Xing Hu and Zhe Jiang and Dawei Yang},
      year={2026},
      eprint={2606.13054},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2606.13054}, 
}

💡 Acknowledgements

This work is released under the Apache 2.0 license. The codes are based on ARB-LLM. Please also follow their licenses. Thanks for their awesome works.

About

[ICML'26] TWLA: Achieving Ternary Weights and Low-Bit Activations for LLMs via Post-Training Quantization

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages