Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization

arXiv EMNLP 2026 Main Hugging Face model Project Page GitHub Stars License: MIT

Yipin Guo, Arun M George, Jie Fu, Tareq Mahmoud, Sixue Xing, Siddharth Joshi

University of Notre Dame · EMNLP 2026 (Main)


News

  • 2026-08-21: QTEA is accepted to the EMNLP 2026 Main Conference.

Abstract

QTEA is a post-training quantization method that quantizes the linear layers of a decoder-only LLM into effectively 1.7 bits per weight using just a single pass over 256 calibration sequences.

We treat salient weights as error compensators. QTEA first quantizes every weight into a compact ternary base {-1, 0, +1}, then spends a small residual budget on the columns where ternarization causes the largest drop in accuracy. This allows for keeping one FP8 value per four rows inside those columns, facilitating a column-semi sparse layout that recovers most of the accuracy of unstructured salient weights while staying GPU-friendly. We also implement two further changes to the GPTQ-style column-by-column sweep: we use a per-column rescale factor that is jointly optimized with the ternary assignments, and we introduce an error decay term that attenuates error propagation so late columns are not over-compensated. A lookup-table CUDA kernel then evaluates the result ensuring that the hardware can take advantage of the ternarization.

QTEA overview

Highlights

Quality. QTEA achieves the best average zero-shot accuracy among the methods we evaluated, with the advantage increasing with model size.

  • Qwen3-14B: average zero-shot accuracy rises from 45.11% to 52.65%, a 16.7% relative gain, with 1.40× lower WikiText-2 perplexity (16.48 → 11.78) and 2.61× lower C4 perplexity (68.13 → 26.14).
  • Llama3-8B: accuracy improves from 37.79% to 40.29% (6.6% relative), with 1.34× and 1.95× lower WikiText-2 and C4 perplexity.
  • The 1:4 semi-sparse residual costs only 0.9 accuracy points against an unstructured salient-weight upper bound, at 4× lower residual storage.

Efficiency. The lookup-table kernel delivers practical speedups.

  • 7.2× faster per-token generation than FP16 with CUDA Graphs on Llama2-70B (41.13 → 5.70 ms/token), and 13.3× against FP16 without CUDA Graphs; 3.62× on Qwen3-14B.
  • Latency matches a ternary-only kernel to within 0.01–0.03 ms/token.
  • Speedup is stable at 3.07–3.20× for batch size 1 across 512–4096 tokens of context, and still 2.54× at batch size 4.
  • On a commercially available TSMC 22nm-based implementation, a co-designed accelerator delivers 3.83× lower latency and 69.4% lower energy than dense FP16 matrix multiplication.

Comparison with sub-2-bit / ternary PTQ baselines (paper Table 1)

Method Qwen3-14B Wiki2 PPL ↓ Qwen3-14B C4 PPL ↓ Qwen3-14B 0-shot avg ↑ Llama3-8B Wiki2 PPL ↓ Llama3-8B C4 PPL ↓ Llama3-8B 0-shot avg ↑
FP16 6.38 9.68 68.23 6.14 9.45 65.59
GPTQ 37.90 74.50 37.31 1480.43 394.74 33.30
Slim-LLM 22.85 68.38 44.13 38.21 390.02 34.52
PB-LLM 2.89e4 2.44e4 32.50 73.08 104.15 36.25
PT²-LLM 16.48 68.13 45.11 32.19 129.83 37.79
QTEA (1.7 bit) 11.78 26.14 52.65 24.09 66.45 40.29

WikiText-2 perplexity vs. model size on Qwen3


Installation

conda create -n qtea python=3.12 -y
conda activate qtea
pip install -r requirements.txt

If you already maintain a PyTorch CUDA environment, install the packages from requirements.txt there instead of creating a new one. lm_eval is only needed for zero-shot accuracy, and a CUDA toolchain (nvcc) is only needed for the lookup-table inference kernel.

Quantization

To quantize a model yourself:

python quantize.py --model Qwen/Qwen3-8B-Base --output packed/qwen3-8b.pt

This writes a packed checkpoint holding the ternary codes, with the residuals and the scales compressed. Only one decoder block is resident on the GPU at a time, so the footprint is set by the calibration activations and the largest block rather than by the model. Lower --nsamples if OOM;

A QTEA-quantized Qwen3-14B-Base checkpoint is also published on the Hugging Face Hub as ims-lab/Qwen3-14B-base-QTEA, so you can skip this step and go straight to Evaluation:

huggingface-cli download ims-lab/Qwen3-14B-base-QTEA packed_qwen3_14b.pt --local-dir packed/

Evaluation

python eval/evaluate.py --model Qwen/Qwen3-8B-Base --checkpoint packed/qwen3-8b.pt

Reports WikiText-2 and C4 perplexity plus zero-shot accuracy on the seven tasks above. Drop --checkpoint for the FP16 baseline, and add --backend lut to run the packed weights through the lookup-table CUDA kernel instead of rebuilding FP16 weights. See eval/README.md for the full set of options and for how the kernel works.

Reproducing the Paper

bash scripts/quantize_qwen3.sh          # Qwen3-Base, 0.6B to 14B
bash scripts/quantize_llama3.sh         # Llama3-8B

Both scripts quantize and then evaluate, writing JSON results to results/. The defaults in quantize.py are the values used for paper. Note that We apply an additional ternary-fitting round for Llama3, this is captured in the script.

Reproduced here on one H200 with torch 2.8 / transformers 4.56:

model WikiText-2 C4 zero-shot avg
Qwen3-0.6B-Base 63.58 172.66 34.37
Qwen3-14B-Base 11.78 26.14 52.67
Llama3-8B 24.09 66.39 40.13

Quantization is deterministic for a given GPU and library stack, but numerics are impacted by differences between environments: the ternary assignment is a hard threshold inside a sequential error-feedback loop, which can be impacted by rounding differences between hardware.

Llama-3 and Qwen3 are supported out of the box. Any decoder stack that exposes model.model.layers with the usual seven projections works unchanged; other families need their projection names added to qtea/sequential.py.


FAQ

What is the best way to quantize an LLM below 2 bits without retraining? QTEA is a post-training method: one calibration pass, no gradient training and no QAT. It keeps a ternary base for every weight and adds a small, hardware-friendly sparse residual only where ternarization hurts most. Among the sub-2-bit PTQ methods we evaluated (PT²-LLM, PB-LLM, Slim-LLM, GPTQ), it gives the best accuracy.

How is QTEA different from BitNet b1.58? BitNet trains ternary models from scratch. QTEA ternarizes an existing pretrained checkpoint after training.

How is it different from GPTQ? QTEA builds on the GPTQ column-by-column sweep and adds three things: a ternary quantizer with per-column rescale factors optimized jointly with the ternary assignments, a salient-column sparse residual, and an error-decay term that stops late columns from being over-compensated.

Does it actually run faster? Yes. Five ternary values are packed per byte and evaluated with a lookup-table CUDA GEMV kernel: 7.2× faster per-token generation than FP16 on Llama2-70B with CUDA Graphs, and 3.62× on Qwen3-14B.

Which models are supported? Llama-3 and Qwen3 (0.6B–14B) out of the box, plus any decoder that exposes model.model.layers with the standard seven projections. A pre-quantized Qwen3-14B-Base checkpoint is on Hugging Face.

Repository Contents

Quantization

  • quantize.py: command-line entry point — quantizes one model and saves the packed checkpoint.
  • qtea/qtea.py: the QTEA algorithm for a single linear layer. Sweeps the columns left to right, picking the salient columns, refining their rescale factors, fitting the sparse residual and propagating the quantization error.
  • qtea/quantizer.py: fits the ternary scale and centre that every group of 128 columns shares.
  • qtea/sequential.py: applies the quantizer to the whole decoder stack, one block at a time, so only one block ever sits on the GPU.
  • qtea/pack.py: the checkpoint format — packs and unpacks the ternary codes (five per byte), the FP8 residuals and the scales.
  • qtea/data.py: WikiText-2 and C4 loaders for calibration and perplexity.
  • qtea/model.py: HuggingFace model loading.

Evaluation

  • eval/evaluate.py: measures perplexity and zero-shot accuracy for a packed checkpoint, or for the FP16 baseline.
  • eval/perplexity.py: perplexity, computed one decoder block at a time to keep memory low.
  • eval/packed_linear.py: loads a packed checkpoint into a HuggingFace model.
  • eval/ternary_kernel.py and eval/csrc/ternary_gemv.cu: the lookup-table GEMV kernel — the Python wrapper and the CUDA source behind the paper's latency numbers.
  • eval/check_kernel.py: checks the kernel against a dense FP16 matmul.

Other

  • scripts/: one reproduction script per model family.
  • tests/: round-trip tests for the checkpoint format (python -m pytest tests).
  • figs/: figures used by this README and by eval/README.md.

Citation

QTEA appears at the EMNLP 2026 Main Conference. If you find it useful in your research, please cite:

@misc{guo2026qteaternaryllmssparse,
      title={QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization}, 
      author={Yipin Guo and Arun M George and Jie Fu and Tareq Mahmoud and Sixue Xing and Siddharth Joshi},
      year={2026},
      eprint={2609.00224},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2609.00224}, 
}

License

This project is released under the MIT License. See LICENSE for details. The evaluation datasets and the models retain their own licences.