Generate
The compressed student produces a response using reduced visual inputs.
On-Policy Self-Distillation for
Extreme Visual Token Reduction
Shanghai Jiao Tong University · University of Cambridge
Xidian University · Xi’an Jiaotong University
visual tokens retained
full-model performance retained
less prefill FLOPs
Paper results · Qwen3.5-4B · averaged over nine benchmarks
October 5, 2026 — New checkpoint with learnable token merging: 84.5% retained performance, as reported in the repository. Separate from the paper’s 82.3% result.
Explore release ↗Extreme compression changes both what a model sees and the responses it produces.
LT-OPD trains a low-token student on its own generated prefixes, supervised by a frozen full-token copy. A decreasing token-budget curriculum supports adaptation. At inference, the compressed student runs alone.
Read the full abstract and paper ↗Low-Token On-Policy Distillation
The compressed student produces a response using reduced visual inputs.
A frozen teacher supplies next-token distributions on those same prefixes.
Token-level Jensen–Shannon divergence aligns the student with its teacher.
175 updates · 25% → 5% visual tokens
Qwen3.5-4B · 5% visual-token budget
100% denotes the full-token Qwen3.5-4B reference.
| Method | V*Bench | HR-4K | GQA | MMMU | MMBench | MME | POPE | TextVQA | OCRBench |
|---|---|---|---|---|---|---|---|---|---|
| Full tokens | 84.3 | 84.4 | 60.7 | 77.6 | 86.4 | 2300.8 | 89.3 | 81.2 | 850 |
| CDPruner · 5% | 52.9 | 68.4 | 50.4 | 45.7 | 74.0 | 1531.8 | 70.2 | 43.2 | 404 |
| LT-OPD · 5% | 74.9 | 73.0 | 54.8 | 50.0 | 78.9 | 2077.0 | 86.3 | 56.0 | 535 |
Retained performance = mean of each benchmark score divided by its full-token reference score. It is not an overall accuracy. Source: paper, Tables 2–5 ↗
Code, weights, and training data
Training, compression runtime, checkpoint export, and benchmark evaluation.
Apache-2.002 / HUGGING FACE ↗Current Qwen3.5-4B release with CDPruner and learnable token merging.
Use the project’s compression runtime03 / HUGGING FACE ↗14,000 training examples with packaged image shards, provenance, and checksums.
Browse the dataset and release filesPython 3.12 + CUDA. Follow the model card for environment setup and compressed inference.
Full model quick start ↗git clone https://github.com/Yrxxxxxxxx1007/LT-OPD.git cd LT-OPD # Download the current checkpoint hf download yyy051007/LT-OPD --local-dir models/LT-OPD # Prepare the training data after installation lt-opd data --dataset yyy051007/LT-OPD-14K \ --output data/LT-OPD-14K
@article{li2026fewer,
title={Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation
for Extreme Visual Token Reduction},
author={Li, Junxian and Yang, Ruixuan and Zhang, Tianao and
Xu, Tiange and Dong, Weisheng and Zhang, Yulun},
journal={arXiv preprint arXiv:2609.32353},
year={2026}
}