LT—OPDRESEARCH PROJECT
arXiv 2026 / Visual token efficiency

Fewer Tokens,
More Self-Teaching.

On-Policy Self-Distillation for
Extreme Visual Token Reduction

Shanghai Jiao Tong University · University of Cambridge
Xidian University · Xi’an Jiaotong University

View on Hugging Face Daily Papers ↗
5%

visual tokens retained

82.3%

full-model performance retained

85.4%

less prefill FLOPs

Paper results · Qwen3.5-4B · averaged over nine benchmarks

RELEASE NOTE

October 5, 2026 — New checkpoint with learnable token merging: 84.5% retained performance, as reported in the repository. Separate from the paper’s 82.3% result.

Explore release ↗
01 / THE IDEA

Teach the model
where it actually goes.

Extreme compression changes both what a model sees and the responses it produces.

LT-OPD trains a low-token student on its own generated prefixes, supervised by a frozen full-token copy. A decreasing token-budget curriculum supports adaptation. At inference, the compressed student runs alone.

Read the full abstract and paper ↗
02 / METHOD

Same model. Different visual budgets.

Low-Token On-Policy Distillation

Image + questionFrozen full-token teacher100% visual tokens · no updatesLow-token studentCurriculum: 25% → 5% tokensStudent rolloutsSelf-generated prefixesSame prefixesToken-level JSDUpdate the student
LT-OPD framework · Redrawn from the paper and code repository.
01

Generate

The compressed student produces a response using reduced visual inputs.

02

Self-teach

A frozen teacher supplies next-token distributions on those same prefixes.

03

Adapt

Token-level Jensen–Shannon divergence aligns the student with its teacher.

A curriculum for compression

175 updates · 25% → 5% visual tokens

25%Warm start / 14 updates
25% ↘ 5%Cosine decay / to update 100
5%Target budget / final 75 updates
03 / RESULTS

Recover capability.
Keep the efficiency.

Qwen3.5-4B · 5% visual-token budget

Average retained performance

CDPruner
68.6%
EPIC
73.7%
DivPrune
77.6%
LT-OPD
82.3%

100% denotes the full-token Qwen3.5-4B reference.

See all nine benchmark scores +
Paper Table 2 · Qwen3.5-4B
MethodV*BenchHR-4KGQAMMMUMMBenchMMEPOPETextVQAOCRBench
Full tokens84.384.460.777.686.42300.889.381.2850
CDPruner · 5%52.968.450.445.774.01531.870.243.2404
LT-OPD · 5%74.973.054.850.078.92077.086.356.0535

Retained performance = mean of each benchmark score divided by its full-token reference score. It is not an overall accuracy. Source: paper, Tables 2–5 ↗

Evaluated across model scales and familiesQwen3.5-9BGLM-4.6V-9BLLaVA-OV-1.5-4B
04 / OPEN RESOURCES

From paper to practice.

Code, weights, and training data

Get started

Python 3.12 + CUDA. Follow the model card for environment setup and compressed inference.

Full model quick start ↗
Evaluation guide ↗
git clone https://github.com/Yrxxxxxxxx1007/LT-OPD.git
cd LT-OPD

# Download the current checkpoint
hf download yyy051007/LT-OPD --local-dir models/LT-OPD

# Prepare the training data after installation
lt-opd data --dataset yyy051007/LT-OPD-14K \
  --output data/LT-OPD-14K
05 / CITATION

BibTeX

@article{li2026fewer,
  title={Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation
         for Extreme Visual Token Reduction},
  author={Li, Junxian and Yang, Ruixuan and Zhang, Tianao and
          Xu, Tiange and Dong, Weisheng and Zhang, Yulun},
  journal={arXiv preprint arXiv:2609.32353},
  year={2026}
}