Skip to main content
Back to Newswire
AI Science

Preprint describes VLM inference method with reported 3.1x faster first token

An arXiv preprint describes PACE, a training-free framework that combines input downsampling before vision encoding with selective visual-token retention after encoding. Integrated into Qwen2.5-VL-7B, its authors report retaining 93.8% of the model's original performance while using 10% of visual tokens and achieving a 3.1x time-to-first-token speedup. The paper says code is available. These are author-reported experimental results, not an announced product release.
Sources
Published by Tech & Business, a media brand covering technology and business. This story was sourced from cs.AI updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire