Arxiv 2509 26625 Learning To See Before Seeing Demystifying Llm Visual Priors
ASecurity
LLMs develop rich visual priors despite text-only training. We reveal that visual priors are composed of separable perception and reasoning priors with unique scaling trends and origins. Visual reasoning ability is predominantly developed by pre-training on reasoning-centric data (code, math, academia). We propose a data-centric recipe for pre-training vision-aware LLMs verified in 1T token scale pre-training across 100+ controlled experiments consuming 500,000 GPU-hours.
Installs into .claude/skills of the current project.
Are you the author of Arxiv 2509 26625 Learning To See Before Seeing Demystifying Llm Visual Priors?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-arxiv-2509-26625-learning-to-see-before-seeing-dem)
---
title: "Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training"
authors: "Junlin Han, Shengbang Tong, David Fan, Yufan Ren, Koustuv Sinha, Philip Torr, Filippos Kokkinos"
arxiv_id: "2509.26625"
categories: "cs.LG; cs.AI; cs.CV; cs.MM"
utility: 0.9
date_added: "2026-09-29"
category: "vision-generative"
---
# Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
## Abstract
LLMs develop rich visual priors despite text-only training. We reveal that visual priors are composed of separable perception and reasoning priors with unique scaling trends and origins. Visual reasoning ability is predominantly developed by pre-training on reasoning-centric data (code, math, academia). We propose a data-centric recipe for pre-training vision-aware LLMs verified in 1T token scale pre-training across 100+ controlled experiments consuming 500,000 GPU-hours.
## Key Contributions
- Novel approach in vision generative domain
- Utility score: 0.9
- Published on arXiv: 2509.26625
## Potential Applications
- Research reference for vision generative
- Building block for related systems