Pixels Into Tokens
Know what an image actually costs and what the model sees after your file is resized.
$87
this phase
- Open1.1
What The Model Sees
Free previewVision encoder, MLP merger, and why a 4000px screenshot is downsampled before anything reads it.
- 1.2
Counting Visual Tokens
Gemini's 258 tokens per PDF page against Claude's 28x28 patch formula, priced over a thousand pages.
- 1.3
Native Resolution & M-RoPE
Dynamic resolution, NaFlex, and the 2x2 patch merge that turns a 224x224 image into 66 tokens.
- 1.4
Current, Legacy, Retired
Pixtral, Llama 3.2 Vision and Qwen2.5-VL are legacy, and video support rarely means a video file.
- 1.5
Benchmarks After Saturation
MMMU published its test answers in February 2026, DocVQA sits at ceiling, and what to read instead.