PUBLISHED EVALUATIONS
FastVLM benchmarks
Compare the three official model sizes across eleven published vision-language evaluations. Select a model to inspect its score profile.
Select model
Benchmark score profile
FastVLM-0.5BHigher is better
AI2D
68.0
ScienceQA
85.2
MMMU
33.9
VQAv2
76.3
ChartQA
76.0
TextVQA
64.5
InfoVQA
46.4
DocVQA
82.5
OCRBench
63.9
RealWorldQA
56.1
SeedBench-Img
71.0
Full comparison table
| Benchmark | 0.5B | 1.5B | 7B |
|---|---|---|---|
| AI2D | 68.0 | 77.4 | 83.6 |
| ScienceQA | 85.2 | 94.4 | 96.7 |
| MMMU | 33.9 | 37.8 | 45.4 |
| VQAv2 | 76.3 | 79.1 | 80.8 |
| ChartQA | 76.0 | 80.1 | 85.0 |
| TextVQA | 64.5 | 70.4 | 74.9 |
| InfoVQA | 46.4 | 59.7 | 75.8 |
| DocVQA | 82.5 | 88.3 | 93.2 |
| OCRBench | 63.9 | 70.2 | 73.1 |
| RealWorldQA | 56.1 | 61.2 | 67.2 |
| SeedBench-Img | 71.0 | 74.2 | 75.4 |
How to read these numbers
These are evaluation scores published on Apple’s FastVLM Hugging Face model cards. They do not include a single universal latency value because runtime depends on resolution, precision, hardware and implementation. Compare scores within a benchmark row; do not compare the numeric scale of different benchmarks as if they were identical.