Skip to content

PUBLISHED EVALUATIONS

FastVLM benchmarks

Compare the three official model sizes across eleven published vision-language evaluations. Select a model to inspect its score profile.

Select model

Benchmark score profile

FastVLM-0.5BHigher is better
AI2D
68.0
ScienceQA
85.2
MMMU
33.9
VQAv2
76.3
ChartQA
76.0
TextVQA
64.5
InfoVQA
46.4
DocVQA
82.5
OCRBench
63.9
RealWorldQA
56.1
SeedBench-Img
71.0

Full comparison table

Benchmark0.5B1.5B7B
AI2D68.077.483.6
ScienceQA85.294.496.7
MMMU33.937.845.4
VQAv276.379.180.8
ChartQA76.080.185.0
TextVQA64.570.474.9
InfoVQA46.459.775.8
DocVQA82.588.393.2
OCRBench63.970.273.1
RealWorldQA56.161.267.2
SeedBench-Img71.074.275.4

How to read these numbers

These are evaluation scores published on Apple’s FastVLM Hugging Face model cards. They do not include a single universal latency value because runtime depends on resolution, precision, hardware and implementation. Compare scores within a benchmark row; do not compare the numeric scale of different benchmarks as if they were identical.