Skip to content

BROWSER AI BENCHMARK · FASTVLM-0.5B

Benchmark your device

Find out how fast a real vision-language model runs on your hardware. The test downloads FastVLM-0.5B once, runs a fixed image and question twice, and reports the warm run: load time, first token and decode speed. Nothing leaves your device except the numbers you choose to share.

Your device

Browser

Operating system

GPU adapter

Checking…

WebGPU

Checking…

shader-f16

Checking…

Reported memory

Method: FastVLM-0.5B ONNX (q4, fp16 embeddings) through Transformers.js and WebGPU; fixed 1024 px receipt sample and prompt; one warm-up run, one measured run; tokens/s = generated tokens ÷ (image→answer − first token). Numbers depend on GPU, driver, browser and other open tabs.

The image and answer stay on your device. Usage events record step names and timings; GPU name and share text are only included in the link you copy.

What to do with the number

  • Under 10 tokens/s: use FastVLM through Python, MLX or Core ML instead of the browser, or pick a smaller model such as SmolVLM-256M.
  • 10–60 tokens/s: fine for single-image questions; expect a few seconds per answer.
  • Over 60 tokens/s: interactive use is comfortable; try the 1.5B model natively for better accuracy. A recent Apple Silicon Mac measures well above 100.