measuredlocally.com
Local AI Last updated 11 October 2026

Local AI, timed on machines I own

I run AI models at home on hardware I bought myself, and I publish what I measure: which model, on which machine, how fast, and how it was timed.

The benchmarks so far

MachineModelWriting codeReading (prefill)Tested
Four DGX Sparksgraphics chips capped at 2,200MHz to cut power useGLM 5.3 FlashNVFP4 at 4 bits, dense layers at 8104.1 tok/s3,159 tok/s11 Oct 2026
Two DGX Sparksgraphics chips capped at 2,200MHz to cut power useQwen 3.8 27BNVFP4 at 4 bits44.7 tok/s2,034 tok/s11 Oct 2026
One DGX Sparkgraphics chips capped at 2,200MHz to cut power useQwen 3.8 27BNVFP4 at 4 bits26.7 tok/s1,188 tok/s11 Oct 2026
RTX 4090 desktopQwen 3.8 27BQ4_K_M, GGUF62.4 tok/s1,994 tok/s11 Oct 2026
HP ZBook Ultra G1aStrix Halo laptop, Ryzen AI Max+ 395, 64GBQwen 3.8 27BQ4_K_M, GGUF23.5 tok/s194 tok/s11 Oct 2026

One user, the model's thinking switched off, the same prompts and the same script on each machine. My Sparks would run a little faster than the figures on this page. I cap their graphics chips at 2,200MHz, against about 2,400 uncapped, to cut their power use. When I tested the cap on one box in August, writing speed stayed within 1 per cent, reading was 2 to 4 per cent slower, and the chip's peak power fell from 55.5W to 43.3W. Every figure, the runs with thinking on, the hardware and how I test.

The weekly video

Local AI, This Week is a short video each week on what changed for people who run AI models at home. These are the latest episodes, and all of them are on the YouTube channel.

What is coming to this site