measuredlocally.com
Benchmarks Last updated 11 October 2026

Local AI benchmarks: what runs on which machine, and how fast

AI models timed on the machines I own, with the same prompts and the same script on each. Speeds are tokens a second for one user unless it says otherwise.

What I have found so far

The machines

Four DGX Sparks4 x 128GB, joined at 200Gb/s. Graphics chips capped at 2,200MHz to cut power use.6 results below
Two DGX Sparks2 x 128GB, joined at 200Gb/s. Graphics chips capped at 2,200MHz to cut power use.3 results below
One DGX Spark128GB, one box on its own. Graphics chips capped at 2,200MHz to cut power use.1 result below
RTX 4090 desktop24GB of graphics memory, Ryzen 9 7950X, 96GB of memory1 result below
HP ZBook Ultra G1aA Strix Halo laptop: Ryzen AI Max+ 395 with Radeon 8060S graphics, and 64GB of memory shared between them1 result below

The same test on each machine

MachineModelWriting codeWriting proseFirst wordReading (prefill)Tested
Four DGX SparksvLLM across all four boxes, version 0.1.dev20051graphics chips capped at 2,200MHz to cut power use4 at once: 112.3 tok/s in totalGLM 5.3 FlashNVFP4 at 4 bits, dense layers at 8104.1 tok/s100.0 to 104.2 over 3 runs62.7 tok/s62.6 to 63.4 over 3 runs0.15 s3,159 tok/sabout 16,000 tokens11 Oct 2026
Two DGX SparksvLLM across two boxes, version 0.1.dev20003graphics chips capped at 2,200MHz to cut power use4 at once: 96.4 tok/s in totalQwen 3.8 27BNVFP4 at 4 bits44.7 tok/s43.5 to 45.5 over 3 runs28.7 tok/s26.1 to 29.0 over 3 runs0.17 s2,034 tok/sabout 19,000 tokens11 Oct 2026
One DGX SparkvLLM on one box, version 0.1.dev20003graphics chips capped at 2,200MHz to cut power use4 at once: 60.9 tok/s in totalQwen 3.8 27BNVFP4 at 4 bits26.7 tok/s25.9 to 27.1 over 3 runs16.9 tok/s16.8 to 17.3 over 3 runs0.16 s1,188 tok/sabout 19,000 tokens11 Oct 2026
RTX 4090 desktopLM Studio, llama.cpp on CUDA, version 2.55.0Qwen 3.8 27BQ4_K_M, GGUF62.4 tok/s61.3 to 62.5 over 3 runs45.0 tok/s44.5 to 46.5 over 3 runs0.63 s1,994 tok/sabout 19,000 tokens11 Oct 2026
HP ZBook Ultra G1aStrix Halo laptop, Ryzen AI Max+ 395, 64GBLM Studio, llama.cpp on Vulkan, version 2.55.0on mains power, Windows power mode BalancedQwen 3.8 27BQ4_K_M, GGUF23.5 tok/s22.7 to 23.7 over 3 runs14.6 tok/s14.6 to 15.4 over 3 runs1.30 s194 tok/sabout 19,000 tokens11 Oct 2026

Each figure is the middle of three runs with the model's thinking switched off, and the range is under it. First word is the wait after a short question. Reading is how fast the machine takes in a long document before it starts to answer, which other sites call prefill or prompt processing. The LM Studio rows were run with LM Studio's defaults, which include the model's own multi-token prediction. My Sparks would run a little faster than the figures on this page. I cap their graphics chips at 2,200MHz, against about 2,400 uncapped, to cut their power use. When I tested the cap on one box in August, writing speed stayed within 1 per cent, reading was 2 to 4 per cent slower, and the chip's peak power fell from 55.5W to 43.3W.

The same test with thinking switched on

MachineModelCode promptProse promptWait before the codeWait before the proseTested
Four DGX SparksvLLM across all four boxes, version 0.1.dev20051graphics chips capped at 2,200MHz to cut power useGLM 5.3 FlashNVFP4 at 4 bits, dense layers at 8106.1 tok/s106.0 to 111.1 over 3 runs61.8 tok/s61.7 to 63.0 over 3 runs0.5 s0.4 to 0.8 s over 3 runsabout 40 tokens of thinking1.3 s1.3 to 1.8 s over 3 runsabout 80 tokens of thinking11 Oct 2026
Two DGX SparksvLLM across two boxes, version 0.1.dev20003graphics chips capped at 2,200MHz to cut power useQwen 3.8 27BNVFP4 at 4 bits40.9 tok/s38.5 to 41.5 over 3 runs36.7 tok/s35.6 to 38.5 over 3 runs14.2 sone run onlyabout 540 tokens of thinkingover 55 sstill thinking at 2,000 tokens11 Oct 2026
One DGX SparkvLLM on one box, version 0.1.dev20003graphics chips capped at 2,200MHz to cut power useQwen 3.8 27BNVFP4 at 4 bits25.2 tok/s25.1 to 26.3 over 3 runs23.3 tok/s21.6 to 23.5 over 3 runs17.3 sone run onlyabout 410 tokens of thinkingover 86 sstill thinking at 2,000 tokens11 Oct 2026
RTX 4090 desktopLM Studio, llama.cpp on CUDA, version 2.55.0Qwen 3.8 27BQ4_K_M, GGUF58.5 tok/s57.2 to 59.6 over 3 runs54.7 tok/s54.5 to 55.7 over 3 runs6.7 s6.5 to 7.8 s over 3 runsabout 350 tokens of thinkingover 37 sstill thinking at 2,000 tokens11 Oct 2026
HP ZBook Ultra G1aStrix Halo laptop, Ryzen AI Max+ 395, 64GBLM Studio, llama.cpp on Vulkan, version 2.55.0on mains power, Windows power mode BalancedQwen 3.8 27BQ4_K_M, GGUF21.1 tok/s20.5 to 21.3 over 3 runs19.5 tok/s19.1 to 19.8 over 2 runs22.2 s19.1 to 34.8 s over 3 runsabout 410 tokens of thinkingover 104 sstill thinking at 2,000 tokens11 Oct 2026

Thinking is the model working through the problem before it answers. For GLM 5.3 Flash, thinking on means high, the middle of the model's three thinking levels, on every machine. For Qwen 3.8 27B, thinking on means medium, the middle of the model's three thinking levels, on every machine. The speeds in this table count the tokens it thinks with and the tokens of the answer. The wait is from sending the prompt to the first word of the answer, and it is the middle of three runs, because the model thinks for a different length each time. The cap of 2,000 tokens covers the thinking and the answer together. A wait marked "one run only" was timed in a single separate run with room for 8,000 tokens, because the three runs of that row were made before the script timed the answer. Where a wait begins with "over", the model was still thinking when the 2,000 tokens ran out. The speed beside it is then the speed of its thinking, and the real wait is longer than the one shown.

The models, exactly

The hardware, exactly

GLM 5.3 Flash on four DGX Sparks, before and after a recipe change

Test5 October
new recipe
2 October
old recipe
Change
Writing codeone user94.2 tok/s58.4 tok/s1.6 times
A long answerone user89.8 tok/s45.6 tok/s2.0 times
A short answerone user60.4 tok/s28.2 tok/s2.1 times
Four at onceall four added together110.7 tok/s65.2 tok/s1.7 times

The same script on both days. The only change was a recipe another owner published. The first word of an answer arrived after 0.26 seconds in a session recorded on 5 October. The change column is my own maths.

Three models, the same eight prompts

ModelRuns onWriting speed
middle of 8 prompts
First word
short prompt
Reading a 345,000-token
document from cold
GLM 5.3 Flash4 DGX Sparks46.3 tok/s0.19 s257 s
Qwen 3.8 27B2 DGX Sparks42.8 tok/s0.16 s810 s
DeepSeek V4 Flash2 DGX Sparks40.9 tok/s0.29 snot measured

Measured in August 2026 with each model thinking at its highest setting, on different prompts from the test at the top of this page. These were also taken before October's recipe change, so they should not be set beside the other tables.

How I test