Local AI benchmarks: what runs on which machine, and how fast
AI models timed on the machines I own, with the same prompts and the same script on each. Speeds are tokens a second for one user unless it says otherwise.
What I have found so far
- My Sparks would run a little faster than the figures on this page. I cap their graphics chips at 2,200MHz, against about 2,400 uncapped, to cut their power use. When I tested the cap on one box in August, writing speed stayed within 1 per cent, reading was 2 to 4 per cent slower, and the chip's peak power fell from 55.5W to 43.3W.
- The four DGX Sparks wrote code at 104.1 tokens a second with GLM 5.3 Flash (NVFP4 at 4 bits, dense layers at 8), and read a document of about 16,000 tokens at 3,159 tokens a second. That is a wait of about 5 seconds before the first word (the size of the document divided by the speed).
- With GLM 5.3 Flash thinking before it answers, the first word of the code answer came after about 0.5 seconds on the four Sparks. With thinking off it came after 0.15 seconds.
- With Qwen 3.8 27B on all of them (each machine in the 4-bit format its own software uses), the RTX 4090 desktop wrote code at 62.4 tokens a second, two of my DGX Sparks at 44.7, one DGX Spark on its own at 26.7 and the HP ZBook Ultra G1a (a Strix Halo laptop with the Ryzen AI Max+ 395) at 23.5.
- Two of my DGX Sparks read a document of about 19,000 tokens at 2,034 tokens a second, the RTX 4090 desktop at 1,994, one DGX Spark on its own at 1,188 and the HP ZBook Ultra G1a (a Strix Halo laptop with the Ryzen AI Max+ 395) at 194. That is a wait of about 9 seconds on the two Sparks, 9 seconds on the desktop, 16 seconds on the single Spark and 97 seconds on the laptop before the first word (the size of the document divided by each speed).
- With Qwen 3.8 27B thinking before it answers, the first word of the code answer came after about 7 seconds on the desktop, 14 seconds on the two Sparks (one run), 17 seconds on the single Spark (one run) and 22 seconds on the laptop. With thinking off it came after 1.3 seconds at most.
- On the prose prompt with thinking on, Qwen 3.8 27B was still thinking when the 2,000 tokens ran out, which took 37 seconds on the desktop, 55 seconds on the two Sparks, 86 seconds on the single Spark and 104 seconds on the laptop.
- Given room for 8,000 tokens, the prose answer began after 52 seconds in one run on the two Sparks and after 271 seconds in one run on the single Spark. The thinking before it ran to about 1,800 and 7,400 tokens, so most of that gap is the model thinking for longer.
- A recipe change alone doubled GLM 5.3 Flash on four DGX Sparks. A short answer went from 28.2 to 60.4 tokens a second between 2 and 5 October, on the same boxes with the same script.
- On the same eight prompts in August, three models came within 6 tokens a second of each other (40.9 to 46.3), whether they ran on two Sparks or four.
The machines
The same test on each machine
| Machine | Model | Writing code | Writing prose | First word | Reading (prefill) | Tested |
|---|---|---|---|---|---|---|
| Four DGX SparksvLLM across all four boxes, version 0.1.dev20051graphics chips capped at 2,200MHz to cut power use4 at once: 112.3 tok/s in total | GLM 5.3 FlashNVFP4 at 4 bits, dense layers at 8 | 104.1 tok/s100.0 to 104.2 over 3 runs | 62.7 tok/s62.6 to 63.4 over 3 runs | 0.15 s | 3,159 tok/sabout 16,000 tokens | 11 Oct 2026 |
| Two DGX SparksvLLM across two boxes, version 0.1.dev20003graphics chips capped at 2,200MHz to cut power use4 at once: 96.4 tok/s in total | Qwen 3.8 27BNVFP4 at 4 bits | 44.7 tok/s43.5 to 45.5 over 3 runs | 28.7 tok/s26.1 to 29.0 over 3 runs | 0.17 s | 2,034 tok/sabout 19,000 tokens | 11 Oct 2026 |
| One DGX SparkvLLM on one box, version 0.1.dev20003graphics chips capped at 2,200MHz to cut power use4 at once: 60.9 tok/s in total | Qwen 3.8 27BNVFP4 at 4 bits | 26.7 tok/s25.9 to 27.1 over 3 runs | 16.9 tok/s16.8 to 17.3 over 3 runs | 0.16 s | 1,188 tok/sabout 19,000 tokens | 11 Oct 2026 |
| RTX 4090 desktopLM Studio, llama.cpp on CUDA, version 2.55.0 | Qwen 3.8 27BQ4_K_M, GGUF | 62.4 tok/s61.3 to 62.5 over 3 runs | 45.0 tok/s44.5 to 46.5 over 3 runs | 0.63 s | 1,994 tok/sabout 19,000 tokens | 11 Oct 2026 |
| HP ZBook Ultra G1aStrix Halo laptop, Ryzen AI Max+ 395, 64GBLM Studio, llama.cpp on Vulkan, version 2.55.0on mains power, Windows power mode Balanced | Qwen 3.8 27BQ4_K_M, GGUF | 23.5 tok/s22.7 to 23.7 over 3 runs | 14.6 tok/s14.6 to 15.4 over 3 runs | 1.30 s | 194 tok/sabout 19,000 tokens | 11 Oct 2026 |
Each figure is the middle of three runs with the model's thinking switched off, and the range is under it. First word is the wait after a short question. Reading is how fast the machine takes in a long document before it starts to answer, which other sites call prefill or prompt processing. The LM Studio rows were run with LM Studio's defaults, which include the model's own multi-token prediction. My Sparks would run a little faster than the figures on this page. I cap their graphics chips at 2,200MHz, against about 2,400 uncapped, to cut their power use. When I tested the cap on one box in August, writing speed stayed within 1 per cent, reading was 2 to 4 per cent slower, and the chip's peak power fell from 55.5W to 43.3W.
The same test with thinking switched on
| Machine | Model | Code prompt | Prose prompt | Wait before the code | Wait before the prose | Tested |
|---|---|---|---|---|---|---|
| Four DGX SparksvLLM across all four boxes, version 0.1.dev20051graphics chips capped at 2,200MHz to cut power use | GLM 5.3 FlashNVFP4 at 4 bits, dense layers at 8 | 106.1 tok/s106.0 to 111.1 over 3 runs | 61.8 tok/s61.7 to 63.0 over 3 runs | 0.5 s0.4 to 0.8 s over 3 runsabout 40 tokens of thinking | 1.3 s1.3 to 1.8 s over 3 runsabout 80 tokens of thinking | 11 Oct 2026 |
| Two DGX SparksvLLM across two boxes, version 0.1.dev20003graphics chips capped at 2,200MHz to cut power use | Qwen 3.8 27BNVFP4 at 4 bits | 40.9 tok/s38.5 to 41.5 over 3 runs | 36.7 tok/s35.6 to 38.5 over 3 runs | 14.2 sone run onlyabout 540 tokens of thinking | over 55 sstill thinking at 2,000 tokens | 11 Oct 2026 |
| One DGX SparkvLLM on one box, version 0.1.dev20003graphics chips capped at 2,200MHz to cut power use | Qwen 3.8 27BNVFP4 at 4 bits | 25.2 tok/s25.1 to 26.3 over 3 runs | 23.3 tok/s21.6 to 23.5 over 3 runs | 17.3 sone run onlyabout 410 tokens of thinking | over 86 sstill thinking at 2,000 tokens | 11 Oct 2026 |
| RTX 4090 desktopLM Studio, llama.cpp on CUDA, version 2.55.0 | Qwen 3.8 27BQ4_K_M, GGUF | 58.5 tok/s57.2 to 59.6 over 3 runs | 54.7 tok/s54.5 to 55.7 over 3 runs | 6.7 s6.5 to 7.8 s over 3 runsabout 350 tokens of thinking | over 37 sstill thinking at 2,000 tokens | 11 Oct 2026 |
| HP ZBook Ultra G1aStrix Halo laptop, Ryzen AI Max+ 395, 64GBLM Studio, llama.cpp on Vulkan, version 2.55.0on mains power, Windows power mode Balanced | Qwen 3.8 27BQ4_K_M, GGUF | 21.1 tok/s20.5 to 21.3 over 3 runs | 19.5 tok/s19.1 to 19.8 over 2 runs | 22.2 s19.1 to 34.8 s over 3 runsabout 410 tokens of thinking | over 104 sstill thinking at 2,000 tokens | 11 Oct 2026 |
Thinking is the model working through the problem before it answers. For GLM 5.3 Flash, thinking on means high, the middle of the model's three thinking levels, on every machine. For Qwen 3.8 27B, thinking on means medium, the middle of the model's three thinking levels, on every machine. The speeds in this table count the tokens it thinks with and the tokens of the answer. The wait is from sending the prompt to the first word of the answer, and it is the middle of three runs, because the model thinks for a different length each time. The cap of 2,000 tokens covers the thinking and the answer together. A wait marked "one run only" was timed in a single separate run with room for 8,000 tokens, because the three runs of that row were made before the script timed the answer. Where a wait begins with "over", the model was still thinking when the 2,000 tokens ran out. The speed beside it is then the speed of its thinking, and the real wait is longer than the one shown.
The models, exactly
- GLM 5.3 Flash on the four DGX Sparks. NVIDIA's own NVFP4 export of GLM 5.3 Flash (4-bit, 204.5GB on disk), with the dense layers re-encoded at 8 bits, which the recipe calls lossless8. A small draft model, DFlash2 converted to FP8, guesses several tokens ahead. It runs on vLLM split across all four boxes (tensor parallel 4) under another owner's published recipe, with a window of 1,048,576 tokens. The model is about 300 billion parameters with about 15 billion active for each token, which is my estimate from its config file. See NVIDIA's weights, the draft model and the recipe.
- Qwen 3.8 27B on two of my DGX Sparks. Unsloth's NVFP4 build of Qwen 3.8 27B (4-bit, 23.4GB on disk), on vLLM split across two boxes through the 200Gb/s switch (tensor parallel 2). The model's own multi-token prediction was on, guessing 3 tokens ahead. The window was 351,232 tokens, the engine kept its working memory of the conversation at 8 bits (FP8), and sampling was the model's own default, a temperature of 1.0. See the weights.
- Qwen 3.8 27B on one DGX Spark on its own. The same weights, engine and settings as the two-box row, on one box by itself (tensor parallel 1). See the weights.
- Qwen 3.8 27B on the RTX 4090 desktop. The Q4_K_M GGUF file of Qwen 3.8 27B from LM Studio's own catalogue (4-bit, 16.8GB), loaded whole onto the graphics card with a window of 20,480 tokens. LM Studio's settings were left as they come, so the model's own multi-token prediction was on. In this run it kept 78 per cent of its guesses. See the file.
- Qwen 3.8 27B on the HP ZBook Ultra G1a (a Strix Halo laptop with the Ryzen AI Max+ 395). The same Q4_K_M GGUF file as the RTX 4090 ran (4-bit, 16.8GB), loaded whole onto the graphics chip with a window of 20,480 tokens. LM Studio's settings were left as they come, so the model's own multi-token prediction was on. In this run it kept 69 per cent of its guesses. See the file.
The hardware, exactly
- The DGX Sparks (the same boxes, used one, two or four at a time). The boxes: four NVIDIA DGX Sparks in two makes, GIGABYTE AI TOP ATOM and ASUS GX10, which behave the same here. Chip: an NVIDIA GB10 in each box, which is a Blackwell graphics processor and a 20-core Arm processor (10 cores at about 3.9GHz and 10 at about 2.8GHz) on one package. Memory: 128GB in each box, shared between the processor and the graphics, with a bandwidth of 273GB a second. Network: two ConnectX-7 ports of 200Gb/s on each box, joined through a MikroTik CRS812 switch. System: NVIDIA DGX OS. Standing setting: graphics chips capped at 2,200MHz to cut power use.
- RTX 4090 desktop. Graphics card: NVIDIA GeForce RTX 4090 with 24GB of its own memory. Processor: AMD Ryzen 9 7950X, 16 cores and 32 threads. Memory: 96GB of DDR5 as two 48GB modules running at 6,000MT/s (Corsair CMK96GX5M2B6000Z30). Motherboard: ASRock X670E Taichi. System: Windows 11 Pro.
- HP ZBook Ultra G1a. Machine: a 14 inch laptop, timed on mains power. Processor: AMD Ryzen AI Max+ 395 (Strix Halo), 16 cores and 32 threads. Graphics: Radeon 8060S, built into the processor. Memory: 64GB of LPDDR5X soldered to the board, running at 8,000MT/s. Windows sees 47.8GB of it, because 16GB is set aside for the graphics. System: Windows 11 Pro.
GLM 5.3 Flash on four DGX Sparks, before and after a recipe change
| Test | 5 October new recipe | 2 October old recipe | Change |
|---|---|---|---|
| Writing codeone user | 94.2 tok/s | 58.4 tok/s | 1.6 times |
| A long answerone user | 89.8 tok/s | 45.6 tok/s | 2.0 times |
| A short answerone user | 60.4 tok/s | 28.2 tok/s | 2.1 times |
| Four at onceall four added together | 110.7 tok/s | 65.2 tok/s | 1.7 times |
The same script on both days. The only change was a recipe another owner published. The first word of an answer arrived after 0.26 seconds in a session recorded on 5 October. The change column is my own maths.
Three models, the same eight prompts
| Model | Runs on | Writing speed middle of 8 prompts | First word short prompt | Reading a 345,000-token document from cold |
|---|---|---|---|---|
| GLM 5.3 Flash | 4 DGX Sparks | 46.3 tok/s | 0.19 s | 257 s |
| Qwen 3.8 27B | 2 DGX Sparks | 42.8 tok/s | 0.16 s | 810 s |
| DeepSeek V4 Flash | 2 DGX Sparks | 40.9 tok/s | 0.29 s | not measured |
Measured in August 2026 with each model thinking at its highest setting, on different prompts from the test at the top of this page. These were also taken before October's recipe change, so they should not be set beside the other tables.
How I test
- One user at a time unless the row says otherwise. A total across several users is always labelled as a total.
- The machine is idle. The script waits for 8 quiet seconds before each run, and a run that had other work arriving part way through is thrown away.
- The same prompts and the same script on every machine, so two rows can be compared.
- Three runs, the middle one reported. The lowest and highest are under each figure.
- Thinking off, and thinking on. Where a model can think before it answers, I run the writing test both ways, and each way has its own table.
- A laptop is tested on mains power, and the row says which Windows power mode it was in.
- Every row has its date. Recipes and software move quickly, and an old figure stays labelled as old.
- Nothing here is a maker's claim. Every figure on this page is one I measured.
- If a figure looks wrong, tell me under the latest video and I will run it again.