Writing · June 2026

Running Gemma-4 26B at 124 tokens/sec on a CPU, no GPU

Gemma-4 26B, a mixture-of-experts model, runs at about 40 tokens per second single-stream, lossless, and about 124 aggregate at batch 32, on an i9-13900K with 64 GB of DDR5 and no graphics card.

I wanted to see how fast a 26B mixture-of-experts model runs on a normal desktop with no graphics card. Just the CPU: an i9-13900K, 64GB of plain DDR5. I went looking because of a question that keeps deciding the architecture of the robot I'm building: how fast can a model run on hardware you own rather than rent. The answer turned into its own project.

About 40 tokens/sec single-stream, lossless, or about 124 if you batch a few requests. For a 26B model with no GPU, that's a little wild.

One thing up front, because it surprised me. The instinct with a mixture of experts is to quantize the experts, that's where the parameters are. But I counted the bytes you actually read per token, and the experts are 16% of them. The output head, the projection to the 262K-token vocabulary, is 32%. For this model you compress the head, not the experts.

The run itself, on the i9, one stream generating live, then the batched sweep filling in past 120 tok/s. No GPU, and the closing numbers are measured from this very run.

The setup: a 26B model that uses 3.8B parameters per token

Gemma-4-26B-A4B is 26B parameters but only ~3.8B are used per token, because it's a mixture of experts: 128 of them, 8 used at a time. That sparsity is the only reason this fits on a CPU at all. And "lossless" here is literal, every trick either skips work the model would have thrown away, or guesses ahead and has the full model check the guess. The tokens that come out are exactly Q4_0's tokens, just faster. (One lever, running fewer experts, is an approximation; I'll show how I checked it.)

Where the time goes: memory bandwidth, not cores

Plain Q4_0, one token at a time: 25 tokens/sec, and it didn't move whether I gave it 8 threads or 24. That's the sign. If cores don't help, you're waiting on memory.

tokens/sec = memory bandwidth ÷ bytes read per token

To make one token you read the active model out of RAM, once. Everything below is about changing those two numbers: faster RAM, or fewer bytes.

Two speedups that cost almost nothing

Speculative decoding. Gemma ships a small official drafter that guesses the next few tokens; the big model verifies them all in a single pass instead of one at a time. Good guesses mean several tokens per pass. That's 25 → 40 tok/s, still exactly lossless, the big model checks every token.

Running 3 of the 8 experts. This one nearly made me drop it: perplexity on raw Wikipedia jumped 1.6x. But raw-text perplexity is meaningless for a chat model, it scores terribly there no matter what, and in that setting small differences look much bigger than they are. So I read the actual outputs instead, and top-3 and top-8 give the same answers on real prompts. Free speed, confirmed by looking rather than trusting the number.

The surprise: compress the head, not the experts

Then I stopped guessing and counted the bytes per token, straight from the model file. What people miss is that bytes-per-token is not size-on-disk. For a mixture of experts they're wildly different: you only read the experts that are used, 3 of 128, but you read the whole head on every token.

share of all weights · on disk
experts · ~88%rest
bytes actually read · per token
always-on · 52%head · 32%experts · 16%
Experts: most of the model on disk, smallest slice per token. The head is the inverse, tiny on disk, read in full every step.
Part of the modelShare of weights on diskShare of bytes read per tokenWhy
Experts (128, 3 used)about 88%16%only the experts in use are read
Always-on: attention and dense layersmost of the rest52%read in full every token
Output head (262K-token vocabulary)small32%read in full every token; 6.5 bits per weight down to 2.4 with no change in output, 606 MB to 225 MB

Counted from the model file, not estimated.

So the head is the part to compress. It sits at 6.5 bits per weight; I dropped it to 2.4 and couldn't find any damage, same answers on every prompt. That's 606 MB down to 225 MB on the single most-read tensor in the model. 2.4 is the lowest setting that works: at 1.75 bits it got slower, more time unpacking the tighter format than saved reading it, and the output started looping in the arithmetic.

What did not work

Two failed attempts, included because they're the useful part. First, quantizing harder is supposed to add to the other gains. A 32%-smaller model is 18% faster on plain decode. But stack it with speculative decoding and it buys exactly nothing, both are limited by memory bandwidth and speculative decoding already took that gain. (I logged a 43.6 once, got excited, re-ran it, noise around 41.) Second, you can't shrink the experts much anyway, their down-projection width doesn't fit the low-bit block formats, so they fall back to 4 bits. Doesn't matter, they're 16% of the bytes; forcing it would buy about 5%.

The memory limit, and the way around it

Is 40 a hard hardware limit, or is the software inefficient? I measured the RAM's actual bandwidth and how much decode uses of it.

DDR5-4800
on paper
76.8 GB/s
actually
achievable
64.5 GB/s
decode
uses
~48 GB/s
Decode already runs at ~78% of what the RAM can deliver. Pinning threads, core counts, more quant, none of it moved that. The gap is how a MoE reads memory, scattered, not an inefficiency in the software.
Memory bandwidth, i9-13900K with DDR5-4800GB/s
On paper76.8
Measured achievable (compiled STREAM triad)64.5
Used by single-stream decode, repacking onabout 48

So single-stream is close to the hardware limit. But the equation is per token, for one stream. Serve a few at once and you read each weight once for all of them, the matmul turns from a vector into a matrix, and the bottleneck moves from memory to compute, the cores that sat idle finally have work.

28
69
81
114
124
batch 1481632
Aggregate tok/s by batch size. It crosses 100 at batch 16 and climbs to 124. Single-stream, meanwhile, never moves off ~40.
Batch sizeAggregate tokens/sec, 8 threadsAggregate tokens/sec, 24 threads
128
469
881
1693114
32124

It crosses 100 at batch 16. And where threads did nothing for single-stream, here they help, batch-16 goes from 93 at eight threads to 114 at twenty-four, because now it's compute-bound. Same chip, two different limits. This is continuous batching, the trick vLLM made famous on GPUs; it just isn't usually used on a CPU. The catch is it's aggregate, not per-stream, at batch 32 each request gets about 4 tok/s. Right for a server with concurrent load, wrong for one person waiting on one answer.

Reproduce it

Everything runs on public models and two public forks of llama.cpp: atomic-llama for the speculative-decode path and the batched benchmark, pinned to commit d86eb0b, and ik_llama.cpp for the low-bit quantization kernels, pinned to f96eadd. Both built for AVX2, since the 13900K has AVX-512 fused off. The model is Google's Gemma-4-26B-A4B instruction-tuned release as a Q4_0 GGUF, with the official MTP drafter alongside it.

git clone https://github.com/arun-prasath2005/gemma4-cpu-moe
cd gemma4-cpu-moe
bash scripts/reproduce.sh      # builds both engines, downloads the two GGUFs, runs both recipes

# the single-stream recipe: top-3 experts plus MTP speculative decoding, 8 threads (bandwidth-bound)
llama-cli -m models/gemma-4-26B_q4_0-it.gguf \
  --override-kv gemma4.expert_used_count=int:3 \
  --model-draft models/mtp-q4.gguf --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-threads 8 \
  -t 8 -f prompt.txt -n 200 --temp 0 --seed 1 -st --simple-io --no-display-prompt

# the batched recipe: read the S_TG column for aggregate generation tok/s, 24 threads (compute-bound)
llama-batched-bench -m models/gemma-4-26B_q4_0-it.gguf \
  --override-kv gemma4.expert_used_count=int:3 \
  -c 8192 -npp 64 -ntg 128 -npl 1,2,4,8,16,32 -t 24

Tokens per second drift with turbo and temperature: the careful single-stream number on this box is 37.7, and an idle, cool machine turbos to about 41. What should reproduce is the shape: threads flat for latency, threads helping for throughput, the byte-budget proportions, and the head compressing with no change in output. If the shape does not hold on your machine, that is a bug, and the repository takes issues and rows for its community results file.

The result: two limits on one chip

single-stream
latency
~40 tok/s
aggregate
throughput
~124 tok/s

100 on a no-GPU desktop is done, as throughput. Single-stream sits at 40 because that's the most the memory bus can deliver, and pushing past it there is a hardware question, faster RAM or more memory channels, not a software one. Two limits, 3x apart, on the same chip.

None of the pieces are mine. Speculative decoding off Google's drafter, the low-bit kernels from ik_llama.cpp, all on llama.cpp. What I did was measure where the limits are, including the failed attempts, and write it down.

It all runs on public models, and the recipe, every number, and the scripts to reproduce it on your own machine are on GitHub: arun-prasath2005/gemma4-cpu-moe.

Common questions

How fast does Gemma-4 26B run on a CPU with no GPU?

About 40 tokens per second for a single stream, identical output to plain Q4_0, and about 124 aggregate with 32 requests batched, on an i9-13900K with 64 GB of DDR5-4800. Plain Q4_0 with no tricks: about 25.

Why is 124 a batched number and not a single-stream one?

Single-stream decode is bound by memory bandwidth. Batching reads each weight once for all requests in flight, so the bottleneck moves to compute. At batch 32 the aggregate is 124 and each request gets about 4. Single-stream never moves off about 40.

What limits CPU inference for a mixture-of-experts model?

Memory bandwidth, not cores. Tokens per second is bandwidth divided by bytes read per token. Thread count from 8 to 24 changed nothing single-stream. Measured bandwidth was 64.5 GB/s against 76.8 on paper, and decode used about 48 of it.

Which part of the model should be quantized hardest?

The output head. The experts are about 88 percent of the weights on disk but 16 percent of the bytes read per token. The head is read in full every token and is 32 percent. From 6.5 to 2.4 bits per weight changed no outputs; 1.75 got slower and looped.

Is the result lossless?

Speculative decoding is exactly lossless; the full model verifies every token. Running 3 of 8 experts is an approximation, checked on real prompts against top-8, and the answers matched.

How do I reproduce it?

Clone the repository and run the reproduce script. It builds the two pinned forks, downloads the GGUFs, and runs both recipes. The exact commands are above.

Cite this

Arun Prasath E G, "Running Gemma-4 26B at 124 tokens/sec on a CPU, no GPU," apeg.dev, published 30 June 2026, updated 15 September 2026. https://apeg.dev/writing/running-gemma4-26b-on-a-cpu/

The scripts and results are archived on Zenodo with a DOI, so the record is fixed: doi.org/10.5281/zenodo.22762963 (version 1.0, September 2026). Source: github.com/arun-prasath2005/gemma4-cpu-moe

Arun Prasath E G

Agentic AI architect. Twenty years in enterprise architecture, seven companies built and sold, more than 8,000 commits in the last twelve months across a retail platform, an ad-acquisition platform, a transactional email service, a video product, robotics and CPU inference research. Every number on this site was measured on real hardware or taken from a real engagement's records.

Working with me

I take on one or two engagements at a time, inside the company, and I hand over when it runs. If you have a problem in your own processes that no single product solves, and you need it built and running, not just a recommendation, that is the kind of work I do.