← writing

Decode speed is one division, and it contradicts most hardware advice

8 October 2026 · by Syamjith NK

I wanted a faster local model, so I measured the machine I already had instead of reading reviews. The answer turned out to be a single division, and it ruled out almost everything I was about to buy.

tokens/second  ~=  memory bandwidth / bytes read per token

Generating one token streams the weights through the memory bus once. That is very nearly the whole story for decode, and it means most of the headline numbers on a spec sheet do not decide the thing you are waiting for.

What I measured

Base M4 Mac mini, 16 GB, 10 cores, running a 7B at q4:

measurementvalue
7B q4 decode21.3 to 21.5 tok/s (two runs, cold and warm)
first token, cold18.47 s
first token, warm0.01 s
one 7B resident4.31 GB
swap after two queries5.65 GB used, 494 MB free

Three conclusions, none of them what I expected.

21.3 tok/s is roughly what this chip can ever do on a 7B. At the published 120 GB/s, the equation predicts a ceiling of 120 / 4.31 = 27.8 t/s. I measured 21.3, which is 77% of the bus. No flag fixes that, because it is not a misconfiguration. It is also the sanity check worth doing before you trust any of the rest of this: a model of your own hardware that lands within a quarter of your own stopwatch is a model you can use to rule out a purchase.

The 18.47 s cold start is a swap symptom, not a loading problem. Keeping the model resident cannot help when there is no RAM to keep it in. What I actually felt as slowness was paging.

Capacity caps the model class, and that costs more than speed. A 32B at q4 is about 20 GB and simply will not load. The choice was never between a smarter model and a faster one; memory decides which models exist for you at all.

Where the division bites

boxbandwidthdense 70B q4
(~40 GB/token)
30B MoE
(~2 GB active)
M4 mini 16 GB~120 GB/s (published)cannot loadcannot load
Strix Halo 128 GB256 GB/s spec, ~215 measured by third parties~5-6 t/s66 t/s measured
Mac Studio M4 Max 128 GB546 GB/s (published)~13 t/s (projected)very fast
Raspberry Pi 517.1 GB/sno1-2 t/s on an 8B

A 128 GB box can be arithmetically pinned at 5 to 6 t/s on a dense 70B while loading it perfectly well. Capacity and speed are different purchases, and a spec sheet blurs them.

Then the reversal that makes the whole thing interesting: the same box reaches 66 t/s on a mixture-of-experts model. An MoE reads only its active experts per token, so bytes-per-token collapses by an order of magnitude while the bandwidth sits exactly where it was. The hardware did not change. The divisor did.

Which means "is this box fast enough" is not answerable without naming the architecture. The same machine on the same afternoon is slow for dense models and excellent for sparse ones.

Two caveats, because the arithmetic cuts both ways

The 21.3 tok/s, the 18.47 s, the 4.31 GB and the swap figures are mine, measured. The 120, 546 and 17.1 GB/s figures are published vendor specs, the 215 GB/s is a third-party measurement, and any projected token rate scaling off those is an estimate. I have labelled them rather than quietly mixing them in with the things I ran.

Second, and more useful: I found a published benchmark claiming about 12 t/s on a dense 70B on the 215 GB/s box, and I do not believe it. Forty GB per token into 215 GB/s is 5.4 t/s. A number that beats the memory bus by roughly two times needs an explanation, whether that is a smaller quant, a draft model or speculative decoding. If no explanation is offered, trust the division. The arithmetic is the floor, and the floor is the part you can check yourself.

What I would tell anyone sizing a box

One thing that has nothing to do with speed and cost me more time than any of it

No Apple-silicon Mac powers on from shutdown over the network. There is no IPMI, and wake-on-LAN only works from sleep, not from off. An x86 desktop does wake from S5. If a machine has to be recoverable remotely with nobody in the room, that is a hard architectural difference, and it was not on any spec comparison I read.