WebAI Studio

Consumer hardware · local inference

Local models run at the speed of the memory bus.

Generating a token means reading the whole model out of memory, so how fast a machine answers comes down to how fast its memory bus is. Two things stand between the spec sheet and reality: a machine only realises part of its rated peak, and the engine only keeps part of what it gets , because fixed per-token overhead swamps a small model's streaming time. Both are tunable in Advanced settings. Pick a year to see what each kind of consumer hardware averaged on each size of model.