How the finder estimates work

Updated 2026-10-06

In short: The finder works out, from published specifications, which open AI models fit in a laptop's memory and roughly how fast they write. We did not test those laptops. Every row says Estimated from specs. A row says Measured only when we ran the model on that laptop. Our scorecards stay measured-only. See how we test.

What is estimated#

For each laptop, each memory size and each model, we give you:

  • Fit. Does the model fit in the memory it can use?
  • Memory needed. How much memory the model wants, in GB.
  • Writing speed. A range, in tokens per second. A token is a small piece of a word.
  • A speed band. Fast, Moderate or Slow, or Unclear when the range is too wide to say.

These words are not the verdicts on our scorecards. They come from specifications, not from a test.

What is not estimated#

  • Reading speed. How fast a model reads your prompt is not estimated in this version. We say so on each page.
  • Quality. Quality belongs to the model, not to the laptop. We show a quality score only for models we have measured. Otherwise the table says "not measured yet". We never invent one.
  • Heat, battery and noise. We do not estimate these.
  • Long documents. All figures assume a context of 16,384 tokens. This is the same setting as our tests.

The memory rule#

Memory the model can use. This depends on the kind of graphics:

  • Unified memory (Apple silicon). The graphics chip may use only part of the memory. Apple does not publish one fixed share for this. Our default is 0.70 of the memory. It is a number we chose, and a Mac owner can change the real limit.
  • A graphics card with its own memory. We use the size of that video memory.
  • Integrated graphics. We use the system memory minus 4 GB that we keep for the operating system.

Memory the model needs. We take the size of the model file for the Q4_K_M version on Ollama. We add 5% for loading overhead. We add the memory for the context, called the KV cache, which we work out from the model's published design on Hugging Face. We then add 1 GB for the runtime.

The fit label:

  • Fits. The model needs at most 0.90 of the usable memory.
  • Tight. It needs more than that, but still fits.
  • Partial offload (slow). Only for a graphics card. The model does not fit in video memory, but it fits when part of it runs from system memory. This is much slower.
  • Won't fit. It does not fit even then.

The speed rule#

Writing a token means reading the model's weights from memory. So the speed depends mostly on the memory bandwidth, which is how many GB per second the memory can move. Apple lists the bandwidth for each of its chips. Other makers list it for their graphics cards.

  1. We find how many GB of weights the laptop reads for each token. For an ordinary model this is the whole file. For a "mixture of experts" model, only a few parts work at a time, so we use the file size times the active share of the parameters.
  2. The middle estimate is the bandwidth times the efficiency, divided by that number of GB. A real laptop never reaches the full bandwidth. The default efficiency is 0.50 for ordinary (dense) models and 0.30 for mixture-of-experts models, which do extra work for each token.
  3. The range runs from 0.7 times the middle estimate to 1.3 times it.
  4. If part of the model runs from system memory, we multiply the speed by 0.15.

The band comes from the whole range. Fast means even the low end reaches 10 tokens per second. Moderate means the whole range sits between 5 and 10. Slow means even the high end is below 5. Unclear means the range crosses one of those lines, and we say which bands are possible. These limits are the same as the writing-speed limits in our test rules.

If we do not have a sourced bandwidth for a laptop, we show the fit only and say that speed is not known.

Some models write out their reasoning first. They take longer to finish a job than the speed alone suggests. We mark them "thinks first".

Error bars and honest limits#

The defaults are our first guesses, set from spot measurements. The range of 0.7 to 1.3 is not a measured error bar yet. We replace a default with a fitted value only after we have at least 3 measured pairs of that kind of model. Until then the page How close are our estimates? says "not enough measurements yet".

All calibration so far comes from one machine: a MacBook Pro with an M1 Max. We took the two default efficiencies from spot measurements of two models on it: one dense model and one mixture-of-experts model. Other chips and other makers may differ, so we keep the ranges wide. Read the speed ranges as a rough guide, not a promise. The calibration page always shows the current pairs and errors.

Other limits:

  • Specifications can differ. Makers sell several versions of one laptop. We list the version named in the row and link its source.
  • Software changes. A new version of the runtime can make a model faster or slower.
  • Memory settings. A different context size, or other programs running, changes how much memory is free.
  • No prices. The finder shows no prices.

Reuse the data#

The finder data is free to reuse under CC BY 4.0. Download finder.json and credit Findra. Each laptop and model row lists its source links.

Sources