How we test laptops: the Findra method

Updated 2026-10-04

In short: We run six small-business jobs on each laptop, using AI that runs on the laptop itself. We time every job, check every answer, and give each job a plain verdict. Every figure is measured on the real machine. None is estimated.

This page is the public version of our test protocol, version 1. Scorecards appear in the laptop list as laptops are tested. Nothing on this page is a result.

Our principles#

  1. Measured, never estimated. Every figure is timed on the actual laptop.
  2. Jobs, not tokens. Results read as "2 minutes, quality 8/10". Raw speed figures sit one click deeper, for technical readers.
  3. Real conditions. Each job runs three times, back to back, while the laptop is plugged in. This shows whether the laptop slows down when it gets hot. Two jobs are repeated on battery.
  4. Our own test files. Every input is a fictional document that we write. No real company or person appears in them. They carry no copyright problems, and no AI model has seen them before.
  5. Fixed settings. The same prompt, the same input and the same settings for every machine.
  6. Verdict first. Each job gets one of five verdicts. The numbers sit underneath.
  7. Versioned. Every result records the protocol version, the model, the app version and the test date. When a new model version arrives, we test again and archive the old result.

The six jobs#

Each job is a task a small business really does. Click a job to see the laptops ranked for it.

Job What we give the AI What it must produce How we check it
Contract summary A fictional services contract of about 20 pages, with 10 key facts planted in it A one-page summary: the parties, the term, the fees, how to end it, and the top 5 risks We count how many of the 10 planted facts it captured correctly
Customer replies 10 fictional customer emails and a one-page refund and shipping policy A reply to each email A written rubric: right policy applied, no invented promises, good tone
Marketing pack A one-page brief for a fictional product A product description of 150 words and 5 social posts A written rubric: every claim traceable to the brief, easy to read, steady brand voice
Invoice extraction 10 fictional invoices as text, in mixed layouts Structured data: vendor, date, invoice number, subtotal, tax, total and line items Field by field against an answer key
Sales data questions A fictional sales spreadsheet of 200 rows Answers to 5 fixed questions Exact match against an answer key
Handbook Q&A A fictional staff handbook of about 30 pages Answers to 10 questions, each naming the section it came from Correct answer and correct section, question by question

We left call transcription out of version 1, because it needs a speech model. We plan an audio test for a later version.

Hardware tiers#

We group laptops by the memory that is available to the AI. Each tier has its own page.

Tier AI memory Typical machines Model sizes we test
Tier A 8 GB of memory Base laptops, budget "AI PCs" 1B to 4B
Tier B 16 GB of memory Mainstream laptops, base MacBooks 7B to 9B
Tier C 24 to 36 GB of shared memory, or a graphics card with 8 to 12 GB MacBook Pro, gaming laptops 12B to 20B
Tier D 48 GB or more of shared memory, or a graphics card with 16 to 24 GB The largest Apple chips, laptops with the strongest graphics cards 30B and up

The "B" in a size like 7B means billions of parameters. It is a rough measure of how big a model is. Bigger models need more memory.

For each tier, we test the two most downloaded general-purpose models in the Ollama library that fit the tier. We use a setting called Q4_K_M. It shrinks a model file so it fits in less memory, and it gives up a little quality in return. The model list for each round of testing is recorded with the results.

We use Ollama as the only AI runner in version 1. A result tells you how the laptop performs with Ollama. Another runner could give different numbers. Our guide on LM Studio and Ollama explains the difference between the two.

Fixed settings#

These are the same on every run.

  • The temperature is 0 and the seed is fixed. This keeps the AI's answers as steady as possible from one run to the next.
  • The context window is 16,384 tokens. The context window is how much text the AI can hold in mind at once. A token is a small piece of a word.
  • Streaming is off, so we time the whole answer.
  • Other apps are closed. Wi-Fi is on, and no downloads are running.
  • Before each timed run, the AI reads the whole input from scratch. A repeated prompt cannot get a free speed-up.
  • The very first run starts cold, with the model not loaded. This measures the first-use wait. Every later run starts with the model loaded. Before each timed run the tool also clears the model's memory of the previous prompt, so every run reads the document from scratch and speeds are not flattered by caching.
  • The full set runs plugged in. The contract summary and the handbook Q&A are then repeated on battery.

What we measure#

Measure What you learn
Task time How long the job takes. It is the average of the two warm runs, runs 2 and 3.
First-use wait How long you wait before the AI starts, the first time.
Reading speed How fast the AI digests a long document.
Writing speed How fast the answer appears.
Fit Whether the model fits in the graphics memory or shared memory.
Peak memory Whether other apps can stay open while the AI works.
Heat slowdown How much longer run 3 takes than run 2. A big gap means the laptop slows down when hot.
Battery cost The share of the battery one job uses when unplugged. It answers "can I do this on a train?"
Quality Whether the output is usable.

How we judge quality#

Four jobs have an answer key: the contract summary, invoice extraction, sales data and the handbook Q&A. Quality is the share of answers that are right, scaled to a score out of 10. This part is objective.

Two jobs have no single right answer: customer replies and the marketing pack. A strong cloud AI model, called the judge, scores each output against a written rubric. The judge follows these rules:

  • It comes from a different model family than any model we test.
  • It does not know which laptop or model produced the output.
  • It scores each output on its own, never side by side with another.

The judge reads saved outputs after the test. The laptop under test never calls a cloud service during a test. All test documents are fictional.

We also check the judge. We plan to read a random 20% of the judged outputs ourselves and compare our scores with the judge's. We will publish the judge agreement rate on this page after the first hand-check is finished. There is no number yet.

What the verdicts mean#

Each job gets one verdict. A job gets the first verdict that fits, checked in this order: Won't run, Not reliable, Great, Usable. A job that fits none of those is Painful.

Verdict In plain words The rule
Won't run The laptop cannot do the job. The model fails to load, runs out of memory, the input does not fit, or the model cannot finish its answer within the fixed reply limit and the 16,384-token window.
Not reliable It runs, but you cannot trust the answers. Quality is below 6/10, at any speed.
Great Good answers, fast. Quality is 8/10 or more, it finishes within the Great time limit, and the writing is fast.
Usable Good enough answers, at a fair pace. Quality is 6/10 or more, it finishes within the Usable time limit, and the writing is at a comfortable pace.
Painful Good enough answers, but too slow. Quality is 6/10 or more, but it is slower than the Usable time limit, or the writing is too slow.

Technical details. The writing-speed rules are measured in tokens per second. Below 5 tokens per second, text appears more slowly than a person reads. At 10 tokens per second or more, it feels instant. Great needs 10 or more. Usable needs 5 or more. Scorecards show these figures only inside a collapsed "Technical details" section.

Time limits#

These are the time limits for each job. They are starting values. We will recalibrate them after we have tested 10 machines, and we will record any change here.

Job Great within Usable within
Contract summary 2 minutes 5 minutes
Customer replies (10) 3 minutes 8 minutes
Marketing pack 1 minute 3 minutes
Invoice extraction (10) 2 minutes 5 minutes
Sales data questions 1 minute 3 minutes
Handbook Q&A (10) 3 minutes 8 minutes

How the scores work#

  • Quality runs from 0 to 10.
  • Speed points are 10 for Great, 6 for Usable, 3 for Painful and 0 for Won't run.
  • Task score = 6 × quality + 4 × speed points. It runs from 0 to 100.
  • Business AI Score is the average of the six task scores, using the best model that fits the laptop.
  • Price per point is the laptop price divided by the Business AI Score. The price is the public market price recorded when we test. It is not a quote from us.

What this method cannot tell you#

  • Small models that run on a laptop are weaker than the best cloud models. Cloud AI is still stronger and easier for most jobs. Local AI wins on privacy, cost per use, and working offline.
  • Our test files are fictional. Your own documents may be harder or easier.
  • A result is for one laptop, one model, one version of Ollama and one test date. A new version can change it.
  • We test one runner, Ollama. We do not yet test how other AI software performs, including software that uses the chips built for AI that some laptops have, called NPUs.

How we publish#

  • Every scorecard shows the protocol version and the test date. Our raw run logs stay private.
  • Public scores are open data. They are licensed CC BY 4.0, so you may copy and reuse them if you give us a credit link. You can download them from our open data page.
  • Every scorecard carries machine-readable markup, so search engines and tools can read the data.
  • A new protocol version, or a new model version in a tier, means we test again. We archive the old scorecard and mark it as old.
  • Affiliate links never change a score. See our affiliate disclosure.

Read how the tiers work in our memory guide, or open the laptop list.

Check our work. Every score we publish is free to download. Get the open data.

Sources