Local AI 101: models, Hugging Face, runners and workflows

Updated 2026-10-04

Short answer: Local AI is an AI model that runs on your own computer. You need three things: a model, an app that runs it, and a routine that puts it to work. Cloud AI is still stronger and easier for most jobs. Local AI wins when your documents must stay private, when you want no fee for each use, and when you need to work offline.

What "local" means#

When you use a cloud AI service, your text travels to the company's computers. When you use local AI, the model sits on your laptop and the work happens there.

Both of the main apps say this plainly. Ollama says it does not see your prompts or data when you run locally. LM Studio says nothing you type into a chat leaves your device. LM Studio also says it works offline once you have the model files. See the Ollama FAQ and the LM Studio offline page.

Two cautions:

  • Both apps can also reach the cloud. Ollama offers cloud models. They run on Ollama's servers, not on your laptop, and you must sign in to use them. In the Ollama app and command line, their names carry a cloud label. See the Ollama cloud page. Before you rely on privacy, check that the model you chose is running on your own machine.
  • Downloading still needs the internet. LM Studio makes network requests when you search for models, download them, and check for updates. After that, chatting does not need a connection.

Models#

A model is a very large file of numbers. It was built by training on huge amounts of text. When you give it text, it predicts a useful reply.

  • Open models are ones you can download and run yourself. Each one has a licence. Some are free for business use, and some are not. Check before you build a business routine on one.
  • Size is counted in parameters. A label like 7B means 7 billion. Bigger models are usually better at hard jobs, and they need more memory. Our memory guide explains how much.
  • Compression makes a model file smaller. It is called quantisation. LM Studio describes it as compressing the file while giving up some quality. A common setting is Q4_K_M, and it is the one we use in our tests.

Hugging Face#

Hugging Face is a website where people share models, datasets and small demo apps. Its documentation describes the Hub as a collection of version-controlled folders that hold these files. See the Hub documentation.

What to look at on a model page:

  • The model card. It describes the model, what it is meant for, and its limits and biases. See model cards.
  • The licence. It is shown on the model page. Hugging Face asks you to respect a project's licence. See licences.
  • The file format. For local use, look for GGUF. It is a file format built to load quickly for local running. See GGUF.

Anyone can upload to the Hub. Hugging Face lists malware scanning among its security features, but you should still prefer well-known publishers.

Apps can pull models straight from the Hub. With Ollama, the command takes the form ollama run hf.co/{username}/{repository}, as the Hugging Face Ollama page shows. LM Studio has a "Use this model" button and a built-in search, as the LM Studio page shows.

Runners#

A runner is the app that loads a model and answers your questions. Two are popular with non-experts:

  • Ollama is open source under the MIT licence. It runs on macOS, Windows and Linux. Other programs can talk to it through a local address on your own machine.
  • LM Studio is a desktop app for finding, downloading and chatting with models. LM Studio announced on 8 July 2025 that it is free to use at home and at work. See the announcement.

Our guide on LM Studio and Ollama compares them. Our tests use Ollama only, for now.

Workflows#

A chat box is not a workflow. A workflow is a routine you can repeat: the same instructions, the same kind of input, the same shape of output. It turns a clever tool into a dependable one.

Good first workflows for a small business match jobs we test:

  • Summarise a document. Drop in a contract or a policy and get a one-page summary. See contract summaries.
  • Draft replies. Paste a customer email and your policy, and get a reply to edit and send. See customer replies.
  • Pull data out of invoices. Turn invoice text into a spreadsheet row. See invoice extraction.

Always read what the AI wrote before you use it. Small models make mistakes, and they can state wrong things with confidence.

If you would rather have this done for you, see our setup service.

Cloud or local?#

Cloud AI Local AI
Strength of the best models Stronger Weaker, and limited by your laptop
Getting started Easier More steps, or a setup service
Privacy Your text goes to the provider Your text stays on your machine
Cost for each use Often paid by use or by subscription No fee for the software or each job once the model is on your machine. You still pay for the laptop and the power it uses
Working offline No Yes, once the models are downloaded

Many businesses will use both. Use local AI for private documents and routine jobs. Use the cloud for the hardest jobs, with documents you are happy to share.

What this means for your laptop#

How well local AI runs depends mostly on your laptop's memory. Start with the memory tiers and the guide on how much memory you need. Then look at the laptop list and the ranking for the job you care about, such as contract summaries. We measure every result. We do not guess.

Want more like this? See the newsletter page.

Sources