Apple MacBook Pro 14-inch (2026, M5 Pro): which AI models can it run?
Estimated from the specifications below. Choose the memory size you have or plan to buy, then read the table for that size.
Specifications
- Brand
- Apple
- Year
- 2026
- Chip
- Apple M5 Pro
- Processor
- 15-core or 18-core CPU
- Graphics
- 16-core or 20-core GPU
- Memory type
- unified memory
- Memory options
- 24 GB, 64 GB
- Memory bandwidth
- 307 GB/s
Sources: https://support.apple.com/en-us/126318, https://www.apple.com/newsroom/2026/03/apple-introduces-macbook-pro-with-all-new-m5-pro-and-m5-max/
Apple lists 307GB/s and up to 64GB for M5 Pro. Intermediate memory options are not listed here because the Apple pages conflict.
With 24 GB of memory
About 16.8 GB of it can hold a model.
| Model | Fit | Needs (GB) | Writing speed (tokens per second) | Speed | Quality | Basis |
|---|---|---|---|---|---|---|
| Llama 3.2 3B llama3.2:3b |
Fits | 5 | 54 to 100 | Fast | not measured yet | Estimated |
| Qwen3 4B qwen3:4b · thinks first |
Fits | 6 | 43 to 80 | Fast | not measured yet | Estimated |
| Gemma 3 4B gemma3:4b |
Fits | 6.7 | 33 to 60 | Fast | not measured yet | Estimated |
| Mistral 7B v0.3 mistral:7b |
Fits | 7.8 | 24 to 45 | Fast | not measured yet | Estimated |
| Llama 3.1 8B llama3.1:8b |
Fits | 8.3 | 22 to 41 | Fast | not measured yet | Estimated |
| Qwen3 8B qwen3:8b · thinks first |
Fits | 8.9 | 21 to 38 | Fast | not measured yet | Estimated |
| Gemma 3 12B gemma3:12b |
Tight | 15.9 | 13 to 25 | Fast | not measured yet | Estimated |
| DeepSeek-R1 Distill Qwen 14B deepseek-r1:14b · thinks first |
Fits | 13.7 | 12 to 22 | Fast | not measured yet | Estimated |
| Phi-4 14B phi4:14b |
Fits | 13.9 | 12 to 22 | Fast | not measured yet | Estimated |
| Qwen3 14B qwen3:14b · thinks first |
Fits | 13.4 | 12 to 21 | Fast | not measured yet | Estimated |
| gpt-oss 20B gpt-oss:20b · thinks first |
Tight | 16.5 | 27 to 50 | Fast | not measured yet | Estimated |
| Gemma 3 27B gemma3:27b |
Won't fit | 27.2 | not known | n/a | not measured yet | Estimated |
| Qwen3 30B (A3B) qwen3:30b |
Won't fit | 22.6 | not known | n/a | 9.3/10 | Estimated |
| Qwen3 32B qwen3:32b |
Won't fit | 26.3 | not known | n/a | not measured yet | Estimated |
| Llama 3.1 70B llama3.1:70b |
Won't fit | 51.5 | not known | n/a | not measured yet | Estimated |
| gpt-oss 120B gpt-oss:120b |
Won't fit | 70.5 | not known | n/a | not measured yet | Estimated |
Models marked "thinks first" can write out their reasoning before they answer, so a job takes longer than the speed alone suggests.
With 64 GB of memory
About 44.8 GB of it can hold a model.
| Model | Fit | Needs (GB) | Writing speed (tokens per second) | Speed | Quality | Basis |
|---|---|---|---|---|---|---|
| Llama 3.2 3B llama3.2:3b |
Fits | 5 | 54 to 100 | Fast | not measured yet | Estimated |
| Qwen3 4B qwen3:4b · thinks first |
Fits | 6 | 43 to 80 | Fast | not measured yet | Estimated |
| Gemma 3 4B gemma3:4b |
Fits | 6.7 | 33 to 60 | Fast | not measured yet | Estimated |
| Mistral 7B v0.3 mistral:7b |
Fits | 7.8 | 24 to 45 | Fast | not measured yet | Estimated |
| Llama 3.1 8B llama3.1:8b |
Fits | 8.3 | 22 to 41 | Fast | not measured yet | Estimated |
| Qwen3 8B qwen3:8b · thinks first |
Fits | 8.9 | 21 to 38 | Fast | not measured yet | Estimated |
| Gemma 3 12B gemma3:12b |
Fits | 15.9 | 13 to 25 | Fast | not measured yet | Estimated |
| DeepSeek-R1 Distill Qwen 14B deepseek-r1:14b · thinks first |
Fits | 13.7 | 12 to 22 | Fast | not measured yet | Estimated |
| Phi-4 14B phi4:14b |
Fits | 13.9 | 12 to 22 | Fast | not measured yet | Estimated |
| Qwen3 14B qwen3:14b · thinks first |
Fits | 13.4 | 12 to 21 | Fast | not measured yet | Estimated |
| gpt-oss 20B gpt-oss:20b · thinks first |
Fits | 16.5 | 27 to 50 | Fast | not measured yet | Estimated |
| Gemma 3 27B gemma3:27b |
Fits | 27.2 | 6.3 to 12 | Unclear could be Moderate or Fast | not measured yet | Estimated |
| Qwen3 30B (A3B) qwen3:30b · thinks first |
Fits | 22.6 | 31 to 58 | Fast | 9.3/10 | Estimated |
| Qwen3 32B qwen3:32b · thinks first |
Fits | 26.3 | 5.4 to 10 | Moderate | not measured yet | Estimated |
| Llama 3.1 70B llama3.1:70b |
Won't fit | 51.5 | not known | n/a | not measured yet | Estimated |
| gpt-oss 120B gpt-oss:120b |
Won't fit | 70.5 | not known | n/a | not measured yet | Estimated |
Models marked "thinks first" can write out their reasoning before they answer, so a job takes longer than the speed alone suggests.
Speed is the speed of writing an answer, shown as a range. Reading speed is not estimated. How the estimates work · How close they are · Back to the finder