Apple MacBook Pro 16-inch (2024, M4 Max 16-core CPU, 40-core GPU): which AI models can it run?
Estimated from the specifications below. Choose the memory size you have or plan to buy, then read the table for that size.
Specifications
- Brand
- Apple
- Year
- 2024
- Chip
- Apple M4 Max
- Processor
- 16-core CPU (12 performance, 4 efficiency)
- Graphics
- 40-core GPU
- Memory type
- unified memory
- Memory options
- 48 GB, 64 GB, 128 GB
- Memory bandwidth
- 546 GB/s
Sources: https://support.apple.com/en-us/121554, https://www.apple.com/newsroom/2024/10/apple-introduces-m4-pro-and-m4-max/
Apple lists 546GB/s for this variant.
With 48 GB of memory
About 33.6 GB of it can hold a model.
| Model | Fit | Needs (GB) | Writing speed (tokens per second) | Speed | Quality | Basis |
|---|---|---|---|---|---|---|
| Llama 3.2 3B llama3.2:3b |
Fits | 5 | 96 to 177 | Fast | not measured yet | Estimated |
| Qwen3 4B qwen3:4b · thinks first |
Fits | 6 | 76 to 142 | Fast | not measured yet | Estimated |
| Gemma 3 4B gemma3:4b |
Fits | 6.7 | 58 to 108 | Fast | not measured yet | Estimated |
| Mistral 7B v0.3 mistral:7b |
Fits | 7.8 | 43 to 81 | Fast | not measured yet | Estimated |
| Llama 3.1 8B llama3.1:8b |
Fits | 8.3 | 39 to 72 | Fast | not measured yet | Estimated |
| Qwen3 8B qwen3:8b · thinks first |
Fits | 8.9 | 37 to 68 | Fast | not measured yet | Estimated |
| Gemma 3 12B gemma3:12b |
Fits | 15.9 | 24 to 44 | Fast | not measured yet | Estimated |
| DeepSeek-R1 Distill Qwen 14B deepseek-r1:14b · thinks first |
Fits | 13.7 | 21 to 39 | Fast | not measured yet | Estimated |
| Phi-4 14B phi4:14b |
Fits | 13.9 | 21 to 39 | Fast | not measured yet | Estimated |
| Qwen3 14B qwen3:14b · thinks first |
Fits | 13.4 | 21 to 38 | Fast | not measured yet | Estimated |
| gpt-oss 20B gpt-oss:20b · thinks first |
Fits | 16.5 | 48 to 88 | Fast | not measured yet | Estimated |
| Gemma 3 27B gemma3:27b |
Fits | 27.2 | 11 to 21 | Fast | not measured yet | Estimated |
| Qwen3 30B (A3B) qwen3:30b · thinks first |
Fits | 22.6 | 56 to 104 | Fast | 9.3/10 | Estimated |
| Qwen3 32B qwen3:32b · thinks first |
Fits | 26.3 | 9.6 to 18 | Unclear could be Moderate or Fast | not measured yet | Estimated |
| Llama 3.1 70B llama3.1:70b |
Won't fit | 51.5 | not known | n/a | not measured yet | Estimated |
| gpt-oss 120B gpt-oss:120b |
Won't fit | 70.5 | not known | n/a | not measured yet | Estimated |
Models marked "thinks first" can write out their reasoning before they answer, so a job takes longer than the speed alone suggests.
With 64 GB of memory
About 44.8 GB of it can hold a model.
| Model | Fit | Needs (GB) | Writing speed (tokens per second) | Speed | Quality | Basis |
|---|---|---|---|---|---|---|
| Llama 3.2 3B llama3.2:3b |
Fits | 5 | 96 to 177 | Fast | not measured yet | Estimated |
| Qwen3 4B qwen3:4b · thinks first |
Fits | 6 | 76 to 142 | Fast | not measured yet | Estimated |
| Gemma 3 4B gemma3:4b |
Fits | 6.7 | 58 to 108 | Fast | not measured yet | Estimated |
| Mistral 7B v0.3 mistral:7b |
Fits | 7.8 | 43 to 81 | Fast | not measured yet | Estimated |
| Llama 3.1 8B llama3.1:8b |
Fits | 8.3 | 39 to 72 | Fast | not measured yet | Estimated |
| Qwen3 8B qwen3:8b · thinks first |
Fits | 8.9 | 37 to 68 | Fast | not measured yet | Estimated |
| Gemma 3 12B gemma3:12b |
Fits | 15.9 | 24 to 44 | Fast | not measured yet | Estimated |
| DeepSeek-R1 Distill Qwen 14B deepseek-r1:14b · thinks first |
Fits | 13.7 | 21 to 39 | Fast | not measured yet | Estimated |
| Phi-4 14B phi4:14b |
Fits | 13.9 | 21 to 39 | Fast | not measured yet | Estimated |
| Qwen3 14B qwen3:14b · thinks first |
Fits | 13.4 | 21 to 38 | Fast | not measured yet | Estimated |
| gpt-oss 20B gpt-oss:20b · thinks first |
Fits | 16.5 | 48 to 88 | Fast | not measured yet | Estimated |
| Gemma 3 27B gemma3:27b |
Fits | 27.2 | 11 to 21 | Fast | not measured yet | Estimated |
| Qwen3 30B (A3B) qwen3:30b · thinks first |
Fits | 22.6 | 56 to 104 | Fast | 9.3/10 | Estimated |
| Qwen3 32B qwen3:32b · thinks first |
Fits | 26.3 | 9.6 to 18 | Unclear could be Moderate or Fast | not measured yet | Estimated |
| Llama 3.1 70B llama3.1:70b |
Won't fit | 51.5 | not known | n/a | not measured yet | Estimated |
| gpt-oss 120B gpt-oss:120b |
Won't fit | 70.5 | not known | n/a | not measured yet | Estimated |
Models marked "thinks first" can write out their reasoning before they answer, so a job takes longer than the speed alone suggests.
With 128 GB of memory
About 89.6 GB of it can hold a model.
| Model | Fit | Needs (GB) | Writing speed (tokens per second) | Speed | Quality | Basis |
|---|---|---|---|---|---|---|
| Llama 3.2 3B llama3.2:3b |
Fits | 5 | 96 to 177 | Fast | not measured yet | Estimated |
| Qwen3 4B qwen3:4b · thinks first |
Fits | 6 | 76 to 142 | Fast | not measured yet | Estimated |
| Gemma 3 4B gemma3:4b |
Fits | 6.7 | 58 to 108 | Fast | not measured yet | Estimated |
| Mistral 7B v0.3 mistral:7b |
Fits | 7.8 | 43 to 81 | Fast | not measured yet | Estimated |
| Llama 3.1 8B llama3.1:8b |
Fits | 8.3 | 39 to 72 | Fast | not measured yet | Estimated |
| Qwen3 8B qwen3:8b · thinks first |
Fits | 8.9 | 37 to 68 | Fast | not measured yet | Estimated |
| Gemma 3 12B gemma3:12b |
Fits | 15.9 | 24 to 44 | Fast | not measured yet | Estimated |
| DeepSeek-R1 Distill Qwen 14B deepseek-r1:14b · thinks first |
Fits | 13.7 | 21 to 39 | Fast | not measured yet | Estimated |
| Phi-4 14B phi4:14b |
Fits | 13.9 | 21 to 39 | Fast | not measured yet | Estimated |
| Qwen3 14B qwen3:14b · thinks first |
Fits | 13.4 | 21 to 38 | Fast | not measured yet | Estimated |
| gpt-oss 20B gpt-oss:20b · thinks first |
Fits | 16.5 | 48 to 88 | Fast | not measured yet | Estimated |
| Gemma 3 27B gemma3:27b |
Fits | 27.2 | 11 to 21 | Fast | not measured yet | Estimated |
| Qwen3 30B (A3B) qwen3:30b · thinks first |
Fits | 22.6 | 56 to 104 | Fast | 9.3/10 | Estimated |
| Qwen3 32B qwen3:32b · thinks first |
Fits | 26.3 | 9.6 to 18 | Unclear could be Moderate or Fast | not measured yet | Estimated |
| Llama 3.1 70B llama3.1:70b |
Fits | 51.5 | 4.4 to 8.3 | Unclear could be Slow or Moderate | not measured yet | Estimated |
| gpt-oss 120B gpt-oss:120b · thinks first |
Fits | 70.5 | 40 to 75 | Fast | not measured yet | Estimated |
Models marked "thinks first" can write out their reasoning before they answer, so a job takes longer than the speed alone suggests.
Speed is the speed of writing an answer, shown as a range. Reading speed is not estimated. How the estimates work · How close they are · Back to the finder