ASUS ROG Strix G16 (2025) G614: which AI models can it run?
Estimated from the specifications below. Choose the memory size you have or plan to buy, then read the table for that size.
Specifications
- Brand
- ASUS
- Year
- 2025
- Chip
- AMD Ryzen 9 9955HX3D, 9955HX or 8940HX
- Processor
- AMD Ryzen 9 9955HX3D, 9955HX or 8940HX (16 cores)
- Graphics
- NVIDIA GeForce RTX 5060 Laptop GPU
- Memory type
- DDR5 SO-DIMM
- Memory options
- 16 GB, 32 GB, 64 GB
- Graphics memory
- 8 GB
- Memory bandwidth
- 384 GB/s
Sources: https://rog.asus.com/us/laptops/rog-strix/rog-strix-g16-2025-g614/spec/, https://www.nvidia.com/en-us/geforce/laptops/50-series/
The spec page lists 16GB and 8GB x2 and 16GB x2 SO-DIMM kits with a 64GB maximum (two slots). GPU bandwidth from NVIDIA's RTX 50 laptop page.
With 16 GB of memory
About 8 GB of it can hold a model.
| Model | Fit | Needs (GB) | Writing speed (tokens per second) | Speed | Quality | Basis |
|---|---|---|---|---|---|---|
| Llama 3.2 3B llama3.2:3b |
Fits | 5 | 67 to 125 | Fast | not measured yet | Estimated |
| Qwen3 4B qwen3:4b · thinks first |
Fits | 6 | 54 to 100 | Fast | not measured yet | Estimated |
| Gemma 3 4B gemma3:4b |
Fits | 6.7 | 41 to 76 | Fast | not measured yet | Estimated |
| Mistral 7B v0.3 mistral:7b |
Tight | 7.8 | 31 to 57 | Fast | not measured yet | Estimated |
| Llama 3.1 8B llama3.1:8b Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 8.3 | 4.1 to 7.6 | Unclear could be Slow or Moderate | not measured yet | Estimated |
| Qwen3 8B qwen3:8b · thinks first Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 8.9 | 3.9 to 7.2 | Unclear could be Slow or Moderate | not measured yet | Estimated |
| Gemma 3 12B gemma3:12b Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 15.9 | 2.5 to 4.6 | Slow | not measured yet | Estimated |
| DeepSeek-R1 Distill Qwen 14B deepseek-r1:14b · thinks first Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 13.7 | 2.2 to 4.2 | Slow | not measured yet | Estimated |
| Phi-4 14B phi4:14b Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 13.9 | 2.2 to 4.1 | Slow | not measured yet | Estimated |
| Qwen3 14B qwen3:14b · thinks first Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 13.4 | 2.2 to 4 | Slow | not measured yet | Estimated |
| gpt-oss 20B gpt-oss:20b · thinks first Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 16.5 | 5 to 9.3 | Moderate | not measured yet | Estimated |
| Gemma 3 27B gemma3:27b |
Won't fit | 27.2 | not known | n/a | not measured yet | Estimated |
| Qwen3 30B (A3B) qwen3:30b |
Won't fit | 22.6 | not known | n/a | 9.3/10 | Estimated |
| Qwen3 32B qwen3:32b |
Won't fit | 26.3 | not known | n/a | not measured yet | Estimated |
| Llama 3.1 70B llama3.1:70b |
Won't fit | 51.5 | not known | n/a | not measured yet | Estimated |
| gpt-oss 120B gpt-oss:120b |
Won't fit | 70.5 | not known | n/a | not measured yet | Estimated |
Models marked "thinks first" can write out their reasoning before they answer, so a job takes longer than the speed alone suggests.
With 32 GB of memory
About 8 GB of it can hold a model.
| Model | Fit | Needs (GB) | Writing speed (tokens per second) | Speed | Quality | Basis |
|---|---|---|---|---|---|---|
| Llama 3.2 3B llama3.2:3b |
Fits | 5 | 67 to 125 | Fast | not measured yet | Estimated |
| Qwen3 4B qwen3:4b · thinks first |
Fits | 6 | 54 to 100 | Fast | not measured yet | Estimated |
| Gemma 3 4B gemma3:4b |
Fits | 6.7 | 41 to 76 | Fast | not measured yet | Estimated |
| Mistral 7B v0.3 mistral:7b |
Tight | 7.8 | 31 to 57 | Fast | not measured yet | Estimated |
| Llama 3.1 8B llama3.1:8b Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 8.3 | 4.1 to 7.6 | Unclear could be Slow or Moderate | not measured yet | Estimated |
| Qwen3 8B qwen3:8b · thinks first Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 8.9 | 3.9 to 7.2 | Unclear could be Slow or Moderate | not measured yet | Estimated |
| Gemma 3 12B gemma3:12b Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 15.9 | 2.5 to 4.6 | Slow | not measured yet | Estimated |
| DeepSeek-R1 Distill Qwen 14B deepseek-r1:14b · thinks first Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 13.7 | 2.2 to 4.2 | Slow | not measured yet | Estimated |
| Phi-4 14B phi4:14b Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 13.9 | 2.2 to 4.1 | Slow | not measured yet | Estimated |
| Qwen3 14B qwen3:14b · thinks first Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 13.4 | 2.2 to 4 | Slow | not measured yet | Estimated |
| gpt-oss 20B gpt-oss:20b · thinks first Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 16.5 | 5 to 9.3 | Moderate | not measured yet | Estimated |
| Gemma 3 27B gemma3:27b Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 27.2 | 1.2 to 2.2 | Slow | not measured yet | Estimated |
| Qwen3 30B (A3B) qwen3:30b · thinks first Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 22.6 | 5.9 to 11 | Unclear could be Moderate or Fast | 9.3/10 | Estimated |
| Qwen3 32B qwen3:32b · thinks first Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 26.3 | 1 to 1.9 | Slow | not measured yet | Estimated |
| Llama 3.1 70B llama3.1:70b |
Won't fit | 51.5 | not known | n/a | not measured yet | Estimated |
| gpt-oss 120B gpt-oss:120b |
Won't fit | 70.5 | not known | n/a | not measured yet | Estimated |
Models marked "thinks first" can write out their reasoning before they answer, so a job takes longer than the speed alone suggests.
With 64 GB of memory
About 8 GB of it can hold a model.
| Model | Fit | Needs (GB) | Writing speed (tokens per second) | Speed | Quality | Basis |
|---|---|---|---|---|---|---|
| Llama 3.2 3B llama3.2:3b |
Fits | 5 | 67 to 125 | Fast | not measured yet | Estimated |
| Qwen3 4B qwen3:4b · thinks first |
Fits | 6 | 54 to 100 | Fast | not measured yet | Estimated |
| Gemma 3 4B gemma3:4b |
Fits | 6.7 | 41 to 76 | Fast | not measured yet | Estimated |
| Mistral 7B v0.3 mistral:7b |
Tight | 7.8 | 31 to 57 | Fast | not measured yet | Estimated |
| Llama 3.1 8B llama3.1:8b Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 8.3 | 4.1 to 7.6 | Unclear could be Slow or Moderate | not measured yet | Estimated |
| Qwen3 8B qwen3:8b · thinks first Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 8.9 | 3.9 to 7.2 | Unclear could be Slow or Moderate | not measured yet | Estimated |
| Gemma 3 12B gemma3:12b Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 15.9 | 2.5 to 4.6 | Slow | not measured yet | Estimated |
| DeepSeek-R1 Distill Qwen 14B deepseek-r1:14b · thinks first Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 13.7 | 2.2 to 4.2 | Slow | not measured yet | Estimated |
| Phi-4 14B phi4:14b Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 13.9 | 2.2 to 4.1 | Slow | not measured yet | Estimated |
| Qwen3 14B qwen3:14b · thinks first Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 13.4 | 2.2 to 4 | Slow | not measured yet | Estimated |
| gpt-oss 20B gpt-oss:20b · thinks first Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 16.5 | 5 to 9.3 | Moderate | not measured yet | Estimated |
| Gemma 3 27B gemma3:27b Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 27.2 | 1.2 to 2.2 | Slow | not measured yet | Estimated |
| Qwen3 30B (A3B) qwen3:30b · thinks first Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 22.6 | 5.9 to 11 | Unclear could be Moderate or Fast | 9.3/10 | Estimated |
| Qwen3 32B qwen3:32b · thinks first Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 26.3 | 1 to 1.9 | Slow | not measured yet | Estimated |
| Llama 3.1 70B llama3.1:70b Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 51.5 | 0.5 to 0.9 | Slow | not measured yet | Estimated |
| gpt-oss 120B gpt-oss:120b |
Won't fit | 70.5 | not known | n/a | not measured yet | Estimated |
Models marked "thinks first" can write out their reasoning before they answer, so a job takes longer than the speed alone suggests.
Speed is the speed of writing an answer, shown as a range. Reading speed is not estimated. How the estimates work · How close they are · Back to the finder