Choosing the Right LLM for Your Mac: A Practical Guide
Pick the model before you buy the Mac. The model decides how much memory you need, and memory is the spec you cannot change after purchase.
This guide lists the open-weight models we recommend as of October 2026, which Mac memory tier each one needs, and what has changed since earlier versions of this post. Model names turn over every few months. The method below does not.
Start with the work
Match the smallest model that does your hardest task well. A larger model than the work needs costs memory and speed and returns nothing you will notice. A smaller one makes local AI look worse than it is.
As a rough guide:
- Short drafting, email, summaries of a page or two: a small model such as Gemma 4 E4B or 12B.
- Writing a professional would sign, analysis of client documents, coding help: a 27B to 35B model such as Qwen3.8-27B or Qwen3.6-35B-A3B.
- Long, judgment-heavy documents, or one machine serving several people: Qwen3.5-122B-A10B or larger.
Some work still calls for a frontier cloud model. We draw that line per workload, so private documents stay on your machine and the rest goes wherever it fits.
The memory rule
A model at 4-bit needs about 0.55 to 0.6 GB of memory per billion parameters, plus 10 to 30 percent headroom for the context window and runtime. These are estimates. macOS and your apps need memory too, so leave room beyond the model itself.
For a mixture-of-experts (MoE) model, total parameters set the memory and active parameters set the speed. Qwen3.6-35B-A3B has 35 billion parameters in total but uses about 3 billion for each token. It needs the memory of a 35B model and runs closer to the speed of a small one.
Picks by memory tier, as of October 2026
| Model | Size | Memory for the model (estimate) | Licence | Mac that holds it |
|---|---|---|---|---|
| Gemma 4 E4B, 12B | 4B-class and 12B | under 10 GB | check before commercial use | Mac mini M6, 16 GB ($899) |
| Qwen3.8-27B | 27B dense | about 16 to 17 GB | Apache 2.0 (reported) | Mac mini M6, 24 GB ($1,099) |
| Qwen3.6-35B-A3B | 35B MoE, 3B active | about 20 GB | Apache 2.0 | Mac mini M6, 32 GB ($1,299), or M5 Pro with 48 GB or more |
| Gemma 4 26B-A4B, 31B | 26B MoE and 31B | about 15 to 19 GB | check before commercial use | Mac mini M6, 24 to 32 GB, or M5 Pro |
| Qwen3.5-122B-A10B | 122B MoE, 10B active | about 70 GB | Apache 2.0 | Mac Studio M5 Max, 128 GB |
| DeepSeek V4 Flash | 284B MoE, 13B active | about 165 GB | MIT | Mac Studio M5 Ultra, 256 GB ($9,499) |
| GLM-5.3 (full) | about 744B | about 430 GB | check before commercial use | Mac Studio M5 Ultra, 512 GB (late October 2026, price not yet published), or two machines clustered |
Prices are Apple US retail as of October 2026. We buy hardware at that price and never mark it up. Higher-memory Mac mini M5 Pro and Mac Studio M5 Max configurations cost more than their starting prices ($1,699 and $2,499); we quote them from Apple's configurator at the time of order.
Qwen3.8-27B succeeds Qwen3.6-27B and is reported to be the most downloaded new open model of 2026. Gemma 4 is strong for its size. Its licence terms are reported inconsistently, so we confirm them for your use before installing it. The same goes for GLM-5.3.
Previous generation
Llama 3.x, Llama 4, gpt-oss, Gemma 3, Qwen2.5 and Qwen3 are the previous generation as of October 2026. Several hosting providers retired gpt-oss-120b and Llama 3.3 70B in September 2026. These models still run on a Mac. We no longer recommend them as a first pick for new setups, and earlier versions of this guide that named them are out of date.
Too large for any Mac
Some open-weight models are built for datacenters. Kimi K3 (2.8 trillion parameters), Qwen3.8-2.4T-A95B and DeepSeek V4 Pro (1.6 trillion total, 49 billion active) are open, but no Mac holds them. When a task needs a model in that class, it goes to a hosted service.
What changed in Ollama
Ollama is the runtime we install on most setups. In March 2026, starting with a preview in version 0.19, it moved its Apple silicon engine from llama.cpp to Apple's MLX. Ollama reported decode speed on Qwen3.5-35B-A3B rising from about 58 to about 112 tokens per second on its own benchmark. The MLX path applies to models in the safetensors format; GGUF models still use the older path. An update on June 30, 2026 (0.31.1) added multi-token prediction for Gemma 4 on Apple silicon.
The gains are largest on M5 and M6 chips, whose GPU cores include Neural Accelerators that MLX uses. That is one reason we recommend safetensors builds of current models on new hardware.
Speed: measured where it exists, estimated where it does not
Maai® Machines has not run its own benchmarks. These figures come from published reviews.
- Qwen3.8-27B on a Mac Studio M5 Ultra: 48 tokens per second at an 8K context, 32 at 128K (MacStories, September 2026). BGR measured about 55 in LM Studio with MLX.
- Qwen3.5-122B-A10B on an M5 Ultra: about 80 tokens per second (BGR, September 2026).
- Qwen3.5-122B-A10B on an M5 Max with 128 GB: 65.9 tokens per second at a 4K context, in a MacBook Pro (Hardware Corner, March 11, 2026).
No independent measurements of the Mac mini M6 or M5 Pro had been published as of October 7, 2026. As an estimate, generation speed is about memory bandwidth × 0.6 to 0.8, divided by the gigabytes read per token: the full model size for a dense model, the active parameters only for an MoE model. By that method, Qwen3.8-27B runs at about 11 to 14 tokens per second on a Mac mini M5 Pro and about 5 to 8 on a Mac mini M6. Estimates, not measurements. The sizing guide covers the hardware side in detail.
Key takeaways
- Choose the smallest model that does your hardest task well, then buy one memory tier above what it needs.
- Budget about 0.55 to 0.6 GB per billion parameters at 4-bit, plus 10 to 30 percent. Estimates. For MoE models, total parameters set memory and active parameters set speed.
- As of October 2026 we recommend Qwen3.8-27B, Qwen3.6-35B-A3B, Gemma 4, Qwen3.5-122B-A10B and DeepSeek V4 Flash, depending on memory.
- Llama 3.x, gpt-oss and Gemma 3 are the previous generation. They run, but they are no longer our first pick.
- Ollama now runs on MLX on Apple silicon, with its largest gains on M5 and M6 chips.
Matching the model to the work is the core of what we do. Book a free 15-minute call and describe the work; we will tell you which model fits and what it needs. If the answer needs a closer look, the $500 Assessment covers it and is waived with any setup. Setup tiers are on our pricing. If what you want is marketing done for you rather than a machine, our sister service MOCO™ handles that.