Local AI Roundup: What Changed and Where Things Stand
This page is updated every month. The top section covers what changed since the last update; the rest is where things stand now. The date above is the last update.
What changed this month (October 2026)
- New Macs. The Mac mini M6 starts at $899 and the Mac mini M5 Pro at $1,699. The Mac Studio M5 Max starts at $2,499 and the M5 Ultra at $5,499. All Apple US retail.
- Higher prices. A memory shortage pushed Mac prices up twice in 2026, and non-Apple AI hardware rose further.
- Ollama runs on MLX on Apple silicon, with its largest gains on M5 and M6 chips.
- New models. Qwen3.8-27B, Qwen3.6-35B-A3B, Gemma 4, Qwen3.5-122B-A10B and DeepSeek V4 Flash are our picks. Llama 3.x, gpt-oss and Gemma 3 are the previous generation.
- Coming in late October: a 512 GB Mac Studio M5 Ultra, price not yet published.
Where things stand
The sections below are current as of the last update.
Apple's new Mac mini and Mac Studio
Apple announced the new lineup by press release on August 25, 2026 and put it on sale on September 22.
Mac mini M6, from $899. Apple's first 2 nm chip, with 16 GB of memory as standard and options up to 32 GB. Memory bandwidth is 153 GB/s, or 170 GB/s on higher-memory configurations. It has three Thunderbolt 4 ports, which matters below.
Mac mini M5 Pro, from $1,699 (15-core) or $1,899 (18-core). 24 GB as standard, up to 64 GB, with 307 GB/s of bandwidth and three Thunderbolt 5 ports. Apple markets it for clustering several machines over Thunderbolt 5 to run large models. The M6 lacks Thunderbolt 5, so it is not a cluster node.
Mac Studio M5 Max, from $2,499. Memory options of 36, 48, 64 or 128 GB, and up to 614 GB/s of bandwidth.
Mac Studio M5 Ultra, from $5,499 (96 GB). The upgrade to 256 GB adds $4,000. Bandwidth is 1.2 TB/s. Apple claims up to 4x faster prompt processing in LM Studio than the M3 Ultra; that is Apple's figure. Both Mac Studios have 10 Gb Ethernet and Thunderbolt 5 as standard.
Every GPU core on the M6 and M5 Pro includes a Neural Accelerator, and Apple's MLX engine uses the M5 generation's accelerators. Our sizing guide covers which machine fits which work.
Prices went up
AI datacenter demand drove a global memory shortage through 2026. Apple raised prices across several Mac lines on June 25, 2026, and the base Mac mini has gone from $599 (M4, 2024) to $899 as of October 2026. NVIDIA and AMD hardware for local AI rose further. Our post on the 2026 memory shortage has the details.
Ollama moved to MLX
Ollama, the runtime we install on most setups, switched its Apple silicon engine from llama.cpp to Apple's MLX, starting with a preview in version 0.19 on March 30, 2026. Ollama's own benchmark showed decode speed on Qwen3.5-35B-A3B rising from about 58 to about 112 tokens per second. The MLX path applies to models in the safetensors format; GGUF models still use the older path. An update on June 30 (0.31.1) added multi-token prediction for Gemma 4 on Apple silicon.
A new generation of open models
As of October 2026, these are the open-weight models we recommend for Macs:
- Qwen3.8-27B, successor to Qwen3.6-27B, reported to be the most downloaded new open model of 2026.
- Qwen3.6-35B-A3B and Qwen3.5-35B-A3B, mixture-of-experts models that use about 3 billion parameters per token. Qwen3.5-35B-A3B was Ollama's MLX launch model.
- Gemma 4, in E4B, 12B, 26B-A4B and 31B sizes. Its licence terms are reported inconsistently, so we confirm them before installing it for commercial use.
- Qwen3.5-122B-A10B, which fits a 128 GB Mac Studio M5 Max at 4-bit.
- DeepSeek V4 Flash, 284 billion parameters with 13 billion active, which fits a 256 GB M5 Ultra at 4-bit (estimate).
Llama 3.x, Llama 4, gpt-oss, Gemma 3, Qwen2.5 and Qwen3 are now the previous generation. Several hosting providers retired gpt-oss-120b and Llama 3.3 70B in September 2026.
At the top end, the largest open-weight releases are datacenter-scale: Kimi K3 (2.8 trillion parameters, weights released July 27, 2026), Qwen3.8-2.4T-A95B (weights released August 12, 2026) and DeepSeek V4 Pro (1.6 trillion total, 49 billion active). No Mac holds them. Our model guide maps the current picks to memory.
NVIDIA and AMD alternatives
- NVIDIA DGX Spark: 128 GB of unified memory and 273 GB/s, Linux only. NVIDIA announced a 64 GB version at $4,999 on October 2, shipping October 23 through Acer, ASUS, Dell, Gigabyte, HP and MSI (reported). The 128 GB model was reported near $6,950 in early October.
- NVIDIA RTX Spark: Windows-on-Arm PCs with up to 128 GB, from ASUS, Acer and MSI, expected in October 2026. No pricing has been announced.
- AMD Strix Halo (Ryzen AI Max+ 395): up to 128 GB and about 256 GB/s, on Windows or Linux. Street prices for 128 GB systems ran roughly $3,000 to $3,450 in mid-2026 (reported).
We default to Mac and say when another platform fits better. Mac vs NVIDIA vs AMD for Local AI compares them.
What is coming
- 512 GB Mac Studio M5 Ultra: late October 2026, confirmed by Apple. It requires the 36-core CPU and 80-core GPU chip. Price not yet published.
- MacBook Pro with M6 Pro and M6 Max: reported by Bloomberg's Mark Gurman to bring a redesign with an OLED touch display, expected between late 2026 and early 2027. Laptops matter less for always-on local AI.
- AMD: AMD said at Computex 2026 that a Strix Halo refresh, Gorgon Halo, will offer up to 192 GB of unified memory (reported), with no date. Its Medusa chip family is confirmed for 2027.
No further Mac mini or Mac Studio update is expected in the near term.
What it means if you are buying now
Size by memory first, then bandwidth, and buy one memory tier above the model you plan to run. Expect prices to be dated: we quote hardware at Apple retail at the time of order, and we never mark it up. Whether a machine pays for itself depends on how much AI work you run; the cost comparison shows where it does and where cloud is cheaper.
Book a free 15-minute call to talk through what the new lineup means for your work.