Ollama v0.40.0 makes the MLX runtime the default on Apple Silicon
Announced October 6, 2026
What happened
Starting with Ollama v0.40.0, model architectures supported by the MLX runtime run on MLX automatically on Apple Silicon devices. Supported models include qwen3.8, gemma4, qwen3.6 and qwen3.5. Decision models such as Nimble tev1, clef and clef-flash, and the embedding model embeddinggemma-2, can also run on MLX.
Why it matters
Practitioners running local LLMs on a Mac will use MLX without any extra setup.
Sources
- Official v0.40.0 Ollama