Everything Ollama does, free — model downloads, the native llama.cpp engine, multi-model, the local OpenAI-compatible API. No account, no caps. Paid tiers cover what no model runner does: dictation, speech, vision, system control, and developer tools.
Explore plans
Free for everyone. Pro for the full toolkit. Studio for creators, Max for developers. Team for your whole office on one Mac.
Free
Everything Ollama does, free. No account, no caps.
$0
Native llama.cpp engine — run any GGUF model
Unlimited model downloads, multi-model switching
Local OpenAI-compatible API + embeddings
MCP server with 359 free tools
All speech: transcription, text-to-speech, voice cloning, dictation
Chat with bundled model
Apple Intelligence on the Neural Engine
Full browser automation, vision, and system control
The free engine is llama.cpp, embedded directly — currently llama-server b10068. The version ships in the app's About panel and updates with every upstream bump.