omlx
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
oMLX is an LLM inference server optimized for Apple Silicon Macs that implements continuous batching and tiered KV caching with SSD persistence. The server provides OpenAI and Anthropic-compatible APIs for running local language models, vision-language models, embeddings, and rerankers with a native macOS menu bar interface. It features automatic model discovery, multi-model serving with LRU eviction, and a web-based admin dashboard for real-time monitoring and configuration. The tiered cache system keeps frequently-accessed blocks in RAM while offloading cold blocks to SSD, enabling efficient context management even across server restarts.
Key Features
Use Cases
- 01Running local LLMs on Apple Silicon Macs with efficient memory management and persistent caching
- 02Serving multiple models simultaneously with automatic loading and unloading based on usage patterns
- 03Integrating local inference into development tools like Claude Code, Cursor, and Copilot
- 04Processing vision-language tasks with multi-image chat and OCR capabilities
- 05Benchmarking model performance with realistic prefix cache hit testing
- 06Managing LLM inference from the macOS menu bar without terminal access
Related MCP Servers
View moremarkitdown
Python tool for converting files and office documents to Markdown.
firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
prompts.chat
f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-host for your organization with complete privacy.
langflow
Langflow is a powerful tool for building and deploying AI-powered agents and workflows.
omlx — FAQ
What is oMLX?+
oMLX is an LLM inference server optimized for Apple Silicon that provides OpenAI and Anthropic-compatible APIs for running local language models, vision models, and embeddings with continuous batching and tiered KV caching.
How do I install oMLX on macOS?+
Download the .dmg from the GitHub releases page and drag it to Applications, or install via Homebrew with 'brew tap jundot/omlx' followed by 'brew install jundot/omlx/omlx'. You can also install from source using pip.
Does oMLX support MCP (Model Context Protocol)?+
Yes, MCP support is available as an optional feature. Install it with pip using 'pip install mcp' or include it during installation with 'pip install -e ".[mcp]"' when building from source.
Which AI clients work with oMLX?+
oMLX works with any OpenAI or Anthropic-compatible client, including Claude Code, OpenCode, Codex, Cursor, Copilot, and Hermes Agent. The admin dashboard provides one-click integration setup for these clients.
What are the system requirements for oMLX?+
oMLX requires macOS 15.0 or later (Sequoia), Python 3.11–3.13, and Apple Silicon (M1/M2/M3/M4/M5). No API keys are needed for core functionality.
Is oMLX free to use?+
Yes, oMLX is open source and released under the Apache 2.0 license, making it free for both personal and commercial use.
How do I install omlx?+
Open the source repository on GitHub and follow its README. omlx is a mcp server — MCP Agents Market links you directly to the official repo.
Is omlx free?+
omlx is an open-source project hosted on GitHub. Check the repository for its license and any usage requirements.