</>MCP Agents Market
MCP Server

omlx

by jundot20.2kPythonUpdated 2026-08-21

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

Claude CodeClaude DesktopCursorCopilotCodex

oMLX is an LLM inference server optimized for Apple Silicon Macs that implements continuous batching and tiered KV caching with SSD persistence. The server provides OpenAI and Anthropic-compatible APIs for running local language models, vision-language models, embeddings, and rerankers with a native macOS menu bar interface. It features automatic model discovery, multi-model serving with LRU eviction, and a web-based admin dashboard for real-time monitoring and configuration. The tiered cache system keeps frequently-accessed blocks in RAM while offloading cold blocks to SSD, enabling efficient context management even across server restarts.

Key Features

Tiered KV cache with hot (RAM) and cold (SSD) tiers that persist across restarts and enable prefix reuse
Continuous batching for concurrent request handling with configurable parallelism
Multi-model serving supporting LLMs, vision-language models, OCR models, embeddings, and rerankers
Native macOS menu bar app with one-click start/stop, persistent stats, and auto-update
Web admin dashboard with real-time monitoring, model management, chat interface, and benchmark tools
OpenAI and Anthropic API compatibility with streaming, tool calling, and structured output
Automatic model pinning, per-model TTL, and memory-aware LRU eviction
Built-in model downloader with HuggingFace integration and one-click integrations for Claude Code, Cursor, and other clients

Use Cases

  • 01Running local LLMs on Apple Silicon Macs with efficient memory management and persistent caching
  • 02Serving multiple models simultaneously with automatic loading and unloading based on usage patterns
  • 03Integrating local inference into development tools like Claude Code, Cursor, and Copilot
  • 04Processing vision-language tasks with multi-image chat and OCR capabilities
  • 05Benchmarking model performance with realistic prefix cache hit testing
  • 06Managing LLM inference from the macOS menu bar without terminal access

Related MCP Servers

View more

omlx — FAQ

What is oMLX?+

oMLX is an LLM inference server optimized for Apple Silicon that provides OpenAI and Anthropic-compatible APIs for running local language models, vision models, and embeddings with continuous batching and tiered KV caching.

How do I install oMLX on macOS?+

Download the .dmg from the GitHub releases page and drag it to Applications, or install via Homebrew with 'brew tap jundot/omlx' followed by 'brew install jundot/omlx/omlx'. You can also install from source using pip.

Does oMLX support MCP (Model Context Protocol)?+

Yes, MCP support is available as an optional feature. Install it with pip using 'pip install mcp' or include it during installation with 'pip install -e ".[mcp]"' when building from source.

Which AI clients work with oMLX?+

oMLX works with any OpenAI or Anthropic-compatible client, including Claude Code, OpenCode, Codex, Cursor, Copilot, and Hermes Agent. The admin dashboard provides one-click integration setup for these clients.

What are the system requirements for oMLX?+

oMLX requires macOS 15.0 or later (Sequoia), Python 3.11–3.13, and Apple Silicon (M1/M2/M3/M4/M5). No API keys are needed for core functionality.

Is oMLX free to use?+

Yes, oMLX is open source and released under the Apache 2.0 license, making it free for both personal and commercial use.

How do I install omlx?+

Open the source repository on GitHub and follow its README. omlx is a mcp server — MCP Agents Market links you directly to the official repo.

Is omlx free?+

omlx is an open-source project hosted on GitHub. Check the repository for its license and any usage requirements.

Related searches