llama-swap
Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc
llama-swap is an MCP server that provides dynamic model-switching capabilities for local AI inference servers, along with documentation and configuration assistance via MCP. Built in Go, it acts as a reverse proxy that automatically hot-swaps between multiple generative AI models running on llama.cpp, vllm, stable-diffusion.cpp, and other OpenAI/Anthropic-compatible servers. The server exposes MCP tools at /api/mcp that allow AI clients to query llama-swap's configuration and documentation, enabling conversational assistance about setup and usage. It combines infrastructure management with agent-accessible knowledge tools in a single binary with zero external dependencies.
Key Features
Use Cases
- 01Enable AI assistants to query llama-swap configuration and documentation through MCP tools
- 02Run multiple LLMs locally and automatically switch between them based on incoming API requests
- 03Manage concurrent text, image, and audio models with custom resource allocation rules
- 04Provide unified OpenAI/Anthropic API interface for heterogeneous local inference backends
- 05Monitor and debug local AI workflows with real-time metrics and log streaming
- 06Test and compare different models through the integrated playground interface
Related MCP Servers
View moremarkitdown
Python tool for converting files and office documents to Markdown.
firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
prompts.chat
f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-host for your organization with complete privacy.
langflow
Langflow is a powerful tool for building and deploying AI-powered agents and workflows.
llama-swap — FAQ
What is llama-swap?+
llama-swap is an MCP server and reverse proxy that automatically hot-swaps between multiple local AI models running on OpenAI/Anthropic-compatible servers like llama.cpp and vllm. It exposes MCP tools at /api/mcp that allow AI clients to query its configuration and documentation.
How do I install llama-swap?+
llama-swap can be installed via Docker, Homebrew (macOS/Linux), MacPorts (macOS), WinGet (Windows), or by downloading pre-built binaries from the GitHub releases page. After installation, run it with a YAML configuration file specifying your model commands.
Which MCP clients work with llama-swap?+
Any MCP client can connect to llama-swap's /api/mcp endpoint to access documentation and configuration query tools. The README mentions using it with tool-capable local models in the built-in Playground.
Do I need API keys to use llama-swap?+
No external API keys are required since llama-swap manages local inference servers. You can optionally configure API keys in llama-swap's own configuration to restrict access to its endpoints.
What inference servers does llama-swap support?+
llama-swap works with any OpenAI or Anthropic API-compatible server, including llama.cpp (llama-server), vllm, tabbyAPI, stable-diffusion.cpp, audio.cpp, whisper.cpp, and ComfyUI. Python-based servers are recommended to run via Docker/Podman.
Is llama-swap free and open source?+
Yes, llama-swap is open source software available on GitHub under the mostlygeek/llama-swap repository. It is free to use and can be built from source or installed via package managers.
How do I install llama-swap?+
Open the source repository on GitHub and follow its README. llama-swap is a mcp server — MCP Agents Market links you directly to the official repo.
Is llama-swap free?+
llama-swap is an open-source project hosted on GitHub. Check the repository for its license and any usage requirements.