MLX 1

Native macOS server for fast local AI model inference
oMLX is a native macOS inference server for running local AI models on Apple Silicon using Apple's MLX framework. It provides paged SSD KV caching to reduce time-to-first-token for coding agents, continuous batching for higher throughput, and OpenAI- and Anthropic-compatible APIs. oMLX supports LLM, vision-language, embedding, and reranker models, along with tool calling, MCP integration, multi-model serving, and a native menu bar application with a web dashboard.