{"repo":"waybarrios/vllm-mlx","free":true,"listed":false,"github":"https://github.com/waybarrios/vllm-mlx","clone":"git clone https://github.com/waybarrios/vllm-mlx.git","description":"High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.","language":"Python","stars":1514,"topics":["apple-silicon","llm","macos","mlx","multimodal-ai","speech-to-text","text-to-speech","vision-language-model","vllm","anthropic"],"license":"Apache-2.0","category":"ai-agents","readme_excerpt":"vllm-mlx Continuous batching + OpenAI + Anthropic APIs in one server. Native Apple Silicon inference. Read this in other languages: English · Español · Français · 中文 --- What is vllm-mlx? A vLLM-style inference server for Apple Silicon Macs. Unlike Ollama or mlx-lm used directly, it ships continuous batching, paged KV cache, prefix caching, and SSD-tiered cache , and exposes both OpenAI /v1/ and Anthropic /v1/messages from a single process. Run LLMs, vision models, audio, and embeddings on Metal with unified memory, no conversion step. Quick start (30 seconds) OpenAI SDK: Anthropic SDK / Claude Code: Features APIs - OpenAI-compatible : /v1/chat/completions , /v1/completions , /v1/embeddings , /v1/rerank , /v1/responses - Anthropic-compatible : /v1/messages (streaming, tool use, system prompts) - MCP Tool Calling : 12 parsers (OpenAI, Anthropic, Gemini, Qwen, DeepSeek, Gemma, and more) - Structured output : JSON Schema via response format (lm-format-enforcer) Throughput & memory - Continuous batching : high throughput for concurrent requests - Paged KV cache : memory-efficient with prefix sharing - SSD-tiered KV cache : spill prefix cache to disk for long-context agents ( --ssd-cache-dir ) - Warm prompts : preload popular prefixes at startup ( --warm-prompts ) for 1.3-2.25x TTFT - Prefix cache : trie-based, shared across requests Multimodal - Text + image + video + audio from one server - Vision models: Gemma 3, Gemma 4, Qwen3-VL, Pixtral, Llama vision - Audio input in chat ( ","default_branch":null,"files":null,"tree":[],"storefront":"/r/waybarrios","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/waybarrios/vllm-mlx/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}