Vllm Mlx
未认领High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.
anthropicanthropic-apiapple-siliconclaude-codecontinuous-batchinginference-serverllmlocal-llmmacosmcpmlxmultimodal-aiopenaiopenai-apiopenai-compatiblespeech-to-texttext-to-speechtool-callingvision-language-modelvllm
安装
$
pip install vllm-mlx