Skip to content
#

inference-server

Here are 208 public repositories matching this topic...

⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.

  • Updated Aug 30, 2026
  • Rust
Rapid-MLX

Rapid-MLX is an open-source (Apache 2.0) OpenAI- and Anthropic-compatible LLM inference server and Mac app for Apple Silicon, built on MLX, focused on reliable tool calling for coding agents. Release-gated with Claude Code, Codex CLI, Aider, Hermes and DeepSeek Harness.

  • Updated Sep 29, 2026
  • Python

High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.

  • Updated Sep 29, 2026
  • Python

Add this topic to your repo

To associate your repository with the inference-server topic, visit your repo's landing page and select "manage topics."

Learn more