Open-source AI coding assistant for VS Code. Code completion, chat, edits and reviews with local or hosted models. Your models, your infrastructure.
-
Updated
Sep 28, 2026 - TypeScript
Open-source AI coding assistant for VS Code. Code completion, chat, edits and reviews with local or hosted models. Your models, your infrastructure.
A universal OpenCode plugin for dynamic model discovery with flexible configuration for OpenAI-compatible providers.
Self-hosted Qwen3.8-27B (FP8) inference with vLLM, KServe and Envoy AI Gateway on RTX 6000 PRo or 2× RTX 4080 Super
Benchmark how well any model performs real-world, long-horizon coding tasks across coding agents (harnesses) -- Claude Code and pi. Run models three ways -- Anthropic on Amazon Bedrock, open-weight on Bedrock via a LiteLLM proxy, or self-hosted on EC2 with vLLM -- score them with an LLM judge, and plot the cost/quality Pareto frontier.
Terraform setup for deploying a private coding LLM on Vast.ai with vLLM, Qwen3 Coder, and OpenCode.
Free ChatGPT API & Free LLM API Key Alternative. 100% local drop-in OpenAI replacement supporting System Prompts, Tool Calling, and Think Mode for AI Agents.
Enterprise-grade Sovereign AI Stack optimized for NVIDIA Blackwell (sm_120) & vLLM. Features 256K context window, 5.8k tok/s prefill, and integrated observability via Langfuse.
The control plane for self-hosted AI inference. Warm-state GPU routing, multi-runtime orchestration across Ollama, vLLM, llama.cpp, TGI and MLX . Single Go binary. Apache-2.0.
A platform-agnostic format + method for sizing local-LLM hardware from your real agent sessions, calibrated with capability-oracle runs.
Windows utility for using Codex Desktop GUI with Ollama and other local/self-hosted LLM backends via profiles.
Alternatively prompt, an LLM-based plugin for the IntelliJ product family. Answer questions and generate code with self-hosted models.
Production-grade multimodal RAG assistant using open-source LLMs and vector databases.
Air-gapped pre-deployment network change validation against a real containerlab digital twin, with sealed PCI/SOC2/NIST evidence. Zero egress.
A university-scale LLM serving platform in miniature: vLLM on cloud GPU, LiteLLM gateway with per-faculty governance, Prometheus/Grafana SLOs, k3d/ArgoCD GitOps. All numbers measured, all failures documented.
Self-hosted AI coding platform — provisions GPU compute on Vast.ai and deploys open-weight LLMs via llama.cpp as a bring-your-own-key backend for GitHub Copilot (also auto-configures Continue.dev/Cline).
Self-hosted, typed inference API for open-weight models. OpenAI-compatible, runs airgapped once your models are cached.
Self-hosted vLLM inference stack with an OpenAI-compatible API, Docker Compose templates, Caddy reverse proxy, and NVIDIA GPU thermal guard.
GLM-5.2 744B at 4-bit on Modal 4x H200 via vLLM, plus a static streaming chat UI.
AI-powered herbal remedies chatbot based on "The Little Handbook of Natural Remedies" by Michael Martin. RAG system for natural medicine research.
convert instagram reels to markdown recipe using self hosted LLMs on VLLM
To associate your repository with the self-hosted-llm topic, visit your repo's landing page and select "manage topics."