claudeers.
// MCP Servers

smg

Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-firs…

// MCP Servers[ cli ][ api ][ claude ]#claude#anthropic#anthropic-api#chat#gemini#inference-gateway#lightseek#llm#mcp-servers◷ Apache-2.0$open-sourceupdated about 1 month ago
Actively maintained
100/100
last commit 10 days ago
last release 13 days ago
releases 17
open issues 59
// star history

Install with your AI

Paste into Claude Code, Cursor, or any agent — it reads the repo and wires the tool into your project.

Install and set up smg (release-binary project) into my current project.
Found on https://claudeers.com/smg
Repo: https://github.com/smg-project/smg
Homepage/docs: https://lightseek.org/smg
Detected install method: release-binary → inspect the README
Category: mcp-servers. Platforms: cli, api.
Read the repo's README for exact setup and env vars, then install it and wire it into my project.

Claudeers Health Verdict:
active; community-verified: false. Confirm the source before running anything.
// or install directly (release-binary)

Grab the latest release asset from GitHub.

# download a build from https://github.com/smg-project/smg/releases
// or clone
git clone https://github.com/smg-project/smg

// compatibility

Platformscli, api
Operating systems—
AI compatibilityclaude
LicenseApache-2.0
Pricingopen-source
LanguageRust

Get your FREE $2.50 API credits to access TickAtlas financial data ↗

SMG Logo

Shepherd Model Gateway

Engine-agnostic, high-performance model-routing gateway for large-scale LLM deployments. SMG centralizes worker lifecycle management, balances traffic across self-hosted engines and cloud providers, and gives you enterprise-grade control over multi-tenancy, chat-history storage, MCP tooling, and observability — behind one unified endpoint.

SMG architecture: clients flow through the gateway layer and router layer to gRPC workers, HTTP workers, and external APIs

Why SMG?

🚀 Maximize GPU UtilizationCache-aware routing tracks each worker's KV-cache state in radix trees to reuse prefixes across SGLang, vLLM, TensorRT-LLM, TokenSpeed, and MLX — with load modeling that accounts for queued token work and KV pressure.
🔌 One API, Any BackendRoute to self-hosted engines over HTTP or gRPC, or to OpenAI, Anthropic, Gemini, and xAI — plus any OpenAI-compatible endpoint — through a single unified gateway.
⚡ Built for SpeedNative Rust with streaming gRPC pipelines, cached tokenization with zero-copy cache hits, prefill/decode disaggregation (including a separate encode stage for vision), and DP-aware routing for data-parallel engines.
🔒 Enterprise ControlPriority admission scheduling with preemption and per-tenant controls, API-key auth with OIDC on the control plane, WebAssembly plugins for custom logic, and chat history that never leaves your infrastructure.
📊 Full Observability90+ Prometheus metrics, OpenTelemetry tracing with W3C trace context propagated into the engines over both HTTP and gRPC, and structured JSON logs with request correlation.

API Coverage: OpenAI Chat Completions, Completions, Embeddings, Rerank, and Classify; Responses and Conversations APIs for agents; Anthropic Messages; Gemini Interactions; Realtime over WebSocket and WebRTC; audio transcription; tokenize/detokenize; and MCP tool execution with approval policies in the Responses and Messages APIs.

Quick Start

Install — pick your preferred method:

# Docker
docker pull lightseekorg/smg:latest

# Kubernetes (Helm)
helm install smg oci://ghcr.io/smg-project/charts/smg

# Python
pip install smg

# Rust (needs protoc)
cargo install smg

Run — point SMG at your inference workers:

# Single worker
smg launch --worker-urls http://localhost:8000

# Multiple workers with cache-aware routing
smg launch --worker-urls http://gpu1:8000 http://gpu2:8000 --policy cache_aware

# With high availability mesh
smg launch --worker-urls http://gpu1:8000 --enable-mesh \
  --mesh-advertise-host 10.0.0.1 --mesh-peer-urls 10.0.0.2:39527

Use — send requests to the gateway:

curl http://localhost:30000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "llama3", "messages": [{"role": "user", "content": "Hello!"}]}'

That's it. SMG is now load-balancing requests across your workers.

Supported Backends

Self-Hosted EnginesvLLM · SGLang · TokenSpeed · TensorRT-LLM · MLX (Apple Silicon) · any OpenAI-compatible server (e.g. Ollama)
Cloud ProvidersOpenAI · Anthropic · Google Gemini · xAI · OCI Generative AI · AWS Bedrock · Azure OpenAI · any OpenAI-compatible provider (Groq, Together, …)

Features

FeatureDescription
10 Routing Policiescache_aware, least_load, power_of_two, consistent_hashing, prefix_hash, bucket, round_robin, random, manual, passthrough
gRPC PipelineNative streaming gRPC to the engines with prefill/decode and encode disaggregation and DP-aware routing
Kubernetes DiscoveryNative pod watchers with label selectors, per-role prefill/decode/encode selectors, and router peer discovery
Model Parsers21 tool-call parsers and 16 reasoning parsers with automatic model detection — DeepSeek, Qwen, Kimi, GLM, Llama, Mistral, Command, Nemotron, and more
MCP IntegrationTool discovery and execution over stdio, SSE, and streamable HTTP, with approval policies and audit logging
High AvailabilityMesh networking with SWIM gossip and CRDT-replicated state for multi-node deployments
Chat HistoryPluggable storage with schema migrations: PostgreSQL, Oracle, Redis, or in-memory
WASM PluginsExtend request and response handling with custom WebAssembly middleware
ResilienceCircuit breakers, retries with backoff and jitter, rate limiting, and priority admission scheduling

Documentation

Full documentation lives at lightseek.org/smg.

Getting StartedInstallation and first steps
ArchitectureHow SMG works
ConfigurationCLI reference and options
API ReferenceOpenAI-compatible endpoints
Kubernetes SetupIn-cluster discovery and production setup

Contributing

We welcome contributions! See the Contributing Guide for details.

// faq

What is smg?

Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.. It is open-source on GitHub.

Is smg free to use?

smg is open-source under the Apache-2.0 license, so it is free to use.

What category does smg belong to?

smg is listed under mcp-servers in the Claudeers registry of Claude-compatible tools.

12 views
★ 543 stars
unclaimed
updated about 1 month ago

// embed badge

smg on Claudeers
[![Claudeers](https://claudeers.com/api/badge/smg.svg)](https://claudeers.com/smg)

// retro hit counter

smg hit counter
[![Hits](https://claudeers.com/api/counter/smg.svg)](https://claudeers.com/smg)

// reviews

// guestbook

0/500

// related in MCP Servers

🔓

f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-host for your organization with complete…

// mcp-serversf/⟨HTML⟩★ 172,096◷ NOASSERTION[ claude ]
🔓

A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Gemini CLI & Hermes Agent. Only official website: ccswitch.io

// mcp-serversfarion1231/⟨Rust⟩★ 140,512◷ MIT[ claude ]
🔓

🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman

// mcp-serversJuliusBrussee/⟨JavaScript⟩★ 107,719◷ MIT[ claude ]
🔓

An open-source AI agent that brings the power of Gemini directly into your terminal.

// mcp-serversgoogle-gemini/⟨TypeScript⟩★ 107,167◷ Apache-2.0[ claude ]
→ see how smg connects across the ecosystem