
codex-router
External-model router for Codex with guided Kimi OAuth/API, DeepSeek, safe migration, and rollback.
Install with your AI
Paste into Claude Code, Cursor, or any agent — it reads the repo and wires the tool into your project.
Install and set up codex-router (git-clone project) into my current project. Found on https://claudeers.com/codex-router Repo: https://github.com/duolahypercho/codex-router Homepage/docs: — Detected install method: git-clone → git clone https://github.com/duolahypercho/codex-router Category: uncategorized. Platforms: cli, api, desktop, web. Read the repo's README for exact setup and env vars, then install it and wire it into my project. Claudeers Health Verdict: active; community-verified: false. Confirm the source before running anything.
git clone https://github.com/duolahypercho/codex-router
// compatibility
| Platforms | cli, api, desktop, web |
|---|---|
| Operating systems | — |
| AI compatibility | claude |
| License | MIT |
| Pricing | open-source |
| Language | JavaScript |
Codex Router
Use Anthropic, Kimi, DeepSeek, xAI, GitHub Copilot, opencode Go, Command Code, and future external models inside the Codex App and CLI — or inside DeepSeek Harness or Gemini CLI — through one local, credential-isolating router. The integration speaks the Responses API and merges external entries into Codex's native model catalog, so routed models appear in the normal picker next to the native GPT models. The same routed models publish into the harness as one provider route, so they appear in its Models page too, and into Gemini CLI through a Gemini-shaped endpoint the router serves for it.
Every client shares one installation: one background service, one gateway, one set of provider credentials, one provider selection. Installing a second or third integration does not ask for a single key again.
The router is also the source of truth for routed model policy. Provider/model
selection and external picker visibility are stored locally in the router state
directory (model-picker.json is an explicit allowlist: only router models you
show or select during curation are published), then republished to every
installed client. A signed-in Codex
installation keeps its native GPT catalog and native visibility client-owned;
the router never lets an external overlay erase that original picker. Codex's
active task remains in Codex configuration. Its default model does too unless
you explicitly opt into a router-owned routed default; that choice is saved
locally, survives rebuilds, and can be restored to the prior Codex default.
Codex Router is an independent community project. It is not affiliated with or endorsed by OpenAI, GitHub, Anthropic, Moonshot AI, DeepSeek, OpenRouter, opencode, Google, or the referenced opencodex project.
Give the link to your agent
Paste this into a Codex task:
Install the router from this public repository:
https://github.com/duolahypercho/codex-router
Follow AGENTS.md. Preserve my existing Codex models, profiles, settings, and
ChatGPT login. Use only the provider authentication I choose, safely migrate
only recognized older versions, run the Codex doctor, and leave the final app
restart to me. Never ask me to paste a token or API key into chat.
If compatible authentication already exists, an agent can finish everything except the final app restart. Provider credentials are entered only through a hidden local terminal prompt.
Install
Homebrew
If you already use Homebrew, install Codex Router from this repository's tap:
brew tap duolahypercho/codex-router https://github.com/duolahypercho/codex-router
brew install codex-router
codex-router setup --guided
The tap URL is needed only once. Homebrew installs the formula's Node.js,
Python, and build dependencies; codex-router setup --guided performs the
one-time provider selection, credential-safe authentication, background
service installation, and Codex integration. When setup finishes, fully quit
and reopen Codex, create a new task, and choose a routed model from the picker.
Upgrade an existing Homebrew installation with:
brew upgrade codex-router
Homebrew command equivalents
A Homebrew install puts a single codex-router command on your PATH instead
of this repository's bin/ directory. Wherever the rest of this README shows
./bin/model-router codex <command> or ./bin/<command>, run:
codex-router <command>
List everything the packaged build exposes with:
codex-router help
To add a custom provider's models — the packaged equivalent of
./bin/curate-models <provider> — run:
codex-router curate-models <provider>
codex-router install is deliberately unavailable: a Homebrew install has no
writable checkout to rewrite, and brew upgrade codex-router performs that
step itself.
Before removing the formula, remove the per-user service and managed Codex configuration that Homebrew does not own:
codex-router uninstall
brew uninstall codex-router
The first Homebrew install can take considerably longer than the guided
installer below because the formula builds the locked Python dependencies from
source. The release workflow generates Formula/codex-router.rb from
requirements/python.txt and refreshes it for each release.
Maintainers preparing the eventual homebrew/core submission should follow
docs/HOMEBREW_CORE.md.
Guided installer
macOS or Linux:
curl -fsSL https://raw.githubusercontent.com/duolahypercho/codex-router/main/install.sh \
| sh -s -- --target codex --guided
Windows PowerShell:
$installer = Join-Path $env:TEMP "codex-router-install.ps1"
Invoke-WebRequest https://raw.githubusercontent.com/duolahypercho/codex-router/main/install.ps1 -OutFile $installer
powershell.exe -NoProfile -ExecutionPolicy Bypass -File $installer -Target codex -Guided
The setup selects providers, detects existing authentication, can run the
official kimi login, prompts invisibly for provider credentials, installs a per-user
background service, and verifies every local layer. It never makes a paid test
request unless --smoke-test is explicitly selected.
To validate the install and uninstall lifecycle before trusting the router
with any credential, pass --no-provider --no-discovery: the router installs
idle, reads no credential from anywhere, and answers Codex traffic with a
local error. See docs/INSTALL.md.
Requirements:
- The Codex App or CLI.
- Node.js 22.19 or newer; Node.js 24 LTS is recommended.
uv, or Python 3.10+ withvenv.- Git for the managed one-command checkout and rollback.
Linux installations support the Codex CLI.
Models and authentication
| Picker label | Model ID | Authentication |
|---|---|---|
| K2.7 Coding Highspeed (OAuth) | kimi-oauth/kimi-for-coding-highspeed | Existing Kimi Code CLI OAuth session |
| K2.7 Coding (OAuth) | kimi-oauth/kimi-for-coding | Existing Kimi Code CLI OAuth session |
| Kimi K3 (OAuth) | kimi-oauth/k3 | Existing Kimi Code CLI OAuth session |
| Kimi K3 (API) | kimi-api/kimi-k3 | Separately billed Kimi Platform API key |
| Kimi K3 (China API) | kimi-api-cn/kimi-k3 | Separately billed Moonshot China platform key |
| DeepSeek V4 Flash (API) | deepseek/deepseek-v4-flash | DeepSeek API key |
| DeepSeek V4 Pro (API) | deepseek/deepseek-v4-pro | DeepSeek API key |
| Grok 4.5 (OAuth) | grok-oauth/grok-4.5 | Official Grok CLI OAuth session |
| Grok 4.5 (API) | grok-api/grok-4.5 | Separately billed xAI API key |
| Claude Opus 4.8 (API) | anthropic-api/claude-opus-4.8 | Separately billed Anthropic API key |
| GLM-5.2 (Ollama Cloud) | ollama-cloud/glm-5.2 | Ollama Cloud API key |
| Kimi K2.7 Code (Ollama Cloud) | ollama-cloud/kimi-k2.7-code | Ollama Cloud API key |
| MiniMax M3 (Ollama Cloud) | ollama-cloud/minimax-m3 | Ollama Cloud API key |
| DeepSeek V4 Pro (Ollama Cloud) | ollama-cloud/deepseek-v4-pro | Ollama Cloud API key |
| DeepSeek V4 Flash (Ollama Cloud) | ollama-cloud/deepseek-v4-flash | Ollama Cloud API key |
| MiniMax M3 | minimax-token-plan/minimax-m3 | MiniMax Token Plan API key |
| MiMo-V2.5 (Xiaomi API) | xiaomi-mimo/mimo-v2.5 | Xiaomi MiMo API key |
| MiMo-V2.5-Pro (Xiaomi API) | xiaomi-mimo/mimo-v2.5-pro | Xiaomi MiMo API key |
| Qwen3.8 Max (Plan) | qwen-plan/qwen3.8-max | Alibaba Model Studio plan API key |
| Qwen3.8 Max Preview (Plan) | qwen-plan/qwen3.8-max-preview | Alibaba Model Studio plan API key |
| Qwen3.7 Max (Plan) | qwen-plan/qwen3.7-max | Alibaba Model Studio plan API key |
| Qwen3.7 Plus (Plan) | qwen-plan/qwen3.7-plus | Alibaba Model Studio plan API key |
| Qwen3.6 Flash (Plan) | qwen-plan/qwen3.6-flash | Alibaba Model Studio plan API key |
| DeepSeek V4 Pro (Qwen Plan) | qwen-plan/deepseek-v4-pro | Alibaba Model Studio plan API key |
| DeepSeek V4 Flash (Qwen Plan) | qwen-plan/deepseek-v4-flash-0731 | Alibaba Model Studio plan API key |
| GLM-5.2 (Qwen Plan) | qwen-plan/glm-5.2 | Alibaba Model Studio plan API key |
| GLM-5.3 (Coding Plan) | zai-coding/glm-5.3 | Z.ai GLM Coding Plan API key |
| GLM-5.2 (Coding Plan) | zai-coding/glm-5.2 | Z.ai GLM Coding Plan API key |
| GLM-5-Turbo (Coding Plan) | zai-coding/glm-5-turbo | Z.ai GLM Coding Plan API key |
| GLM-5.3 (Z.ai API) | zai-api/glm-5.3 | Separately billed Z.ai platform API key |
| GLM-5.2 (Z.ai API) | zai-api/glm-5.2 | Separately billed Z.ai platform API key |
| GLM-4.7 (Z.ai API) | zai-api/glm-4.7 | Separately billed Z.ai platform API key |
| Muse Spark 1.2 (Meta) | meta/muse-spark-1.2 | Meta Model API key |
| Muse Spark 1.2 Contributor (Meta) | meta/muse-spark-1.2-contributor | Meta Model API key |
| Muse Spark 1.1 (Meta) | meta/muse-spark-1.1 | Meta Model API key |
| GLM-5.2 (ClinePass) | clinepass/glm-5.2 | ClinePass API key |
| Kimi K3 (ClinePass) | clinepass/kimi-k3 | ClinePass API key |
| Kimi K2.7 Code (ClinePass) | clinepass/kimi-k2.7-code | ClinePass API key |
| Kimi K2.6 (ClinePass) | clinepass/kimi-k2.6 | ClinePass API key |
| DeepSeek V4 Pro (ClinePass) | clinepass/deepseek-v4-pro | ClinePass API key |
| DeepSeek V4 Flash (ClinePass) | clinepass/deepseek-v4-flash | ClinePass API key |
| MiMo-V2.5 (ClinePass) | clinepass/mimo-v2.5 | ClinePass API key |
| MiMo-V2.5-Pro (ClinePass) | clinepass/mimo-v2.5-pro | ClinePass API key |
| MiniMax M3 (ClinePass) | clinepass/minimax-m3 | ClinePass API key |
| Qwen3.7 Max (ClinePass) | clinepass/qwen3.7-max | ClinePass API key |
| Qwen3.7 Plus (ClinePass) | clinepass/qwen3.7-plus | ClinePass API key |
| Qwen3.8 Max (ClinePass) | clinepass/qwen3.8-max | ClinePass API key |
Kimi has two API platforms and they are not interchangeable. kimi-api is the
global console at platform.moonshot.ai; kimi-api-cn is the mainland console at
platform.moonshot.cn. Accounts, billing, and keys are separate — a key minted on
one platform is rejected by the other — so each is enabled and credentialed on
its own, and both can be active at once. Pick the one matching where your key
was created. (kimi-oauth is a third, distinct thing: the Kimi Code
subscription reused through the official CLI's session.)
The Codex catalog is credential-aware. It includes models only from enabled
external providers with a stored credential or valid OAuth session. Native GPT
models are included only when codex login status confirms an OpenAI login.
Qwen is key-only. Alibaba discontinued the Qwen Code OAuth free tier on
2026-04-15, so the Model Studio plan key is the sole Qwen surface; qwen-plan
points at the token-plan endpoint. Set QWEN_PLAN_BASE_URL to
https://dashscope-intl.aliyuncs.com/compatible-mode/v1 to bill a
pay-as-you-go DashScope key through the same provider. Alibaba publishes no
quota or balance API on either endpoint, so the tray shows router-observed
traffic and links to the console for actual spend.
ClinePass uses Cline's OpenAI-compatible API at
https://api.cline.bot/api/v1. An API key alone does not grant access to the
cline-pass/* models: the account also needs an active ClinePass subscription.
Create the key under Cline Settings > API Keys, then store it with
./bin/model-router codex provider-key clinepass set.
Grok OAuth reuses the official CLI credential at ~/.grok/auth.json and sends
it only to xAI's documented Grok CLI inference proxy. On that path the router
also attaches bare hosted web_search and x_search tools, the same agentic
surface Grok Build uses. xAI's backend chooses when to search and how to filter
results; the router does not take search env knobs or request-side filter
config. Install the official CLI and authenticate before enabling the route:
Other routed providers can use Codex's client-side (standalone) web search when
the selected model has been verified for it. DeepSeek V4 Flash is enabled on
its direct API and opencode Go routes. A compatible model declares
"searchTool": { "mode": "standalone" } in its registry or user-model
metadata, and the managed Codex provider table advertises
supports_standalone_web_search = true. This is intentionally opt-in per
model; the router does not infer search compatibility from an OpenAI-compatible
endpoint.
npm install -g @xai-official/grok
grok login --oauth
Antigravity OAuth uses the router's own browser sign-in and the Google AI Pro/Ultra entitlement on the signed-in account. It needs neither a Gemini API key nor a separate Antigravity CLI. Signing in and enabling are separate so a re-authentication never replaces the rest of the provider selection:
Antigravity OAuth requires an integration client secret. Set
ANTIGRAVITY_CLIENT_SECRET in the environment used for installation and
sign-in; the generated background-service definition preserves it for token
refreshes. The command fails before opening Google consent when it is absent.
export ANTIGRAVITY_CLIENT_SECRET='your-integration-client-secret'
./bin/model-router codex providers login antigravity-oauth
./bin/model-router codex providers enable antigravity-oauth
On Windows PowerShell, use the matching wrapper:
$env:ANTIGRAVITY_CLIENT_SECRET = 'your-integration-client-secret'
.\model-router.ps1 codex providers login antigravity-oauth
.\model-router.ps1 codex providers enable antigravity-oauth
The credential stays in the router's owner-only state directory. This is an unofficial compatibility route over Google's internal Antigravity service, not a public Gemini API contract, so availability and wire behavior can change.
MiMo (Xiaomi API) uses Xiaomi's official OpenAI-compatible endpoint at
https://api.xiaomimimo.com/v1. Unlike MiMo reseller routes, the direct API
serves mimo-v2.5 and mimo-v2.5-pro through the standard
/chat/completions surface, so requests never touch the Responses gateway.
mimo-v2.5 is verified for text/image input and Codex standalone web search;
mimo-v2.5-pro is text-only. Store the key with
./bin/model-router codex provider-key xiaomi-mimo set.
Native GPT models continue to use Codex directly. There is no separate GPT or ChatGPT OAuth provider in the router.
GitHub Copilot
github-copilot routes account-visible models that explicitly advertise the
Responses API, streaming, and tool calls. The catalog is plan- and
policy-specific, so this provider ships no hard-coded models: store a
fine-grained GitHub PAT with the Copilot Requests permission, then curate
from the live catalog. This initial integration targets GitHub.com; GitHub
Enterprise Cloud data-residency hosts are not yet configured by the router.
./bin/model-router codex provider-key github-copilot set
./bin/curate-models github-copilot
The hidden prompt stores the GitHub token in protected router state. For a
foreground process, COPILOT_GITHUB_TOKEN, GH_TOKEN, and GITHUB_TOKEN are
checked in that order. Classic ghp_ tokens are not supported by Copilot;
create a fine-grained github_pat_ token
at GitHub personal access tokens.
The router deliberately does not read or copy the official Copilot CLI's
credential store.
At request time the GitHub credential is validated through the Copilot account endpoint, which also selects the account's inference host. That host is accepted only when it is GitHub-owned. The tray reads the account's AI-credit or legacy request quota when GitHub exposes a per-user meter; organization-managed plans that expose no per-seat quota fall back to router-observed traffic.
GitHub documents the PAT permission and Copilot clients, while the inference interface may continue to evolve. Requests consume the user's Copilot allowance; use it within the GitHub Copilot terms and acceptable use policies.
Kimi Code OAuth and Kimi Platform API access are separate authentication and billing systems. The two Kimi entries intentionally coexist. Older DeepSeek aliases remain hidden compatibility routes and are not advertised to new users.
The Ollama Cloud entries bill through an ollama.com account and can host the
same model families as other providers under a separate quota. Matching entries
(for example DeepSeek V4 Pro) intentionally coexist with the vendor-direct
providers because credentials and billing differ.
The Qwen plan entries cover every chat model the Individual Plan serves,
including the cross-vendor models it resells (DeepSeek V4 and GLM-5.2) under
the same plan key and quota. The cross-vendor entries use DashScope's
compatible-mode request profile because DashScope rejects each vendor's native
thinking parameters.
The Qwen entries default to the Alibaba Model Studio Token Plan endpoint in
the Singapore region. Coding Plan subscribers or other regions can point
QWEN_PLAN_BASE_URL at their dashboard-issued base URL. Plan keys use the
sk-sp- prefix and are separate from pay-as-you-go Model Studio keys; Alibaba
reserves plan endpoints for interactive coding tools.
The zai-coding entries use the GLM Coding Plan's dedicated endpoint and its
subscription API key. That key is not interchangeable with general Z.ai
platform keys, and Z.ai reserves the coding endpoint for interactive coding
tools. The metered platform is therefore a separate provider, zai-api, on
https://api.z.ai/api/paas/v4 with its own key file and its own environment
variable (ZAI_PLATFORM_API_KEY, never the plan's ZAI_API_KEY) — connecting
one does not connect the other. GLM-5.3 ships on both routes with Z.ai's
documented low/high/max reasoning tiers and a one-million-token context
window. The [1m] model suffix that circulated for GLM-5.3 does not exist on
either Z.ai endpoint -- both the OpenAI-compatible coding route and the
Anthropic route reject glm-5.3[1m] with error 1214 -- and it was never
needed: a live run accepted 990,020 prompt tokens on the plain glm-5.3
code.
Beyond the built-in models, each API-key provider's live catalog can be
curated interactively: ./bin/curate-models PROVIDER lists the models the
provider currently advertises that are not in the registry, lets you toggle
the ones you want, and stores them as user models in protected state
(surviving updates, editable in place, and removable by re-running the
command and deselecting). Curation asks for each new model's context window,
image support, and reasoning efforts — so curated models get the effort
switcher in the picker — and everything defaults conservatively when
unanswered. The context window is not guessed when the provider publishes one:
the context_length its catalog advertises for the model is offered as the
default and stored by both curation forms, so a million-token model is not
filed as a 131K one and told to compact at 110K. The non-interactive
--models id1,id2 form is additive: it keeps
existing curated entries and their metadata while adding the named models;
--efforts minimal,low,medium,high,xhigh sets the new entries' ladder. Remove
entries explicitly with --remove id1,id2. Every value stays editable in
user-models.json. Curation also asks whether the model rejects a forced
tool_choice: a few upstreams call tools happily when the choice is auto
but answer HTTP 400 when one is required, which fails the compatibility check
and the routed-subagent handoff even though tool calling works. Answering yes
stores "requestProfile": "auto-tool-choice", and the router downgrades the
forced choice for that model only (--request-profile auto-tool-choice in the
--models form). The provider's own /v1/models endpoint always decides
which models exist. Curated models are local to your machine and are not
vetted by the repository's compatibility tests.
opencode (Go subscription and Zen)
The opencode provider family covers both of opencode's endpoints with one
stored API key (OPENCODE_API_KEY or OPENCODE_GO_API_KEY in the
environment): the flat-rate Go subscription at
https://opencode.ai/zen/go/v1, whose tested models ship in the registry
below, and the pay-per-use Zen endpoint at https://opencode.ai/zen/v1,
whose larger catalog is available through local curation
(./bin/curate-models opencode-zen). Everything appears as a single
"opencode Go/Zen" provider; internally the catalog is split across provider
IDs by
endpoint and by the protocol each model speaks upstream. Set the key once and
enable the family:
./bin/model-router codex provider-key opencode-go set
./bin/model-router codex providers enable opencode-go
The desktop panel and macOS tray Settings tab provide both per-model controls and provider-level Select all / Unselect all actions for which registry-proven v2 models can run as subagents and which models appear in installed client pickers. Local settings cannot promote an unverified model. Fully quit and reopen Codex after changing either list; DeepSeek Harness hot-reloads its route, and the next Gemini CLI invocation reads the new environment.
| Picker label | Model ID |
|---|---|
| Grok 4.5 (opencode Go) | opencode-go/grok-4.5 |
| GLM-5.3 (opencode Go) | opencode-go/glm-5.3 |
| GLM-5.2 (opencode Go) | opencode-go/glm-5.2 |
| GLM-5.1 (opencode Go) | opencode-go/glm-5.1 |
| Kimi K3 (opencode Go) | opencode-go/kimi-k3 |
| Kimi K2.7 Code (opencode Go) | opencode-go/kimi-k2.7-code |
| Kimi K2.6 (opencode Go) | opencode-go/kimi-k2.6 |
| DeepSeek V4 Pro (opencode Go) | opencode-go/deepseek-v4-pro |
| DeepSeek V4 Flash (opencode Go) | opencode-go/deepseek-v4-flash |
| MiMo-V2.5 (opencode Go) | opencode-go/mimo-v2.5 |
| MiMo-V2.5-Pro (opencode Go) | opencode-go/mimo-v2.5-pro |
| Hy3 (opencode Go) | opencode-go/hy3 |
| MiniMax M3 (opencode Go) | opencode-go-messages/minimax-m3 |
| MiniMax M2.7 (opencode Go) | opencode-go-messages/minimax-m2.7 |
| Qwen3.8 Max (opencode Go) | opencode-go-messages/qwen3.8-max |
| Qwen3.7 Max (opencode Go) | opencode-go-messages/qwen3.7-max |
| Qwen3.7 Plus (opencode Go) | opencode-go-messages/qwen3.7-plus |
| Qwen3.6 Plus (opencode Go) | opencode-go-messages/qwen3.6-plus |
| GPT 5.6 Luna (opencode Go) | opencode-go-responses/gpt-5.6-luna |
opencode-go carries the Chat Completions models, opencode-go-messages the
Anthropic Messages models, opencode-go-responses the Responses models, and
opencode-zen the pay-per-use Zen endpoint (no preselected models — curate
the ones you want). All four are one selectable family: they share a single
stored key, and enabling or disabling any of them toggles all of them
together.
Entries that duplicate a vendor-direct provider (for example DeepSeek V4 Pro)
intentionally coexist because the subscription bills separately. Point
OPENCODE_GO_BASE_URL (or OPENCODE_ZEN_BASE_URL) elsewhere to override the
endpoints.
Anonymous free model gateways
Two additional entries use providers' documented free-model exceptions. Neither asks for an API key, neither is ever selected on your behalf, and each is pinned in code to its official endpoint.
| Picker label | Provider ID | Endpoint | Free-model rule |
|---|---|---|---|
| OpenCode Free | opencode-free | https://opencode.ai/zen/v1 | big-pickle and IDs ending in -free |
| Kilo Free | kilo-free | https://api.kilo.ai/api/gateway | IDs ending in :free |
Neither ships its free subset as checked-in metadata, with a single exception:
Ox Alpha on OpenCode Free is checked in — see Ox Alpha below.
Everything else comes from the provider's live /models response, filtered to
the free subset and then added locally with ./bin/curate-models. OpenCode Free
curation routes muse-spark-1.2-contributor-free through its internal Responses
sibling while keeping Ox Alpha Free (x-preview-f-free) and the other free IDs
on Chat Completions; the provider remains one selection in setup and the picker.
An existing Chat-routed copy of that one Muse model is migrated only when the
operator explicitly runs curate-models; install, update, and catalog reads do
not rewrite the user model or picker state. Zen's /models response publishes
no context limits, so those two IDs are sized from OpenCode's own published
per-free-ID metadata instead of the conservative 131K fallback, and each stored
entry's description records where its window came from. Every other free ID
keeps the conservative default, and any window is editable in
user-models.json.
./bin/model-router codex providers enable opencode-free
./bin/curate-models opencode-free
./bin/model-router codex providers enable kilo-free
./bin/curate-models kilo-free
OpenCode Console documents that free chat models can omit the bearer header;
the paid Console models still require a key. Kilo documents anonymous access
only for :free models and limits anonymous traffic to 200 requests per hour
per IP. Both catalogs and limits are provider-controlled and can change, so
the router refuses paid IDs and shows traffic-only usage when no quota header
has been observed. Kilo's general SDK setup guide still asks external SDK
users for an API key; this entry intentionally covers only the gateway's
documented anonymous :free path.
Custom: one provider, many endpoints
Every other provider owns one address. custom owns none — each of its models
names its own endpoint, its own auth, and its own metadata, so a single picker
entry can hold a free community endpoint, a friend's self-hosted server, and a
paid API you have a key for, all at once.
./bin/model-router codex providers enable custom
Enabling it costs nothing and asks for nothing: a model that needs a key says so on its own row. It is never selected for you and never part of the default set, because what it holds is whatever somebody put in it.
| Model | Endpoint | Auth |
|---|---|---|
| Qwen3.8-27-free-victor | https://g9hnto0u7lvbu837.us-east-2.aws.endpoints.huggingface.cloud/v1 | none |
That first model is a free community Hugging Face Inference
Endpoint for
Qwen/Qwen3.8-27B, published by an individual rather than by Qwen or Hugging
Face: BF16 on one H200 behind vLLM, 262,144-token context, image input, tool
calling, and a thinking budget you dial with the normal effort picker. It is
shared and rate limited to roughly 30 requests per minute per IP, and its owner
says it will be retired once launch interest fades — so treat it as a model to
try, not one to depend on.
An endpoint reached with no credential is the one thing a registry fragment
cannot introduce on its own. Its address has to be allowlisted in
src/model-registry.mjs, exactly as an anonymous provider's is, because
otherwise adding a JSON file under config/custom/ would be enough to send your
prompts to any host on the internet with nothing to authenticate them. An
endpoint that carries a key, or one that stays on loopback, needs no allowlist
entry — the key or the address is already the boundary.
Use these at your own risk. The two gateways above, and any
custommodel whose endpoint carries no credential, are the only routes here that reach an upstream with no account behind them, and that changes what "supported" can mean. Nobody has agreed to serve you: access is a published exception, not an entitlement, and it can be narrowed, rate-limited, or withdrawn without notice. On the two reseller gateways the naming rule is a heuristic rather than a promise — their catalogs carry no pricing field to check, so a model whose ID saysfreecan still answer401 Paid inference requests require an Authorization bearer token, and the router cannot tell in advance. Anonymous traffic is identified by IP, so a router fanning out parallel subagents spends a budget shared with everyone behind that address. Treat these as a way to try a model, not as something to depend on: nothing in this repository can keep them working, and a failure here is not a bug the project can fix.
Command Code
Command Code's official Provider API is an OpenAI-compatible chat completions
surface plus an Anthropic Messages surface at https://api.commandcode.ai/provider/v1
(COMMAND_CODE_API_KEY or COMMANDCODE_API_KEY in the environment, or store
the key once). Every plan except Go has API access; GOAT, Pro, Max, Team, and
Provider accounts use the API. Everything appears as one
"Command Code" provider; internally the catalog is split between
commandcode for Chat Completions models and commandcode-messages for
models that require the Messages protocol (Claude).
The Go plan is the exception. A Go-plan account is refused by /provider/v1
with Your Go plan doesn't include API access. That is an entitlement, not a
credential problem: no key or reinstall changes it. Check the plan
at commandcode.ai/billing before enabling
this provider.
Store an API key. Create one in Command Code Studio and save it here:
./bin/model-router codex provider-key commandcode set
./bin/model-router codex providers enable commandcode
When multiple API-key sources exist, the exported environment variable wins,
then the key stored here, then the macOS Keychain. doctor names whichever
source is live. The router does not install, launch, or read a Command Code
CLI session.
| Picker label | Model ID |
|---|---|
| Ox Alpha (Command Code) | commandcode/ox-alpha |
| DeepSeek V4 Flash (Command Code) | commandcode/deepseek-v4-flash |
| DeepSeek V4 Pro (Command Code) | commandcode/deepseek-v4-pro |
| GLM-5.2 (Command Code) | commandcode/glm-5.2 |
| Kimi K3 (Command Code) | commandcode/kimi-k3 |
| Kimi K2.7 Code (Command Code) | commandcode/kimi-k2.7-code |
| Qwen3.8 Max (Command Code) | commandcode/qwen3.8-max |
| Qwen3.7 Max (Command Code) | commandcode/qwen3.7-max |
| Qwen3.7 Plus (Command Code) | commandcode/qwen3.7-plus |
| MiniMax M3 (Command Code) | commandcode/minimax-m3 |
| MiniMax M2.7 (Command Code) | commandcode/minimax-m2.7 |
| MiMo-V2.5-Pro (Command Code) | commandcode/mimo-v2.5-pro |
| Grok 4.5 (Command Code) | commandcode/grok-4.5 |
| GPT 5.6 Luna (Command Code) | commandcode/gpt-5.6-luna |
| GPT 5.5 (Command Code) | commandcode/gpt-5.5 |
| Gemini 3.5 Flash (Command Code) | commandcode/gemini-3.5-flash |
| Hy3 (Command Code) | commandcode/hy3-paid |
| Step 3.7 Flash (Command Code) | commandcode/step-3.7-flash |
| Claude Sonnet 5 (Command Code) | commandcode-messages/claude-sonnet-5 |
| Claude Opus 4.8 (Command Code) | commandcode-messages/claude-opus-4.8 |
| Claude Fable 5 (Command Code) | commandcode-messages/claude-fable-5 |
| Claude Haiku 4.5 (Command Code) | commandcode-messages/claude-haiku-4.5 |
Both entries are one selectable family that shares a single stored key;
enabling or disabling either toggles the whole family together. The live
catalog is available without authentication from
https://api.commandcode.ai/provider/v1/models, and additional models can be
added per machine with ./bin/curate-models commandcode. Point
COMMANDCODE_BASE_URL elsewhere to override the endpoint — both routes follow
it, so a redirected provider stays coherent. The tray reports the plan's
remaining credits and its 5-hour and weekly windows from the same undocumented
billing route the official CLI polls, and links to Command Code Studio when
that route is unavailable.
Ox Alpha
Ox Alpha is a stealth reasoning model for coding and long-horizon agentic work: a 1,048,576-token context window, 131,072 tokens of output, text and image input, and tool calling. Six of this repository's routes resell the same model, and it is priced at zero on all of them during the preview, so the entries carry a Free badge in the control center.
| Picker label | Model ID | Needs a key |
|---|---|---|
| Ox Alpha (OpenCode Free) | opencode-free/ox-alpha | no |
| Ox Alpha (opencode Go) | opencode-go/ox-alpha | opencode |
| Ox Alpha (OpenRouter) | openrouter/ox-alpha | OpenRouter |
| Ox Alpha (Command Code) | commandcode/ox-alpha | Command Code |
| Ox Alpha (Nous Research) | nousresearch/ox-alpha | Nous Portal |
| Ox Alpha (Venice) | venice/ox-alpha | Venice |
Reasoning effort is low · high · max on every route, defaulting to max.
Only three rungs exist because the model always thinks and its upstream says so
outright — anything else comes back as 400 — This model always engages in thinking and cannot be disabled; please use low, high, or max. Codex has more
rungs than that, and a Codex older than 0.143 has no max at all, so the router
clamps whatever effort you pick onto the three the model accepts. Switching
effort in the picker is safe on all six routes.
The quickest route needs nothing at all:
./bin/model-router codex providers enable opencode-free
For the credentialed routes, store the key and enable the provider:
./bin/model-router codex provider-key venice set
./bin/model-router codex providers enable venice
The free preview is a preview. No lab has claimed this model, the routes that serve it can narrow or withdraw it without notice, and the retention terms differ per provider — OpenCode advertises zero data retention, Venice anonymizes, and other resellers say less. Treat it as a way to try a model, not as something to depend on.
Meta Model API
Meta's Muse Spark models speak the Responses protocol at
https://api.meta.ai/v1 (META_API_KEY in the environment, or store the key
once):
./bin/model-router codex provider-key meta set
./bin/model-router codex providers enable meta
Three Muse Spark models ship in the registry: 1.2 and its cheaper
Contributor tier (whose inputs and outputs Meta may use for training) with a
1M context window, reasoning efforts from minimal to xhigh, and reasoning
summaries enabled, plus the previous-generation 1.1. Additional Meta models
can be added per machine with ./bin/curate-models meta. Point
META_BASE_URL elsewhere to override the endpoint.
Catalog-only providers
These OpenAI-compatible providers are registered for routing and credential isolation but ship no preselected models, because their catalogs change too often for the repository to pin and live-verify individual entries:
| Provider | Provider ID | Base URL |
|---|---|---|
| Groq | groq | https://api.groq.com/openai/v1 |
| Together AI | together | https://api.together.xyz/v1 |
| Fireworks AI | fireworks | https://api.fireworks.ai/inference/v1 |
| Cerebras | cerebras | https://api.cerebras.ai/v1 |
| Mistral AI | mistral | https://api.mistral.ai/v1 |
| NVIDIA NIM | nvidia-nim | https://integrate.api.nvidia.com/v1 |
| SiliconFlow | siliconflow | https://api.siliconflow.cn/v1 |
| Hugging Face Router | huggingface | https://router.huggingface.co/v1 |
| Google Gemini API | gemini-api | https://generativelanguage.googleapis.com/v1beta/openai |
| GitHub Copilot | github-copilot | Account-specific GitHub Copilot endpoint |
| Chutes | chutes | https://llm.chutes.ai/v1 |
| OrcaRouter | orca | https://api.orcarouter.ai/v1 |
Three more providers work the same way but arrive with the single checked-in Ox Alpha entry, so their picker is not empty once a key is stored:
| Provider | Provider ID | Base URL | Key from |
|---|---|---|---|
| OpenRouter | openrouter | https://openrouter.ai/api/v1 | openrouter.ai/settings/keys |
| Venice | venice | https://api.venice.ai/api/v1 | venice.ai/settings/api |
| Nous Research (Hermes) | nousresearch | https://inference-api.nousresearch.com/v1 | portal.nousresearch.com |
Venice API access is an entitlement, not just a key: a free Venice account has none. A Pro subscription (the low-rate-limit Explorer tier), a funded USD balance, or staked VVV that grants VCU is what makes the key usable, and the router prints that requirement wherever you connect the provider rather than letting it arrive as a 403 inside Codex. Nous Research keys are Nous Portal API keys and authenticate the same endpoint the Hermes agent uses.
Add a key, then pick the models you want from the provider's live catalog:
./bin/model-router codex provider-key groq set
./bin/curate-models groq
OrcaRouter's public catalog includes paid models and concrete zero-price model
deployments. Inference still requires an OrcaRouter API key, including for free
models. The moving orcarouter/free meta-router is intentionally not curated:
the picker shows the concrete model identity with a Free badge instead. To
add every currently advertised free OpenAI-compatible model without pinning
that changing list in the repository:
./bin/model-router codex provider-key orca set
./bin/curate-models orca --free-only --apply
The free list is read live from OrcaRouter's /models response. Re-run the
command when its catalog changes, and verify a curated model with
./bin/test-model 'orca/MODEL_ID' --live --yes before relying on it for
tool-driven work.
Curated entries use the context window, image support, and reasoning efforts you provide during curation — the context window falling back to the one the provider's catalog advertises, and to a conservative default only when it advertises none — and are local to your machine. Verify a model before relying on it:
./bin/test-model 'groq/MODEL_ID' --live --yes
Each base URL is overridable through the provider's baseUrlEnv variable, so a
regional endpoint or a self-hosted gateway can reuse the same provider entry.
Quota cards work for these providers without any extra configuration. Most
OpenAI-compatible services report the caller's remaining window on every
response through x-ratelimit-* headers, and Anthropic reports the same facts
under an anthropic-ratelimit-* prefix. The router reads those headers as
traffic passes through, so a provider starts showing real request and token
limits after its first request — no balance endpoint, no extra API call, and no
separate credential. Providers that publish no such headers, including Google
Gemini, keep showing router traffic only.
Gemini is routed through Google's OpenAI-compatible surface rather than the
native Gemini protocol, so it shares the existing forwarder and needs no
separate adapter.
Only explicitly selected router models from enabled providers appear in installed client pickers. Adding a model during curation selects it for the picker; merely enabling a provider does not flood the list:
./bin/model-router codex providers
./bin/model-router codex providers enable deepseek
./bin/model-router codex provider-key deepseek set
./bin/model-router codex provider-key anthropic-api set
On Windows, use ./model-router.ps1 codex with the same commands.
Router-owned default model (optional)
In a normal signed-in Codex installation, you can opt into an external router model as the default for new tasks. The model must already be selected for the picker. The router snapshots the prior Codex default, reapplies your router choice after an update or repair, and restores that prior default when cleared:
./bin/control router-default set deepseek/deepseek-v4-flash
./bin/control router-default clear
This is separate from login-free mode, which has always owned its routed default. Fully quit and reopen Codex after changing either default.
The API-key prompt disables terminal echo. Protected files use mode 600 on
POSIX and an inheritance-disabled, current-user ACL on Windows. Diagnostics
report credential presence and source, never the value.
Make models appear in Codex
After setup:
- Run
./bin/model-router codex doctorand resolve anyFAILline. - Confirm
providerssaysSHOWandreadyfor the intended provider. - Fully quit Codex, reopen it, and create a new task.
- Open the normal model picker.
Codex loads model_catalog_json only at app startup. If models are still
missing, run ./bin/refresh-catalog, fully quit Codex, and reopen it.
Large compressed Codex contexts use separate safety limits for bytes received
on the loopback socket and bytes produced after decompression. The defaults are
64 MiB encoded and 256 MiB decoded. Override them with
MODEL_ROUTER_MAX_BODY_BYTES and MODEL_ROUTER_MAX_DECODED_BODY_BYTES
respectively when a deliberately larger local workload requires it.
For routed external models, old textual tool results larger than 32 KiB are compacted after the model has acted on them. The four newest tool results stay intact, and each compacted result keeps a hash, head/tail evidence, and an exact rerun instruction.
This is off by default. It rewrites what the model sees mid-conversation, so it is opted into rather than discovered after it has already altered a session. Turning it on is remembered: a stored answer is kept verbatim and is never re-defaulted by a later release.
Toggle Compact old tool results in the router Settings;
the next external-model request sees the change without restarting Codex or the
router. The equivalent CLI commands are ./bin/control tool-result-aging on,
off, and status.
When the estimated request reaches 70% of that model's auto-compact budget, the same switch automatically enters token maxxing for the turn. It applies a small deterministic output shaper inspired by RTK: terminal progress rewrites, exact repeated lines, blank runs, and deep boilerplate are collapsed while error-bearing lines stay visible. The newest-result frontier remains intact below that pressure threshold. Under pressure, every shaped result carries its original byte count, SHA-256 digest, and an exact rerun instruction, and the router adds a terse execution overlay inspired by Caveman so the model favors targeted reads, bounded command output, and concise prose. Routed compaction requests use the same dense shaping because they are already at the context boundary. No second toggle or restart is required.
Native OpenAI traffic is unchanged by default. ./bin/control tool-result-aging native on extends the same compaction to native GPT models;
native off restores the default. It is opt-in because it changes what is sent
to OpenAI's own endpoint, and an install that has never run it keeps the
pre-existing behavior. Set CODEX_ROUTER_TOOL_RESULT_AGING=0 for a hard
environment-level override that disables both the routed and the native path.
Where compaction parks the exact original bytes of a result it rewrote, they go
to an owner-private store at <state dir>/retained-tool-results (override with
MODEL_ROUTER_TOOL_RESULT_RETENTION_DIR). Nothing evicts that store, so both a
way to see it and a way to empty it are part of the feature:
./bin/doctor # count, size, oldest entry, TTL
./bin/control tool-result-aging purge # says what it would remove
./bin/control tool-result-aging purge --yes # removes it
./bin/control tool-result-aging purge --expired # only what the TTL outlived
./bin/control tool-result-aging ttl 30 # keep retained results 30 days
./bin/control tool-result-aging ttl off # keep them until purged
./bin/control tool-result-aging ttl default # back to 7 days
The doctor row appears whether or not the store exists, because an install that
has never retained anything is the answer most people should see and seeing it
is how the directory becomes discoverable at all. The purge is a report by
default: without --yes it prints what it would remove and removes nothing, and
--dry-run says the same thing explicitly and outranks --yes. It removes only
files this store wrote, only inside that one directory, never recursing and
never following a symlink out of it; anything else that ends up there is left in
place and named.
Retained results expire after 7 days. Nothing ever reads those bytes back
into a turn — the receipt tells the model to repeat the tool call — so a
retained original's only reader is you, and only while the session that produced
it still matters. A week is also what keeps the store's caps from becoming
permanent: at 512 files or 512 MiB retention stops accepting new results, and
with a TTL that state drains by itself instead of waiting for somebody to notice
it. Nothing sweeps on a timer: the store expires when it is next written to, and
purge --expired runs the same sweep by hand, with the same --yes consent and
the same containment as a full purge. The key that binds the store to this
install is never expired, only purged. ttl off keeps everything until an
explicit purge and is remembered verbatim, and the
CODEX_ROUTER_TOOL_RESULT_AGING=0 kill switch does not disable expiry — it
stops the router rewriting context, while expiry is disk hygiene for bytes that
are already written.
To estimate the effect without spending provider quota, run:
node scripts/measure-tool-result-aging.mjs /path/to/rollout.jsonl
The report compares each observed compaction boundary and the latest history
before and after aging; this is an estimate and spends no provider quota.
node scripts/aging-benchmark.mjs reports the savings already recorded in
usage-events.jsonl — measured turns rather than an estimate. For a
live check, leave the setting on and inspect usage-events.jsonl after a routed
turn; events that compacted history include toolResultsAged and
toolResultBytesSaved. Pressure-shaped turns additionally include
toolResultsShaped and toolResultShapeBytesSaved. Those counters measure
serialized context bytes, while provider-billed token counts remain the
authoritative cost measurement.
For a reproducible provider-reported A/B, see
docs/tool-result-aging-benchmark.md.
The integration preserves the built-in OpenAI provider, native GPT models, ChatGPT sign-in, profiles, MCP settings, project trust, and reasoning defaults. It adds one marked root block and one inert custom-provider table to the user's Codex config:
# BEGIN codex-router-managed
openai_base_url = "http://127.0.0.1:4202/_codex-router/<generated-capability>/v1"
model_catalog_json = "/absolute/path/to/.codex/codex-router/merged-models.json"
# END codex-router-managed
# BEGIN codex-router-provider-managed
[model_providers.codex-router]
name = "Codex Router (external models)"
base_url = "http://127.0.0.1:4202/_codex-router/<generated-capability>/v1"
wire_api = "responses"
supports_standalone_web_search = true
# END codex-router-provider-managed
The generated path is local caller authentication. Do not paste the complete managed URL into an issue.
Run GPT-5.6 Sol at its documented 1M context window
OpenAI documents GPT-5.6 Sol at 1,050,000 tokens. The catalog Codex ships
declares 272,000, and it has moved more than once
(openai/codex#31860,
#32806). The single-install
answer is model_context_window and model_auto_compact_token_limit in
~/.codex/config.toml; the router's answer is a second entry in the picker, so
the choice is per task rather than per machine:
| Picker label | Model ID | Context window | Auto-compaction |
|---|---|---|---|
| GPT-5.6-Sol (1M context) | gpt-5.6-sol-1m | 1,000,000 | 900,000 |
It is the same upstream model. Everything else in the entry — instructions,
reasoning ladder, image input, subagent behavior — is copied from
gpt-5.6-sol, and the router rewrites the slug back before the turn leaves for
chatgpt.com, so OpenAI only ever sees the model it published.
It ships switched off, because it costs more than the model it shadows: a turn resends the whole conversation, and a request above 272,000 input tokens is billed at a higher rate in full. Switch it on under OpenAI in the router Settings model list, or:
./bin/control picker set gpt-5.6-sol-1m show # and `hide` to put it back
Your answer is remembered. Later catalog rebuilds never re-apply the default to a model you have already decided, in either direction. Fully quit and reopen Codex afterwards — the picker is read at startup.
A login-free install does not get this entry: signed-out Codex only displays native slugs from a server-supplied allowlist, and a slot spent on a synthesized slug is a slot a routed model does not get.
Windows Codex Desktop running through WSL
When Codex Desktop runs on Windows while commands are executed through WSL, there may be two different Codex home directories:
C:\Users\<WindowsUser>\.codex
and:
/home/<LinuxUser>/.codex
Router commands use the Codex home selected by CODEX_HOME. Running them inside
WSL without overriding that variable may update the Linux CLI configuration
instead of the configuration used by Windows Codex Desktop.
To target the Windows Desktop configuration from WSL:
export CODEX_HOME=/mnt/c/Users/<WindowsUser>/.codex
export CODEX_ROUTER_STATE_DIR="$CODEX_HOME/codex-router"
Then run the router command normally. For example, to return to authenticated mode with native GPT models and enabled external providers in the merged catalog:
_…[view the full README on GitHub](https://github.com/duolahypercho/codex-router)._
// faq
What is codex-router?
External-model router for Codex with guided Kimi OAuth/API, DeepSeek, safe migration, and rollback.. It is open-source on GitHub.
Is codex-router free to use?
codex-router is open-source under the MIT license, so it is free to use.
What category does codex-router belong to?
codex-router is listed under uncategorized in the Claudeers registry of Claude-compatible tools.
// embed badge
[](https://claudeers.com/codex-router)
// retro hit counter
[](https://claudeers.com/codex-router)
// reviews
// guestbook
// related in Uncategorized / Others
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
The agent engineering platform.
FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Q…
100+ AI Agent & RAG apps you can actually run — clone, customize, ship.