A switch on every
coding request,
flipped intelligently.

Stay in complete control while every request routes to the cheapest model capable of the task, across only the providers you allow.

Connects to Claude Code, Cursor, and Codex.

your coding agentTokenSwitchclassify · routeDeepSeek V4 Flashtrivial · quick editsKimi K3economy · cheapest capableGLM 5.2balanced · most tasksSonnet 5strong · complex workClaude Opus 4.8frontier · escalate when needed
One-command setup

Implement with one command

No SDK, no code changes. One command points Claude Code or Codex at TokenSwitch, and every request is routed, capped, and logged from your next prompt on.

$ terminal
# Point Claude Code at TokenSwitch  (Codex: add --tool codex)
npx tokenswitch-setup@latest ts_setup_your_code

 Redeemed your setup code and provisioned a token
 Claude Code now routes every request through TokenSwitch

$ 
Control without slowing teams

Governance for every
AI coding dollar

Set the rules once. Every developer’s agent inherits your policy for providers, budgets, privacy, and visibility.

See the controls →
Data residency

Route through your own Bedrock or OpenRouter account, so data stays in a boundary you control.

Per-developer visibility

Spend, usage, and model mix per developer and repo, with no agent overhead.

Privacy by default

Prompts and source are never stored. Only privacy-safe metadata leaves your environment.

Budgets & spend caps

Hard caps per org, team, or developer. Over-budget requests are refused before a provider is called.

How it works

Control every dollar your coding agents spend

01
Classify

Every task is scored by complexity, cost sensitivity, and risk, so routing decisions are auditable, not a black box.

02
Route under policy

The request goes to the cheapest capable model your policy allows, across only the providers you approve, like OpenRouter or your own Amazon Bedrock account.

03
Escalate & log

If a cheaper model falls short, TokenSwitch escalates automatically, and every hop is logged and attributed to a developer, team, and repo.

Where TokenSwitch sends a coding workload

Only 1 in 5 requests needs a frontier model

per 100 requests
52 Free & open-source
28 Cheaper & mid-tier
20 Frontier

TokenSwitch escalates only when a task calls for it.

Speed for your team.
Control for you.

Point your team’s agents at TokenSwitch. Every AI coding dollar routed, capped, and logged under your policy.