3  Models & Budget

3.1 Available models

The gateway gives you access to Claude models and several open-source models. The default is Claude Haiku 4.5 — fast, capable for most coding tasks, and the most budget-friendly Claude option.

Name Model Relative cost
claude-haiku-4-5-20251001 Claude Haiku 4.5 (default) 1x
claude-sonnet-4-6 Claude Sonnet 4.6 ~3x Haiku
claude-opus-4-6 Claude Opus 4.6 (strongest) ~5x Haiku
qwen3-coder-480b Qwen3 Coder 480B ~0.4x Haiku
qwen3-coder-30b Qwen3 Coder 30B (small, fast) ~0.15x Haiku
devstral-2 Mistral Devstral 2 ~0.35x Haiku
kimi-k2.5 Kimi K2.5 ~0.55x Haiku
glm-5 GLM 5 ~0.9x Haiku
minimax-m2.5 MiniMax M2.5 ~0.25x Haiku
gpt-oss-120b OpenAI GPT-OSS 120B ~0.15x Haiku

3.2 When to switch models

  • Haiku — your everyday model. Use it for most tasks.
  • Sonnet or Opus — reach for these when Haiku struggles: complex reasoning, large refactors, subtle bugs. Switch back when you’re done.
  • Open models — interesting to try, but they use file-reading and editing tools less reliably than Claude. If one answers without looking at your files, ask it explicitly (e.g., “use the read tool on notes.txt”).

3.3 Switching models

Inside Claude Code, type /model followed by a name from the table, for example /model claude-sonnet-4-6.

3.4 Checking your budget

In any terminal on the hub:

claude-agent-coders --budget

It shows what you have spent, your budget, and when your key expires. Spending shows up about a minute after each request.

3.5 Making your budget last

  • Stay on Haiku for routine tasks. It’s the best cost-to-capability ratio.
  • Switch to Sonnet/Opus only when you need to, then switch back.
  • Use /clear between tasks — this resets the conversation, which means the agent sends less context with each request and each request costs less.
  • The gateway caches repeated context for Claude models, so later requests in a session cost less than the first one.
Tip

Running /usage in Claude Code shows what the current session has spent.