Pricing & Limits

Plain pricing. No surprises. Credits that never expire. See full plans and side-by-side comparison on the pricing page.

  • Resets monthly: Your request budget refreshes at the start of every billing cycle.
  • Rolling 5-hour & weekly limits: Each plan also has shorter rolling windows so a single burst can't drain your whole month. See usage limits.
  • Top up automatically: Buy extra credits at model cost. They roll over. They never expire.
  • Your code stays yours: Command Code never trains on it. Never stores it.

  • Pick a plan or top up credits: Head to Studio > Billing.
  • Watch every request, in real time: The Usage page in Studio tracks tokens and request history as it happens.
  • Cards. Invoices. Auto-reload: Manage them all from the billing portal at Studio > Billing.

For teams that need their code and data to stay on their own infrastructure. Enterprise gives you the security, compliance, and control your organization requires - without compromise.

Talk to our team at support@commandcode.ai to get started.

PlanPrice/moCredits/moIncluded LLM UsageModels
Go$1$10~15K requests
GOAT$10$70~75K requestsOpen models, some closed per-model allowances
Pro$20$80~100K requests
Provider$15Pay as you goProvider API accessOpen-source + premium models
Max 10×$100$150~230K requestsOpen-source + premium models
Max 20×$200$300~370K requestsOpen-source + premium models

PlanPrice/moIncluded LLM UsageModels
Team Pro$40~35K requestsOpen-source + premium models
Enterprise$5,000+CustomCustom pool

How it works.

  • Monthly subscription credits. Reset at the start of every cycle.
  • Top up at model cost. Credits roll over. Forever.
  • Enterprise? Email support@commandcode.ai for custom terms, SLAs, and dedicated support.

On top of your monthly credits, every subscription plan has two rolling usage limits - a 5-hour window and a weekly window. They smooth out bursts so a single busy session can't drain a whole month in an afternoon, while your full monthly pool stays available across the cycle.

Each window caps a set amount of your plan's monthly credit allocation:

PlanYour costMonthly credits5-hour limitWeekly limit
Go$1$10$3$6
GOAT$10$70 of usage$14$35
Pro$20$80 of usage$16$40
Max 10×$100$150$45$90
Max 20×$200$300$90$180
Team Pro$40$40$12$24

Limits are measured in credit value, not request count - so how many requests you get depends on the models you run. Cheaper models like DeepSeek V4 Flash stretch much further than premium models like Claude Opus. The usage estimates below and the calculator help you gauge your own mix.

These windows apply to your plan's included monthly credits. Usage with no subscription (free or balance-only) and custom Enterprise pools have no rolling window caps - they're gated only by your credit balance. On-demand credits are never throttled either, on any plan (details below).

How the windows roll

Both windows roll from first use, not on a fixed clock:

  • A window opens on the first request you make after the previous one elapsed, then resets exactly one window-length later (5 hours or 7 days).
  • Usage never carries between windows - each new window starts at zero.
  • Switching plans resets both windows, so you always start a new plan with a fresh allowance.

Usage beyond your limit

Rolling limits throttle only your included monthly credits. On-demand credits are never throttled:

  • Top-up and pay-as-you-go credits are exempt: When you have on-demand credits available - or once your monthly pool is spent - requests skip the window check entirely. Adding credits with /extra unblocks you immediately.
  • While you're over a window limit, Command Code spends your on-demand credits first, preserving your monthly allocation for when the window resets.

So you're never hard-stopped: keep working past a window on on-demand credits, or wait for it to reset to use your included credits again.

Tracking your limits

Run /usage in the CLI to see live 5-hour and weekly meters - each shows how much of the window you've used and when it resets:

Usage limits 5-hour ███░░░░░░░ 32% · resets in 3h 12m Weekly ████░░░░░░ 41% · resets in 2d 4h

If you do hit a limit, the CLI tells you which window you reached and when it frees up:

You've reached your 5-hour usage limit. Resets in 2h 41m (3:00 PM). Run /upgrade for a higher plan or /extra for on-demand credits.

The numbers above are estimates, not ceilings. Your actual usage depends on the models you pick, the length of your prompts, and how long your conversations run.

  • A typical request: ~700–1K input tokens, ~125–200 output tokens, plus ~42K-56K cache reads on average. The calculator below has a cache slider to model your own session.
  • A range, not a single number: Cheaper models stretch further. On Go, DeepSeek V4 Flash runs ~32K requests with no cache - ~15K once the typical 50K cache reads are included. Our plan estimates are the mix average across each plan's typical models.
  • A balanced mix: Even split across models within each tier. Pro, Max 10×, and Max 20× assume ~35% open-source / ~65% premium.
  • Long conversations cost more: Every new turn re-reads the full context, so a long-running session burns through credits faster than a fresh one. Use /clear or start a new session for unrelated tasks to keep things efficient.
  • Track the real number: The Usage page in Studio always reflects the actual price charged per request.
Interactive

Usage calculator

On the Go plan, you pay $1/mo
and get $10 in model credits
effectively ~$20 with deals
Plan
what you pay
Model
34 on this plan
Input tokensfresh prompt
Output tokensmodel reply
Cache read tokensre-read context
Total / 30 days
K
on Muse Spark 1.2 Contributor($10 ÷ $0.0002 per request)
caching varies
Total
2.3B
Cache
2.3B
In
36.4M
Out
9.1M
Plan mix average:K(across Go's typical model mix)
Cost / request
$.
Muse Spark 1.2 Contributor
Rate / 1M tokens
In
$0.100
Out
$0.200
Cache
$0.0020
Quick fixes
.K
one-shot edits
Bug fixes
.K
patch + verify
Small CLIs
end-to-end builds
Feature PRs
multi-file change
Side projects
weekend builds
What you can build / month — rough estimates
on Muse Spark 1.2 Contributor with K reqs / month

Pick a model. Adjust the sliders. See how far your credits go. These are estimates based on typical patterns - your actual mileage may vary.

Command Code runs on the best models from the open-source community, Anthropic, OpenAI, Google, Sakana, and xAI. Each one chosen for what it does best.

Checkout all our available models. Switch any time with /model in interactive mode. Per-token rates for each model are listed below.

Plans include usage at model API rates. You pay for what you use, by model and by token. All prices below are per million tokens. Subscriptions include 2-10x more credits for the same cost as API pricing.

// Model pricingper 1M tokens
Model
Context
Input/M
Output/M
Cache Read
Cache Write
Caps
1M
Free
Free
Free
⚡ Free while the stealth preview lasts
256K
Free
Free
Free
⚡ Free while capacity lasts
Ling 3.0 Flash
256K
Free
Free
Free
Tencent Hy3
262K
$0.14
$0.58
$0.035
Kimi K3
1M
$3.00
$15.00
$0.30
Kimi K2.7 Code
256K
$0.95
$4.00
$0.19
Kimi K2.7 Code HighSpeed
262K
$1.90
$8.00
$0.38
Kimi K2.6
256K
$0.95
$4.00
$0.16
Kimi K2.5
256K
$0.60
$3.00
$0.10
GLM-5.3
1M
$1.40
$4.40
$0.26
GLM-5.2
1M
$1.40
$4.40
$0.26
GLM-5.2 Fast
1M
$3.00
$10.25
$0.50
GLM-5.1
$1.40
$4.40
$0.26
GLM-5
200K
$1.00
$3.20
$0.20
1M
$0.60$0.30
$2.40$1.20
$0.12$0.06
⚡ 50% off
MiniMax M2.7
$0.30
$1.20
$0.06
MiniMax M2.5
200K
$0.30
$1.20
$0.03
DeepSeek V4 Pro (latest)
1M
$0.66
$1.98
$0.022
Off-peak rate shown (17h/day) · peak $1.32 / $3.96 · 01–04 & 06–10 UTC (7h/day)
DeepSeek V4 Flash (latest)
1M
$0.22
$0.66
$0.007
Off-peak rate shown (17h/day) · peak $0.44 / $1.32 · 01–04 & 06–10 UTC (7h/day)
DeepSeek V4 Flash Vision (exp)
1M
$0.22
$0.66
$0.01
Off-peak rate shown (17h/day) · peak $0.44 / $1.32 · 01–04 & 06–10 UTC (7h/day)
Qwen 3.8 Max
1M
$2.00
$6.00
$0.25
$2.50
Qwen 3.8 27B
262K
$0.40
$3.00
$0.04
Qwen 3.6 Max Preview
$1.30
$7.80
$0.26
$1.63
Qwen 3.6 Plus
$0.50
$3.00
$0.10
Qwen 3.7 Max
1M
$2.50
$7.50
$0.50
$3.13
Qwen 3.7 Plus
1M
$0.40
$1.60
$0.08
$0.50
Qwen 3.7 Flash
1M
$0.03
$0.13
$0.006
$0.038
Step 3.7 Flash
256K
$0.20
$1.15
$0.04
Step 3.5 Flash
1M
$0.10
$0.30
$0.02
1M
$2.00$0.435
$6.00$0.87
$0.40$0.0036
⚡ 99% off
1M
$0.80$0.14
$4.00$0.28
$0.16$0.0028
⚡ 98% off
Nemotron 3 Ultra
1M
$0.60
$2.40
$0.12
Claude Fable 5
1M
$10.00
$50.00
$1.00
$12.50
Claude Opus 5
1M
$5.00
$25.00
$0.50
$6.25
Claude Opus 4.8
1M
$5.00
$25.00
$0.50
$6.25
Claude Opus 4.7
1M
$5.00
$25.00
$0.50
$6.25
Claude Opus 4.6
1M
$5.00
$25.00
$0.50
$6.25
Claude Sonnet 5
1M
$2.00
$10.00
$0.20
$2.50
Claude Sonnet 4.6
1M
$3.00
$15.00
$0.30
$3.75
Claude Haiku 4.5
200K
$1.00
$5.00
$0.10
$1.25
GPT-5.6 Sol
1.1M
$5.00
$30.00
$0.50
$6.25
Available on GOAT and above.
GPT-5.6 Terra
1.1M
$2.00
$12.00
$0.20
$2.50
GPT-5.6 Luna
1.1M
$0.20
$1.20
$0.02
$0.25
Available on every plan, including Go.
GPT-5.5
400K
$5.00
$30.00
$0.50
GPT-5.4
400K
$2.50
$15.00
$0.25
GPT-5.4 Mini
400K
$0.75
$4.50
$0.075
GPT-5.3 Codex
400K
$2.00
$8.00
$0.50
1M
$1.50$0.75
$7.50$3.75
$0.15$0.075
$0.04167
⚡ 50% off · reverts to $1.50 in / $7.50 out · ends December 31, 2026
Available on GOAT and above.
Gemini 3.6 Flash
1M
$1.50
$7.50
$0.15
Gemini 3.5 Flash
1M
$1.50
$9.00
$0.15
Gemini 3.5 Flash Lite
1M
$0.30
$2.50
$0.03
Gemini 3.1 Flash Lite
1M
$0.25
$1.50
$0.03
Fugu Ultra
1M
$5.00
$30.00
$0.50
Muse Spark 1.2
1M
$1.25
$4.25
$0.15
Available on GOAT and above.
Muse Spark 1.2 Contributor
1M
$0.10
$0.20
$0.002
Available on every plan, including Go.
Muse Spark 1.1
1M
$1.25
$4.25
$0.15
Grok 4.6
500K
$2.00
$6.00
$0.50
Available on GOAT and above.
Grok 4.5
500K
$2.00
$6.00
$0.50
Available on every plan, including Go.
Inkling
256K
$1.00
$4.05
$0.17
Inkling Small
1M
$0.50
$1.20
$0.10
Claude Sonnet 4.5
1M
$3.00
$15.00
$0.30
$3.75
61/61 Models · Prices per 1M tokens · USD
Text Vision Reasoning
Note

Open-source models are routed across multiple upstream providers to maintain high availability. The values above are the mean per-provider price. Actual cost on a given request may vary slightly based on which upstream serves it.

A context window is the maximum span of tokens (text, images, and code) a model can hold in mind at once. Input prompt plus everything it generates back.

Every Command Code session has its own context. As you prompt, read files, and respond, that context fills up. Command Code compresses older messages automatically, so you can keep going without starting over.

Context window size depends on the model. Most current models from Anthropic, OpenAI, xAI, Google, DeepSeek and more support up to 1 million tokens. Check the models page for a specific model's exact window.

How do I get started?

Install npm i -g command-code. Sign in. Code. Pick a plan from Studio > Billing or visit the pricing page. Need more? Buy extra credits any time.

What happens when I run out of credits?

Once your included credits reach zero, requests stop until your subscription renews or you add credits. To keep going, buy more from Studio > Billing, or enable auto top-up. Extra credits are at model cost. They roll over month to month. They never expire.

How do I buy more credits?

You can upgrade to one of our plans or purchase extra credits at API pricing. Go to Studio > Billing to purchase extra credits manually or enable auto-reload to top up automatically when your balance is low.

Where can I check my usage?

Visit the Usage page in Studio. It shows per-request cost, token counts, and a full history of your usage. Learn more in the usage docs. Run /usage in the CLI to see your credit balance plus live 5-hour and weekly limit meters.

What are the 5-hour and weekly limits?

On top of your monthly credits, subscription plans have two rolling windows - a 5-hour limit and a weekly limit - so a single burst can't drain your whole month. They apply only to your included monthly credits: on-demand and top-up credits are never throttled, so you can keep working past a window with /extra credits or wait for it to reset. See Usage limits for details.

Do you train on my code?

No. Command Code does not train on your code or store your code snippets. Taste processing runs on your codebase and stores learning data in your project and on your local machine only. See our Privacy Policy for details.

Can I switch plans?

Yes. You can upgrade or downgrade anytime from Settings > Billing in Studio. Upgrades take effect immediately - the new plan replaces your current one and you're billed for it right away. Downgrades take effect at the start of your next billing cycle.

Can I cancel my subscription?

Yes. Your subscription can be set to cancel at any time from Settings > Billing in Studio. Cancel before your renewal date and it stays active until the end of the current billing cycle. Cancel after the renewal date has already passed and it is set to cancel at the end of the next subscription cycle. See the cancellation FAQs for billing edge cases.

How does team billing work?

Team plans are billed per seat. Credits are pooled at the team level, so power users and occasional users share from the same allocation. Organization admins can manage seats and billing from the org settings.

What is the right plan for me?

Go for international users. Pro for active agent users. Provider for pay-as-you-go API access to all models. Max 10× for power users running Command Code all day. Max 20× for maxed-out daily usage at the highest request volume. Team for shared usage and pooled billing. Enterprise for custom terms, SLAs, and dedicated support.

What are my payment options?

Self-serve plans support all major credit and debit cards via Settings > Billing in Studio. For invoice-based billing or wire transfers, contact us at support@commandcode.ai to discuss Enterprise plans.

How does usage-based pricing work?

Every plan includes a set of credits. Usage is charged at model API rates, by model and by token. Run out? Buy more, or enable auto top-up. Credits roll over. They never expire.

What happens when a model provider changes its prices?

We charge you what they charge us. No markup. If a model gets cheaper, you pay less. If it gets more expensive, you pay more.

Your subscription plan never changes because of this. Same monthly price, same usage credits. The only thing that changes is how far your usage credits go on that one model. Check any model in the calculator to see the numbers.

Why do DeepSeek V4 rates change during the day?

DeepSeek now charges by the hour. From 16:00 UTC on August 16, 2026, V4 Pro, V4 Flash, and V4 Flash Vision (exp) cost:

  • Peak - 01:00-04:00 and 06:00-10:00 UTC. 7 hours a day. Full price.
  • Off-peak - the other 17 hours. Half price.

We show the off-peak price as the main number, because that is what you pay 17 hours out of 24. Hover it to see the peak price. The Usage page shows what each request actually cost you.

Does the DeepSeek price change affect my plan or credits?

No. Your subscription plan is the same. Same monthly price, same usage credits, and they still roll over.

Your usage credits are an amount of money, not a number of requests. DeepSeek costs more now, so the same credits buy fewer DeepSeek requests than before. During peak hours, about half as many as off-peak.

How do I pay the lower DeepSeek rate?

Run outside 01:00-04:00 and 06:00-10:00 UTC. That is it - nothing to turn on, no setting to change. Everything outside those hours is half price. Big jobs like long refactors or batch runs cost half as much if you run them then.

How can I see and manage usage in my organization?

Organization admins can view per-seat usage, manage billing, and monitor team-wide consumption from org settings in Studio. Individual and Org usage is available on the Usage page.

Can I buy Command Code from a reseller or third party?

No. Command Code subscriptions are only sold directly through commandcode.ai. We do not authorize any resellers or third-party sellers. Subscriptions purchased from any other source are unauthorized and may be suspended or terminated. Purchase only through our official website.

Where are models hosted?

Commercial models are hosted by Anthropic, OpenAI, Google, and Azure on their respective US-based infrastructure (EU on demand). Command Code does not train on your code or store code snippets. When processing requests, data is sent to the provider's API and handled per their privacy policy. Taste data is stored locally in your project directory. See our Privacy Policy for details. For more options, contact us at support@commandcode.ai to discuss Enterprise plans.

What about data and privacy for open-source models?

The $1 Go plan is meant for international developers. We want coding agents to work for everyone. Open-source models are available globally with infrastructure in the US, EU, and Singapore for reliable access worldwide. For more options, contact us at support@commandcode.ai to discuss Team & Enterprise plans.

Is Command Code open source?

Currently Command Code is not open source.

Can I opt out of telemetry?

Yes. See Telemetry for what's collected and how to disable it.

Do you support zero data retention (ZDR)?

Yes. 99% of our models route through ZDR-capable upstreams. Run the CLI with CMD_ZDR=1 (for example, CMD_ZDR=1 cmd) to enforce both zero data retention and no prompt training on every request across all plans. The same opt-in is available on the Provider API by sending the header x-cmd-zdr: 1 on any request. A small handful of models don't yet have a ZDR-capable upstream; under CMD_ZDR=1 those requests will either pass to a ZDR-capable provider or fail rather than route through a non-zdr provider. We're actively adding ZDR coverage to the remaining models. ZDR can change which upstream provider that serves your request, and so it may charge more depending on the provider. Need ZDR on a specific model that doesn't have it today? Email support@commandcode.ai.

ZDR capacity is priced per provider. We route each request to whichever provider can serve it at the lowest cost with capacity at that moment and pass through exactly what we are charged. Committed team and enterprise plans bill the rates in their contract. That is why a ZDR request can cost more than the same model without it.

Ox Alpha is free. What is the catch?

The price is genuinely $0: input, output and cache reads are all free on stealth/ox-alpha, on every plan, and free requests draw down neither your credit balance nor your usage limits. What you should know before switching is who serves it. Ox Alpha is a stealth model, developed and operated by a third-party provider who has chosen to stay anonymous for the duration of the preview.

The provider retains prompts and completions, and states they are not used for training. That is a weaker guarantee than the rest of our catalog holds, so Ox Alpha is deliberately not served under zero data retention: a session running with CMD_ZDR=1 (or the x-cmd-zdr: 1 header on the Provider API) will refuse to send to it rather than quietly fall back to a provider that retains your prompts. Keep sensitive work on a ZDR-capable model. See the ZDR question above.

The preview can end at any time. If the model disappears from /model, switch to another and carry on; nothing else about your plan changes.

Where can I ask more questions?

Join our Discord community to ask questions, share feedback, and connect with other developers. Follow us on 𝕏 @CommandCodeAI. For private inquiries, email us at support@commandcode.ai.

Command Code deals. Apply automatically. No codes. No toggles. Discounts work on every plan and even extra pay-as-you-go credits.

DEALGemini 3.7 Flash
Every $1 of credit on this model
2×usage
50% off listauto-applied when you pick gemini-3.7-flash
Model
Gemini 3.7 Flash
Discount
50%
Ends
December 31, 2026

Details

Gemini 3.7 Flash is 50% off through December 31, 2026.

Because the model's per-token cost is half of list price for the duration of this deal, every dollar of credit you spend on Gemini 3.7 Flash is worth 2× the requests it normally would be. Practical implications:

  • Available on GOAT and above (GOAT, Pro, Max, Ultra, Team Pro, Provider API): credits routed to Gemini 3.7 Flash stretch 2× further.
  • On the $10 GOAT plan and $20 Pro plan: applied automatically. Gemini 3.7 Flash's monthly allowance ($40 on GOAT, $60 on Pro) is denominated at the discounted rates, so it's effectively up to ~$80 of list-price usage on GOAT and ~$120 on Pro.
  • On Max, Ultra, and Team Pro: it bills against your full credit balance at half list price, so those credits go 2× further on Gemini 3.7 Flash.
  • On top-up credits: extra credits you purchase from Studio > Billing get the same treatment. Buying $20 of credits and using Gemini 3.7 Flash is equivalent to $40 of usage.
  • No code, no toggle: pick Gemini 3.7 Flash from /model in the CLI and the discounted rate applies automatically. The Usage page reflects the discounted price per request in real time.
Gemini 3.7 Flash, per 1M tokensWasNowOff
Input$1.50$0.7550%
Output$7.50$3.7550%
Cache read$0.15$0.07550%

Discount expires December 31, 2026. When it ends, the model returns to full list price ($1.50 input / $7.50 output / $0.15 cache read per 1M tokens).

DEALMiniMax M3
Every $1 of credit on this model
2×usage
50% off listauto-applied when you pick minimax-m3
Model
MiniMax M3
Discount
50%

Details

MiniMax M3 is 50% off.

Because MiniMax M3's per-token cost is half of list price, every dollar of credit you spend on MiniMax M3 is worth 2× the requests it normally would be. Practical implications:

  • On any subscription (Go, GOAT, Pro, Max, Ultra, Team Pro): credits routed to MiniMax M3 stretch 2× further. A $1 Go plan with $10 credits effectively has up to $20 of MiniMax M3 usage if you stay on this model.
  • On the $10 GOAT plan and $20 Pro plan: applied automatically. The deal is baked into MiniMax M3's boosted per-model allowance: $47 of monthly usage on GOAT, $57 on Pro. Nothing to enable.
  • On top-up credits: extra credits you purchase from Studio > Billing get the same treatment. Buying $20 of credits and using MiniMax M3 is equivalent to $40 of usage.
  • No code, no toggle: pick MiniMax M3 from /model in the CLI and the discounted rate applies automatically. The Usage page reflects the discounted price per request in real time.
Per 1M tokens (context ≤512K)WasNowOff
Input$0.60$0.3050%
Output$2.40$1.2050%
Cache read$0.12$0.0650%
Per 1M tokens (context >512K)WasNowOff
Input$0.60$0.3050%
Output$2.40$1.2050%
Cache read$0.12$0.0650%
DEALMiMo V2.5 Pro
Price cut
99%off
auto-applied when you pick mimo-v2.5-pro
Model
MiMo V2.5 Pro
Discount
99%

Details

MiMo V2.5 Pro is up to 99% off.

Command Code has slashed MiMo V2.5 pricing in collab with Xiaomi. Unified across all context lengths. Because the model's per-token cost is a fraction of the previous price, every dollar of credit you spend on MiMo V2.5 Pro goes dramatically further.

Practical implications:

  • On $1 Go plan ($10 credits): your $10 in credits effectively delivers up to ~$50 of MiMo V2.5 Pro usage.
  • On $10 GOAT plan: applied automatically. MiMo V2.5 Pro's $20 monthly allowance is denominated at the deal rates, effectively up to ~$100 of usage at the old list price.
  • On $20 Pro plan: the allowance is $30, effectively up to ~$150 of MiMo V2.5 Pro usage. Nothing to enable.
  • On top-up credits: extra credits you purchase from Studio > Billing get the same treatment. $20 of credits on MiMo V2.5 Pro is equivalent to ~$100 of usage.
  • On every plan + pay-as-you-go: Go, GOAT, Pro, Max, Ultra, Team Pro, and even standalone top-up credits all benefit from the same discount automatically.
  • No code, no toggle: pick mimo-v2.5-pro from /model in the CLI and the discounted rate applies automatically. The Usage page reflects the discounted price per request in real time.
Per 1M tokensWasNowOffMultiplier
Output$6.00$0.8786%7x
Input$2.00$0.43578%4.6x
Cache read$0.40$0.003699%111x

Xiaomi has reduced their API pricing with unified rates across all context lengths.

DEALMiMo V2.5
Price cut
98%off
auto-applied when you pick mimo-v2.5
Model
MiMo V2.5
Discount
98%

Details

MiMo V2.5 is up to 98% off.

Command Code has slashed MiMo V2.5 pricing in collab with Xiaomi. Unified across all context lengths. Because the model's per-token cost is a fraction of the previous price, every dollar of credit you spend on MiMo V2.5 goes dramatically further.

Practical implications:

  • On $1 Go plan ($10 credits): your $10 in credits effectively delivers up to ~$100 of MiMo V2.5 usage.
  • On $10 GOAT plan: applied automatically. MiMo V2.5 carries a boosted $30 monthly allowance at the deal rates, effectively up to ~$300 of usage at the old list price.
  • On $20 Pro plan: the allowance grows to $40, effectively up to ~$400 of MiMo V2.5 usage. Nothing to enable.
  • On top-up credits: extra credits you purchase from Studio > Billing get the same treatment. $20 of credits on MiMo V2.5 is equivalent to ~$200 of usage.
  • On every plan + pay-as-you-go: Go, GOAT, Pro, Max, Ultra, Team Pro, and even standalone top-up credits all benefit from the same discount automatically.
  • No code, no toggle: pick mimo-v2.5 from /model in the CLI and the discounted rate applies automatically. The Usage page reflects the discounted price per request in real time.
Per 1M tokensWasNowOffMultiplier
Output$4.00$0.2893%14x
Input$0.80$0.1483%5.7x
Cache read$0.16$0.002898%57x

Xiaomi has reduced their API pricing with unified rates across all context lengths.

DEALOx Alpha
Free while the stealth preview lasts
100%off
auto-applied when you pick ox-alpha
Model
Ox Alpha
Discount
100%
Term
while the stealth preview lasts

Details

Ox Alpha is a free stealth preview.

Ox Alpha is a reasoning model for coding and long-horizon agentic work, served as an anonymous stealth preview.

  • Requests cost no credits: Input, output, and cache reads are all $0.00 per million tokens.
  • Available to all customers: Go, GOAT, Pro, Max 10×, Max 20×, and Team Pro. Need to have $1 of credits in your account to start a session.
  • On GOAT and Pro: applied automatically. Free requests cost $0, so they never count against your usage limits. Nothing to enable.
Per 1M tokensPrice
Input$0.00
Output$0.00
Cache read$0.00

A stealth preview can change or end without notice. If it goes offline, switch models with /model.

DEALLaguna S 2.1
Free while capacity lasts
100%off
auto-applied when you pick laguna-s-2.1-free
Model
Laguna S 2.1
Discount
100%
Term
while capacity lasts

Details

Laguna S 2.1 is free while capacity lasts.

Laguna S 2.1 is Poolside's open-weight model for agentic coding and long-horizon work.

  • Requests cost no credits: Input, output, and cache reads are all $0.00 per million tokens.
  • Available to all customers: Go, GOAT, Pro, Max, Ultra, and Team Pro. Need to have $1 of credits in your account to start a session.
  • On GOAT and Pro: applied automatically. Free requests cost $0, so they don't draw down your per-model allowances. Nothing to enable.
Per 1M tokensPrice
Input$0.00
Output$0.00
Cache read$0.00

How deals work.

  • Automatic: Pick the discounted model from /model in the CLI or in Studio. The rate drops for the length of the deal.
  • On every plan: Go, GOAT, Pro, Max 10×, Max 20×. Credits routed to the discounted model stretch further on every tier. On GOAT and Pro, deals are baked into the per-model allowances automatically, with nothing to enable.
  • Stacks with your plan's credits: Your plan's credit bonus is a discount of its own - $100 on Max 10× buys $150 of credits - and any active deal then applies to those credits. On GOAT and Pro, deals are already baked into the per-model allowances. A model at 50% off makes every credit go twice as far, so the two savings compound. How much that adds up to depends on your plan and which deal is running - each deal card above spells out its own math.
  • Discounted upstream only: The deal rate applies only when the discounted upstream serves the request.
  • On top-up credits: Extra credits you buy from Studio > Billing get the same discount when spent on the discounted model.
  • Real-time billing: The Usage page reflects the discounted price per request as it happens.
  • Auto-expires: When a deal ends, the model returns to its full list price across the docs, the calculator, and Studio. Nothing for you to do.