DeepFrugal Early Preview Find the cheapest way to run any model

Command Code Flash Fast vs Flash: the cache trap

Command Code gives both tiers the same quota, but Flash Fast bills cached reads at more than 5× the rate. A cache-heavy month runs out far sooner.

9 min read DeepFrugal
Command Code Flash Fast vs Flash: the cache trap

Command Code lists DeepSeek V4.1 Flash Fast next to the standard tier of the same model, and gives both the same monthly quota. Fresh input costs about the same on either tier. Cached reads cost more than 5× as much on Fast, and for a coding agent that is most of the bill.

The short version

  • Small month, the tier does not matter. Below Fast's quota limit you pay the same monthly fee either way, so the Fast rate buys nothing on the bill.
  • Big cache-heavy month, the standard tier wins. The gap grows with volume: the overflow is billed at listed rates, and the listed cached-read rate is where the tiers diverge.
  • Fast is not a trap, just a different product. Its rates are not outrageous next to pay-as-you-go; they are expensive next to the same tokens on the standard tier. Buy it where throughput is what you are paying for, and leave it out of a batch job that rereads a large context.

The four Command Code rows

The model publishes four rows in the plan catalog: one per tier, one per time window. Read the cached-read column next to the input column. That is the whole story. The table opens on listed rates; switch to effective rates to see what a monthly fee spread over its quota really costs.

DeepSeek V4.1 Flash on Command Code — Command Code GOAT
VariantInput /1MOutput /1MCached read /1MThroughputMonthly quota
DeepSeek V4.1 Flash (Off-Peak)$0.15$0.6$0.003—$60
DeepSeek V4.1 Flash Fast (Off-Peak)$0.16$0.58$0.016—$60
DeepSeek V4.1 Flash (Peak)$0.3$1.2$0.006—$60
DeepSeek V4.1 Flash Fast (Peak)$0.32$1.16$0.032—$60

Command Code publishes the same window hours for both tiers, so the peak rows carry the same shape as the off-peak ones. It publishes no throughput for any of them: the column reads — because a rate nobody published is unmeasured, not because it is slow. Benchmark the two tiers yourself if speed decides. The full plan is in the live table.

Everything above is written against the catalog as it stood when the article was published. Rates, quotas and what each provider publishes change, so re-check the table before you decide.

Where the two tiers diverge

Input moves by a few percent between the tiers, and output moves the other way by about as much. Those differences wash out over a month. The cached-read rate does not: it is several times higher on Fast, and it applies to the largest slice of a cache-heavy agent's traffic.

A quota is spent at listed rates, so a higher cached-read rate is spent faster. That is the trap: the tier with the faster responses empties the same budget in fewer tokens.

A typical month against both tiers

The calculator below runs the same workload against the plan once per tier, at identical volumes.

The volumes are a typical month on DeepSeek V4.1 Flash driven by a coding agent: tens of millions of fresh input tokens against billions of cached reads, with output a fraction of the input. The observed ratio is what matters: reused context outweighs fresh input by an order of magnitude, and that is where a cache-heavy month actually spends.

The horizontal axis is total tokens per month on a log scale. The vertical axis is the average cost per million tokens. Each tier keeps its own quota-limit line, and the verdict names the cheaper side at the volumes you enter. Start from those volumes and overwrite them to match your own usage, including how much of the month lands in the peak window.

Variant burn rate — Command Code GOAT · Experimental

One workload, two tiers on Command Code GOAT: DeepSeek V4.1 Flash vs DeepSeek V4.1 Flash Fast. Open the article to edit the volumes.

ModelInputOutputCached read
DeepSeek V4.1 Flash (Off-Peak) 121M 18M 4.1B
DeepSeek V4.1 Flash (Peak) 121M 18M 4.1B
DeepSeek V4.1 Flash (Peak)15% of the month at the peak rate
Total4.2B tokens
Cheaper for this mix
DeepSeek V4.1 Flash$49.376/mo less
Usage: 4.2B tokens · same volumes on both sides, only the tier differs.
DeepSeek V4.1 Flashcheaper
Quota used78%
Covered by the quota$47.093
Above the quota$0
Average /1M$0.002
Monthly cost$10
DeepSeek V4.1 Flash Fast
Quota used100%
Covered by the quota$60
Above the quota$49.376
Average /1M$0.014
Monthly cost$59.376
Variant burn details
$0.002$0.004$0.011$0.026500M10B4.24B900M2.05B8.02B5.4BTotal tokens per month (M tokens/mo, log scale)Average cost per 1M tokens your mix
  • DeepSeek V4.1 Flash
  • DeepSeek V4.1 Flash Fast
  • DeepSeek V4.1 Flash listed rate
  • DeepSeek V4.1 Flash Fast listed rate
  • Cheapest pay-as-you-go route
  • DeepSeek V4.1 Flash quota limit
  • DeepSeek V4.1 Flash Fast quota limit
  • Cheaper than the cheapest
  • Quota usage break-even
  • Plan break-even / loses
  • Average at your mix

Cached reads cost 5.3× more on DeepSeek V4.1 Flash Fast. Both tiers share the same fee, so the curves share the segment inside the quota.

Hover the curve to read the values.
▾ How the graph is calculated

Each plan draws one curve: the average cost per 1M tokens as the mix scales up. Inside the quota the average is that plan's monthly subscription fee divided by the tokens used, so it falls and reaches the plan's effective rate at its quota limit. Above the limit the extra tokens are paid outside the plan at the overflow base chosen in Options — each plan's own fallback route, the pay-as-you-go market its gateway publishes (the listed rate where there is none), its cheapest pay-as-you-go rate, or the cheaper of the offered bases — so both curves follow that base.

Each plan grants one monthly pool: each model's allowance is that same pool restated at its rates — dollars for a dollar-credit plan, credits for a token-credit plan — and it is allocated in the model order shown, so a different order or blend moves the curve. A token type a plan does not price still counts in the volume but adds no cost. The token axis is logarithmic, so a wide range of usage fits on one chart.

Each plan's listed rate is a horizontal reference, drawn once as a shared neutral line when the two rates match. The pay-as-you-go line is the shared cheapest real rate per model (fee included; sales tax excluded) — an exact variant match when a pay-as-you-go gateway publishes it, otherwise the model's nearest published row — so it can combine more than one gateway. When it equals the cheaper listed rate the chart drops it.

Only the cheaper plan for the mix carries the markers: the quota usage break-even is where its average meets the cheaper reference, the plan break-even is where it starts to beat it, and loses to the cheapest is where it stops. The shaded band is where the cheaper plan beats the cheapest reference, split at that plan's quota limit: solid up to it, fainter beyond. The your mix guide is shared by both plans at your stated usage, labelled with its monthly token volume, with a horizontal guide at the average cost of the cheaper plan per 1M. Each quota limit is a dashed vertical line.

Experimental. Confirm every figure against the provider's own pricing before you rely on it. DeepFrugal is not responsible for calculation errors.

Conclusions

For a cache-heavy workload the standard tier is the cheaper one, and the reason is not the input rate. Both tiers spend from the same monthly budget; the standard tier simply gets more tokens out of it, because reused context is where the budget actually goes. The chart shows the mechanism: the Fast quota-limit line sits to the left of the standard one, so it runs out at a lower monthly volume, and everything after that point is billed outside the plan.

The check to run before choosing: how much of your traffic is a cache hit. On a cache-heavy agent that one number decides the bill. A month that barely touches the cache barely feels the rate difference, and one that rereads context all day feels all of it.

Next steps

Frequently asked questions

What is the cache trap?

Both tiers carry the same monthly quota, so the fee is identical. But the Fast tier lists a much higher rate for cached reads, so a cache-heavy month spends that quota on far fewer tokens.

Is this only about Command Code?

Today, yes: Command Code is the only provider announcing the Fast tier, so this page values both tiers inside the same Command Code plan. The comparison itself is not Command Code specific: any provider that splits a model into a cached-read-cheap tier and a cached-read-dear one needs the same arithmetic.

Do peak hours matter here?

Yes. Each tier lists a peak and an off-peak rate, and the calculator splits your month across the two windows.

When is the Fast tier worth it?

When throughput is the thing you are paying for: a latency-sensitive workflow, an interactive loop, or a month so small the quota never runs out. Its rate is not outrageous on its own; it is just expensive relative to the same tokens on the standard tier, so it has to be used where the speed is worth paying for.

Does Command Code publish throughput for these tiers?

Not at the time of writing. The table keeps a throughput column and shows a dash for all four rows, because a provider that publishes no rate has not been measured, not because it is slow. Treat the speed claim as unverified and benchmark it yourself.

How do I test my own workload?

Edit the volumes and the peak share in the burn-rate widget, or sweep the usage scale, then confirm the live rates in the table. Always compare the two tiers on your own mix rather than on the listed rate alone.

How current is this?

Written against the catalog as it stood when the article was published. Rates, quotas and what each provider publishes change, so re-check the live table before deciding.