DeepFrugal Early Preview Find the cheapest way to run any model

LLM cost per token: listed vs effective

The listed price is not the real cost per token. How subscriptions, reseller fees, cached tokens and peak hours change the cost per million tokens.

10 min read DeepFrugal
LLM cost per token: listed vs effective

A price per token looks simple: a number of dollars for a million tokens. That number is a list price. What you pay depends on how you buy — a subscription spreads a flat fee over a quota, a reseller adds a service fee, and caching and peak hours change the rate again. This guide explains the difference between the listed price and the effective price per token, and how to put every model and plan on the same unit.

Why the list price is not the real price

Gateways publish prices in three shapes. To compare them, DeepFrugal normalises all three to $ per 1M tokens.

Model 1 · one listed rate, three billing shapeslisted $1.00/1M
Pay as you go × 1 $1.00/1M the listed rate — the price you see
Subscription × 10 ÷ 60 $0.17/1M Plan A: $10/month buys a $60 quota
Reseller × 1 + 10% $1.10/1M a 10% service fee, added on top

The listed rate is the anchor: $1.00/1M. Pay as you go leaves it untouched. A subscription scales it by subscription ÷ quota — Plan A pays $10 for a $60 quota, so each million tokens costs $0.17. A reseller scales it by 1 + fee: a 10% service fee makes it $1.10. Normalising all three to the same unit is what makes them comparable.

Sample plans and models — figures are illustrative. Plan A is a sample: a $10/month plan with a $60 quota. Sales tax is not included.

Billing Real price
Pay as you go The listed price.
Subscription Listed × (subscription ÷ quota).
Reseller Listed × (1 + service fee). Sales tax is not included.

A subscription does not charge by the token. It charges a flat fee for a quota of included usage. Spread that fee over the tokens the quota buys and you get a lower rate per token — the effective rate. The more of the quota you use, the closer the effective rate falls to its floor.

The effective rate is a best case: it assumes you use the whole quota. Use less and the fee is spread over fewer tokens, so the real rate is higher.

The effective cost formula

The listed rate is the anchor. Each billing shape scales it:

  • Pay as you goeffective = listed.
  • Subscriptioneffective = listed × (subscription ÷ quota).
  • Resellereffective = listed × (1 + service fee).

The subscription term is the one that moves. A quota is a budget of dollars, and each model spends it at its own listed rates, so the same fee buys a different number of tokens per model. A cheap model stretches the quota further; an expensive one drains it faster. For the full picture, see how AI subscription credits work.

Here is one model's listed and effective rates in a real plan:

Listed vs effective rates — OpenCode Go
Subscription $10.00/month · effective rates are published per token type (quota: a per-model dollar allowance) · DeepSeek V4.1 Flash (Off-Peak): 0 input · 0 cached · 1 output credits per 1M tokens
Input /1MOutput /1MCached read /1M
$0.025$0.15$0.1$0.6$0.001$0.003

Turn on the effective column to see the rate with the plan applied. The toggle in the live table does the same for every model.

Reseller fee and tax

Some gateways resell another provider's models and add a service fee on top. The fee is a percentage of the listed price, so the effective rate is listed × (1 + fee). A fee is not tax: sales tax is separate and is never included in a DeepFrugal price. Confirm the tax treatment with the provider.

Add the fee before you compare. A reseller that looks a few percent dearer on the listed rate can widen once the fee is in.

Cached tokens

Most providers bill cached input at a lower rate than fresh input, and some charge a separate cache write fee. Four token types therefore drive the real cost:

  • Input — fresh prompt tokens.
  • Cached read — prompt tokens served from cache.
  • Cached write — tokens written into the cache, when the provider charges.
  • Output — generated tokens.

A cached read pulls the average down; a cache write pushes it up. The rates table above lists the types this plan publishes, and the live table carries the same columns for every model.

Peak, off-peak and tiers

Some providers charge less at quiet hours and more at peak hours, and some price long-context or priority requests as a separate tier. These are not a discount on the listed price; they are separate rows, each with its own listed rate and its own effective rate. Compare them as different models.

How to compare models and plans

Put every candidate on the same unit — dollars per million tokens, effective — and the comparison becomes arithmetic:

Run your numbers

Estimate your monthly input, cached read, cached write and output tokens, then put them against a plan. The calculator values your usage at the listed rates, spreads the plan's fee over its quota, and shows where the subscription starts to win.

Plan break-even — OpenCode Go · Experimental
Cheaper
Subscriptionsave $24.19/month
You pay $10.00/month instead of $34.19 of usage.
Usage: 663M tokens · $50.35 at list rates · $34.19 cheapest pay-as-you-go.
Subscriptioncheaper$10.00/month
Quota used84%
Covered by the quota$50.35
Above the quota$0.00
Monthly cost$10.00
Cheapest pay-as-you-go route

Real rates include the service fee; sales tax is not included.

GLM 5.3 Flash via Ozore API

DeepSeek V4.1 Flash via DeepSeek API

Monthly cost$34.19
Break-even: $10.00 of usage (≈ 132M tokens at this mix)
At your mix: subscription $10.00/mo · cheapest metered $34.19/mo · at list rates $50.35/mo — subscription is cheapest by $24.19/mo.
Plan break-even details
$0.013$0.015$0.052$0.076500M1B663M132M194M2.05B790MTotal tokens per month (M tokens/mo, log scale)Average cost per 1M tokens usage above the plan paid at the listed rate your mix
  • Average cost per token
  • Listed rate
  • Cheapest pay-as-you-go route
  • Quota limit
  • Cheaper than the cheapest
  • Quota usage break-even
  • Plan break-even / loses
  • Average at your mix
Hover the curve to read the values.
How the graph is calculated

Inside the quota the average is the monthly commitment divided by the tokens used, so it falls as the mix scales up and reaches the mix's effective rate at the quota limit. Above the limit the pool is spent and the extra tokens are paid outside the plan at the overflow base chosen in Options — the model's listed rate (default), the cheapest pay-as-you-go rate found for it, or the cheaper of the two — so the average climbs towards that blended rate.

The pool is one shared budget: each model's allowance is that same pool restated at its rates — dollars for a dollar-credit plan, credits for a token-credit plan — and it is allocated in the model order shown, so a different order or blend moves the curve. A token type a model does not price still counts in the volume but adds no cost. The token axis is logarithmic, so a wide range of usage fits on one chart.

The pay-as-you-go line is the cheapest real rate per model (fee included; sales tax excluded) — an exact variant match when a pay-as-you-go gateway publishes that variant, otherwise the model's nearest published row — so it can combine more than one gateway. When that rate equals the plan's own listed rate the chart shows the listed line only.

The markers: the quota usage break-even is where the average meets the plan's list rate, so below it you pay more per token than list; the plan break-even is where the average starts to beat the cheaper of the two references; loses to the cheapest is where it stops doing so (it exists only when the overflow base is dearer than that reference); the your mix point is your stated usage, labelled with its monthly token volume, where the horizontal guide marks the mix's average cost per 1M, and the shaded band is where the subscription wins. With Promo prices off, every figure uses the pre-promo list rates and quotas, and the commitment reverts to the list subscription.

How the plan quota is consumed

A subscription grants one monthly pool. The pool is allocated to the models in the order listed here: each model spends its usage divided by its quota, capped by what is left, so the total never exceeds 100%. Reorder the models to approximate your own pattern. Usage above a model's allowance, or beyond the pool, is paid outside the plan at the listed rate.

ModelUsageMonthly quota% of QuotaPay as you go
GLM 5.3 Flash$29.10$60.0048.5%
DeepSeek V4.1 Flash (Off-Peak)$15.71$60.0026.2%
DeepSeek V4.1 Flash (Peak)$5.54$60.009.2%
Total$50.3584%$0.00

Budget used = 0.839 (84%) · Covered = $50.35 · Pay as you go = $0.00

Effective rates (listed vs effective)
ModelInput /1MOutput /1MCached read /1M
GLM 5.3 Flash$0.03$0.15$0.099$0.5$0.006$0.03
DeepSeek V4.1 Flash (Off-Peak)$0.03$0.15$0.119$0.6$0.001$0.003
DeepSeek V4.1 Flash (Peak)$0.06$0.3$0.238$1.2$0.001$0.006

Effective = listed × 0.199 at this usage · 84% of the plan budget used

Promo active · ends 2026-09-27

These figures use promotional rates or quotas. They change when the promo ends — confirm current pricing before you decide.

  • DeepSeek V4.1 Flash (Off-Peak) — 4× usage promo ($15 → $60 quota) · ends 2026-09-27
  • DeepSeek V4.1 Flash (Peak) — 4× usage promo ($15 → $60 quota) · ends 2026-09-27

Experimental. Confirm every figure against the provider's own pricing before you rely on it. DeepFrugal is not responsible for calculation errors.

Not sure how to read it? How the break-even calculator works.

Summary

The listed price is the price per token on pay as you go. A subscription and a reseller change it, and cached tokens and peak hours change it again. Normalise every model to the same unit — dollars per million tokens, effective — and you can compare them. Check the live comparison table for current values, or run the break-even calculator on your own usage.

Frequently asked questions

What is an effective price per token?

The listed rate scaled by how you buy it: a subscription spreads its fee over the plan's quota, and a reseller adds a service fee. It assumes the whole quota is used, so it is a best case.

Why is the listed price not the real price?

Because most plans do not charge by the token. A flat fee, a reseller fee, cached rates and peak hours all change what a million tokens costs.

How does a subscription change the cost per token?

It scales the listed rate by the subscription divided by the quota. The more of the quota you use, the lower the average rate falls.

What is a reseller service fee?

A percentage a gateway adds on top of another provider's listed price. It is not sales tax, which is never included in a DeepFrugal price.

Do cached tokens change the cost per token?

Yes. Cached reads are usually cheaper than fresh input, and a cache write may cost extra. The average depends on your mix of input, cached read, cached write and output.

Do peak and off-peak hours change the price?

They are separate rates, not a discount. A quiet-hours window has its own listed and effective rate, so compare it as a different row.

How do I compare two models on cost per token?

Put both on the same unit — dollars per million tokens, effective — and compare like for like. The live table does it for every model and provider.

How do I find my own effective rate?

Estimate your monthly tokens, then use the break-even calculator, which spreads a plan's fee over its quota.