Foxnut Studios
Format
Statistics
Territory
AI consulting
Family
Readiness and measurement
Basis
First-hand

Reviewed

Statistics

LLM cost per task: the published prices and what drives real bills over them

What a unit of LLM work costs, from the vendors' own price lists: Anthropic, OpenAI and Google rates read at source, the discounts they publish, and the two overrun causes the studio sees in practice.

Reviewed by Ameya Sahasrabudhe and Swati Thakur,

What this page records

What a unit of LLM work costs is public: every major model vendor publishes a per-token price list, and this page reads them at source rather than through aggregators. The pattern across the three lists below is consistent - a frontier-to-small price gap of 5x to 25x for the same volume of work, half price for work that can wait, a 90% discount on context the model has already read, and at least one rate that changes with the shape of the prompt itself. What no price list states is why real bills drift past estimates built on these exact numbers, and there this page switches sources: in the studio’s own engagement records the two usual causes are added use cases after the system proves itself, and upgrades to frontier models the task never needed. Token consumption against the rate card is what this studio measures on a live AI system when an estimate meets reality; the records below are the public half of that arithmetic, dated, so any row can be re-checked against its list.

5x

The published gap between Anthropic's frontier tier and its small tier: Claude Opus 5 at $5 per million input tokens and $25 per million output tokens, against Claude Haiku 4.5 at $1 and $5 - five times the price on both sides for the same volume of work

Method Both rows read from Anthropic's published pricing page on 5 August 2026Sample Anthropic's current model price list, all modelsPeriod As published on 5 August 2026Source Anthropic published price list

$2 then $3

A published price with an expiry date printed on it: Claude Sonnet 5 input is $2 per million tokens through 31 August 2026 and $3 from 1 September 2026, with output moving from $10 to $15 - the price list itself states when its own numbers stop being true

Method The introductory-pricing note on Anthropic's pricing page, quoted dates and rates as printedSample One model's published rate schedulePeriod Read 5 August 2026; change effective 1 September 2026Source Anthropic published price list

25x

The span across OpenAI's current GPT-5.6 models: the flagship gpt-5.6-sol at $5.00 per million input tokens and $30.00 per million output tokens, against the small gpt-5.6-luna at $0.20 and $1.20 - twenty-five times the price on both sides

Method Both rows read from OpenAI's published API pricing page on 5 August 2026, and re-checked against the same page in a second read the same daySample OpenAI's current standard-tier price listPeriod As published on 5 August 2026Source OpenAI published price list

$2.00 or $4.00

What the same input token costs at Google depending on how long the prompt is: Gemini 3.1 Pro Preview input is $2.00 per million tokens up to 200,000 tokens of context and $4.00 above it, with output moving from $12.00 to $18.00 - the task's shape, not just the model, sets the rate

Method The over-200,000-token tier read from Google's Gemini Developer API pricing page on 5 August 2026, confirmed in a second read the same daySample Google's paid-tier Gemini price listPeriod As published on 5 August 2026Source Google Gemini API published price list

Half price

What all three vendors publish for asynchronous batch processing - the discount for not needing the answer immediately: Anthropic's batch table halves every standard rate, OpenAI's Batch tier prices are half its standard rates ($2.50 and $15.00 against $5.00 and $30.00 on its flagship), and Google states a 50% cost reduction

Method Each vendor's batch pricing read from its own pricing page on 5 August 2026Sample The batch tiers of all three price listsPeriod As published on 5 August 2026Source Anthropic, OpenAI and Google published price lists

One tenth

What repeated context costs once cached: Anthropic prices a cache read at 0.1 times its base input rate, and OpenAI's cached-input rates are one tenth of its standard input rates ($0.50 against $5.00 on its flagship) - two vendors independently pricing re-read context at 90% off

Method Cache-read and cached-input rows read from the two vendors' pricing pages on 5 August 2026; Google prices caching on a different basis (per-token rates plus hourly storage) and is excluded from this recordSample The caching rows of the Anthropic and OpenAI price listsPeriod As published on 5 August 2026Source Anthropic and OpenAI published price lists

New use cases

The most common reason a real AI bill comes in higher than its estimate, in the studio's experience: scope creep. Once a client sees the system working well, they insist on adding new use cases, and token consumption grows past what was estimated - the rate card never moved, the volume did

Method The founder's observed pattern across the systems the studio has built and handed over; stated as operating experience, not a cited market statisticSample Client systems the studio has built and handed overPeriod As of August 2026Source Foxnut Studios engagement records (not externally retrievable)

Unnecessary upgrades

The second common overrun cause the studio sees: client-insisted upgrades to state of the art models when the initially configured, cheaper model handles the task fine - the work does not change, but every token of it is billed at the frontier rate

Method The founder's observed pattern on running client systems; stated as operating experience, not a cited market statisticSample Client systems the studio has built and handed overPeriod As of August 2026Source Foxnut Studios engagement records (not externally retrievable)

Every number on this page, with what it measures, its period and its source
NumberWhat it measuresPeriodSource
5xThe published gap between Anthropic's frontier tier and its small tier: Claude Opus 5 at $5 per million input tokens and $25 per million output tokens, against Claude Haiku 4.5 at $1 and $5 - five times the price on both sides for the same volume of workAs published on 5 August 2026Anthropic published price list
$2 then $3A published price with an expiry date printed on it: Claude Sonnet 5 input is $2 per million tokens through 31 August 2026 and $3 from 1 September 2026, with output moving from $10 to $15 - the price list itself states when its own numbers stop being trueRead 5 August 2026; change effective 1 September 2026Anthropic published price list
25xThe span across OpenAI's current GPT-5.6 models: the flagship gpt-5.6-sol at $5.00 per million input tokens and $30.00 per million output tokens, against the small gpt-5.6-luna at $0.20 and $1.20 - twenty-five times the price on both sidesAs published on 5 August 2026OpenAI published price list
$2.00 or $4.00What the same input token costs at Google depending on how long the prompt is: Gemini 3.1 Pro Preview input is $2.00 per million tokens up to 200,000 tokens of context and $4.00 above it, with output moving from $12.00 to $18.00 - the task's shape, not just the model, sets the rateAs published on 5 August 2026Google Gemini API published price list
Half priceWhat all three vendors publish for asynchronous batch processing - the discount for not needing the answer immediately: Anthropic's batch table halves every standard rate, OpenAI's Batch tier prices are half its standard rates ($2.50 and $15.00 against $5.00 and $30.00 on its flagship), and Google states a 50% cost reductionAs published on 5 August 2026Anthropic, OpenAI and Google published price lists
One tenthWhat repeated context costs once cached: Anthropic prices a cache read at 0.1 times its base input rate, and OpenAI's cached-input rates are one tenth of its standard input rates ($0.50 against $5.00 on its flagship) - two vendors independently pricing re-read context at 90% offAs published on 5 August 2026Anthropic and OpenAI published price lists
New use casesThe most common reason a real AI bill comes in higher than its estimate, in the studio's experience: scope creep. Once a client sees the system working well, they insist on adding new use cases, and token consumption grows past what was estimated - the rate card never moved, the volume didAs of August 2026Foxnut Studios engagement records (not externally retrievable)
Unnecessary upgradesThe second common overrun cause the studio sees: client-insisted upgrades to state of the art models when the initially configured, cheaper model handles the task fine - the work does not change, but every token of it is billed at the frontier rateAs of August 2026Foxnut Studios engagement records (not externally retrievable)

From a price list to a cost per task

The published unit is a million tokens, but nobody buys a million tokens - they run tasks, and a task’s cost is its input tokens at the input rate plus its output tokens at the output rate. The arithmetic on the cited rates, labelled as arithmetic and not as a measurement: a task that reads 2,000 tokens and writes 500 costs about 2.3 cents on a $5-and-$25 frontier model, about 0.45 cents on a $1-and-$5 small model, and about a tenth of a cent on the cheapest current small model on these lists. That is the whole method. What makes real per-task costs vary a hundredfold is not the rate card but the multipliers on top of it: how many tokens the task actually consumes (a retrieval-heavy task can read fifty times what it writes), which tier the task genuinely needs, whether the work can run in a batch at half price, and how much of the context is repeated and therefore cacheable at a tenth of the input rate. A cost estimate that starts from the task’s token profile and applies the published rates is checkable to the cent; an estimate that starts from a circulating monthly average is not checkable at all.

What most cost comparisons get wrong

The first mistake is comparing input prices alone. Output tokens are five to six times the input price on every row cited here, and one vendor’s list states outright that output prices include the model’s thinking tokens - so a reasoning-heavy task can bill most of its cost on the side of the rate card a comparison ignored. The second mistake is treating today’s number as durable. One cited rate carries its own expiry date on the page, another carries a Preview suffix, and this page’s own retrieval dates exist because any of these rows can be different next quarter. The third is ignoring the published discounts, which are not fine print: half price for batch work and 90% off cached context are on the same pages as the headline rates, and for high-volume repetitive workloads they move the bill more than the choice of vendor does. The last is stopping at the rate card entirely - which is where the next section, and the studio’s own experience, comes in.

Why real bills drift from the estimate

Nothing in the records above explains an overrun, because in the studio’s experience the overruns do not come from the prices. Two patterns account for most of the drift, and both are stated here plainly as the founder’s observed record, not as industry statistics. The first is scope creep: once a client sees the system working well, they insist on adding new use cases, and each one adds token consumption the original estimate never contained. The estimate was right about the system it priced; the system it priced is not the system that ends up running. The second is unnecessary model upgrades: clients sometimes insist on state of the art models even when the initially configured, cheaper model does the task perfectly well - and on the rate cards above, that single decision multiplies the same volume of work by five to twenty-five times. Both causes are decisions, not surprises, which is why the studio treats them as scoping conversations rather than billing incidents. What a running system costs its owner beyond the model bill - monitoring, people, the weeks after deployment - is a separate question, answered separately on this site.

What this page cannot show you

This page publishes no per-task token measurements from the studio’s own systems: client token bills are client-confidential, and a per-task figure without its workflow context would be a misleading benchmark rather than a useful one. It also does not rank vendors - the lists cited price different model families with different capabilities, and a cheaper token that needs more attempts is not cheaper. Three classes of figure are deliberately absent: circulating monthly averages for what an LLM costs that trace to no primary list, Google’s context-caching rates (priced on a different basis - per-token charges plus hourly storage - and not comparable row-for-row with the other two vendors), and any prediction of where prices go next. The one dated claim this page does make beyond the lists - the two overrun causes - is asserted from the studio’s engagement records and labelled as exactly that.

How these figures were compiled

Every price on this page was fetched and read from the vendor’s own live pricing page on 5 August 2026: Anthropic’s Claude pricing page, OpenAI’s API pricing page, and Google’s Gemini Developer API pricing page, each linked in the sources with its retrieval date. The OpenAI and Google pages were each read twice the same day and the rows quoted here were consistent across reads. Aggregator comparisons and secondary roundups were not used, even where they quote the same lists - prices change often enough that only the primary page is evidence. The worked cost-per-task figures are arithmetic on the cited rates and are labelled as arithmetic in place. The two overrun causes come from one source only, the studio’s own engagement records as summarised by the founder on 5 August 2026, edited for publication without adding numbers, and are presented as operating experience rather than dressed up as market data. Where this page asserts, it says it asserts.

Sources

  1. Anthropic, Claude pricing page: per-million-token input and output rates for every current Claude model, batch rates at half the standard rates, cache reads at 0.1x the base input price, and the Claude Sonnet 5 introductory-pricing note, read on the page Retrieved
  2. OpenAI, API pricing page: per-million-token input, cached-input and output rates for the GPT-5 family through the GPT-5.6 models, with a Batch tier priced at half the standard rates, read on the page Retrieved
  3. Google, Gemini Developer API pricing page: paid-tier per-million-token rates for the current Gemini models, including the higher input rate above 200,000 tokens of context and the 50% batch cost reduction, read on the page Retrieved

Foxnut Studios works on briefs like this one from Bengaluru and Paris. If you want the shape of that before you talk to anyone, here is what an AI engagement costs before you ask.