Subscription Usage vs. API Token Economics.
A same-dollar comparison of how much AI model usage consumer subscriptions provide relative to spending the same monthly dollar amount through corresponding APIs.
Summary
This study compares two ways of paying for access to advanced AI models: purchasing a consumer subscription and purchasing usage directly through an API. The governing rule is equal monetary value. Every subscription is paired with an API budget equal to that plan's actual monthly price. A $7.99 plan is therefore compared with $7.99 of API usage, a $20 plan with $20, a $30 plan with $30, a $100 plan with $100, and a $200 plan with $200. Standardized $20 charts appear only as market-reference visuals; they are not used in place of the plan-specific equal-dollar comparison.
The principal research question is how much comparable model usage the user receives through each channel for the same out-of-pocket amount. The API side can usually be calculated precisely from published input and output token prices. The subscription side is more difficult because most providers meter access through rolling windows, weekly compute pools, model-dependent allowances, relative multipliers, fair-use policies, or dynamic limits rather than a fixed monthly token balance. This report therefore expresses the usage difference as an exact difference when public data permits it, as a scenario range when published message or credit limits can be converted responsibly, or as a same-dollar break-even threshold when the subscription allowance is not published in token terms.
Raw model usage is kept separate from bundled product value. Consumer subscriptions may include search, memory, file handling, voice, image generation, research tools, connectors, agents, or other capabilities that can make the subscription attractive even when raw token throughput is not superior. APIs provide different advantages, including metering precision, automation, model pinning, prompt caching, batching, routing, orchestration, and granular spend control. The central comparison in this paper is therefore model usage received for the same dollars; non-token features are discussed separately so they do not obscure that result.
The central finding of this study is methodological but directly useful: subscription usage and API usage can be compared fairly only when the dollar budget, model, effort level, context conditions, and workload shape are held as constant as the products allow. For every paid plan, the API budget is set equal to the subscription's monthly price. The analysis then asks how much matched-model API usage that amount purchases and whether the subscription's included usage is greater than, approximately equal to, or less than that amount under a comparable workload.
This equal-dollar rule is essential because a $100 or $200 subscription should not be judged against a $20 API example. If API pricing is linear, a $200 API budget buys ten times the token capacity of a $20 budget at the same model, effort, context tier, and input/output mix. Higher-priced subscriptions therefore face proportionally higher API comparison thresholds. The only material departures from linearity occur when the effective API rate itself changes through caching, Batch or Flex discounts, premium service tiers, promotional prices, long-context surcharges, bundled credits, or separate search and tool charges.
The report therefore produces three kinds of answers. Where a subscription exposes a directly convertible allowance, the paper can calculate an actual usage difference. Where a provider publishes message limits or relative multipliers, the paper can estimate scenarios or compare how included usage scales with price. Where the provider exposes only a dynamic compute pool, the paper calculates the exact same-dollar API capacity and uses it as the subscription's break-even point: a subscription provides more raw model usage only if the user can sustainably consume more comparable work than that threshold. This is not a retreat from the comparison; it is the most accurate way to compare products whose subscription-side metering is intentionally not expressed in tokens. [S01, S10, S18, S24, S27]
The break-even threshold is the point at which the two access methods deliver the same amount of comparable model usage for the same dollars. If a $20 subscription supports more than the number of matched-model tokens or standardized tasks that $20 buys through the API, the subscription has the raw-usage advantage for that workload. If it supports less, the API has the raw-usage advantage. When the subscription allowance is unpublished, the paper states this relationship conditionally and provides a measurement protocol so a reader can convert their own sustained subscription use into an observed advantage or disadvantage rather than relying on an invented universal quota.
Core findings
The provider-specific findings should be read as matched pairs. The first side of each pair describes the subscription: its price, model access, reset cadence, usage pool, published multipliers, and any other limit information that determines how much work a user can sustain. The second side converts exactly the same dollar amount into API capacity for the matched model and workload. The comparison then states whether the difference is directly measurable, scenario-estimable, or conditional on the user's observed subscription throughput. This structure keeps the paper focused on the amount of usable AI access received through each payment method rather than treating API pricing as an isolated rate-card exercise.
A useful way to interpret every result is to think in terms of comparable tasks rather than tokens alone. If the subscription and API are given the same prompt class, similar context, the same named model where possible, the same reasoning objective, and the same output requirement, the economic question becomes how many such tasks each payment method can sustain before the subscription reaches its allowance or the API reaches the plan-equivalent dollar ceiling. Tokens remain the billing denominator for API arithmetic, but standardized task counts make the result easier to understand and reduce the risk of treating one million tokens from different models as one million units of identical productive work.
ChatGPT Plus at $20 vs GPT-5.6 Sol API
At current promotional Sol pricing, the same $20 buys 5.00M input-only tokens, 1.00M output-only tokens, or 2.50M total tokens at the 75/25 baseline. ChatGPT Plus therefore becomes cheaper on raw token throughput only if its included monthly Sol usage exceeds the relevant workload-specific API threshold. OpenAI does not currently publish a fixed Plus token pool.
The practical implication is that the 2.50M-token figure is not an estimate of what Plus contains; it is the amount of baseline Sol token throughput that $20 could purchase directly. If an individual Plus user's comparable monthly use is materially above that level, the subscription has a raw-token advantage for that workload. If the user's measured use is below it, the API may be cheaper before assigning value to ChatGPT's bundled interface and tools.
Claude Pro at $20 vs Claude Sonnet 5 API
The same $20 buys 10.00M input-only, 2.00M output-only, or 5.00M total baseline tokens. Claude Pro has rolling/session and weekly limits whose consumption depends on message length, files, model and features; a precise monthly token total is not public.
Because Claude's subscription allowance is multi-dimensional, two Pro users can reach limits after very different amounts of tokenized work. A user with long document-heavy conversations can consume capacity faster than one sending short prompts, even if the visible message count is similar. The five-million-token API threshold is therefore best used as a measurement target against the user's own observed monthly activity.
Google AI Pro at $19.99 vs Gemini API
Using the closest current Pro-family API SKU, Gemini 3.1 Pro Preview, $19.99 buys about 4.44M baseline tokens for <=200K prompts. If the workload is served by Gemini 3.7 Flash, the same spend buys about 13.33M baseline tokens. The Gemini app does not expose a pinned API SKU and meters compute rather than tokens, so these are family-crosswalk benchmarks, not exact backend equivalence.
The large difference between Pro-class and Flash-class API capacity shows why model identity matters more than the subscription's sticker price. A single consumer plan can expose multiple model families whose direct API economics differ by several times. Any attempt to describe Google AI Pro with one universal token-equivalent number would therefore hide the model mix that actually drives cost.
SuperGrok at $30 vs Grok 4.6 API
$30 buys 15.00M input-only, 5.00M output-only, or 10.00M baseline Grok 4.6 tokens. SuperGrok draws from a shared weekly compute allowance across Grok products; xAI does not publish the pool as tokens.
For a user who consumes Grok features beyond text chat, the shared allowance means the economic comparison should consider the entire paid-product workload. Image, voice, build, or other compute-intensive activities can reduce the capacity remaining for chat. A text-only API-equivalent benchmark therefore provides a clean model-level reference but should not be interpreted as a complete substitute for the bundled consumer experience.
Mistral Vibe Pro
The unusual economics are the bundled API credits. Mistral's current pricing page displays $30/month in API credits with Pro while the normal Pro list price is $14.99. If those credits are available to the user under the displayed offer, their face value alone is roughly 2.0x the plan's monthly price before assigning any value to Vibe chat/coding usage.
This creates a two-part value proposition: a fair-use consumer product plus a separately metered API balance. The correct accounting treatment is to keep those components distinct, then evaluate their combined economic value. Counting the credit balance as if it were chat usage would overstate the subscription's token quota; ignoring the credit balance would understate the plan's total economic benefit.
Perplexity
Perplexity subscriptions expose multiple third-party models rather than one provider model. The appropriate API-equivalent benchmark is the direct first-party rate for the exact named model, or the Perplexity Agent API pass-through rate where it is available and explicitly matched. A $20 Pro plan therefore has materially different API break-even thresholds depending on whether the user selects Terra, Sonnet 5, Gemini 3.1 Pro, Grok 4.5, or another included model.
This means that a Perplexity user's own model-selection distribution is part of the cost equation. A month dominated by premium models can represent more direct-API value than a month dominated by lower-cost models, even if the number of visible user interactions is identical. The model-by-model API-equivalent comparison table later in the report is designed to make that variation explicit.
Perplexity's economics are therefore portfolio economics. The subscription can be valuable precisely because it lets a user move among providers without opening separate accounts, but that convenience means the underlying marginal cost varies with every model choice. A robust valuation should record the distribution of usage across models and distinguish native Sonar workflows from third-party model calls. Without that decomposition, a single average token-equivalent figure can obscure the premium-model access that may be the primary reason a user purchases Max rather than Pro.
DeepSeek
DeepSeek now offers V4 Pro in web/app Expert Mode and via API, but the first-party web/app remains free. Because there is no paid subscription price to match, a subscription-versus-equal-dollar-API comparison cannot be constructed on the same basis used for paid plans. DeepSeek is therefore included as a provider-coverage case rather than being assigned an artificial $20 subscription comparison. [S04, S10, S18, S25, S27, S30, S35]
The separate $20 API reference shown later is therefore a market-normalization example only. It answers how much DeepSeek API capacity $20 could buy; it does not represent the price of a DeepSeek subscription. Maintaining that distinction prevents standardized market charts from being mistaken for the core equal-dollar methodology.
DeepSeek also demonstrates why a comprehensive market report can contain useful API analysis even when the subscription comparison is undefined. For developers, the relevant questions include peak versus off-peak scheduling, cache reuse, and model selection; for consumer users, the relevant fact is that the first-party app does not create a paid monthly benchmark to match. Preserving those two facts side by side is more informative than excluding the provider or inventing a hypothetical subscription price solely to make the tables look symmetrical.
How value is interpreted in this report
Raw token value and product value are deliberately separated. A subscription may be economically preferable even before it beats the API token threshold because it bundles web search, connectors, file handling, memory, voice, code execution, research workflows, storage, image/video generation, browser agents, collaboration, and consumer UX. Conversely, an API can be preferable even when its raw token volume is lower because it offers deterministic metering, automation, batching, prompt caching, model pinning, programmatic orchestration, and workload-level cost controls. The report therefore presents token economics as one decision axis, not a universal product-value score.
For procurement purposes, raw token economics should be combined with a total-cost-of-ownership view. API use can require engineering, observability, retry logic, security controls, storage, retrieval infrastructure, and user-interface development that a subscription product already bundles. A subscription can likewise impose workflow constraints, weaker automation, less deterministic model selection, or limited integration with internal systems. The report therefore treats token cost as a measurable economic component inside a broader product decision rather than as a substitute for evaluating reliability, capability, governance, and operational fit.
Figure 1
Figure 1. API token capacity obtainable for the same monthly dollars as representative subscriptions. Core assumption: uncached text, short-context tier, 75% input / 25% billable output.
Figure 2
Figure 2. Market reference: total uncached tokens purchasable for $20 at a 75/25 input/output workload. DeepSeek uses off-peak pricing; Gemini 3.7 Flash uses its current 2026 promotional rate.
Methodology for Comparing Subscription Usage with Same-Dollar API Usage
A rigorous comparison begins with a matched-use definition. The subscription and API should be evaluated at the same monetary value, on the same named model where a public mapping exists, at the closest available reasoning or effort level, and under the same prompt, context, output, and tool conditions. If one of those variables changes, the result no longer isolates the difference between subscription access and API access. A cheaper API model may be a better engineering choice, but it is a model-routing comparison rather than a delivery-channel comparison.
The monetary threshold is the subscription's current monthly list price unless a row explicitly analyzes another billing basis. Taxes, regional price differences, annual-prepayment discounts, introductory offers, and negotiated enterprise contracts are excluded from the core comparison because they change the cash amount being matched. When a promotional API rate is current and first-party documented, the report uses that rate and labels the time condition because the same subscription price can purchase materially different API capacity when the API denominator changes.
The methodology distinguishes three levels of evidence on the subscription side. A published token, credit, or message allowance can sometimes be converted directly or through a clearly labeled workload scenario. A published relative multiplier can quantify how usage scales inside one provider's plan family even when the absolute baseline is unknown. A dynamic or undisclosed compute pool cannot be converted into a defensible universal token total from public documentation alone; in that case the exact same-dollar API capacity becomes the comparison threshold and the subscription difference must be established by observed use. This hierarchy prevents false precision while still answering the economic question as far as the evidence permits.
Unit of comparison
For a plan with monthly price P, the report asks: if the user spends exactly P dollars on the same model through an API, how many tokens can the user buy? Where the consumer product does not disclose its token allowance, the resulting API quantity is the subscription break-even threshold. If the subscription delivers more comparable billable model tokens than that threshold over the billing period, subscription chat is cheaper on raw token throughput; if it delivers less, direct API usage is cheaper on raw tokens.
Using the monthly subscription price as the API budget also prevents the analysis from favoring lower-cost plans by construction. The comparison scales automatically: a higher-priced plan must clear a higher API-equivalent threshold because the same dollars could have purchased proportionally more direct model usage. This makes the method suitable for comparing $8, $20, $30, $100, $200, and enterprise-priced tiers without changing the governing equation. It also means that annual discounts should be converted to an effective monthly cost only when the analysis is explicitly evaluating the annual commitment rather than the month-to-month plan.
Core workload scenarios
| Scenario | Token mix | Representative workload |
|---|---|---|
| Input-heavy | 90% input / 10% output | RAG, document analysis, long-context Q&A, codebase reading |
| Baseline | 75% input / 25% output | General developer/knowledge-work interaction; principal comparison |
| Output-heavy | 50% input / 50% output | Long-form generation, reasoning-heavy answers, code generation |
Figure 3
Figure 3. $20 API purchasing power under three input/output mixes. Higher output share reduces total-token purchasing power because output tokens are generally priced at a premium.
How the Usage Difference Is Reported. When both sides can be expressed in comparable units, the usage advantage is reported as the percentage difference between sustainable subscription throughput and the amount purchasable through the API at the same monthly dollar value. A positive result means the subscription provides more standardized model consumption for the money; a negative result means the API provides more. When only the API side is precisely quantifiable, the same-dollar API amount becomes a break-even threshold rather than a claim about the subscription's actual token capacity. This distinction is maintained throughout the paper so that readers can separate measured differences, defensible scenario estimates, and thresholds that still require account-level observation.
When comparable subscription usage can be measured as Tsub and the same-dollar API capacity is Tapi, the absolute usage difference is Tsub - Tapi. The subscription usage advantage is ((Tsub / Tapi) - 1) × 100 percent when Tsub exceeds Tapi. If the API quantity is larger, the API advantage can be expressed as ((Tapi / Tsub) - 1) × 100 percent. These formulas are intentionally symmetric with respect to the economic question: they compare how much matched work is obtained for the same cash outlay rather than comparing unit prices in isolation.
A direct difference is reported only when the subscription-side allowance can be measured or converted with reasonable confidence. This can occur when the provider exposes a token or credit balance, when a business plan publishes a message range that can be paired with a clearly defined token-per-task scenario, or when an organization has its own sustained telemetry from the subscription. In those cases the paper can state both the same-dollar API capacity and the subscription's observed or scenario-based capacity, then calculate the percentage advantage.
When the subscription uses a dynamic five-hour window, a weekly compute pool, or a fair-use system without an absolute public balance, the paper reports a decision threshold instead of pretending that the missing value is known. The threshold is still actionable: it specifies exactly how much comparable subscription usage must be achieved before the fixed subscription becomes more economical than spending the same dollars through the API. A reader can then compare their own real usage against that threshold.
Relative multipliers are treated as a separate result type. If a $100 tier provides 5x the usage of a $20 baseline, price and usage scale together and the published included-usage efficiency is unchanged relative to the baseline. If a $200 tier provides 20x the usage of the same $20 baseline, price rises 10x while published usage rises 20x, implying 2x the relative included usage per dollar inside that provider's subscription family. This does not reveal absolute tokens, but it does reveal whether upgrading improves the subscription side faster or slower than an API budget, which normally scales linearly with spend.
The three workload shapes are intended to bracket common text-use patterns, not to describe universal user behavior. A 90/10 mix approximates input-heavy work such as classification, retrieval-assisted analysis, or summarization where the model produces a relatively compact answer. The 50/50 case represents generation-heavy work with substantial output or reasoning. The 75/25 case is retained as a neutral planning midpoint. Because output tokens are usually priced above input tokens, moving toward an output-heavy workload lowers the number of total tokens that a fixed API budget can support even when the model rate card is unchanged. The baseline is not a claim that all users produce exactly a 75/25 mix; it is a neutral normalization. [S04, S11, S21, S26, S28]
Standardized Task Equivalents
Token capacity is the most auditable API billing measure, but readers generally care about completed work. For that reason, the report treats a standardized task as a repeatable prompt-and-response workload with a defined input size, expected output size, context condition, model, and reasoning objective. The subscription and API versions of a task should be as similar as the product surfaces allow. The number of comparable tasks per dollar can then be calculated by dividing the equal-dollar API budget by the API cost of one standardized task and comparing that count with how many of the same tasks the subscription can sustain before its applicable limit is reached.
Standardized tasks are especially useful when subscription plans are described in messages rather than tokens. A message cap is not itself a token allowance because one message can contain a short question while another can carry a long document, extensive conversation history, tool results, and a large reasoning trace. By defining representative short-interaction, professional, coding, document-analysis, and long-context tasks, the analyst can convert a published message range into a scenario range without claiming that every message has the same computational weight.
Task-equivalent comparisons remain model-specific. A task completed by a small fast model and the same task completed by a premium reasoning model may differ in quality, latency, and tokenization even when the prompt text is identical. The report therefore does not use standardized tasks to claim that all models produce equal value per task. Their purpose is narrower: to compare the subscription and API access methods for the same model or the closest documented model mapping under a consistent workload.
Reasoning and effort normalization
Reasoning modes create a second comparability problem. OpenAI, Anthropic, Google, xAI, Mistral and DeepSeek expose effort controls, but the names are not semantically interchangeable and most are behavioral signals rather than fixed token budgets. A "high" response can use different hidden or visible reasoning volumes from one request to the next. Accordingly, the primary token-capacity tables do not pretend that an effort label fixes the number of output tokens. Instead, they keep the model and effort condition constant conceptually and report how much billable token volume the equal-dollar API budget can fund. A separate scenario model illustrates task-count sensitivity as reasoning spend rises. [S12, S22, S25, S29, S37]
For the illustrative effort-sensitivity figure only, a standardized task is defined as 2,000 input tokens plus 750 visible output tokens. Billable-output multipliers are 1.0x for none/minimal, 1.5x low, 2.0x medium, 4.0x high, 8.0x xhigh and 16.0x max. These multipliers are deliberately labeled as scenario assumptions; they are not vendor guarantees, not measured model averages, and not used to claim a subscription token quota.
Figure 4
Figure 4. Illustrative effect of increasing billable reasoning/output on standardized tasks per $20. The curve is a sensitivity model, not an official effort-to-token mapping.
What Counts as Comparable Subscription Usage
For a subscription-versus-API comparison, the relevant subscription quantity is the amount of model work that would have been billable if the same workload had been sent through the matched API. Where possible, this includes user input, carried conversation or document context, visible model output, and reasoning or thinking tokens when the provider exposes them. Tool calls and retrieval results should be included only when the API comparison prices the corresponding tool or context in the same way. The goal is not to reconstruct the provider's private infrastructure cost; it is to estimate the public API-equivalent usage represented by the user's subscription workload.
Consumer applications can add hidden system instructions, summarize history, cache repeated context, route requests among model variants, or inject search and tool content that the user cannot fully observe. Those hidden operations mean a subscription-side token count is often an estimate of comparable public-model usage rather than a literal accounting of every internal token processed by the provider. The report treats this uncertainty explicitly and does not inflate the subscription estimate with guessed system or routing overhead.
Rolling reset windows should be normalized from sustained use rather than multiplied into a theoretical twenty-four-hour maximum. A user who can consume a five-hour allowance several times per day may still encounter a weekly cap, and a weekly pool may be consumed faster by long contexts, higher effort, or tool-heavy work. For public-facing comparisons, a representative measurement period of several weeks is more defensible than a one-day stress test because it captures both local resets and higher-level limits.
The paper also separates raw text-model usage from bundled consumer features. If a subscription includes image generation, voice, storage, browser actions, research workflows, or connectors, those features may improve the total value proposition but they should not be converted into imaginary text tokens. The primary comparison remains the amount of matched language-model work received for the same dollars; bundled capabilities are discussed as additional product value after the raw-usage comparison is established.
Scope of the Analysis
The principal comparisons in this study use standard, uncached text-token pricing and normal-speed API inference unless a provider offers only one applicable text-inference tier. This establishes a consistent baseline across providers and reduces the risk that implementation-specific optimizations will distort the underlying comparison between subscription pricing and direct API consumption. The resulting figures should therefore be interpreted as standardized cost-equivalent estimates rather than representations of every possible production configuration.
Prompt caching is evaluated separately from the principal calculations because its economic benefit depends on workload architecture, prompt reuse, cache-retention policies, provider-specific pricing, and the proportion of input tokens that qualify for discounted cached-input rates. Consumer AI applications may also employ internal caching or context-management techniques that are not disclosed to users. For these reasons, assuming a universal cache-hit rate would create artificial precision and would reduce comparability between consumer subscriptions and independently operated API workloads.
API delivery options such as Batch, Flex, Fast, Priority, or comparable service tiers are likewise treated independently from model reasoning or effort settings. These mechanisms primarily affect processing priority, latency, availability, scheduling, or price and should not be interpreted as equivalent to a model's reasoning-effort control. Where such service tiers materially alter API economics, their effect is examined separately so that readers can distinguish the underlying model price from the cost implications of a particular delivery configuration.
The core token-capacity calculations focus on text-model inference and therefore do not ordinarily incorporate separate charges associated with web search, retrieval, citations, computer use, image generation, video generation, audio processing, or third-party tool calls. Where such functionality is inseparable from the published price of a model or API product, the associated cost is incorporated or specifically identified. This distinction is necessary because consumer subscriptions frequently bundle capabilities that an API user may need to purchase or operate separately. Consequently, the token comparisons presented in this report should not be interpreted as a complete measure of the total product value provided by a consumer subscription.
Long-context pricing is also treated separately from the standard baseline. Several providers apply different rates once prompts or contexts exceed specified token thresholds. Unless otherwise identified, the principal comparisons assume workloads remain within the provider's standard pricing range. Workloads that regularly process very large documents, repositories, retrieval corpora, or extended conversation histories may therefore experience different effective API costs from those shown in the baseline scenarios.
Context-window capacity should not be interpreted as usage allowance. A model advertised with a context window of hundreds of thousands or one million tokens can potentially address that quantity of information within an individual request or active context, subject to provider restrictions, but this does not imply that a subscription includes an equivalent monthly quantity of usable tokens. Context-window size describes the model's addressable working context, whereas subscription usage limits describe how frequently or extensively the service may be consumed over time.
Finally, token counts are not perfectly interchangeable across model families. Providers may use different tokenization systems, and newer generations of a provider's own models may employ different tokenizers from earlier generations. As a result, one million tokens purchased from two different APIs does not necessarily correspond to an identical quantity of English text, source code, multilingual content, or completed semantic work. Token capacity is therefore used in this report as a standardized billing and consumption measure, not as a claim that an equal number of tokens represents equal model capability, productivity, or task completion.
Provider Comparability Scorecard
The scorecard is an evidence-quality tool as much as a product comparison. High comparability means that a named model can be identified on both the subscription surface and a first-party API rate card. Medium comparability means that a model family is visible to the consumer but the provider may route among versioned backends, so the API figure serves as a family-level economic reference rather than a guaranteed backend match. N/A is used when an equal-dollar comparison cannot be constructed, such as for a free first-party assistant or an open-weight model without a directly matched paid consumer plan.
This distinction matters because false precision can reverse a procurement conclusion. If a subscription UI says only "Pro" or "Flash" while the API contains several versions with materially different prices, selecting one SKU without qualification can make the subscription appear artificially cheap or expensive. The scorecard therefore tells the reader how much confidence to place in each quantitative row before the raw numbers are used for budgeting.
A "High" comparability rating does not mean the subscription token total is known. It means the same model can be identified on both sides, allowing a defensible API break-even benchmark. "Medium" typically reflects family routing or product-layer abstraction. "N/A" means one side of the proposed equal-dollar comparison does not exist.
Provider comparability scorecard
| Provider | Comparability | Subscription↔API identity | Subscription metering | API price transparency | Main caveat |
|---|---|---|---|---|---|
| OpenAI | High | Named GPT-5.6 model matches; Pro mode uses Sol Pro and requires separate treatment | Dynamic on most plans; Business has local-message estimates | Excellent | Promo/API and Business rules change frequently |
| Anthropic | High | Named Claude model matches; Fable 5.1 included only on Max/premium entitlements; Mythos 5.1 is trusted-access only | Dynamic 5h/session + weekly limits | Excellent | Tokenizer and effort affect consumption; Fable 5.1 has model-specific weekly inclusion rules and cheaper cache reads; Mythos 5.1 has no general consumer-plan comparator |
| Medium | App family names; API uses versioned SKUs | Compute-based 5h + weekly pool | Excellent API pricing; medium app parity | Backend family-to-SKU mapping not pinned | |
| xAI | High for Grok 4.6 | Named Grok 4.6 match | Shared weekly compute pool | Good | No published numeric pool |
| Mistral | Medium | Vibe routes latest models; API names exact | Fair-use rather than tokenized | Excellent | Bundled API credits materially change economics |
| Perplexity | High for named third-party model identity; API benchmark varies | Explicit model selector | Plan limits/credits not tokenized | Good | Platform bundles search/orchestration; native Sonar adds request fees |
| DeepSeek | No paid-plan comparison | Named V4 app/API model match | Consumer access is free | Excellent | No paid subscription price to match |
| Meta / Amazon / NVIDIA / others | Usually low or N/A | Varies | No directly matched paid consumer plan | Varies | Audited in coverage appendix |
Subscription Usage Versus Equal-Dollar API Usage
This chapter is the quantitative center of the report because it joins the two sides of the research question. The subscription side records what the provider actually promises or exposes: a fixed allowance, a message range, a relative multiplier, a rolling or weekly compute pool, or a fair-use policy. The API side converts the same monthly dollar amount into matched-model token capacity under the report's standardized workload. The final column explains what can and cannot be concluded about the usage difference from those two pieces of evidence.
The 75/25 input/output workload remains the central numerical reference because it represents a mixed professional workload without assuming that every user is input-heavy or output-heavy. The sensitivity tables that follow show how the API side changes under 90/10 and 50/50 workloads. Subscription consumption can also change with prompt size, reasoning effort, tool use, and chat length, so the same workload definition should be used when measuring the subscription side.
The primary comparison matrix should be read as a set of equal-cost decision points, not as a claim that every subscription has a hidden fixed token balance. For plans with unpublished absolute usage, the API capacity is the amount of comparable subscription work that must be exceeded before the subscription has the raw-throughput advantage. For plans with relative multipliers, the table also notes whether usage scales faster, slower, or approximately in line with the increase in subscription price.
The comparison is most informative when combined with the reader's own sustained workload. A heavy interactive user may clear a same-dollar API threshold and obtain substantially more model work from a fixed subscription fee. A light or intermittent user may consume far less than the threshold, in which case the API can be economically preferable because spending stops when usage stops. The report therefore treats subscription-versus-API value as a utilization problem as well as a pricing problem.
Primary Usage Comparison Matrix
The primary comparison matrix is designed to be read as a usage-value table, not merely a pricing table. The API column states how much standardized text-model consumption the subscription price can purchase at published API rates. The subscription-evidence column states what the provider actually discloses about included usage, such as a fixed allowance, a rolling limit, a weekly pool, or a multiplier relative to another plan. The final column then explains the decision boundary: when the subscription delivers more comparable usage than the equal-dollar API threshold, the subscription has the raw-usage advantage; when it delivers less, the API has the advantage. Where the subscription allowance cannot be converted responsibly into an absolute token quantity, the report preserves that uncertainty instead of presenting an invented figure.
The table below summarizes the main equal-dollar comparison for representative paid plans. API capacity is total uncached text tokens at the 75/25 baseline and is model-specific. Where a plan exposes several materially different models, a range or multiple model references are shown. "Threshold" means the provider does not publish an absolute subscription token pool; the listed API quantity is therefore the amount of comparable subscription usage that must be exceeded for the subscription to provide more raw model throughput at the same monthly cost.
Primary subscription-versus-API comparison
| Provider / plan | Price | Matched model / benchmark | Same-dollar API capacity | Subscription usage evidence | How to read it |
|---|---|---|---|---|---|
| OpenAI ChatGPT Go | $8 | GPT-5.6 Luna | 17.778M | Unlimited everyday text chats; separate tool/feature limits | No fixed token pool. Compare sustained Luna-equivalent text use with 17.778M; tools remain separate. |
| OpenAI ChatGPT Plus | $20 | GPT-5.6 Sol | 2.500M | Dynamic plan-level reasoning allowance; no public monthly token pool | Subscription has the raw-usage advantage only if comparable monthly Sol use exceeds 2.500M at the baseline workload. |
| OpenAI Pro $100 | $100 | GPT-5.6 Sol benchmark | 12.500M | Published as 5x Plus usage | 5x usage at 5x the Plus price implies roughly unchanged relative included-usage efficiency. Pro mode itself uses Sol Pro, for which this report does not assign an unverified public API token rate. |
| OpenAI Pro $200 | $200 | GPT-5.6 Sol benchmark | 25.000M | Published as 20x Plus usage | 20x usage at 10x the Plus price implies 2x the published relative usage per dollar versus Plus. Sol table is a Sol benchmark, not an exact Sol Pro comparison. |
| Anthropic Claude Pro | $20 | Claude Sonnet 5 | 5.000M | 5-hour/session and weekly limits; usage varies by model, length, files, features, and effort | Subscription clears the raw-usage break-even point only when sustained comparable Sonnet use exceeds 5.000M baseline tokens. |
| Anthropic Claude Max 5x | $100 | Sonnet 5 / Opus 5 / Fable 5.1 | 5.000M-25.000M | 5x Pro per session plus weekly limits; Fable models may use up to 50% of the regular weekly limit at no extra cost | 5x usage at 5x price is roughly linear relative to Pro. Fable 5.1 is an included Max model, but its included use is capped within the shared weekly pool. |
| Anthropic Claude Max 20x | $200 | Sonnet 5 / Opus 5 / Fable 5.1 | 10.000M-50.000M | 20x Pro per session plus weekly limits; Fable models may use up to 50% of the regular weekly limit at no extra cost | 20x usage at 10x price implies 2x relative included-usage efficiency versus Pro before model-mix effects. |
| Google AI Pro | $19.99 | Gemini 3.1 Pro / 3.7 Flash | 4.442M / 13.327M | 4x standard Gemini-app usage; compute-based 5-hour refreshes until weekly limit | Model family determines the API threshold. App routing prevents one universal token-equivalent figure for the plan. |
| xAI SuperGrok | $30 | Grok 4.6 | 10.000M | Shared weekly usage pool across Grok products | Text-chat subscription use must exceed 10.000M comparable Grok 4.6 baseline tokens to beat $30 of text API throughput. |
| Mistral Vibe Pro | $14.99 | Mistral Medium 3.5 reference | 4.997M | Fair-use Vibe access plus displayed $30/month API-credit benefit | Vibe chat usage is not tokenized. The API-credit benefit is a separate directly metered component and should not be counted as chat tokens. |
| Perplexity Pro | $20 | Terra / Sonnet 5 / Gemini 3.1 Pro / Grok 4.5 | 4.444M-6.667M | Plan-limited access to multiple third-party models; no unified token pool | The correct threshold is weighted by the user's actual model mix. |
Equal-Dollar API Capacity by Plan
The detailed matrix that follows isolates the API side of the comparison. It is retained because the subscription difference cannot be interpreted without knowing the exact amount of API usage available at the same price.
Same-dollar API capacity matrix across workload mixes
| Provider | Plan | Input only | Output only | 90/10 total | 75/25 total | 50/50 total |
|---|---|---|---|---|---|---|
| OpenAI | ChatGPT Go | 40.00M | 6.67M | 26.67M | 17.78M | 11.43M |
| OpenAI | ChatGPT Plus | 5.00M | 1.00M | 3.57M | 2.50M | 1.67M |
| OpenAI | ChatGPT Pro 100 | 25.00M | 5.00M | 17.86M | 12.50M | 8.33M |
| OpenAI | ChatGPT Pro 200 | 50.00M | 10.00M | 35.71M | 25.00M | 16.67M |
| OpenAI | Business Standard | 6.25M | 1.25M | 4.46M | 3.12M | 2.08M |
| OpenAI | Business Premium | 31.25M | 6.25M | 22.32M | 15.62M | 10.42M |
| Anthropic | Claude Pro - Sonnet 5 | 10.00M | 2.00M | 7.14M | 5.00M | 3.33M |
| Anthropic | Claude Pro - Opus 5 | 4.00M | 0.80M | 2.86M | 2.00M | 1.33M |
| Anthropic | Claude Max 5x - Opus 5 | 20.00M | 4.00M | 14.29M | 10.00M | 6.67M |
| Anthropic | Claude Max 5x - Fable 5.1 | 10.00M | 2.00M | 7.14M | 5.00M | 3.33M |
| Anthropic | Claude Max 20x - Opus 5 | 40.00M | 8.00M | 28.57M | 20.00M | 13.33M |
| Anthropic | Claude Max 20x - Fable 5.1 | 20.00M | 4.00M | 14.29M | 10.00M | 6.67M |
| Google AI Plus - Gemini 3.1 Pro | 4.00M | 0.67M | 2.66M | 1.78M | 1.14M | |
| Google AI Plus - Gemini 3.7 Flash | 10.65M | 2.13M | 7.61M | 5.33M | 3.55M | |
| Google AI Pro - Gemini 3.1 Pro | 9.99M | 1.67M | 6.66M | 4.44M | 2.86M | |
| Google AI Pro - Gemini 3.7 Flash | 26.65M | 5.33M | 19.04M | 13.33M | 8.88M | |
| Google AI Ultra 100 | 50.00M | 8.33M | 33.33M | 22.22M | 14.29M | |
| Google AI Ultra 200 | 100.00M | 16.67M | 66.67M | 44.44M | 28.57M | |
| xAI | SuperGrok | 15.00M | 5.00M | 12.50M | 10.00M | 7.50M |
| xAI | SuperGrok Plus | 50.00M | 16.67M | 41.67M | 33.33M | 25.00M |
| Mistral | Vibe Pro - Medium 3.5 | 9.99M | 2.00M | 7.14M | 5.00M | 3.33M |
| Mistral | Vibe Pro - Small 4 | 99.93M | 24.98M | 76.87M | 57.10M | 39.97M |
| Mistral | Vibe Pro - Large 3 | 29.98M | 9.99M | 24.98M | 19.99M | 14.99M |
| Perplexity | Pro - GPT-5.6 Terra | 10.00M | 1.67M | 6.67M | 4.44M | 2.86M |
| Perplexity | Pro - Claude Sonnet 5 | 10.00M | 2.00M | 7.14M | 5.00M | 3.33M |
| Perplexity | Pro - Gemini 3.1 Pro | 10.00M | 1.67M | 6.67M | 4.44M | 2.86M |
| Perplexity | Pro - Grok 4.5 | 10.00M | 3.33M | 8.33M | 6.67M | 5.00M |
| Perplexity | Max - GPT-5.6 Sol | 50.00M | 10.00M | 35.71M | 25.00M | 16.67M |
| Perplexity | Max - Claude Opus 5 | 40.00M | 8.00M | 28.57M | 20.00M | 13.33M |
Subscription and API Parity Notes
The parity notes record the evidence quality behind each subscription-to-API mapping. They distinguish exact named-model matches from family-level references and cases where the consumer product does not expose a directly comparable API SKU. This qualification is essential when interpreting the threshold tables: the arithmetic can be exact for the API price while the mapping to the subscription remains conditional.
Subscription-to-API parity notes
| Provider | Plan | Parity |
|---|---|---|
| OpenAI | ChatGPT Go | Named-model match |
| OpenAI | ChatGPT Plus | Named-model match |
| OpenAI | ChatGPT Pro 100 / 200 | Sol benchmark only; Pro mode uses Sol Pro |
| OpenAI | Business Standard / Premium | Named-model match |
| Anthropic | Claude Pro / Max - Sonnet 5, Opus 5 | Named-model match |
| Anthropic | Claude Max - Fable 5.1 | Named-model match; shared weekly Fable cap applies |
| All AI plan tiers | Family crosswalk | |
| xAI | SuperGrok / SuperGrok Plus | Named-model match |
| Mistral | Vibe Pro (Medium/Small/Large) | Routing / likely |
| Perplexity | Pro - named third-party models | Named-model reference; direct-provider API benchmark |
| Perplexity | Pro - Gemini family | Named-family reference; direct-provider API benchmark |
| Perplexity | Max - GPT-5.6 Sol, Claude Opus 5 | Named-model reference; direct-provider API benchmark |
OpenAI
OpenAI's product structure illustrates why subscription price alone does not reveal token economics. The consumer tiers differ in model access and usage multipliers, while the API separately prices Sol, Terra, and Luna and can further vary effective cost through cached input, service tier, and long-context rules. The correct comparison is therefore plan-to-model specific: ChatGPT Plus should be benchmarked against the model actually used for the matched task, and $100 or $200 Pro tiers should be compared with $100 or $200 of that same API model rather than with the $20 Plus benchmark.
For developers, OpenAI's API can become materially more efficient than the uncached baseline when prompts contain reusable context. Conversely, high reasoning effort and large output volumes can deplete an API budget rapidly because reasoning is billed on the output side. This makes task-level telemetry essential. A subscription user experiences these effects indirectly as throttling or dynamic usage limits, while an API user sees them directly in token accounting and spend.
OpenAI also illustrates the difference between plan scaling and model scaling. A higher ChatGPT tier may increase access to the same premium model, while the API gives the developer an additional option to move a workload to a smaller or cheaper member of the model family. Those are different economic levers. Subscription upgrades primarily change allowance and product access; API optimization can change the model, cache behavior, service tier, and context strategy for each request. As a result, a team comparing the two channels should evaluate both the matched-model benchmark and the savings available from intelligent API routing.
Plan architecture
OpenAI now separates consumer ChatGPT, ChatGPT Work and Codex availability more explicitly. GPT-5.6 Luna is the default for Free and Go, while eligible paid ChatGPT plans use GPT-5.6 Sol for Instant/Medium/High/Extra High and Sol Pro for the Pro option. Terra and Luna are available to eligible paid users in Work/Codex even though they are not selectable in ordinary ChatGPT conversations. [S01, S02, S04, S05, S06, S07, S08]
The architecture of the plan lineup matters because the user's economic alternative is not always the same model at every price tier. Lower-cost plans can default to a smaller model, while premium tiers may unlock higher-effort variants or additional work surfaces. For a clean comparison, the analyst should first identify which model actually performs the target task on the subscription and then price that same model through the API where possible. Comparing a lower-cost subscription model with a premium API model would measure a capability change rather than a delivery-channel difference.
Subscription metering
The individual plans do not expose a stable token pool. Go advertises unlimited everyday text chats subject to separate limits for tools/features. Plus exposes plan-dependent reasoning limits. Pro 100 and Pro 200 are described as 5x and 20x the usage of Plus. Business is more quantifiable: Standard seats publish local-message estimates per five hours by model, but OpenAI explicitly warns that task size, reasoning effort, cloud execution and model choice affect usage. [S01, S02, S04, S05, S06, S07, S08]
Dynamic metering also means that visible message counts are a weak proxy for sustained capacity. A long conversation can repeatedly carry forward prior context, while tool calls and reasoning can consume resources that are not obvious from the number of user turns. The most defensible evaluation therefore samples ordinary work over multiple reset windows rather than stress-testing a new account with short prompts.
API economics
At current rates, Sol costs $4/M input and $20/M output; Terra $2/$12; Luna $0.20/$1.20. This creates enormous differences in equal-dollar throughput even inside one GPT-5.6 family. At the 75/25 baseline, Sol costs $8/M total, Terra $4.50/M, and Luna $0.45/M. [S01, S02, S04, S05, S06, S07, S08]
The gap between input and output pricing makes prompt design and response control economically material. Applications that can reuse cached context, keep outputs concise, or route simple requests to lower-cost models can increase tasks completed per dollar without changing the nominal subscription comparison. Conversely, applications that request long explanations, extensive reasoning, or repeated regeneration will move toward the output-heavy side of the sensitivity range.
Effort modes
Sol, Terra and Luna expose none, low, medium, high, xhigh and max effort controls through their public APIs. ChatGPT does not map those labels one-to-one: eligible paid plans use GPT-5.6 Sol for Medium, High, and Extra High, while the Pro option uses GPT-5.6 Sol Pro. Work and Codex expose their own model and effort controls. An effort-matched comparison should therefore match the underlying model first and then compare task quality and observed reasoning/output consumption; Pro mode should not be treated as ordinary Sol merely because both appear in the same ChatGPT plan. [S01, S04, S05, S06]
OpenAI subscription plan matrix
| Plan | Price | Cadence | Models | Published usage | Quantifiability |
|---|---|---|---|---|---|
| ChatGPT Go | $8.00 | Monthly | GPT-5.6 Luna | Unlimited everyday text chats; Think uses Luna; separate feature/tool limits | Named-model match; subscription token quota not published |
| ChatGPT Plus | $20.00 | Monthly | GPT-5.6 Sol (Medium/High/Extra High); Work/Codex: Sol/Terra/Luna | Reasoning limits depend on plan; no current public exact token quota | Named-model Sol match; dynamic allowance. Pro mode is not part of Plus |
| ChatGPT Pro 100 | $100.00 | Monthly | GPT-5.6 Sol and Sol Pro (Pro) | 5x higher usage than Plus | Published 5x relative multiplier; Sol Pro API token rate not assigned in this audit |
| ChatGPT Pro 200 | $200.00 | Monthly | GPT-5.6 Sol and Sol Pro (Pro) | 20x higher usage than Plus | Published 20x relative multiplier; Sol Pro API token rate not assigned in this audit |
| Business Standard | $25.00 | Monthly seat | Sol; Work/Codex Sol/Terra/Luna | Local-message estimates per 5h: Sol 10-100; Terra 25-200; Luna 250-2,000. Annual equivalent $20/mo | Only major current plan with model-level local-message estimates; not token quotas |
| Business Premium | $125.00 | Monthly seat | Sol; Work/Codex Sol/Terra/Luna | 5x included usage vs Standard; no 5-hour limit. Annual equivalent $100/mo | Dynamic work/task-based allowance |
OpenAI current text-model / API matrix
| Model | Input | Cached | Output | Context | Effort / thinking | Subscription parity |
|---|---|---|---|---|---|---|
| GPT-5.6 Sol | $4/M | $0.4/M | $20/M | 1.05M | none, low, medium, high, xhigh, max | Named-model match: ChatGPT Medium/High/Extra High; Work/Codex |
| GPT-5.6 Terra | $2/M | $0.2/M | $12/M | 1.05M | none, low, medium, high, xhigh, max | Named-model match: Work/Codex; not standard ChatGPT picker |
| GPT-5.6 Luna | $0.2/M | $0.02/M | $1.2/M | 1.05M | none, low, medium, high, xhigh, max | Named-model match: Go/Free default; Work/Codex |
Equal-dollar API break-even thresholds
Each row converts that plan's own monthly price into API capacity for the named model. The figures are thresholds, not claimed subscription quotas, and the three workload columns show the sensitivity to output share. Interpretation: the subscription has lower raw model-token cost only after its comparable included usage exceeds the API token count shown for the same dollar spend.
OpenAI equal-dollar API capacity by plan
| Plan | Budget | API model | 90/10 | 75/25 | 50/50 | Mapping |
|---|---|---|---|---|---|---|
| ChatGPT Go | $8.00 | GPT-5.6 Luna | 26.667M | 17.778M | 11.429M | Named-model match |
| ChatGPT Plus | $20.00 | GPT-5.6 Sol | 3.571M | 2.500M | 1.667M | Named-model match |
| ChatGPT Pro 100 | $100.00 | GPT-5.6 Sol | 17.857M | 12.500M | 8.333M | Sol benchmark only; Pro mode uses Sol Pro |
| ChatGPT Pro 200 | $200.00 | GPT-5.6 Sol | 35.714M | 25.000M | 16.667M | Sol benchmark only; Pro mode uses Sol Pro |
| Business Standard | $25.00 | GPT-5.6 Sol | 4.464M | 3.125M | 2.083M | Named-model match |
| Business Premium | $125.00 | GPT-5.6 Sol | 22.321M | 15.625M | 10.417M | Named-model match |
Subscription Versus API Usage Interpretation
For ChatGPT Plus, the 75/25 baseline produces a particularly clear decision point: $20 of GPT-5.6 Sol API usage buys 2.50 million total billable tokens. A Plus user who can sustainably complete more than 2.50 million comparable Sol tokens of work during the monthly billing period receives more raw Sol usage from the subscription than from spending the same $20 through the API. A user who consumes materially less than that amount is, on raw model throughput alone, paying for unused subscription capacity that an API account would not have charged. The comparison should be made with the same workload and reasoning objective rather than by counting visible chat turns. [S01, S03, S04, S47]
The Pro tiers require an additional distinction. OpenAI publishes Pro $100 as 5x Plus usage and Pro $200 as 20x Plus usage, while the price rises 5x and 10x respectively. On the published multiplier alone, the $100 tier scales included usage approximately linearly with price, while the $200 tier offers twice the relative included usage per dollar of Plus. That relative result is useful even without an absolute token pool. However, ChatGPT's Pro option uses GPT-5.6 Sol Pro rather than ordinary Sol. Because this audit does not assign an unverified public token-priced API rate to Sol Pro, the $100 and $200 Sol rows should be read as benchmarks for Sol usage available within those plans, not as exact same-model comparisons for Pro mode. [S01, S02, S04]
OpenAI Business provides a different kind of evidence because it publishes local-message ranges for Sol, Terra, and Luna over five-hour windows. Those ranges permit scenario analysis if the analyst defines a representative token cost per message, but they still should not be multiplied into a theoretical monthly maximum without accounting for task size, cloud execution, reasoning, and higher-level policy limits. For organizations with telemetry, Business can therefore support a more empirical subscription-versus-API comparison than Plus or Pro, but the result remains workload-specific rather than a universal token entitlement. [S07, S08]
Anthropic
Anthropic's subscription economics are dominated by rolling session limits and weekly limits rather than a monthly token bank. Pro establishes the baseline, while Max tiers advertise relative session capacity. The September 1, 2026 release of Claude Fable 5.1 materially changes the premium-model portion of this analysis because Fable 5.1 is now the current Fable model, its prompt-cache reads are substantially cheaper than Fable 5, and its subscription inclusion rules differ by plan. Claude Mythos 5.1 shares Fable 5.1's underlying capabilities and API pricing but is restricted to vetted trusted-access programs, so it cannot be treated as a normal consumer-subscription comparator.
The API side is comparatively transparent: Sonnet, Opus, and Fable occupy distinct price bands, while prompt caching, Batch processing, effort, and model-specific behavior can materially alter cost per completed task. Fable 5.1 keeps Fable 5's $10-per-million input and $50-per-million output rates, but reduces cache-read pricing from $1.00 to $0.25 per million tokens. Anthropic estimates this lowers total token-billed cost by roughly 25% for typical workloads and by as much as about 45% for highly agentic, context-heavy workloads. Those savings do not change the report's uncached baseline token threshold, but they can significantly increase the amount of repeated or agentic work that an equal-dollar API budget can complete.
Anthropic's structure makes sustained usage pattern especially important. Session-based and weekly limits can interact in ways that make a short burst of availability look more generous than the amount of work that can be maintained over an entire month. A client-facing economic evaluation should therefore distinguish peak interactive capacity from sustainable recurring capacity.
Plan architecture
Claude Pro is $20 monthly (or $200/year), Max 5x is $100 monthly, and Max 20x is $200 monthly. Team Standard/Premium mirror the $25/$125 monthly seat pattern with $20/$100 annual equivalents. Anthropic's current Fable policy is especially important for a fair subscription-versus-API comparison: Fable 5 and Fable 5.1 are standard included models on Max plans and premium Team or eligible premium Enterprise seats, but combined Fable usage may consume no more than 50% of the plan's regular weekly usage limit at no extra cost. On Pro and standard Team or legacy Enterprise seats, Fable 5 and Fable 5.1 run on pay-as-you-go usage credits from the first Fable request and therefore are not part of the fixed subscription entitlement. The earlier Fable 5 promotion ended July 19 and never applied to Fable 5.1.
Anthropic's plan ladder should be evaluated against both model mix and cadence of use. A Max subscriber who depends on Fable 5.1 can receive included access to a model whose direct API rate is expensive, but the value is constrained by the shared weekly pool and the 50% Fable ceiling. A Pro subscriber cannot count Fable 5.1 usage-credit consumption as subscription value because those credits are incremental spend beyond the $20 plan price. This distinction is essential to the paper's equal-dollar rule: a $20 Pro-versus-$20 API comparison should use models included in Pro limits, such as Sonnet or eligible Opus usage, whereas the appropriate fixed-price Fable 5.1 comparisons are the $100 and $200 Max tiers or qualifying premium seats.
Subscription metering
Claude usage is deliberately dynamic. Pro usage resets on a five-hour session cadence and also has a weekly all-model limit; consumption varies with conversation length, attachments, selected model, features, and effort. Max multiplies the Pro session capacity by 5x or 20x, but weekly limits still apply. Fable 5.1 adds another model-specific constraint: on Max and qualifying premium seats, Fable 5 and Fable 5.1 together can account for up to 50% of the plan's weekly usage limit before additional Fable requests require usage credits. Because Fable models draw down the shared allowance faster than lower-cost Claude models, a message count or simple session multiplier cannot be converted into a fixed monthly Fable token total without empirical usage data.
API economics
Sonnet 5 at $2/$10 yields a $4/M blended 75/25 rate. Opus 5 at $5/$25 yields $10/M. Fable 5.1, like Fable 5, is $10/M input and $50/M output, producing a $20/M uncached 75/25 baseline. At standard uncached rates, $100 therefore buys 5.00 million total Fable 5.1 tokens under the 75/25 workload and $200 buys 10.00 million. The 90/10 equivalents are 7.143 million and 14.286 million; the 50/50 equivalents are 3.333 million and 6.667 million. Claude Mythos 5.1 shares the same published rates and specifications, but because it is available only through trusted-access programs rather than a general consumer subscription, this paper reports its API economics without fabricating a consumer-plan break-even comparison.
The Sonnet-to-Fable price spread makes model routing economically significant, and Fable 5.1 adds a second dimension: cache reuse can substantially change effective task cost without changing headline input/output rates. Routine classification, extraction, and drafting may justify Sonnet, while difficult reasoning or long-horizon work may justify Opus or Fable 5.1 despite the smaller raw-token capacity per dollar. For repeated-context agents, however, Fable 5.1's $0.25/M cache-read rate can make the premium model materially more economical than a simple uncached $10/$50 comparison suggests. Batch-eligible workloads also receive a 50% input/output discount, which doubles the uncached token capacity of an equal-dollar batch budget but is not equivalent to the interactive experience of a consumer subscription and is therefore kept as a separate sensitivity case.
Fable 5.1 and Mythos 5.1 release analysis
Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1, 2026. Anthropic describes them as the same underlying model with different safeguards: Fable 5.1 is generally available, while Mythos 5.1 is reserved for trusted-access programs supporting vetted cybersecurity and life-sciences work. Both use a 1 million-token context window, allow up to 128,000 output tokens, use always-on adaptive thinking, support low, medium, high, xhigh, and max effort, and publish the same $10/M input and $50/M output prices. Fable 5.1 is the normal public model for this research paper's consumer-versus-API analysis; Mythos 5.1 is included as a specialized API-access case rather than treated as a consumer subscription benefit.
Anthropic's launch evaluations indicate that Fable 5.1 improves substantially on Fable 5 across long-horizon coding, scientific research, computer use, knowledge work, and business automation. Reported results include 52.6% on Terminal-Bench-Science 0.1 versus 24.7% for Fable 5, 55.8% on Terminal-Bench 4.0 versus 42.0%, a GDPval-AA v2 score of 1853 versus 1723, 31.4% on AutomationBench versus 17.1%, and 73.4% on CursorBench 3.2.0 versus 70.5%. Mythos 5.1 reached 60.9% on Terminal-Bench 4.0 under Anthropic's evaluation. These are provider-reported benchmarks rather than independent Omnesly measurements, and Anthropic notes that production safeguards can affect some benchmark outcomes, so they are used here to characterize workload capability rather than to convert subscription limits into tokens.
The most economically relevant capability claim is not simply higher benchmark accuracy but improved performance at lower effort. Anthropic states that Fable 5.1 at Low or Medium effort can achieve results similar to or better than Fable 5 at materially lower cost. Fable 5.1 defaults to High effort in Claude Code and the API, while Claude.ai and Cowork default to Medium. Consequently, a fair subscription-versus-API test must explicitly match effort and workload. Comparing a Medium-effort consumer session with a High- or Max-effort API run would confound access method with reasoning depth.
Anthropic launch benchmark comparison
| Benchmark | Fable 5.1 | Fable 5 | Opus 5 |
|---|---|---|---|
| Terminal-Bench-Science 0.1 | 52.6% | 24.7% | 29.0% |
| Terminal-Bench 4.0 | 55.8% (Mythos 60.9%) | 42.0% | 52.3% |
| GDPval-AA v2 | 1853 | 1723 | 1824 |
| OSWorld 2.0 partial / strict | 77.9% / 41.7% | 72.9% / 36.1% | 75.4% / 39.6% |
| Humanity's Last Exam (no tools / tools) | 60.9% / 65.0% | 57.8% / 63.8% | 56.6% / 63.6% |
| AutomationBench | 31.4% | 17.1% | 26.9% |
| CursorBench 3.2.0 | 73.4% | 70.5% | 70.0% |
Cache economics and equal-dollar usage
Fable 5.1 preserves Fable 5's base input and output prices but cuts prompt-cache reads from $1.00 to $0.25 per million tokens, a 75% reduction in the cache-read unit price. Anthropic reports that, when measured over four weeks of real August 2026 default-effort usage, this change reduced the cost of running the same typical workloads by about 25% and reduced highly agentic, context-heavy workload costs by as much as about 45%. If those savings transfer to a particular workload, a fixed API budget can execute roughly 1.33 times as many of the same typical workloads at a 25% cost reduction and up to about 1.82 times as many at a 45% reduction. These are workload-throughput multipliers, not raw-token multipliers: the uncached $10/$50 token rates are unchanged, and actual savings depend on cache hit rate and workload structure.
This distinction directly affects the paper's break-even interpretation. At the standard uncached 75/25 baseline, a $100 Max 5x subscription should be compared with 5.00 million Fable 5.1 API tokens, while a $200 Max 20x subscription should be compared with 10.00 million. If the user's API implementation reuses large prompt prefixes and realizes Anthropic's reported cache savings, the relevant question becomes whether the subscription completes more standardized Fable work than the cache-optimized API can complete for the same $100 or $200.
Fable 5.1 API cost and cache-efficiency scenarios
| Fable 5.1 API cost case | Anthropic-reported cost effect | Same-dollar completed-work multiplier | Interpretation |
|---|---|---|---|
| Fresh-token baseline | 0% change in $10/$50 base rates | 1.00x raw-token baseline | Equal-dollar token thresholds remain 5.00M at $100 and 10.00M at $200 for 75/25 |
| Cache-read unit price | $1.00/M → $0.25/M | 4.00x cache-read tokens per dollar | Applies only to cache reads, not output or fresh input |
| Typical real workload | ~25% lower total cost vs Fable 5 | ~1.33x same workload executions | Anthropic estimate from four weeks of August 2026 default-effort usage |
| Highly agentic workload | Up to ~45% lower total cost | Up to ~1.82x same workload executions | Most relevant when repeated context/cache reads dominate cost |
| Batch-eligible fresh tokens | 50% input/output discount | 2.00x raw-token capacity | Separate async service-tier sensitivity; not equivalent to interactive subscription use |
Subscription inclusion and Mythos access
Fable 5.1 is visible across paid Claude plans, but visibility is not the same as included subscription usage. Max plans and qualifying premium Team or legacy Enterprise seats include Fable 5 and Fable 5.1 within the regular subscription pool, with Fable models collectively limited to up to 50% of the weekly allowance at no extra cost. Pro and standard seats can select Fable 5.1 only through usage credits billed at pay-as-you-go rates from the start. For this reason, the research does not claim that a $20 Claude Pro subscription includes a quantifiable Fable 5.1 allowance. Counting those separately purchased credits as part of the $20 subscription would violate the equal-dollar methodology.
Claude Mythos 5.1 has the same model specifications and published token prices as Fable 5.1, but Anthropic limits access to vetted participants through the Cyber Verification Program, Life Sciences Verification Program, Project Glasswing, and related trusted-access arrangements. That access model has no generally published consumer subscription price that can be paired with an equal-dollar API budget. Mythos 5.1 is therefore analyzed for technical and API-cost completeness but is marked not directly comparable in the subscription tables.
Developer integration, safety, and operational implications
For API developers, Fable 5.1 is not a purely cosmetic model-ID update. It does not support forced tool choice, earlier models cannot read Fable 5.1 thinking blocks, and editing earlier conversation turns can invalidate preserved thinking blocks in contexts where binding is enforced. The release also adds per-message effort, turn-scoped system messages, readable progress updates between tool calls, content provenance, and lower cache-read pricing. These integration differences can change retries, cache reuse, latency, and tokens per completed task, so production cost comparisons should be performed on completed workflows rather than on price-per-million alone.
Anthropic also reports materially improved safeguards and alignment. Fable 5.1's cybersecurity safeguards generated about 60% fewer interventions per session than the previous Fable safeguards while permitting software-vulnerability discovery but continuing to restrict exploit development. Mythos 5.1 demonstrated Anthropic's strongest released cyber capabilities and greater chemical/biological capability than Mythos 5, while remaining below Anthropic's next Responsible Scaling Policy risk tier. These findings do not enter the equal-dollar token equations directly, but they explain why Fable 5.1 and Mythos 5.1 occupy a premium/specialized tier and why access controls differ between the two models.
Fable 5.1 is available through the Claude API and Anthropic's supported partner channels, including Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS. Anthropic's migration documentation identifies Fable 5.1 and Mythos 5.1 as Covered Models with 30-day retention by default and states that zero-data-retention treatment requires authorization; Anthropic separately announced Enterprise Frontier Safeguards, intended to provide safeguard monitoring while keeping protected customer data inside customer-controlled cloud infrastructure. These privacy mechanisms do not change the base token-price equations, but they can affect which deployment path is commercially viable for regulated or sensitive workloads.
Tokenizer effect
Claude 4.7-and-later models, including Fable 5.1 and Mythos 5.1, use a newer tokenizer that produces roughly 30% more tokens for the same text than models older than Claude 4.7. Anthropic states that Fable 5.1 uses the same tokenizer as Fable 5, so direct 5-to-5.1 token comparisons are not distorted by a tokenizer change. Cross-generation "tokens per dollar" comparisons can still mislead, however: the newer model can have a lower effective cost per completed task while representing the same payload with more token units. The workload-normalized comparisons in this report therefore remain more informative than raw cross-generation token counts alone.
Anthropic subscription plan matrix
| Plan | Price | Cadence | Models | Published usage | Quantifiability |
|---|---|---|---|---|---|
| Claude Pro | $20.00 | Monthly | Sonnet/Opus and paid model selector; Fable 5/5.1 via PAYG credits | At least 5x free per session; resets every 5h; weekly all-model limit; Fable 5/5.1 excluded from included Pro limits | Dynamic compute/message pool |
| Claude Max 5x | $100.00 | Monthly | Includes Fable 5 and Fable 5.1 plus paid model set | 5x Pro capacity per session; weekly limits also apply; Fable models collectively may use up to 50% of the regular weekly limit at no extra cost | Published relative multiplier |
| Claude Max 20x | $200.00 | Monthly | Includes Fable 5 and Fable 5.1 plus paid model set | 20x Pro capacity per session; weekly limits also apply; 50% shared Fable ceiling | Published relative multiplier |
| Claude Team Standard | $25.00 | Monthly seat | Claude paid model set | More usage than Pro; annual equivalent $20/mo | Absolute token quota not published |
| Claude Team Premium | $125.00 | Monthly seat | Claude paid model set incl. Fable 5 and Fable 5.1 standard | 5x Standard usage; annual equivalent $100/mo; subject to the shared 50% weekly Fable ceiling | Relative multiplier only |
Anthropic current text-model / API matrix
| Model | Input | Cached | Output | Context | Effort / thinking | Subscription parity |
|---|---|---|---|---|---|---|
| Claude Fable 5.1 | $10/M | $0.25/M | $50/M | 1M | low, medium, high, xhigh, max | Named-model match on Max/premium seats; PAYG credits on Pro; included usage capped within weekly pool |
| Claude Mythos 5.1 | $10/M | $0.25/M | $50/M | 1M | low, medium, high, xhigh, max | No general consumer match; trusted-access / Project Glasswing only |
| Claude Fable 5 | $10/M | $1/M | $50/M | 1M | low, medium, high, xhigh, max | Named-model match on Max; PAYG credits on Pro |
| Claude Opus 5 | $5/M | $0.5/M | $25/M | 1M | low, medium, high, xhigh, max | Named-model match on paid Claude plans where selectable |
| Claude Sonnet 5 | $2/M | $0.2/M | $10/M | 1M | low, medium, high, xhigh, max | Named-model match on paid Claude plans |
| Claude Haiku 4.5 | $1/M | $0.1/M | $5/M | 200K | No current effort parameter | API; selectable availability varies |
Equal-dollar API break-even thresholds
Anthropic equal-dollar API capacity by plan
| Plan | Budget | API model | 90/10 | 75/25 | 50/50 | Mapping |
|---|---|---|---|---|---|---|
| Claude Pro | $20.00 | Claude Sonnet 5 | 7.143M | 5.000M | 3.333M | Named-model match |
| Claude Pro | $20.00 | Claude Opus 5 | 2.857M | 2.000M | 1.333M | Named-model match |
| Claude Max 5x | $100.00 | Claude Opus 5 | 14.286M | 10.000M | 6.667M | Named-model match |
| Claude Max 5x | $100.00 | Claude Fable 5.1 | 7.143M | 5.000M | 3.333M | Named-model match on Max; capped by shared weekly Fable allowance |
| Claude Max 20x | $200.00 | Claude Opus 5 | 28.571M | 20.000M | 13.333M | Named-model match |
| Claude Max 20x | $200.00 | Claude Fable 5.1 | 14.286M | 10.000M | 6.667M | Named-model match on Max; capped by shared weekly Fable allowance |
| Trusted access | N/A | Claude Mythos 5.1 | N/A | N/A | N/A | No general consumer subscription price; API economics shown separately |
Subscription Versus API Usage Interpretation
Claude Pro at $20 should be compared with $20 of the specific Claude model actually included in the user's subscription usage. At the 75/25 baseline, $20 buys 5.00 million Sonnet 5 API tokens or 2.00 million Opus 5 API tokens at standard uncached rates. Fable 5.1 should not be added as a third $20 Pro threshold because Anthropic does not include Fable 5.1 in Pro's fixed usage limits; Pro users access it through separately purchased usage credits. For Fable 5.1, the appropriate fixed-subscription comparisons are Max 5x at $100 versus $100 of Fable 5.1 API usage and Max 20x at $200 versus $200 of API usage. Those API baselines are 5.00 million and 10.00 million total tokens at 75/25 before cache or Batch discounts.
Anthropic's relative Max scaling can be analyzed even though the absolute token pool is not public. Max 5x costs five times Pro and advertises five times the session usage, so its published included-usage efficiency is approximately linear with the baseline. Max 20x costs ten times Pro while advertising twenty times the session usage, implying twice the relative session entitlement per dollar before weekly limits. Fable 5.1 complicates that otherwise simple scaling because Fable models collectively may consume no more than 50% of the regular weekly limit as included usage. A Max user's observed Fable throughput must therefore be measured against the 5.00M or 10.00M uncached API baseline - and, for cache-heavy workloads, against the larger task-equivalent throughput made possible by the $0.25/M cache-read rate.
The five-hour session reset and weekly all-model cap prevent a simple monthly multiplication. A user who performs long document analysis, uses higher effort, selects Fable 5.1, or runs long tool chains can consume the shared subscription budget faster than a user sending short Sonnet requests. The subscription-versus-API difference should therefore be measured from sustained completed work over the billing period, not inferred from the number of five-hour reset opportunities.
Google
Google requires extra caution because the Gemini consumer application exposes family-level choices while the developer platform exposes versioned SKUs with their own pricing, context rules, and thinking behavior. The report therefore treats app-to-API comparisons as family crosswalks unless Google explicitly pins the same model identifier on both surfaces. This is a methodological limitation, not a claim that the consumer application is necessarily using a different backend.
Compute-based subscription metering also prevents a reliable conversion from a published plan multiplier to monthly tokens. Longer chats, more complex requests, premium models, and advanced features can consume the allowance at different rates. As a result, the equal-dollar API threshold is most useful as an empirical target: a user can log actual subscription use over several reset cycles and then determine whether the measured workload exceeds what the plan price would have purchased through the API.
Plan architecture
Google AI Plus is $7.99/month in the U.S.; AI Pro is $19.99/month; AI Ultra has $100 and $200 monthly tiers. Gemini Apps state plan multipliers relative to the standard limit rather than a fixed prompt or token pool: Plus 2x, Pro 4x, Ultra 5x or 20x higher than Pro depending on Ultra tier. [S18, S19, S20, S21, S22, S23]
Subscription metering
Gemini Apps are compute-metered. The provider states that complexity, model/features and chat length affect consumption; limits refresh every five hours until a weekly limit is reached. This makes a token-total conversion impossible without telemetry from an actual account and workload. [S18, S19, S20, S21, S22, S23]
The consumer app exposes Gemini Flash-Lite, Flash and Pro family names, while the developer API exposes versioned SKUs such as 3.7 Flash, 3.5 Flash-Lite and 3.1 Pro Preview. Because Google states model names/versions/availability can change, the app-to-API comparisons in this report are family crosswalks rather than assertions that a consumer prompt was served by one pinned SKU. [S18, S19, S20, S21, S22, S23]
API economics
Current 3.7 Flash promotional pricing is $0.75/$3.75 through December 31, 2026. 3.5 Flash-Lite is $0.30/$2.50. 3.1 Pro Preview is $2/$12 up to 200K prompt tokens and $4/$18 above 200K. Output prices include thinking tokens, so higher thinking effort directly consumes the output-side budget. [S18, S19, S20, S21, S22, S23]
Google subscription plan matrix
| Plan | Price | Cadence | Models | Published usage | Quantifiability |
|---|---|---|---|---|---|
| Google AI Plus | $7.99 | Monthly | Gemini Flash-Lite / Flash / Pro families | 2x standard Gemini app usage; 128K app context | Compute-based 5h pool until weekly limit; no token pool |
| Google AI Pro | $19.99 | Monthly | Gemini Flash-Lite / Flash / Pro families | 4x standard Gemini app usage; 1M app context | Compute-based 5h pool until weekly limit; no token pool |
| Google AI Ultra 100 | $100.00 | Monthly | Gemini families incl. Deep Think | 5x AI Pro usage in Gemini app/Antigravity (=20x standard if multipliers compose) | Compute-based; no token pool |
| Google AI Ultra 200 | $200.00 | Monthly | Gemini families incl. highest access/Deep Think | 20x AI Pro usage (=80x standard if multipliers compose) | Compute-based; no token pool |
Google current text-model / API matrix
| Model | Input | Cached | Output | Context | Effort / thinking | Subscription parity |
|---|---|---|---|---|---|---|
| Gemini 3.7 Flash | $0.75/M | — | $3.75/M | 1M | low, medium, high | Family crosswalk to Gemini Flash |
| Gemini 3.5 Flash-Lite | $0.3/M | $0.03/M | $2.5/M | 1M | minimal, low, medium, high | Family crosswalk to Gemini Flash-Lite |
| Gemini 3.1 Pro Preview | $2/M | $0.2/M | $12/M | 1M | low, medium, high | Family crosswalk to Gemini Pro |
| Gemini 2.5 Pro | $1.25/M | — | $10/M | 1M | low, medium, high (thinkingBudget) | Legacy app/API parity |
| Gemini 2.5 Flash | $0.3/M | — | $2.5/M | 1M | low, medium, high (thinkingBudget) | Legacy app/API parity |
Equal-dollar API break-even thresholds
Google equal-dollar API capacity by plan
| Plan | Budget | API model | 90/10 | 75/25 | 50/50 | Mapping |
|---|---|---|---|---|---|---|
| Google AI Plus | $7.99 | Gemini 3.1 Pro Preview | 2.663M | 1.776M | 1.141M | Family crosswalk |
| Google AI Plus | $7.99 | Gemini 3.7 Flash | 7.610M | 5.327M | 3.551M | Family crosswalk |
| Google AI Pro | $19.99 | Gemini 3.1 Pro Preview | 6.663M | 4.442M | 2.856M | Family crosswalk |
| Google AI Pro | $19.99 | Gemini 3.7 Flash | 19.038M | 13.327M | 8.884M | Family crosswalk |
| Google AI Ultra 100 | $100.00 | Gemini 3.1 Pro Preview | 33.333M | 22.222M | 14.286M | Family crosswalk |
| Google AI Ultra 200 | $200.00 | Gemini 3.1 Pro Preview | 66.667M | 44.444M | 28.571M | Family crosswalk |
Subscription Versus API Usage Interpretation
Google AI Pro demonstrates why one subscription cannot always be assigned one token-equivalent number. At the $19.99 plan price, the 75/25 benchmark is about 4.44 million tokens for Gemini 3.1 Pro within the standard <=200K prompt tier, but roughly 13.33 million tokens for Gemini 3.7 Flash at its current rate. Both can be relevant to a Gemini subscriber, yet they differ by roughly threefold in raw API purchasing power. A correct comparison must therefore identify which model family performs the subscription task before deciding whether the subscription delivered more or less usage than the API alternative. [S18, S21, S23]
Google publishes relative Gemini-app usage multipliers rather than token pools: AI Plus is 2x standard, AI Pro is 4x standard, and Ultra tiers provide 5x or 20x AI Pro usage depending on the subscription. Relative to the $19.99 Pro plan, a $100 Ultra tier is priced at approximately five times as much and is described as five times Pro usage, so the published usage-per-dollar relationship is roughly linear. A $200 Ultra tier is priced at roughly ten times Pro and is described as twenty times Pro usage, implying about twice the relative included usage per dollar. [S18, S20]
Because Gemini limits are compute-based and refresh every five hours only until a weekly cap is reached, the subscription side should be measured over sustained use rather than extrapolated from a single reset window. Prompt complexity, chat length, selected model, and feature choice can all change how quickly the allowance is consumed. The API side remains a precise dollar denominator, while the subscription side is best represented as observed model-specific throughput or as a conditional break-even threshold when telemetry is unavailable. [S18]
xAI
xAI's paid consumer plans use a shared weekly allowance across multiple Grok experiences. This means a user's effective chat capacity can be reduced by other compute-intensive activities that draw from the same pool. Because the provider does not publish the weekly pool as a fixed number of text tokens, the report avoids converting the shared allowance into a synthetic monthly token total and instead prices the direct Grok API at the subscription's exact monthly cost.
The long-context rule is particularly important for Grok economics. At ordinary context lengths, the model's blended cost can be competitive, but requests crossing the long-context threshold can materially reduce equal-dollar token capacity. A user comparing SuperGrok with API spend should therefore consider not only the number of prompts but also whether the workload routinely carries large documents, conversation histories, or retrieved context.
Plan architecture
SuperGrok is $30/month and SuperGrok Plus is $100/month on the current public plan page. Grok's consumer product uses a shared weekly usage allowance across paid products; different activities consume different amounts of compute. [S24, S25, S26, S43]
API economics
Grok 4.6 is $2/M input, $0.50/M cached input and $6/M output below the 200K long-context threshold. At a 75/25 workload the blended rate is $3/M, so $30 buys 10M total tokens and $100 buys 33.33M. At >=200K context the input/output rates double to $4/$12, halving equal-dollar token capacity for the same token mix. [S24, S25, S26, S43]
Effort modes
Grok 4.6 supports low, medium, high (default) and xhigh reasoning effort. Because xAI does not publish a deterministic token allocation for each setting, effort-matched economics must be measured from actual output/reasoning usage rather than inferred from the label. [S24, S25, S26, S43]
xAI subscription plan matrix
| Plan | Price | Cadence | Models | Published usage | Quantifiability |
|---|---|---|---|---|---|
| SuperGrok | $30.00 | Monthly | Grok 4.6 | Higher limits; all paid products draw from one shared weekly allowance | Compute-weighted weekly pool; no token quota |
| SuperGrok Plus | $100.00 | Monthly | Grok 4.6 | Significantly higher usage than SuperGrok across Chat/Imagine/Voice/Build | Compute-weighted weekly pool; no numeric multiplier |
xAI current text-model / API matrix
| Model | Input | Cached | Output | Context | Effort / thinking | Subscription parity |
|---|---|---|---|---|---|---|
| Grok 4.6 | $2/M | $0.5/M | $6/M | 500K | low, medium, high, xhigh | Named-model match: SuperGrok paid plans |
| Grok 4.5 | $2/M | $0.3/M | $6/M | 500K | low, medium, high | Named-model reference via Perplexity; prior Grok app generation |
| Grok 4.3 | $1.25/M | $0.2/M | $2.5/M | 1M | Model-specific reasoning behavior | API catalog |
| grok-build-0.1 | $1/M | $0.2/M | $2/M | 256K | agent build mode | Build product/API |
Equal-dollar API break-even thresholds
xAI equal-dollar API capacity by plan
| Plan | Budget | API model | 90/10 | 75/25 | 50/50 | Mapping |
|---|---|---|---|---|---|---|
| SuperGrok | $30.00 | Grok 4.6 | 12.500M | 10.000M | 7.500M | Named-model match |
| SuperGrok Plus | $100.00 | Grok 4.6 | 41.667M | 33.333M | 25.000M | Named-model match |
Subscription Versus API Usage Interpretation
For SuperGrok at $30, the 75/25 Grok 4.6 API benchmark is 10.00 million tokens under the standard short-context rate. A text-focused subscriber must therefore sustain more than 10.00 million comparable Grok 4.6 tokens of monthly work for the subscription to exceed $30 of direct API throughput on the same workload. Because xAI does not publish the weekly subscription pool as tokens, that statement is a decision threshold rather than an asserted quota. [S24, S25, S43]
The shared weekly pool materially changes how the threshold should be measured. Chat, Imagine, Voice, Build, and other Grok products draw from the same allowance, so a subscriber who uses image or voice features may reach the weekly limit after substantially less text-model work than a text-only user. For a clean text-model comparison, non-text activity should either be excluded from the measurement period or assigned its own economic value rather than being silently treated as Grok 4.6 text tokens. [S24]
SuperGrok Plus is described as providing significantly higher usage than SuperGrok, but xAI does not publish a stable numeric multiplier that would support the same relative-efficiency calculation available for OpenAI, Anthropic, or Google. The $100 API denominator can still be calculated exactly; the subscription advantage remains conditional on the user's observed weekly-pool consumption until xAI publishes a numeric allowance or the account exposes sufficiently detailed telemetry. [S24, S43]
Mistral
Mistral is structurally different from the other major consumer subscriptions in this report because its plan page displays API credits as part of the subscription proposition. Those credits should not be counted as subscription chat tokens; they are a separate, directly monetizable API benefit. Nevertheless, they alter the economic decision because the face value of the credits can equal or exceed the cash price of the plan before any value is assigned to Vibe chat or coding usage.
The underlying Mistral API catalog also spans a very broad price range. Small-class models can deliver extremely high raw token throughput per dollar, while Medium is materially more expensive and potentially more capable for difficult tasks. Batch pricing and cache discounts widen the gap further.
Figure 5
Figure 5. Mistral Pro bundled API-credit face value compared with spending the $14.99 plan price directly on API tokens.
Plan architecture
Mistral's Vibe Free/Pro/Team plans are fair-use products. The current pricing page lists normal Pro at $14.99/month and states verified-student pricing of $5.99; it also displays $30/month in API credits on Pro and $10/month in API credits on Free. Team is $24.99/user/month and the page displays a $50/month total at the two-user level. [S27, S28, S29]
Fair-use language means the chat component cannot be treated as a contractual token bank, while the explicit API credit can be valued directly. This makes Mistral a useful example of why subscription benefits should be decomposed into individually measurable components.
Bundled API credits and plan economics
The bundled API credits are economically important because they can exceed the cash price of the subscription. At the displayed $30 credit value, a Pro user could purchase about 10M baseline Medium 3.5 tokens, 40M Large 3 tokens, or 114.29M Small 4 tokens, before considering any included Vibe usage. This is a face-value calculation and should be rechecked at purchase time because promotions and eligibility can change. [S27, S28, S29]
API economics
Mistral Medium 3.5 is $1.50/$7.50, Large 3 is $0.50/$1.50, and Small 4 is $0.15/$0.60. Batch is generally half price and cached input can reduce repeated input cost by up to 90%, making Mistral especially sensitive to implementation choices. [S27, S28, S29]
Effort modes
The public reasoning guide documents adjustable reasoning on Mistral Small and Medium 3.5 with reasoning_effort none or high. High exposes a full thinking chunk and uses more tokens; none minimizes thinking and omits the thinking chunk. [S27, S28, S29]
Mistral subscription plan matrix
| Plan | Price | Cadence | Models | Published usage | Quantifiability |
|---|---|---|---|---|---|
| Vibe Pro | $14.99 | Monthly | Latest Mistral models / Vibe routing | Fair-use limits; current pricing page displays $30/mo API credits; more messages/search/coding | Chat quota not tokenized; API-credit component is exact face value |
| Vibe Team | $24.99 | Monthly seat | Latest Mistral models / Vibe routing | Fair-use; current UI also shows a $50/mo total at 2-user minimum | Chat quota not tokenized |
Mistral current text-model / API matrix
| Model | Input | Cached | Output | Context | Effort / thinking | Subscription parity |
|---|---|---|---|---|---|---|
| Mistral Medium 3.5 | $1.5/M | $0.15/M | $7.5/M | 256K | none, high | Current Vibe / API |
| Mistral Small 4 | $0.15/M | $0.015/M | $0.6/M | 256K | none, high | Current Vibe / API |
| Mistral Large 3 | $0.5/M | $0.05/M | $1.5/M | 256K | none / standard | Current Vibe / API |
Equal-dollar API break-even thresholds
Mistral equal-dollar API capacity by plan
| Plan | Budget | API model | 90/10 | 75/25 | 50/50 | Mapping |
|---|---|---|---|---|---|---|
| Vibe Pro | $14.99 | Mistral Medium 3.5 | 7.138M | 4.997M | 3.331M | Routing / likely |
| Vibe Pro | $14.99 | Mistral Small 4 | 76.872M | 57.105M | 39.973M | Routing / likely |
| Vibe Pro | $14.99 | Mistral Large 3 | 24.983M | 19.987M | 14.990M | Routing / likely |
Subscription Versus API Usage Interpretation
Mistral's Vibe Pro plan needs to be decomposed into two separate usage channels. The fair-use Vibe allowance is a subscription benefit whose text-token capacity is not publicly fixed, while the displayed API-credit benefit is already a metered API balance. At the $14.99 plan price, the 75/25 Medium 3.5 benchmark is about 4.997 million tokens. If the displayed $30 monthly API credit is actually included for the user, that credit balance funds about 10.0 million Medium 3.5 baseline tokens at current rates, before assigning any value at all to Vibe chat or coding usage. [S27, S28]
The $30 credit should not be relabeled as subscription chat tokens because doing so would double-count different products. A correct value analysis treats the credits as a directly quantifiable API component and the Vibe allowance as a separate fair-use component. This makes Mistral unusual among the providers in the report: part of the subscription-versus-API difference is directly observable through the included API balance, while the consumer-interface portion still requires measured usage or a break-even interpretation.
Vibe also routes among current Mistral models rather than guaranteeing that every task uses one pinned API SKU. A user who wants to measure the raw model-usage advantage should record the model selected or inferred for representative work where the product exposes it, then price the same workload against that model's API rate. If the backend routing is not observable, the Medium, Small, and Large benchmarks should be presented as model-specific scenarios rather than collapsed into one false-precision token number. [S27, S28, S29]
Perplexity
Perplexity is best understood as a multi-model subscription layer rather than a single-model subscription. One monthly fee can expose models from several underlying providers, each with a different direct API rate. The subscription therefore has no single token break-even threshold. The correct threshold changes with the model selected for the task, and Max-only access to premium models can substantially change the economic value of the higher tier for users who regularly require those models.
Native Sonar products introduce another complication because search and research functionality can include request, citation, search, or reasoning charges in addition to token rates. A direct comparison with a simple chat-token rate can therefore understate API cost for research-heavy workflows. The report separates direct-provider API-equivalent benchmarking for named third-party models from native Sonar economics so that readers can distinguish figures representing model-token consumption from those that include a broader retrieval product.
Plan architecture
Perplexity Pro is $20/month, Max $200/month, Enterprise Pro $40/month per seat and Enterprise Max $325/month per seat. The Search model selector currently includes Sonar 2 and multiple third-party models. Sol and Opus 5 are Max-only in Search; Terra, Gemini 3.1 Pro, Sonnet 5, Grok 4.5, Nemotron 3 Ultra and others are available on Pro. [S30, S31, S32, S33, S34, S44, S45, S46]
Direct API-equivalent benchmarking
For third-party model comparisons, the clearest API-equivalent benchmark is the model provider's direct API price because the model identity is explicit. Perplexity's Agent API also states that supported third-party models are passed through at direct first-party token rates with no markup, although its published model list can lag the consumer selector. [S30, S31, S32, S33, S34, S44, S45, S46]
Native Sonar economics
Native Sonar APIs are not pure token-priced chat. Sonar and Sonar Pro add request fees that depend on search context, while Sonar Deep Research adds reasoning-token, citation-token and search-query charges. A "tokens for $20" number that ignores those components would materially overstate usable research capacity, so native Sonar is modeled with its full cost components rather than forced into the simple blended token table. [S30, S31, S32, S33, S34, S44, S45, S46]
Thinking toggle
Perplexity's consumer UI exposes Thinking as optional, always-on or unavailable depending on model. That setting is a platform-level choice and does not necessarily expose every first-party effort level. For example, Claude or OpenAI may have several API effort values while Perplexity presents a binary Thinking toggle. [S30, S31, S32, S33, S34, S44, S45, S46]
Perplexity subscription plan matrix
| Plan | Price | Cadence | Models | Published usage | Quantifiability |
|---|---|---|---|---|---|
| Perplexity Pro | $20.00 | Monthly | Sonar 2; Terra; Gemini 3.1 Pro; Sonnet 5; Kimi K3; GLM 5.2; Grok 4.5; Nemotron 3 Ultra | Advanced-model use is plan-limited; model list subject to change | Third-party model token quota not published |
| Perplexity Max | $200.00 | Monthly | Pro set + GPT-5.6 Sol + Claude Opus 5; Model Council | Higher advanced-model limits; Model Council included within cap | Third-party model token quota not published |
| Enterprise Pro | $40.00 | Monthly seat | Enterprise model set excluding Kimi/GLM | Advanced Search/Computer limits; org controls may restrict models | No token pool |
| Enterprise Max | $325.00 | Monthly seat | Enterprise model set + Sol/Opus + Model Council | Highest enterprise allowances | No token pool |
Perplexity API pricing and parity matrix
| Model | Input | Output | Effort / thinking | Subscription parity |
|---|---|---|---|---|
| Sonar | $1/M | $1/M | No reasoning effort | Sonar-family consumer crosswalk uncertain |
| Sonar Pro | $3/M | $15/M | No model effort; search type affects request fee | Consumer Pro Search crosswalk |
| Sonar Reasoning Pro | $2/M | $8/M | reasoning model | Consumer Thinking crosswalk |
| Sonar Deep Research | $2/M | $8/M | low, medium, high | Deep Research feature/API |
Equal-dollar model benchmark matrix
Perplexity equal-dollar model benchmark matrix
| Plan | Budget | API model | 90/10 | 75/25 | 50/50 | Mapping |
|---|---|---|---|---|---|---|
| Perplexity Pro | $20.00 | GPT-5.6 Terra | 6.667M | 4.444M | 2.857M | Named-model reference; direct-provider API benchmark |
| Perplexity Pro | $20.00 | Claude Sonnet 5 | 7.143M | 5.000M | 3.333M | Named-model reference; direct-provider API benchmark |
| Perplexity Pro | $20.00 | Gemini 3.1 Pro Preview | 6.667M | 4.444M | 2.857M | Named-family reference; direct-provider API benchmark |
| Perplexity Pro | $20.00 | Grok 4.5 | 8.333M | 6.667M | 5.000M | Named-model reference; direct-provider API benchmark |
| Perplexity Max | $200.00 | GPT-5.6 Sol | 35.714M | 25.000M | 16.667M | Named-model reference; direct-provider API benchmark |
| Perplexity Max | $200.00 | Claude Opus 5 | 28.571M | 20.000M | 13.333M | Named-model reference; direct-provider API benchmark |
Subscription Versus API Usage Interpretation
Perplexity Pro is best analyzed as a portfolio of model access rather than as one AI model subscription. At the $20 plan price and the 75/25 baseline, the direct-provider API benchmarks in this report are approximately 4.44 million tokens for GPT-5.6 Terra, 5.00 million for Claude Sonnet 5, 4.44 million for Gemini 3.1 Pro, and 6.67 million for Grok 4.5. The economic value of a user's month therefore depends on how usage is distributed across those models. A month dominated by a more expensive model can represent more direct-API value than the same number of interactions dominated by a cheaper model. [S30, S32]
A weighted model-mix benchmark provides a more faithful comparison. If sm is the share of comparable usage assigned to model m and Cm is that model's blended API cost per million tokens, the portfolio API rate is the sum of sm × Cm across the models used. The same-dollar API capacity is then the subscription price divided by that weighted rate. This method lets a user compare a real Perplexity month with the amount that the same $20 or $200 would have purchased from the underlying providers without pretending that all selected models have the same marginal cost.
Native Perplexity Search, Sonar, Research, Computer, and Model Council workflows can add search, orchestration, citation, reasoning, or credit charges that do not reduce cleanly to ordinary text tokens. Those capabilities should be valued separately or modeled with their native price structure. The multi-model text comparison remains useful, but it should not be stretched into a claim that every Perplexity feature is equivalent to a fixed number of third-party API tokens. [S30, S33, S34, S44]
DeepSeek
DeepSeek demonstrates the importance of preserving the equal-dollar rule even when it prevents a numeric subscription comparison. The first-party consumer app is free, so there is no positive monthly subscription price that can be mirrored as API spend. Assigning an arbitrary $20 or $30 budget would answer a different question. The report therefore includes DeepSeek's API economics as a market reference while labeling the paid-subscription comparison as undefined.
For API users, time-of-day pricing and cache-hit pricing can matter more than the consumer/API distinction. A workload that can be scheduled into off-peak windows or engineered for high cache reuse can achieve substantially greater throughput than one that runs at peak cache-miss rates. This makes DeepSeek a useful example of why operational scheduling and prompt architecture belong in total-cost analysis rather than being treated as minor implementation details.
Plan architecture
DeepSeek's official site advertises free consumer access. V4 Pro is available in app/web Expert Mode and via the API, while V4 Flash is the efficient model. Because the current official materials do not identify a first-party paid consumer subscription price for the matched product, the equal-dollar subscription test used elsewhere in this report is not applicable. [S35, S36, S37, S38]
API economics
DeepSeek uses time-of-day peak/off-peak pricing. V4 Flash costs $0.22/$0.66 off-peak and $0.44/$1.32 peak for cache-miss input/output; V4 Pro costs $0.66/$1.98 off-peak and $1.32/$3.96 peak. Cache hits are dramatically cheaper. Peak periods are 01:00-04:00 and 06:00-10:00 UTC; other hours are off-peak. [S35, S36, S37, S38]
The peak/off-peak schedule makes demand timing a controllable variable. Batchable jobs, evaluations, indexing, or background analysis can potentially be shifted toward lower-priced windows, while interactive traffic may need to accept peak pricing. A realistic monthly model should therefore estimate the percentage of tokens that can be scheduled rather than applying the most favorable rate to all traffic.
Effort modes
V4 Flash and V4 Pro support non-thinking plus low, high and max thinking effort. For compatibility, medium/xhigh map to high. This is an example where forcing effort-name equality across providers would be incorrect. [S35, S36, S37, S38]
DeepSeek consumer access matrix
| Plan | Price | Cadence | Models | Published usage | Quantifiability |
|---|---|---|---|---|---|
| DeepSeek Web/App | $0 / free | Free | V4 Flash Instant; V4 Pro Expert | Official site states free access. No first-party paid consumer subscription price identified | No equal-dollar paid-plan comparison exists |
DeepSeek API pricing and parity matrix
| Model | Input | Cached | Output | Effort / thinking | Subscription parity |
|---|---|---|---|---|---|
| DeepSeek V4 Flash - off-peak | $0.22/M | $0.007/M | $0.66/M | none, low, high, max | Named-model app/API match; no paid subscription |
| DeepSeek V4 Flash - peak | $0.44/M | $0.014/M | $1.32/M | none, low, high, max | Named-model app/API match; no paid subscription |
| DeepSeek V4 Pro - off-peak | $0.66/M | $0.022/M | $1.98/M | none, low, high, max | Named-model Expert Mode/API match; no paid subscription |
| DeepSeek V4 Pro - peak | $1.32/M | $0.044/M | $3.96/M | none, low, high, max | Named-model Expert Mode/API match; no paid subscription |
Why DeepSeek Is Not a Paid Subscription Comparison
DeepSeek is retained because it is a major API provider and because its web/app and API model identities are relevant to the broader market. However, the first-party consumer product is free rather than a paid monthly subscription with a price that can be matched to an equal API budget. Assigning an artificial $20 subscription value would violate the central methodology of this paper. DeepSeek API calculations are therefore presented only as market references and are excluded from rankings that claim to compare paid subscription usage with equal-dollar API usage. [S35, S36, S38]
Effort and Reasoning-Mode Comparison Matrix
The effort matrix is not intended to rank labels across providers. Terms such as low, high, xhigh, max, thinking level, and reasoning effort are interface controls with provider-specific semantics. Two models set to "high" may spend very different amounts of hidden reasoning, visible output, tool calls, or latency. The matrix therefore documents which controls exist and how they are billed, while the economic comparison remains anchored to observed billable tokens or tasks completed per dollar.
This distinction becomes especially important in orchestration systems that dynamically choose a model and effort mode. If an application assumes that all high-effort modes cost the same multiple of a low-effort mode, budgets can drift rapidly. A production system should log input tokens, cached input, reasoning tokens where exposed, visible output, tool calls, latency, and success rate by model-effort pair so that routing decisions can be made on empirical cost-performance data.
This matrix enumerates the current reasoning controls captured in the provider documentation reviewed for this report. It is intentionally model-specific. Labels with the same spelling are not assumed to represent the same compute or token budget across providers.
The matrix is particularly useful for designing evaluation harnesses. Instead of asking whether one provider's "high" setting is cheaper than another provider's "high" setting, a benchmark can define a task set and quality threshold, run each model-effort pair, and record the billable tokens, latency, tool usage, and success rate. The resulting cost-per-success measure is comparable even when the providers use different effort names or expose different amounts of internal reasoning. This converts qualitative controls into an empirical engineering decision.
Effort and reasoning-mode comparison matrix
| Provider | Model | Supported effort/thinking | Input | Output | Context | Important behavior |
|---|---|---|---|---|---|---|
| OpenAI | GPT-5.6 Sol | none, low, medium, high, xhigh, max | $4/M | $20/M | 1.05M | Promotional rate through at least Nov. 21, 2026; >272K input triggers 2x input and 1.5x output pricing |
| OpenAI | GPT-5.6 Terra | none, low, medium, high, xhigh, max | $2/M | $12/M | 1.05M | >272K input triggers long-context surcharge |
| OpenAI | GPT-5.6 Luna | none, low, medium, high, xhigh, max | $0.2/M | $1.2/M | 1.05M | Go uses Luna; >272K input triggers long-context surcharge |
| Anthropic | Claude Fable 5.1 / Mythos 5.1 | low, medium, high, xhigh, max | $10/M | $50/M | 1M | Adaptive thinking; $0.25/M cache reads; newer tokenizer uses ~30% more tokens for the same text |
| Anthropic | Claude Opus 5 | low, medium, high, xhigh, max | $5/M | $25/M | 1M | Thinking on by default |
| Anthropic | Claude Sonnet 5 | low, medium, high, xhigh, max | $2/M | $10/M | 1M | $2/$10 rate made permanent Aug. 10, 2026 |
| Anthropic | Claude Haiku 4.5 | No current effort parameter | $1/M | $5/M | 200K | Fastest Claude tier |
| Gemini 3.7 Flash | low, medium, high | $0.75/M | $3.75/M | 1M | Introductory pricing through Dec. 31, 2026; output includes thinking tokens | |
| Gemini 3.1 Pro Preview | low, medium, high | $2/M | $12/M | 1M | For prompts >200K: $4 input / $18 output | |
| xAI | Grok 4.6 | low, medium, high, xhigh | $2/M | $6/M | 500K | Long context >=200K: $4 input / $12 output; Batch not supported |
| xAI | Grok 4.5 | low, medium, high | $2/M | $6/M | 500K | Long context >=200K: $4/$12 |
| Mistral | Mistral Medium 3.5 / Small 4 | none, high | $1.5/M – $0.15/M | $7.5/M – $0.6/M | 256K | Adjustable reasoning documented at none and high only |
| Perplexity | Sonar Pro | No model effort; search type affects request fee | $3/M | $15/M | Model dependent | Adds request fees: Fast $6/$10/$14 or Pro $14/$18/$22 per 1K, by context |
| Perplexity | Sonar Deep Research | low, medium, high | $2/M | $8/M | Research workflow | Additional reasoning-token, citation-token, and search-query charges |
| DeepSeek | V4 Flash / V4 Pro | none, low, high, max | $0.22/M – $1.32/M | $0.66/M – $3.96/M | 1M | Off-peak is all hours except 01:00-04:00 and 06:00-10:00 UTC; medium/xhigh map to high |
Reasoning Effort: How It Changes Cost Without Changing the Rate Card
A posted token rate answers how much each billable token costs; it does not answer how many tokens the model will choose to consume on a difficult request. Reasoning controls change the second quantity. This is why effort can change cost dramatically even when the provider's dollars-per-million rate remains unchanged. The economic unit that matters to a developer is frequently successful tasks per dollar, not nominal tokens per dollar.
For subscription users, the same phenomenon appears indirectly. Higher-effort modes may consume a rolling allowance faster or trigger stricter limits even though the user never sees the hidden reasoning token count. Consequently, a fair subscription/API comparison should match not only the model name but also the quality and reasoning behavior expected from the task. Comparing low-effort API responses with high-effort subscription responses would understate API cost for the matched quality target.
For most token-priced APIs, changing reasoning effort does not change the posted dollars-per-million-token rate. It changes the number of output, reasoning, tool-call or agent-loop tokens the model chooses to consume. That distinction is essential: "$20 buys 2.5M GPT-5.6 Sol total tokens at a 75/25 mix" remains arithmetically true, but a max-effort task may spend those output tokens much faster than a low-effort task. The correct unit for an effort comparison is therefore tasks-per-dollar or successful-work-per-dollar, not a fictitious fixed number of high-effort messages.
Reasoning-heavy workloads should also be evaluated for variance, not only for averages. A model may answer routine prompts with modest reasoning and occasionally spend far more on difficult cases, creating a long tail of expensive requests. Capacity planning based only on the mean can therefore underestimate budget risk. Logging the distribution of reasoning and output tokens by task category allows developers to set routing rules, effort caps, or escalation policies that preserve quality while limiting unexpectedly expensive outliers.
Provider-specific interpretation
OpenAI exposes API reasoning.effort settings from none through max across the GPT-5.6 model sizes covered in this study. Hidden reasoning tokens are billed as output tokens, while ChatGPT's consumer-facing settings do not disclose a one-to-one public token budget for each named mode. Accordingly, the report treats the API controls as precise billing inputs but does not assume that a similarly named consumer setting consumes an identical quantity of reasoning tokens.
Anthropic describes effort as a behavioral control that can influence total response-token expenditure, tool calls, and thinking. High is the API default, while xhigh and max are available only on specified models. Because Anthropic recommends evaluating effort through observed performance and token consumption, this report avoids assigning fixed token multipliers to the provider's named effort levels except in clearly identified illustrative sensitivity scenarios.
Google's Gemini 3 thinking_level settings represent relative thinking allowances, and published output pricing includes thinking tokens. Gemini 2.5 uses the thinkingBudget mechanism rather than thinkingLevel. These controls are therefore treated as model-specific reasoning mechanisms rather than as universally comparable labels across model generations.
xAI's Grok 4.6 exposes low, medium, high, and xhigh reasoning settings. The public documentation describes capability and latency trade-offs but does not publish fixed token allocations for those levels. The analysis therefore does not convert the labels into exact token budgets unless a measured workload provides that evidence.
Mistral's current reasoning guidance documents none and high settings for the Small and Medium models evaluated here. The high setting surfaces a full thinking segment and consumes additional tokens, making it economically distinct from a non-reasoning request even when the underlying per-token rate remains unchanged.
DeepSeek identifies low, high, and max as effective effort settings, while medium and xhigh map to high for compatibility. This mapping illustrates why effort labels should not be compared by name alone: two interfaces can expose different labels while resolving to the same underlying behavior.
Perplexity Sonar Deep Research exposes low, medium, and high reasoning effort and can charge for reasoning tokens in addition to search and citation components. Higher effort can also increase the number of searches performed. For research-oriented workloads, total API economics therefore depend on more than text-token pricing alone.
API Service-Tier and Caching Sensitivity
API optimization can change the break-even point without changing the subscription's sticker price. Batch or flex processing may reduce token rates in exchange for latency, while fast or priority tiers may increase rates to obtain lower latency or stronger throughput guarantees. Prompt caching can reduce repeated-prefix costs dramatically when an application repeatedly sends the same system instructions, documents, schemas, or long-lived conversational context.
These mechanisms are especially relevant to agentic and retrieval systems because repeated context can dominate input volume. A consumer subscription may benefit from provider-side caching internally, but the user generally cannot audit or price that optimization. An API application, by contrast, can often measure cache-hit tokens directly and redesign prompts around reuse. The uncached baseline in this report is therefore a conservative common reference, not the minimum achievable API cost.
Service-tier and caching decisions are best evaluated after the baseline model choice has been established. The baseline answers what the model costs under ordinary uncached inference; the optimization layer then asks how much of that cost can be exchanged for latency, scheduling flexibility, or prompt reuse. Keeping these stages separate makes the economics easier to audit and prevents a heavily optimized API deployment from being compared against a subscription as though the discount were an inherent property of the model itself.
Why service tier is not "effort"
Batch, Flex, Fast and Priority alter scheduling, latency, throughput or SLA characteristics. They do not mean the model reasons more deeply. Mixing these dimensions would make the subscription comparison incoherent. The core model uses standard service. This section shows how implementation choices can move the API break-even threshold without changing the consumer subscription price.
A useful way to separate the concepts is to treat service tier as a delivery contract and effort as a model-behavior control. Service tier changes when or how quickly the request is processed and what price is charged for that delivery. Effort changes how much reasoning the model may perform to produce the answer. A low-effort request can be sent through a premium latency tier, and a high-effort request can be processed asynchronously through a discounted tier. Conflating the two would obscure both the quality target and the actual source of the cost change.
Prompt caching
Prompt caching can radically change API economics for iterative conversations, agents and RAG. OpenAI GPT-5.6 cached input is generally 10% of uncached input. Anthropic cache hits are typically one-tenth of base input, while cache writes have multipliers. Google and Mistral also offer reduced cached-input rates on supported models. A subscription does not expose an equivalent user-visible cache-hit discount, but the provider may internally optimize context handling. Therefore a production API with high cache reuse can outperform the uncached break-even figures in this report by a large margin. [S04, S11, S21, S28]
Caching produces the greatest benefit when a large, stable prefix is reused across many requests. Typical examples include long system instructions, policy documents, codebase context, tool schemas, or retrieval material that remains unchanged for a series of turns. The engineering trade-off is that applications must structure prompts so reusable content is eligible for the provider's cache rules and must account for cache writes, expiration, and misses. Consequently, the effective savings should be measured from observed cache-hit ratios rather than assuming the headline cached-input discount applies to every input token.
API service-tier and caching sensitivity
| Provider/model | Tier | Price relationship | Effect on equal-dollar comparison |
|---|---|---|---|
| OpenAI GPT-5.6 | Batch/Flex | Typically lower than Standard; Sol promotion applies to these tiers | Raises tokens purchasable per dollar; asynchronous/latency trade-offs |
| OpenAI GPT-5.6 | Fast | Premium to Standard | Lowers tokens per dollar in exchange for latency |
| Anthropic | Batch | 50% input/output discount | Approximately doubles uncached token capacity at identical workload mix |
| Anthropic Opus 5/4.8 | Fast | $10 input / $50 output | About half the raw token capacity of Standard $5/$25 |
| Google Gemini | Batch/Flex | Often ~50% of Standard on supported models | Can roughly double token capacity; model-dependent |
| Google Gemini | Priority | 75-100% premium over Standard | Reduces token capacity in exchange for priority inference |
| Mistral | Batch | 50% discount | Approximately doubles token capacity |
Long-Context Economics
Long context creates both a technical-capacity benefit and an economic penalty. A large context window allows a model to accept more documents or conversation history in a single request, but several providers charge higher rates above specific input thresholds. When the surcharge applies to the entire request, a small increase past the threshold can raise the effective cost of both input and output tokens, reducing the amount of work an equal-dollar API budget can support.
This is one reason token compaction, retrieval, summarization, and context pruning can have first-order economic effects. A subscription user may experience long-context use as faster allowance consumption without seeing the underlying price curve. An API user can model the surcharge explicitly and decide whether to shorten context, route the request to a cheaper model, or accept the higher price in exchange for fidelity.
Long-context cost should therefore be treated as a distribution problem. A system with occasional very large requests can have a different effective monthly rate from one that remains above the surcharge threshold almost continuously. Instrumentation should record prompt length before each call and classify requests by pricing tier so that the weighted monthly cost reflects the actual context distribution. This is particularly important for agents that accumulate conversation history automatically, because a workload can drift into a more expensive tier even when the user's visible task has not changed.
Long-context comparisons are especially prone to a conceptual error: context window size is a per-request capacity limit, while subscription allowance is a rolling or periodic usage limit. They should never be added together or treated as equivalent quantities.
Long-context pricing thresholds and economic consequences
| Model | Threshold | Pricing rule | Economic consequence |
|---|---|---|---|
| OpenAI GPT-5.6 Sol | >272K input | 2x input, 1.5x output for full request | 75/25 blended rate rises from $8/M to $13.5/M; a $20 budget falls from 2.50M to ~1.48M total tokens |
| Google Gemini 3.1 Pro | >200K prompt | $4 input / $18 output vs $2/$12 | 75/25 blended cost rises from $4.50/M to $7.50/M; $19.99 capacity falls from ~4.44M to ~2.67M |
| xAI Grok 4.6 | >=200K context | $4 input / $12 output vs $2/$6 | Rates double; token capacity halves for the same input/output mix |
| Anthropic listed 1M models | Up to 1M | Standard token rates across full context | No long-context per-token surcharge for the listed current models |
| Mistral current flagship models | Up to 256K | No separate long-context surcharge identified | Standard rates used within published context |
Relative Subscription Multipliers: What Can Be Quantified
Relative multipliers are the strongest public evidence available when an absolute token pool is unavailable, but their usefulness is confined to comparisons within the same provider. A 20x tier from one company cannot be compared directly with a 20x tier from another because the baseline product, compute accounting, model mix, and reset logic can all differ. The chart below therefore visualizes internal scaling rather than cross-provider capacity.
Price-to-multiplier efficiency can still reveal plan structure. When usage scales faster than price inside one provider's lineup, the higher tier may offer better marginal capacity per dollar for users who can actually consume the allowance. That does not prove it is the best market-wide value; the absolute baseline remains unknown and should be measured empirically before a large subscription commitment is justified solely on usage economics.
OpenAI and Anthropic use the same 5x/20x nomenclature for their $100/$200 individual tiers, but that does not imply the underlying Plus and Pro baselines are equal. Google's 2x/4x/5x/20x hierarchy compounds from a different "standard limits" baseline and uses compute-based metering. These multipliers are useful inside each provider's product line and unsuitable for cross-provider absolute token arithmetic. [S02, S10, S18, S20]
Relative multipliers can nevertheless support useful break-even reasoning once the baseline is measured. If an organization observes the sustainable workload of a baseline plan under a stable task mix, a published 5x or 20x tier provides a testable expectation for how the higher tier should scale, subject to weekly caps and policy caveats. The measured baseline must remain provider-specific; carrying it across providers would erase the differences in model mix, reset cadence, and compute accounting that make the multipliers non-comparable in the first place.
Figure 6
Figure 6. Published relative subscription usage multipliers. Each provider's baseline is different; bar lengths must not be interpreted as cross-provider absolute usage.
Published subscription usage multipliers vs. price
| Provider | Tier | Monthly $ | Published relative capacity | Relative capacity ÷ relative price vs baseline |
|---|---|---|---|---|
| OpenAI | Plus | 20 | 1x baseline | 1.00 |
| OpenAI | Pro 100 | 100 | 5x Plus | 1.00 |
| OpenAI | Pro 200 | 200 | 20x Plus | 2.00 |
| Anthropic | Pro | 20 | 1x baseline | 1.00 |
| Anthropic | Max 5x | 100 | 5x Pro | 1.00 |
| Anthropic | Max 20x | 200 | 20x Pro | 2.00 |
| AI Plus | 7.99 | 2x standard | N/A | |
| AI Pro | 19.99 | 4x standard | N/A | |
| Ultra 100 | 100 | 5x Pro = 20x standard | N/A | |
| Ultra 200 | 200 | 20x Pro = 80x standard | N/A |
Price-to-multiplier efficiency inside a provider
This ratio is a marginal-value indicator, not a recommendation to purchase the highest tier automatically. A tier can have attractive relative capacity per dollar and still be uneconomic for a user who cannot consume the additional allowance. The relevant decision is whether expected sustained demand is high enough to use the increment before reset windows or monthly billing cycles expire. When demand is intermittent, direct API spend can remain preferable because unused API budget is not committed in advance in the same way as a fixed subscription fee.
For OpenAI and Anthropic, the $100 tiers scale usage approximately linearly with price relative to the $20 baseline (5x price, 5x usage), while the $200 tiers advertise 20x usage for 10x price. On the published multiplier alone, the $200 tier therefore offers twice the relative usage-per-dollar of the baseline. This is an intra-provider ratio only; it does not reveal absolute tokens.
Subscription Usage Limits and Their API-Equivalent Interpretation
A recurring problem in consumer-AI cost analysis is the temptation to convert messages into tokens using one assumed message size and then multiply reset windows into a month. That approach can generate a precise-looking number that the provider never promised. Message size varies, context grows during a conversation, reasoning modes consume different compute, and providers can enforce multiple overlapping limits. This report therefore treats message-to-token conversion only as scenario analysis unless a provider publishes a true token allowance.
The strongest empirical alternative is account-level measurement. A user can record the number and size of representative requests, the models and modes used, when throttles occur, and when allowances reset. Over several weeks, that data can be translated into an observed token-equivalent workload and compared with the API threshold. Such a measurement is specific to the user's behavior, but it is substantially more defensible than a universal token estimate derived from marketing language.
Account-level measurement should also separate attempted demand from completed demand. If a subscription begins throttling, users may reduce prompt frequency, switch models, shorten requests, or postpone work; the observed token total can therefore understate what they would have consumed in the absence of limits. A rigorous study records both the workload completed and the work deferred or rerouted after limits appear. This makes it possible to distinguish a genuinely low-usage account from one whose behavior was constrained by the subscription itself.
The central practical challenge is that consumer AI subscriptions are sold as access products rather than as prepaid token bundles. A provider may describe a tier as having higher limits, five times the usage of a baseline plan, twenty times the usage, or a shared compute pool without publishing the number of input, output, or reasoning tokens represented by that allowance. For this reason, the study treats the subscription side as an observed or inferable quantity rather than assigning a fictional token balance. Where the provider publishes a stable numerical allowance, that allowance is converted under an explicitly defined workload. Where the provider publishes only relative multipliers, the multiplier is used to compare tier efficiency and to scale an empirically measured baseline. Where the provider uses dynamic compute limits, the report identifies the equal-dollar API threshold and explains the amount of comparable subscription use that must be observed for the subscription to provide more raw model consumption per dollar.
This distinction also explains why users can report very different experiences on the same plan without either account necessarily being inconsistent. Long chats carry more context, document uploads increase input volume, higher reasoning settings consume more compute, tool use can draw from separate or shared pools, and some providers enforce both rolling-session and weekly ceilings. A useful comparison therefore cannot equate one message with one fixed number of tokens. It must normalize the workload, record the model and effort setting, and evaluate sustainable usage across the provider's full limiting period. The resulting comparison is more informative than a headline message count because it expresses the practical question in a common economic unit: how much comparable model work can be completed through the subscription before the same amount of money would have been exhausted through the API.
Subscription usage limits and report treatment
| Plan | Published usage statement | Why token conversion is unstable | Treatment in this report |
|---|---|---|---|
| OpenAI Go | Unlimited everyday text chats | Separate tool limits; Think uses Luna | Named-model comparison; no fixed monthly token pool |
| OpenAI Plus | Plan-dependent reasoning limits | No current exact public token quota | Same-dollar API threshold plus observed subscription usage |
| OpenAI Pro | 5x or 20x Plus | Core capabilities same; tier changes allowance | Relative-multiplier analysis; Sol benchmark; Pro-mode caveat |
| OpenAI Business Standard | Sol 10-100; Terra 25-200; Luna 250-2,000 local messages / 5h | Actual task/effort/cloud use varies | Scenario conversion from published message ranges |
| Claude Pro | At least 5x free per session; 5h resets + weekly cap | Message length/files/model/features change usage | Same-dollar API threshold plus observed subscription usage |
| Claude Max | 5x or 20x Pro per session | Weekly/discretionary limits also apply | Relative-multiplier analysis plus model-specific API thresholds |
| Gemini AI plans | Plus 2x; Pro 4x; Ultra 5x/20x Pro | Compute-based; 5h refresh until weekly limit | Relative compute multipliers plus model-family API threshold range |
| SuperGrok | Shared weekly pool | Product actions consume different compute | Shared-pool same-dollar threshold; isolate text-chat use for text comparison |
| Mistral Vibe | Fair use; example feature counts; PAYG extension | Current Pro page displays $30 API credits | Separate fair-use Vibe allowance from included API-credit value |
| Perplexity | Feature/model limits and credits | Search/Computer/Research use separate mechanisms | Weighted model-mix API benchmark; no unified subscription token pool |
OpenAI Business local-message scenario conversion
OpenAI Business is the closest major subscription to a numeric model-level allowance, but even here the published values are local-message estimates rather than tokens. To demonstrate the conversion mechanics without mislabeling them, the table below applies three hypothetical message sizes. It should be read as a sensitivity analysis, not as OpenAI's token quota.
Illustrative five-hour Business usage scenarios
| Model | Published local messages / 5h | Hypothetical billable tokens/message | Implied scenario tokens / 5h |
|---|---|---|---|
| Sol | 10-100 | 1,000 | 0.010-0.100M |
| Sol | 10-100 | 10,000 | 0.100-1.000M |
| Terra | 25-200 | 1,000 | 0.025-0.200M |
| Terra | 25-200 | 10,000 | 0.250-2.000M |
| Luna | 250-2000 | 1,000 | 0.250-2.000M |
| Luna | 250-2000 | 10,000 | 2.500-20.000M |
The hypothetical message sizes are useful because they expose the degree of uncertainty rather than hiding it. A tenfold change in assumed tokens per message creates a tenfold change in the implied token total even before reasoning, tool activity, or cloud execution is considered. The exercise demonstrates why a provider-published message range cannot be collapsed into one token figure without additional evidence. For a real organization, telemetry from representative Business workloads should replace the hypothetical message sizes before any procurement conclusion is drawn.
A monthly extrapolation is deliberately omitted. Running every five-hour window continuously would assume a usage pattern and reset behavior beyond the published local-message estimate and would ignore cloud-task consumption and other plan constraints. The correct empirical method is to measure actual usage on the account over a representative month.
Perplexity Multi-Model Subscription: Model-by-Model API-Equivalent Cost
The multi-model table illustrates why one subscription can represent several different API-equivalent values. A month dominated by a lower-cost model has a much higher raw-token break-even point than a month dominated by a premium model. Therefore, the economic value of Perplexity depends heavily on the user's model-selection behavior, not merely on how many searches or chats are performed.
For procurement or platform design, the appropriate method is to build a weighted model mix. If telemetry shows that 60% of billable tokens would have gone to one provider model, 25% to another, and 15% to a premium reasoning model, the direct-API equivalent cost can be calculated from those weights. The resulting blended threshold is more informative than selecting a single representative model for the entire subscription.
A weighted model mix can also be segmented by task category. Research queries may favor one provider, coding tasks another, and high-stakes synthesis a premium reasoning model. Calculating API-equivalent spend by task reveals whether the subscription's value comes from broad convenience or from concentrated use of a small number of expensive models. This information is useful when deciding whether to retain the subscription, replace part of the workload with direct APIs, or use a hybrid strategy in which the subscription remains the interactive front end while automated volume moves to APIs.
Perplexity multi-model direct API-equivalent costs
| Subscription access | Model | Provider | Direct API rate used | Pro $20 (75/25) | Max $200 (75/25) | Caveat |
|---|---|---|---|---|---|---|
| Pro / Max | GPT-5.6 Terra | OpenAI | $2 / $12 | 4.44M | 44.44M | Thinking optional in Perplexity |
| Max | GPT-5.6 Sol | OpenAI | $4 / $20 | N/A | 25.00M | Max-only in Search |
| Pro / Max | Gemini 3.1 Pro | $2 / $12 <=200K | 4.44M | 44.44M | Thinking always on in Perplexity | |
| Pro / Max | Claude Sonnet 5 | Anthropic | $2 / $10 | 5.00M | 50.00M | Thinking optional |
| Max | Claude Opus 5 | Anthropic | $5 / $25 | N/A | 20.00M | Max-only in Search; Thinking optional |
| Pro / Max | Grok 4.5 | xAI | $2 / $6 | 6.67M | 66.67M | Thinking optional |
| Pro / Max | Kimi K3, GLM 5.2, Nemotron 3 Ultra, Sonar 2 | Moonshot, Z.ai, NVIDIA, Perplexity | Not verified in this audit | N/A | N/A | Kept N/A rather than importing an unverified rate |
Coverage of Major Providers Without a Direct Paid-Subscription Match
Completeness requires documenting non-comparability rather than forcing every provider into the same table. Open-weight model publishers, enterprise-only offerings, preview APIs, and consumer assistants that do not expose a matched paid model can all be economically important without satisfying the equal-dollar subscription test. Marking these cases N/A preserves the meaning of the comparison and makes the report easier to extend when a provider later introduces a compatible paid plan.
For an AI platform that aggregates frontier models, these N/A providers may still be strategically relevant because the API or self-hosted model can be integrated even when there is no consumer subscription benchmark. Their absence from the equal-dollar ranking should therefore be interpreted as a data-structure limitation, not as a judgment about model quality, capability, or deployment value.
The user request calls for every major model provider. A "complete" analysis therefore cannot simply omit providers that fail the equal-dollar test. The correct treatment is to audit them and state why a numeric subscription-vs-API comparison is unavailable.
Non-comparable providers should still remain in the market map because their availability can affect architecture choices. An open-weight or enterprise-only model may have deployment, privacy, latency, or customization advantages that matter more than consumer subscription parity. The N/A classification therefore limits only the specific equal-dollar subscription test; it does not imply that the provider is outside the competitive set for API orchestration, self-hosting, or enterprise procurement. A future update can move a provider into the numeric comparison if a directly matched paid plan becomes publicly available.
This report deliberately does not manufacture comparison budgets for N/A rows. A free assistant, an enterprise contract, an open-weight download, or a routed subscription is economically different from a public fixed-price subscription that exposes the same model as a token-priced API.
Major providers without a direct paid-subscription match
| Provider | Model family | API / developer access | Reason no equal-dollar result |
|---|---|---|---|
| Meta | Muse / Meta AI | Meta AI app + Meta Model API preview | No directly matched paid consumer model subscription price established |
| Amazon | Nova | Amazon Bedrock / Nova APIs; Alexa is a different consumer abstraction | Consumer/service bundles are not the same metered product as the Nova API |
| NVIDIA | Nemotron | NIM / inference providers / open models | No mainstream paid consumer assistant subscription for a named Nemotron model |
| Alibaba | Qwen | Qwen Chat / Model Studio / open weights | No U.S.-comparable paid consumer plan with a stable same-model token allowance verified here |
| Cohere | Command | Enterprise/API products | Enterprise commercial model, not a $20-style consumer plan |
| Moonshot AI | Kimi | Consumer/app and API ecosystem | Included only as a third-party Perplexity model; first-party USD parity not verified here |
| Z.ai | GLM | API/open model ecosystem | Direct first-party subscription parity not established |
| Microsoft | Copilot | Consumer/business subscription routing third-party foundation models | Microsoft is not the underlying provider of the frontier models selected in Copilot |
Consolidated API Rate Catalog
The consolidated catalog is intended as a reusable audit table. It collects the model rates, context limits, effort controls, and subscription-parity notes used elsewhere in the report so that a reader can reproduce the equal-dollar calculations without searching across provider sections. Prices should be treated as point-in-time values and revalidated before budgeting because promotional rates, previews, model replacements, and long-context policies can change rapidly.
For reproducibility, the catalog should be read together with model version and pricing-condition notes. A rate is not fully specified by input and output dollars per million tokens when long-context thresholds, cache pricing, promotional windows, or time-of-day rules can change the effective cost. Teams using this appendix as a budgeting reference should therefore preserve the model identifier and applicable pricing condition in their own records, then revalidate the source page before committing material spend.
Reasoning, Parity, and Implementation Notes
This companion table captures the non-price attributes that determine whether a rate can be used as a true parity benchmark. Context size, effort controls, subscription mapping, and special pricing behavior can materially change both cost and comparability. It is included separately from the rate catalog so that readers can audit the economic assumptions without overloading the core price table, while still retaining the information necessary to reproduce the calculations accurately - see the per-provider model matrices and the effort and reasoning-mode comparison matrix above for those fields.
Consolidated API rate catalog
| Provider | Model | Input | Cached | Output |
|---|---|---|---|---|
| OpenAI | GPT-5.6 Sol | $4/M | $0.4/M | $20/M |
| OpenAI | GPT-5.6 Terra | $2/M | $0.2/M | $12/M |
| OpenAI | GPT-5.6 Luna | $0.2/M | $0.02/M | $1.2/M |
| Anthropic | Claude Fable 5.1 / Mythos 5.1 | $10/M | $0.25/M | $50/M |
| Anthropic | Claude Fable 5 | $10/M | $1/M | $50/M |
| Anthropic | Claude Opus 5 | $5/M | $0.5/M | $25/M |
| Anthropic | Claude Sonnet 5 | $2/M | $0.2/M | $10/M |
| Anthropic | Claude Haiku 4.5 | $1/M | $0.1/M | $5/M |
| Gemini 3.7 Flash | $0.75/M | — | $3.75/M | |
| Gemini 3.5 Flash-Lite | $0.3/M | $0.03/M | $2.5/M | |
| Gemini 3.1 Pro Preview | $2/M | $0.2/M | $12/M | |
| Gemini 2.5 Pro | $1.25/M | — | $10/M | |
| Gemini 2.5 Flash | $0.3/M | — | $2.5/M | |
| xAI | Grok 4.6 | $2/M | $0.5/M | $6/M |
| xAI | Grok 4.5 | $2/M | $0.3/M | $6/M |
| xAI | Grok 4.3 | $1.25/M | $0.2/M | $2.5/M |
| xAI | grok-build-0.1 | $1/M | $0.2/M | $2/M |
| Mistral | Mistral Medium 3.5 | $1.5/M | $0.15/M | $7.5/M |
| Mistral | Mistral Small 4 | $0.15/M | $0.015/M | $0.6/M |
| Mistral | Mistral Large 3 | $0.5/M | $0.05/M | $1.5/M |
| Perplexity | Sonar | $1/M | — | $1/M |
| Perplexity | Sonar Pro | $3/M | — | $15/M |
| Perplexity | Sonar Reasoning Pro | $2/M | — | $8/M |
| Perplexity | Sonar Deep Research | $2/M | — | $8/M |
| DeepSeek | V4 Flash - off-peak | $0.22/M | $0.007/M | $0.66/M |
| DeepSeek | V4 Pro - off-peak | $0.66/M | $0.022/M | $1.98/M |
Worked Calculation Examples
The worked examples expose the arithmetic behind the headline numbers. Each example begins with the subscription's actual monthly price, computes a blended dollars-per-million-token rate from the chosen input/output mix, and divides the budget by that blended rate. The examples also show why changing context tier or workload mix can materially change the answer even when the plan price and model remain fixed.
The examples are deliberately transparent enough to be recomputed in a spreadsheet or a few lines of code. Their purpose is not to privilege the 75/25 workload but to show how the same formula can be parameterized for any observed input/output mix. When a provider adds a surcharge, discount, or different reasoning behavior, the corresponding rate can be substituted directly. This makes the worked examples a reusable calculation pattern rather than a static set of August 2026 answers.
Turning a Break-Even Threshold Into an Observed Usage Difference
The API calculations below become a true subscription-versus-API comparison once the subscription's comparable usage is observed. Consider ChatGPT Plus only as an illustration: the $20 Sol baseline is 2.50 million API tokens. If a particular user measures 3.00 million comparable Sol tokens of subscription work during a representative month, the subscription delivered 20 percent more raw model usage than $20 of API spend because 3.00M ÷ 2.50M - 1 = 0.20. If the same user measures only 2.00 million comparable tokens, the API would provide 25 percent more raw usage because 2.50M ÷ 2.00M - 1 = 0.25. These 3.00M and 2.00M figures are hypothetical examples, not claims about the Plus allowance.
The same logic applies to Claude, Gemini, Grok, Mistral, and other matched plans. The API threshold establishes the denominator at the plan's own price; the user's observed subscription throughput establishes the numerator. When the provider exposes only a relative multiplier, the baseline plan can first be measured empirically and then the higher tier can be tested against the provider's stated scaling. This produces a defensible percentage difference without converting undisclosed provider compute into a fabricated universal token quota.
ChatGPT Plus → GPT-5.6 Sol
The spread between the input-only and output-only bounds demonstrates the sensitivity of Sol economics to response length and reasoning spend. Workloads dominated by short outputs can approach the higher capacity bound, while long-form generation or high-effort reasoning moves the effective throughput toward the lower bound.
Claude Pro → Sonnet 5
For Claude Sonnet, changing only the workload mix moves $20 of API capacity by more than twofold. This is why message counts cannot be converted into a universal cost without knowing how much of each interaction is prompt/context versus model-generated output and thinking.
Google AI Pro → Gemini 3.1 Pro
The long-context surcharge reduces equal-dollar capacity by roughly forty percent in this normalized example. Applications that routinely attach large corpora should therefore model the distribution of requests above and below the threshold rather than applying the short-context rate to the entire month.
SuperGrok → Grok 4.6
This is a particularly clear example of threshold economics: the same monthly dollars, model, and nominal token mix can support only half as many tokens once the long-context pricing tier applies. Context management can therefore be as important as model selection for API cost control.
Mistral Vibe Pro → Medium 3.5
The credit comparison also highlights why subscription benefits should be decomposed before valuation. The cash plan price, the included API-credit face value, and the fair-use Vibe allowance are separate economic components. Their combined value can be attractive, but only the credit portion converts directly into a known number of API tokens.
DeepSeek V4 Pro off-peak reference
There is no paid consumer plan to match. As a market reference only, the calculation below applies a hypothetical $20 API spend to DeepSeek's off-peak rates.
This is not presented as a subscription comparison. The reference remains useful for cross-provider rate awareness because it shows the scale of raw token purchasing power available at DeepSeek's off-peak rates. It should not be used to conclude that a $20 DeepSeek subscription would contain that capacity, because no matched $20 paid plan exists as the basis for that comparison in this analysis.
Decision Framework for Developers and AI Platforms
The decision framework translates the report from a static rate comparison into a repeatable engineering process. The objective is to replace generic assumptions with workload telemetry, separate model cost from tool and infrastructure cost, and revisit the economics whenever providers change pricing or routing. For an API-centric application, this process can also support dynamic model routing based on cost, latency, quality, context length, and user-defined budget constraints.
For a developer deciding between consumer subscriptions and API funding - or for an application such as Omnesly that connects user-owned API keys - the economically relevant question is workload-specific. The following decision sequence avoids the most common errors.
A rigorous evaluation begins by identifying the exact model SKU and effort behavior required by the workload. Provider-family names such as "Claude" or "Gemini" are too broad for cost comparison when several model tiers, versions, and reasoning configurations are available at materially different prices.
The next step is to measure actual input tokens, cached-input tokens, visible output tokens, and reasoning-output tokens on a representative workload. The 75/25 input-to-output assumption used in this report is useful as a planning baseline, but production decisions should replace that assumption with telemetry from the organization's own prompts, documents, agents, and user behavior.
Total workload cost should then incorporate tool charges, search fees, retrieval charges, image or audio processing, and repeated agent-loop execution where applicable. In research and agentic products, these non-token or auxiliary charges can become a larger component of total API expenditure than the underlying language-model tokens.
Implementation-specific economics should also be modeled explicitly. Cache-hit rates, eligibility for Batch or Flex processing, and the distribution of requests that trigger long-context surcharges can materially change effective API capacity. A production architecture with high prompt reuse may therefore reach a substantially different cost conclusion from an uncached baseline.
For subscription benchmarking, actual usage should be recorded across several representative weeks and across the provider's relevant throttle or reset windows. Many plans use rolling five-hour windows, weekly pools, or dynamic compute accounting, so a short stress test cannot reliably establish the amount of sustained monthly capacity available to a typical user.
Bundled product capabilities should be valued separately from raw token throughput. Storage, connectors, browser control, research tools, multi-model coordination, image or video generation, IDE agents, and collaboration features can provide substantial utility, but converting those benefits into an invented token quantity would reduce analytical clarity.
Finally, the comparison should be recalculated whenever a provider changes pricing, model routing, plan limits, or included features. Several rates reviewed for this study changed within weeks of the August 31 research date, which illustrates how quickly a previously accurate cost comparison can become outdated.
For an individual deciding between a fixed subscription and an API account, utilization is often more important than the sticker price. A heavy interactive user who repeatedly clears the same-dollar API threshold can receive more model work from a subscription because the marginal cost of additional included use is effectively zero until the plan limit is reached. A light or irregular user may be better served by the API because there is no need to prepay for unused access. This is why the report avoids declaring one channel universally cheaper and instead calculates the point at which the answer changes for each plan and model.
For developers, the API can also outperform a direct same-model benchmark through architecture. Prompt caching, model routing, shorter outputs, batching, and context management can reduce the effective cost below the standard uncached rate used in the main comparison. A subscription may still provide more gross interactive access, but an optimized API workflow can require fewer paid tokens to complete the same business process. The report therefore distinguishes raw model-usage capacity from engineering efficiency and total product value.
Limitations and Reproducibility Notes
These limitations are substantive constraints on interpretation rather than boilerplate. Subscription backends can change without exposing every infrastructure detail, tokenizers differ across model generations, consumer model names may represent routed families, and provider documentation can change after publication. The reproducible quantity is the calculation method: once current plan prices, API rates, workload mix, and observed subscription usage are known, the same equations can be rerun with updated inputs.
This report is a point-in-time analysis based on information available as of September 1, 2026. AI pricing, model availability, and subscription limits are unusually volatile, so the cited source pages should be rechecked before procurement, budgeting, or architecture decisions are finalized.
Consumer subscription backends may route requests among model variants, quantize workloads, cache repeated context, summarize conversation history, or apply other infrastructure optimizations that are not visible to the user. The API-equivalent comparison in this report prices published billable units rather than attempting to estimate a provider's internal compute cost.
"Same effort" can be treated as exact only when the same model and the same effort control are available on both the subscription and API surfaces. When a consumer interface uses labels such as Think, Extended, Deep Think, or Pro instead of the API parameter, the report identifies the resulting mapping limitation rather than assuming equivalence.
Google family-crosswalk rows should not be interpreted as assertions about a pinned consumer backend. They represent the closest current first-party API economic references to the model-family names exposed in the consumer application, subject to the routing and versioning caveats described in the Google section.
For Perplexity, third-party model rows use direct-provider API rates as the API-equivalent benchmark when the current first-party model identity is explicit. Native Sonar products require additional consideration of request, search, citation, or reasoning charges and therefore cannot always be reduced to a simple text-token comparison.
Mistral's displayed API-credit benefits are reported as observed on the official pricing page during the research period. Eligibility, regional plan availability, localization, or promotional terms may change and should be confirmed directly with the provider before purchase.
Tokenizers are model-specific. Raw token counts are appropriate for billing arithmetic and standardized capacity comparisons, but they should not be treated as a universal unit of text length, semantic work, or task completion across providers.
The analysis does not assign a single dollar value to latency, output quality, tool quality, reliability, privacy controls, storage, or other included ecosystem benefits. Those dimensions can dominate a real purchasing decision even when raw token economics appear to favor a different delivery channel, and they should therefore be evaluated alongside the quantitative results presented here.
Most importantly, the report does not claim an exact monthly subscription token total where the provider does not publish or expose one. In those cases the numerical API side is exact within the stated pricing assumptions, while the subscription difference is conditional on observed throughput. This is a deliberate evidence standard. A publication that converts an opaque weekly compute pool into a precise monthly token number without telemetry would appear more complete but would be less correct.
Sources
- S01OpenAI - GPT-5.6 in ChatGPT
help.openai.com/en/articles/20001354-gpt-56-in-chatgptPlan/model mapping; Sol powers Medium/High/Extra High on eligible paid plans; Sol Pro powers Pro; usage-limit disclosure. - S02OpenAI - About ChatGPT Pro tiers
help.openai.com/en/articles/9793128Pro $100 = 5x Plus usage; Pro $200 = 20x Plus. - S03OpenAI - Introducing ChatGPT Go
openai.com/index/introducing-chatgpt-goU.S. Go $8, Plus $20, Pro $200 reference pricing. - S04OpenAI - GPT-5.6 Sol model
developers.openai.com/api/docs/models/gpt-5.6-solSol pricing, context, effort levels, long-context surcharge and promo. - S05OpenAI - GPT-5.6 Terra model
developers.openai.com/api/docs/models/gpt-5.6-terraTerra pricing, context, effort levels, long-context surcharge. - S06OpenAI - GPT-5.6 Luna model
developers.openai.com/api/docs/models/gpt-5.6-lunaLuna pricing, context, effort levels, long-context surcharge. - S07OpenAI - ChatGPT Business models and limits
help.openai.com/en/articles/12003714Model-level local-message estimates; task/effort-dependent usage. - S08OpenAI - Business Pricing
openai.com/business/pricingBusiness Standard and Premium seat list prices and annual equivalents. - S09OpenAI - Advancing the price-performance frontier with GPT-5.6
openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6Terra/Luna price cuts and subscription-credit impact. - S10Anthropic - Choose a Claude plan
support.claude.com/en/articles/11049762-choose-a-claude-planPro $20; Max $100/$200; 5x and 20x session capacity. - S11Anthropic - Claude pricing
platform.claude.com/docs/en/about-claude/pricingCurrent model, cache, Batch, fast-mode and long-context pricing. - S12Anthropic - Effort
platform.claude.com/docs/en/build-with-claude/effortSupported effort levels by model; behavioral rather than strict token budgets. - S13Anthropic - What's new in Claude Sonnet 5
platform.claude.com/docs/en/docs/about-claude/models/whats-new-sonnet-5Sonnet 5 $2/$10; tokenizer change; API availability. - S14Anthropic - Claude Opus 5 - what is new
platform.claude.com/docs/en/about-claude/models/whats-new-claude-4-6Opus 5 $5/$25, 1M context, thinking. - S15Anthropic - Claude Fable 5 and Mythos 5
platform.claude.com/docs/en/models/fable-5/introducing-claude-fable-5-and-claude-mythos-5Fable 5 $10/$50, context/output, availability. - S16Anthropic - Claude Fable models on your plan
support.claude.com/en/articles/15424964-claude-fable-models-on-your-planFable 5 and Fable 5.1 inclusion on Max/premium seats, shared 50% weekly Fable ceiling, and PAYG usage-credit treatment on Pro/standard seats. - S17Anthropic - Claude Team pricing
claude.com/pricingTeam Standard/Premium prices and relative usage. - S18Google - Gemini Apps limits and upgrades
support.google.com/gemini/answer/16275805Compute-based limits, 5-hour/weekly resets, plan multipliers, model families, context windows. - S19Google - Google AI Plus availability and U.S. price
blog.google/products-and-platforms/products/google-one/google-ai-plus-availabilityAI Plus U.S. list price $7.99. - S20Google - Google AI subscription updates from I/O 2026
blog.google/products-and-platforms/products/google-one/google-ai-subscriptionsAI Ultra $100 and $200; 5x/20x Pro usage. - S21Google - Gemini Developer API pricing
ai.google.dev/gemini-api/docs/pricingCurrent Gemini API Standard/Batch/Flex pricing; thinking included in output. - S22Google - Gemini thinking
ai.google.dev/gemini-api/docs/thinkingThinking-level support by model. - S23Google - Gemini 3.7 Flash release
ai.google.dev/gemini-api/docs/latest-model3.7 Flash pricing, 1M context, low/medium/high thinking; promo end date. - S24xAI - FAQ - Grok Website / Apps
docs.x.ai/grok/faqPaid plans draw from a single weekly usage allowance; different products consume different compute. - S25xAI - Grok 4.6 model
docs.x.ai/developers/models/grok-4-6Grok 4.6 $2/$6, 500K context, low/medium/high/xhigh. - S26xAI - API Pricing
docs.x.ai/developers/pricingCurrent Grok text model rates and long-context surcharges. - S27Mistral - Pricing
mistral.ai/pricingVibe plan prices, fair-use, API credits, PAYG extension. - S28Mistral - API pricing
docs.mistral.ai/inference/pricingMistral model input/cache/output rates. - S29Mistral - Reasoning
docs.mistral.ai/studio/conversations/reasoningAdjustable reasoning on Small and Medium; none/high documented. - S30Perplexity - Advanced AI models included in subscription
perplexity.ai/help-center/en/articles/10354919Current Pro/Max/Enterprise model selector and Thinking availability. - S31Perplexity - Which Perplexity Subscription Plan is right for you?
perplexity.ai/help-center/en/articles/11187416Consumer and enterprise plan comparison and published usage-limit descriptions. - S32Perplexity - Agent API models
docs.perplexity.ai/docs/agent-api/modelsThird-party pass-through pricing policy; model catalog; no markup statement. - S33Perplexity - API pricing
docs.perplexity.ai/docs/getting-started/pricingSonar/Deep Research charge components and examples. - S34Perplexity - Sonar Deep Research
docs.perplexity.ai/docs/sonar/models/sonar-deep-researchLow/medium/high reasoning effort. - S35DeepSeek - Models & Pricing
api-docs.deepseek.com/quick_start/pricingV4 Flash/Pro peak and off-peak API rates, context and max output. - S36DeepSeek - DeepSeek V4 Pro GA Release
api-docs.deepseek.com/news/news260813App/web Expert Mode and API parity; low/high/max reasoning. - S37DeepSeek - Thinking Mode
api-docs.deepseek.com/guides/thinking_modeThinking toggle and effort mapping. - S38DeepSeek - DeepSeek homepage
deepseek.comOfficial free access statement. - S39Meta - Introducing Muse Spark 1.1
research.meta.ai/blog/introducing-muse-spark-meta-model-apiMeta Model API preview and Meta AI app Thinking-mode deployment; used for comparability audit. - S40Meta - Introducing Muse Glimmer
research.meta.ai/blog/introducing-muse-glimmer-open-agentic-modelOpen-model/local-delivery evidence; used for provider coverage audit. - S41Amazon - Amazon Nova pricing
aws.amazon.com/nova/pricingNova usage-based API/service pricing; no same-named consumer assistant plan match. - S42NVIDIA - Nemotron models
nvidia.com/en-us/ai-data-science/foundation-models/nemotronOpen/inference-provider model distribution; used for coverage audit. - S43xAI - Pricing: Compare Grok Plans
x.ai/pricingFree, SuperGrok $30 and SuperGrok Plus $100 consumer plan pricing and features. - S44Perplexity - Perplexity Max
perplexity.ai/help-center/en/articles/11680686Max $200/month and $2,000/year; highest-volume consumer tier. - S45Perplexity - AI for the Curious
perplexity.ai/hubPro $20/month or $200/year public price. - S46Perplexity - Enterprise Pricing and Billing FAQ
perplexity.ai/help-center/en/articles/10352986Enterprise Pro $40/month and Enterprise Max $325/month per seat. - S47OpenAI - What is ChatGPT Plus?
help.openai.com/en/articles/6950777Plus $20/month; higher model limits; API usage is billed separately. - S48Anthropic - How do usage and length limits work?
support.claude.com/en/articles/11647753Usage depends on conversation length and complexity, features, selected model, and effort. - S49Anthropic - Claude Fable 5.1 and Claude Mythos 5.1 launch
anthropic.com/claude-fable-and-mythos-5-1September 1, 2026 release; performance, scientific research, safeguards, trusted access, cache economics, pricing and availability. - S50Anthropic - Claude Fable 5.1 model overview
platform.claude.com/docs/en/models/fable-5-1/overview1M context, 128K output, $10/$50 pricing, $0.25/M cache reads, model IDs, effort defaults. - S51Anthropic - What's new in Claude Fable 5.1
platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1Fable/Mythos 5.1 feature changes, tokenizer, tool-use changes, per-message effort. - S52Anthropic - Migrating to Claude Fable 5.1 and Claude Mythos 5.1
platform.claude.com/docs/en/models/fable-5-1/migration-guideBreaking changes, preserved-thinking behavior, retention requirements, migration guidance. - S53Anthropic - Claude Fable models on your plan
support.claude.com/en/articles/15424964Fable 5/5.1 inclusion on Max and premium seats, 50% weekly Fable ceiling, PAYG treatment, ended promotion. - S54Anthropic - Effort
platform.claude.com/docs/en/build-with-claude/effortLow/medium/high/xhigh/max controls; Fable 5.1/Mythos 5.1 support all five levels. - S55Anthropic - Claude Fable 5.1 and Claude Mythos 5.1 System Card
anthropic.com/claude-fable-5-1-mythos-5-1-system-cardDetailed capability, alignment, cybersecurity, CBRN, safeguards, and deployment evaluation. - S56Anthropic - Preserved thinking API change
support.claude.com/en/articles/16761192Fable 5.1 thinking-block immutability/binding rules and cache-reuse implications. - S57Anthropic - Developing Enterprise Frontier Safeguards with our customers
anthropic.com/news/enterprise-frontier-safeguardsEFS privacy/safeguard architecture and interim Fable 5/5.1 zero-data-retention treatment. - S58Google Cloud - Claude Fable 5.1 on Google Cloud
docs.cloud.google.com/.../claude/fable-5-1Partner-platform GA confirmation, 1M input/128K output limits, supported capabilities and regional availability. - S59Anthropic - Prompting Claude Fable 5.1
platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1Model-specific guidance on effort, progress updates, parallel tool calls, targeted file edits, and token/cost implications. - S60Anthropic - Claude Platform pricing
platform.claude.com/docs/en/about-claude/pricingOfficial Fable 5.1/Mythos 5.1 input, cache-write, cache-read, output, tokenizer, cloud-platform, and Batch pricing. - S61Anthropic - Claude Fable 5.1 system prompts
platform.claude.com/docs/en/release-notes/system-prompts/claude-fable-5-1Release-day consumer-system-prompt documentation confirming Fable 5.1 product identity and availability context. - S62Anthropic - Batch processing
platform.claude.com/docs/en/build-with-claude/batch-processingMessage Batches API mechanics, supported active models, and 50% standard-API price treatment used in the sensitivity analysis.
This publication is provided by Omnesly for informational, educational, and research purposes only and does not constitute legal, financial, technical, regulatory, or other professional advice. Unless otherwise stated, Omnesly owns this publication and all original content contained within it. Properly cited or attributed third-party information, data, trademarks, logos, research, and other materials remain the property of their respective owners. Omnesly makes reasonable efforts to ensure accuracy but does not guarantee that all information is complete, error-free, or current, particularly given the rapidly evolving nature of artificial intelligence technologies, models, products, APIs, pricing, and regulations. References to third-party companies, products, models, or services do not imply endorsement, sponsorship, affiliation, or partnership unless expressly stated. Readers should independently verify material information before relying upon it or making decisions.