INDEPENDENT TOOL / OLLAMA CLOUD

Ollama cost calculator.
Know what you’ll pay.

Turn your AI workload into a monthly bill. Compare pay-as-you-go, Pro, Max and Team—with cache and peak pricing accounted for.

Source-checked pricing2026-09-05A dated snapshot, not live rates
01 / YOUR WORKLOAD
Calculation mode

Synthetic example · Replace with your own assumptions.

02 / THE CASH VIEWUSD / month

Your estimate starts here.

JavaScript runs the calculation in your browser. No prompt, API key or account connection is needed.

Monthly plans only. Before tax. Free starter credits, annual billing and legacy quotas excluded.

Already subscribed? Estimate your remaining-credit runway +

This is a separate, hypothetical remaining-period estimate. It does not read your account.

01

Consumption ≠ payment

See model usage value separately from the subscription fee and overage you pay.

02

One workload. Four plans.

Compare the same usage across plans. Ties stay ties; quality and speed are not ranked.

03

Private by design

Calculations stay in your browser. No account, prompt upload or API key required.

THE INPUTS BEHIND THE ANSWER

Rates, in the open.

Check official pricing ↗

USD per 1 million tokens. Ollama-hosted rates, not another provider’s API prices. Verified 2026-09-05.

ModelInputCached inputOutputPeak pricing
DeepSeek V4 Flash$0.22$0.007$0.662× all categories
DeepSeek V4 Pro$0.66$0.022$1.982× all categories
GLM 5.3$1.4$0.26$4.4No published surcharge
MiniMax M3$0.6$0.12$2.4No published surcharge
GPT-OSS 120B$0.15$0.014$0.6No published surcharge
Gemma 4 Cloud$0.14$0.05$0.4No published surcharge

DeepSeek peak window: weekdays 12:00–18:00 UTC. The calculator weights your share of tokens, not hours. Rates are a dated snapshot, not a live feed.

UNDERSTAND THE NUMBERS

A few minutes of reading.
A more useful estimate.

One workload. Four different cash outcomes.

A model price tells you the cost of consumption, not necessarily the amount you pay for a subscription month. This calculator starts with an Ollama-hosted workload, values fresh input, cached input and output, then compares PAYG, Pro, Max and Team against that same consumption value. The result separates the fixed fee, covered usage, unused allowance and overage so you can see why a plan has a lower estimated bill.

Choose workload mode when you have token assumptions. Choose direct consumption mode when you already have a dollar-valued usage estimate for one monthly billing period on the same current Ollama rate basis. Neither mode connects to your account. Sample values are there to make the tool inspectable; they are not a prediction of what a typical developer will use. Replace them with assumptions for that single billing month, not an annual total.

Start with the provider, not just the model name

The billing provider matters. A model running locally through Ollama is not the same billing scenario as a hosted Ollama request. A coding client can also connect to different backends, so the name on the application's window does not establish who charges for the call. The Ollama cloud documentation explains the local and hosted distinction.

This first version deliberately covers one provider's current monthly credit plans and six selected hosted model presets. It does not compare all AI subscriptions, native provider APIs, local electricity costs or model quality. A dropdown keeps the calculation in one useful place without pretending every model name requires a separate website. Confirm the host and variant before copying rates into a decision.

Put the right quantities into the workload

Total input tokens means all input processed for an average request, including the portion served from cache. The cache percentage splits that total into fresh and cached parts. Output tokens are generated separately. Requests per active day and active days within one monthly billing period determine the total number of requests. Enter one to thirty-one active days, without crossing a credit reset. Fewer active days do not prorate the monthly fee or allowance. Your actual billing month need not begin on the first calendar day.

A session is not a request count. Agent loops, follow-up calls, retries and background operations may add requests. Equally, a model's context-window capacity is not a claim that every call fills that window. Use representative usage information where available. If your sample does not contain a cached-input field, a zero-cache scenario is an assumption, not an observed fact about the provider.

A worked example you can reproduce

Consider a synthetic GLM 5.3 scenario with 10,000 input tokens per request, 80% cached input, 1,000 output tokens and 100 daily requests over thirty days. That is 3,000 requests, containing 6 million fresh input tokens, 24 million cached input tokens and 3 million output tokens. Applying the checked category rates produces $8.40, $6.24 and $13.20 respectively: $27.84 of consumption in total.

Under the current monthly plan assumptions, that consumption fits inside Pro's included usage, giving a modeled $20 cash cost. PAYG consumption expense is $27.84 before any starter allowance. The example is not a recommendation to subscribe: different workload assumptions, account benefits or access requirements could change the decision. Its purpose is to let you follow every step from tokens to consumption to cash.

The lowest price can include overage

A plan with no overage is not always the cheapest plan. For a synthetic $100 consumption estimate, Pro costs $20 plus $40 overage, or $60. Max costs its $100 fee because the consumption fits inside its larger allowance. Paying some overage on Pro is cheaper in this case than buying unused headroom on Max. At a different consumption level, the relationship reverses.

The derived cash boundaries are $20 between PAYG and Pro, $140 between Pro and Max, and $700 between Max and Team. At each boundary, both plans are shown as tied. The tool compares unrounded values internally, so display rounding should not silently decide the winner. Read the Pro versus Max guide for the complete comparison, including where PAYG and Team belong.

Peak share is the share of work at peak rates

For the listed DeepSeek models, the checked official pricing table distinguishes base and peak prices. The published peak window is Monday through Friday, 12:00 to 18:00 UTC. The form's peak percentage represents a share of token volume, not a share of clock hours. If all your batch work happens outside the window, six peak hours in a day do not imply peak consumption.

The simplified calculation assumes the same token-category and cache mix during peak and off-peak work. If those groups differ substantially, run separate scenarios and combine their consumption values before comparing plans. A model without a separate published peak rate is unaffected by the peak slider. The cache and peak guide walks through the field semantics and rate blending.

Use the remaining-credit view for a specific window

A monthly plan comparison answers which modeled plan has the lowest cash cost for a consumption estimate. A remaining-credit forecast answers a different question: how long a stated remaining allowance lasts at an assumed pace before the next reset. Enter the allowance, used amount, expected daily consumption and actual days until reset. The calculator does not discover those values from your account.

For example, a synthetic $40 remaining allowance at $8 a day lasts five days at that pace. If reset is in seven days, projected consumption before reset exceeds the allowance by $16. If reset is in three days, it does not. This forecast is not a promise that service will stop on a particular date; prepaid credit and account behavior are not modeled. Read credits versus cash before treating an allowance as a bill or bank balance.

What the result deliberately leaves out

The figures are in US dollars before taxes and exchange-rate effects. They exclude unknown starter credit, annual purchase cash flow, legacy quota conversions, top-up purchase timing, promotions, mid-cycle changes and separate software or infrastructure costs. They also do not score model quality, response latency, reliability or operational fit. Team is an early-access shared plan, so a low modeled price does not establish that it is available or suitable for you.

Concurrency is another separate constraint: the checked table lists one simultaneous request for Free/PAYG, three for Pro, and ten for Max and Team. A sequential personal workflow and an application serving several people may therefore make different choices at the same monthly consumption. Price is useful evidence, but it is not a replacement for those requirements or for the provider's current account terms.

Inspect, vary, then verify

A sensible budget is a range of defensible scenarios. Change one assumption at a time: requests, input size, output length, cache reuse or peak share. See whether the preferred plan stays the same. That is more informative than presenting one precise number built on an uncertain average. Shared scenarios should be treated as records of assumptions, not proof that the workload will behave that way.

The calculator runs in your browser and does not need prompts, credentials or paid inference calls. Share links use numeric state in a URL fragment; anyone receiving that link can inspect it. See the privacy explanation before sharing. Rates carry a verification date and are not streamed live. Check the official source before purchasing, inspect the methodology and exclusions, and return to the calculator whenever a workload or pricing rule changes.