Skip to content
Where your AI runs/AI cost calculator

AI cost calculator

Rent or own:
what your AI really costs.

See what renting AI by the token costs your business today, what owning the steady share would cost, and the hardware that would run it.

This calculator estimates what your team's AI use costs when you rent it from cloud providers, and what it would cost to run the same work on hardware you own. Tell it how many people use AI and how heavily, which cloud models you pay for, and how much of the work you'd run yourself, and it works out the hardware you'd need. At Claude Opus 5's published prices, light use (assistants and search) costs about $110 per person a month, moderate use (AI coding and knowledge work through the day) about $1,100, and heavy use (AI agents working on their own) about $3,200. With the default desktop system, owned hardware pays for itself in about 13 months for moderate use and about 17 months for light use. All figures are planning estimates in Australian dollars, excluding GST, not quotes.

Start with a few simple choices. Refine the detail when you’re ready.

Calculate my costs Already have a brief? Talk to a specialist
No sign-up to calculateHardware from any supplierCompares on-site, Australian-controlled and fully offline optionsEvery figure shows whether it’s yours, our estimate or a placeholder
  1. Start with your people Say how many use AI and how heavily. Each usage tile shows what it costs per person per month, so you can check it against a bill you already pay.
  2. Then your models Pick the cloud models you pay for today. Grouping is by published price, never by how good they are.
  3. Then the split Choose how much of the work could run on your own hardware, and in what kind of environment.
  4. Then the machines Pick a model to run privately and a system to run it on. The estimate updates as you go, and every figure says where it came from.
YOUR ESTIMATE · STEP 1 OF 4

How does your team use AI?

Choose a starting profile for your team. Each tile shows what that usage costs per person per month on your current cloud mix, so you can sanity-check it against your bills.

Analysis period
Typical AI usage

Typical usage profiles, not measured averages. Moderate: 5.7M input + 574K output tokens / person / day. Output includes Reasoning tokens. Working-out the model does before answering. You are billed for it as output even though you may never see it. .

Your monthly Token. The unit AI providers bill by. A token is roughly three quarters of a word, so a page of text is about 500 tokens. You pay separately for tokens going in, which is everything you send the model, and tokens coming out, which is what it writes back. Output usually costs several times more than input. estimate4.3B input / 430.5M output
Not sure about your usage? Ask Airon to size it
Next: choose your cloud models

YOUR NEXT STEP / TALK TO A SPECIALIST

Turn your estimate
into a practical AI plan.

You don’t need to know which hardware to buy. Tell us what you want AI to do and an Airon specialist will come back with the models, compute and deployment that fit, starting from the numbers you just built.

01

Find the right fit

Discuss open-weight models and hardware options that suit your workload and budget.

02

Keep control of your data

Explore sovereign, on-premises or air-gapped deployment, with no preferred hardware vendor.

03

Know what to validate

Identify the benchmarks, capacity and operating costs to confirm before investing.

Start with a conversation. No obligation to buy.

Talk to an Airon specialist

Leave your details and we’ll start from your estimate.

YOUR ESTIMATE IS INCLUDED25 people · 50% private AIOn-premises AI · 3-year plan
Add models, budget & timing Optional

No obligation. Your estimate travels with the request, so the conversation starts from your numbers.

Already have planner access? Sign in ↗

How the estimate works

How this estimate is calculated

Cloud spend multiplies each person's daily input and output tokens by the published per-million-token prices of the models you allocate, applies caching, batch, context and contract modifiers, and sums the months in your analysis period. Private cost adds the hardware purchase, setup, power at your electricity rate and PUE, support and any hosting, and credits end-of-life value in the final month. Fleet size is the number of whole units needed to process the routed tokens inside each unit's daily operating window at your usable-capacity share, with a floor set by peak concurrent requests. Planning throughput = memory bandwidth ÷ active weight bytes at the chosen precision (plus 20% for scale tensors), reduced by a √(context ÷ 8K) attention penalty and a √(concurrent requests) sharing penalty, taken at 35% of that ceiling. Prompt processing is an editable multiple of generation. Sustained power is 75% of the configured design load. Prices are catalogue planning allowances plus published upgrade allowances. Memory fit adds full weights at the chosen precision, KV cache for the context and concurrent requests, and a runtime allowance against 20% reserved hardware memory.

Worked example · the calculator’s default scenario, On-premises AI, AUD excluding GST
People using AI25
WorkloadModerate · Agentic coding and knowledge work · 5.7M input / 574K output tokens per person per day
Cloud modelClaude Opus 5 · 20% standard, 10% cache write, 70% cache read
Cloud spend$27,103 per month · $1,084 per person
Private systemNVIDIA DGX Spark running OpenAI gpt-oss-120b at 4-bit · 31 output tokens/s per unit (planning value)
Fleet for 100% private46 units · $299,000 hardware · $2,188 per month to run
Totals over 36 monthsCloud $975,711 · Hybrid at 50% $686,231 · Fully private $382,751
Break-evenFully private 13 months · Hybrid 13 months
Load this worked example into the calculator
How much does AI cost per person per month?+

It depends on token volume and model price. Using Claude Opus 5 list prices with a 20/10/70 standard, cache-write and cache-read split, the calculator's tiers work out to about $110 per person per month for light assistant use (574K input and 57K output tokens per person per day), $1,100 for agentic coding and knowledge work (5.7M input and 574K output tokens per person per day) and $3,200 for autonomous agents and reasoning (16.7M input and 1.7M output tokens per person per day). Change the model mix, caching or contract discount to see your own figure.

When does private AI hardware pay for itself?+

Break-even is the month in which cumulative private spend drops below cumulative cloud spend. With the default desktop system and a moderate workload it is about 13 months; heavier workloads break even sooner and light workloads later. The calculator shows the break-even month for a fully private and a hybrid fleet, and it never reports break-even when hardware inputs are incomplete or capacity is short.

How does the cloud / private mix change costs?+

At 0% private, Hybrid equals Cloud. At 100%, Hybrid equals the fully private fleet. Between these points, auto-size buys enough whole systems for the locally routed workload. Cloud token charges fall as the private share rises; new hardware can cause steps in total spend. A manually configured fleet keeps its entered quantities when private work is enabled.

How are workload and model percentages used?+

Workload percentages create a weighted daily input and output token volume per user. Cloud-model percentages create weighted input and output prices. The calculator multiplies those together with users and active days to estimate monthly cloud API spend. Both allocations should total 100%.

How are private costs and capacity calculated?+

Auto-size uses your configured system and adds input processing time to output generation time within its usable daily inference window. Build a fleet combines the configured systems and checks capacity against the final month's demand. Private cost includes hardware, setup, power, support and selected hosting. Hosted accelerators use your hourly compute rate and powered schedule in place of a purchase price and electricity charge.

Where do the hardware prices and throughput figures come from?+

Catalogue systems carry planning allowances, not vendor quotes, and throughput is derived from published memory bandwidth using a stated formula: Planning throughput = memory bandwidth ÷ active weight bytes at the chosen precision (plus 20% for scale tensors), reduced by a √(context ÷ 8K) attention penalty and a √(concurrent requests) sharing penalty, taken at 35% of that ceiling. Prompt processing is an editable multiple of generation. Sustained power is 75% of the configured design load. Prices are catalogue planning allowances plus published upgrade allowances. Every value is labelled as your figure, a catalogue planning value or a placeholder, and you can replace any of them with a quote or a measured benchmark.

How does Airon stay vendor agnostic?+

Every hardware profile uses the same cost and capacity formulas. Product names and published specifications identify a configuration; they do not establish tokens per second. Add your own quoted prices, measured throughput, power and benchmark context for a fair comparison. Quote-required and roadmap platforms use a stated reference placeholder until you enter a quote, and Google Cloud TPUs are treated as hosted compute rather than an air-gapped appliance.

What needs validation before buying hardware?+

Model licence and quality, accelerator compatibility, model and KV-cache memory, context length, measured throughput, latency, availability, networking and storage. The calculator does not assume that local and cloud models provide equal capability. Financing, tax, depreciation, migration and hardware refresh are excluded. The optional end-of-life value is credited in the final month, not deducted from the upfront purchase.

Sources and dates

LAST REVIEWED 2026-09-18 · REVIEWED BY AIRON SYSTEMS ENGINEERING