Find the right fit
Discuss open-weight models and hardware options that suit your workload and budget.
AI cost calculator
See what renting AI by the token costs your business today, what owning the steady share would cost, and the hardware that would run it.
This calculator estimates what your team's AI use costs when you rent it from cloud providers, and what it would cost to run the same work on hardware you own. Tell it how many people use AI and how heavily, which cloud models you pay for, and how much of the work you'd run yourself, and it works out the hardware you'd need. At Claude Opus 5's published prices, light use (assistants and search) costs about $110 per person a month, moderate use (AI coding and knowledge work through the day) about $1,100, and heavy use (AI agents working on their own) about $3,200. With the default desktop system, owned hardware pays for itself in about 13 months for moderate use and about 17 months for light use. All figures are planning estimates in Australian dollars, excluding GST, not quotes.
Start with a few simple choices. Refine the detail when you’re ready.
Calculate my costs Already have a brief? Talk to a specialistChoose a starting profile for your team. Each tile shows what that usage costs per person per month on your current cloud mix, so you can sanity-check it against your bills.
Typical usage profiles, not measured averages. Moderate: 5.7M input + 574K output tokens / person / day. Output includes Reasoning tokens. Working-out the model does before answering. You are billed for it as output even though you may never see it. .
100% allocated — all users assigned.
Pick one model for a quick comparison, or split usage across your model mix. Models are grouped by published price band, not by capability.
Grouped by published price band · lowest input price first, then output price.
100% allocated — all users assigned.
Cloud and local models can have different capabilities. This compares the cost of token workloads, not model quality. Prompt caching. A discount for re-sending the same context. The provider keeps a copy for a few minutes, so repeat sends cost a fraction of the first one. It matters because coding and agent tools send the same project files on every step. This estimate assumes a conservative split: a fifth of input at full price, a tenth writing to the cache, and the rest reading from it. is on by default with a 20 / 10 / 70 standard, cache-write and cache-read split; change it under advanced assumptions.
Model pricing current as at18 September 2026
Check the labs’ published rates and apply them to this estimate. Prices may change.Move the slider to explore the trade-off between cloud token charges and owning your infrastructure. This is the Hybrid. Running some AI work on your own hardware and the rest through cloud providers. Move the slider to choose the split. Hardware is bought in whole machines, so the cost can step up rather than rise smoothly. split.
Run models at your office or site, close to your people and data.
Running OpenAI gpt-oss-120b at 4-bit. 46 units for a fully private fleet, 23 at 50% private. Change the model, system and options in the next step.
Choose an open-weight model and a system. Fit, unit count, price and power update as you go. Every figure says whether it is yours, a catalogue planning value or a placeholder.
What these mean: Open-weight model. A model you can download and run on your own hardware, rather than renting it by the token from a provider. Running one privately means no per-token bill and no data leaving your network, in exchange for buying and operating the machines. Capability is not the same as the cloud models, which is why this tool compares cost rather than quality. · Weight precision. How finely the model's numbers are stored. Lower precision makes the model smaller and faster, with some loss of accuracy. Four-bit roughly quarters the memory needed against the sixteen-bit original. It is the usual way a large model is made to fit on affordable hardware, and it needs validating for your workload before you buy. · Context. How much text the model can hold in mind at once, counted in tokens. Longer context needs more memory for every simultaneous request and slows generation, which is why it changes both the fit and the unit count here.
Compared with the same rules for every vendor. Memory fit. Whether the model actually fits in a machine's memory, with room for the conversations it is handling. If it does not fit, the machine cannot run that model at all, at any speed. That is why a system that does not fit is greyed out rather than shown as slow. is checked at 32.8K context with 4 concurrent requests per unit and 20% of memory held in reserve.
Enter your system quote and throughput measured with the model, quantisation and context you plan to run. A$6,500 is a planning allowance, not a vendor quote. Native ARM runtime and quantisation kernels need validation. Vendor details ↗ (opens in a new tab)
GPUs per system: 1 · Host / unified memory: 128 GB · Storage: 4 TB
Throughput. How fast a machine reads and writes tokens, in tokens per second. It decides how many machines your workload needs. Prompt processing is reading what you sent; generation is writing the reply, and it is the slower of the two. The figures here are derived from published memory bandwidth, not measured, until you enter your own benchmark. figures follow the stated bandwidth method and prices are catalogue allowances, both shown as a Planning value. A figure we derived from published specifications so the estimate can be completed. It is not a price anyone has quoted you. Anything you type replaces it and is labelled as your figure. Get a supplier quote and a benchmark on your own workload before making a purchase decision.. Enter a measured benchmark or a quote to replace any of them.
A day's tokens need ≈ 234 unit-hours (106 prompt processing + 128 generation) in the final month; 46 units × 8 hours at 65% usable capacity provide ≈ 239.
Memory: Fits 71 GB weights + 5 GB context cache + 8 GB runtime of 103 GB usable.
NVIDIA DGX Spark: price, power, throughput from catalogue planning values.
Mixing systems or modelling hosted compute?
See my cost comparisonYOUR NEXT STEP / TALK TO A SPECIALIST
You don’t need to know which hardware to buy. Tell us what you want AI to do and an Airon specialist will come back with the models, compute and deployment that fit, starting from the numbers you just built.
Discuss open-weight models and hardware options that suit your workload and budget.
Explore sovereign, on-premises or air-gapped deployment, with no preferred hardware vendor.
Identify the benchmarks, capacity and operating costs to confirm before investing.
Start with a conversation. No obligation to buy.
How the estimate works
Cloud spend multiplies each person's daily input and output tokens by the published per-million-token prices of the models you allocate, applies caching, batch, context and contract modifiers, and sums the months in your analysis period. Private cost adds the hardware purchase, setup, power at your electricity rate and PUE, support and any hosting, and credits end-of-life value in the final month. Fleet size is the number of whole units needed to process the routed tokens inside each unit's daily operating window at your usable-capacity share, with a floor set by peak concurrent requests. Planning throughput = memory bandwidth ÷ active weight bytes at the chosen precision (plus 20% for scale tensors), reduced by a √(context ÷ 8K) attention penalty and a √(concurrent requests) sharing penalty, taken at 35% of that ceiling. Prompt processing is an editable multiple of generation. Sustained power is 75% of the configured design load. Prices are catalogue planning allowances plus published upgrade allowances. Memory fit adds full weights at the chosen precision, KV cache for the context and concurrent requests, and a runtime allowance against 20% reserved hardware memory.
| People using AI | 25 |
|---|---|
| Workload | Moderate · Agentic coding and knowledge work · 5.7M input / 574K output tokens per person per day |
| Cloud model | Claude Opus 5 · 20% standard, 10% cache write, 70% cache read |
| Cloud spend | $27,103 per month · $1,084 per person |
| Private system | NVIDIA DGX Spark running OpenAI gpt-oss-120b at 4-bit · 31 output tokens/s per unit (planning value) |
| Fleet for 100% private | 46 units · $299,000 hardware · $2,188 per month to run |
| Totals over 36 months | Cloud $975,711 · Hybrid at 50% $686,231 · Fully private $382,751 |
| Break-even | Fully private 13 months · Hybrid 13 months |
It depends on token volume and model price. Using Claude Opus 5 list prices with a 20/10/70 standard, cache-write and cache-read split, the calculator's tiers work out to about $110 per person per month for light assistant use (574K input and 57K output tokens per person per day), $1,100 for agentic coding and knowledge work (5.7M input and 574K output tokens per person per day) and $3,200 for autonomous agents and reasoning (16.7M input and 1.7M output tokens per person per day). Change the model mix, caching or contract discount to see your own figure.
Break-even is the month in which cumulative private spend drops below cumulative cloud spend. With the default desktop system and a moderate workload it is about 13 months; heavier workloads break even sooner and light workloads later. The calculator shows the break-even month for a fully private and a hybrid fleet, and it never reports break-even when hardware inputs are incomplete or capacity is short.
At 0% private, Hybrid equals Cloud. At 100%, Hybrid equals the fully private fleet. Between these points, auto-size buys enough whole systems for the locally routed workload. Cloud token charges fall as the private share rises; new hardware can cause steps in total spend. A manually configured fleet keeps its entered quantities when private work is enabled.
Workload percentages create a weighted daily input and output token volume per user. Cloud-model percentages create weighted input and output prices. The calculator multiplies those together with users and active days to estimate monthly cloud API spend. Both allocations should total 100%.
Auto-size uses your configured system and adds input processing time to output generation time within its usable daily inference window. Build a fleet combines the configured systems and checks capacity against the final month's demand. Private cost includes hardware, setup, power, support and selected hosting. Hosted accelerators use your hourly compute rate and powered schedule in place of a purchase price and electricity charge.
Catalogue systems carry planning allowances, not vendor quotes, and throughput is derived from published memory bandwidth using a stated formula: Planning throughput = memory bandwidth ÷ active weight bytes at the chosen precision (plus 20% for scale tensors), reduced by a √(context ÷ 8K) attention penalty and a √(concurrent requests) sharing penalty, taken at 35% of that ceiling. Prompt processing is an editable multiple of generation. Sustained power is 75% of the configured design load. Prices are catalogue planning allowances plus published upgrade allowances. Every value is labelled as your figure, a catalogue planning value or a placeholder, and you can replace any of them with a quote or a measured benchmark.
Every hardware profile uses the same cost and capacity formulas. Product names and published specifications identify a configuration; they do not establish tokens per second. Add your own quoted prices, measured throughput, power and benchmark context for a fair comparison. Quote-required and roadmap platforms use a stated reference placeholder until you enter a quote, and Google Cloud TPUs are treated as hosted compute rather than an air-gapped appliance.
Model licence and quality, accelerator compatibility, model and KV-cache memory, context length, measured throughput, latency, availability, networking and storage. The calculator does not assume that local and cloud models provide equal capability. Financing, tax, depreciation, migration and hardware refresh are excluded. The optional end-of-life value is credited in the final month, not deducted from the upfront purchase.
LAST REVIEWED 2026-09-18 · REVIEWED BY AIRON SYSTEMS ENGINEERING