Lesson 4 of 5 · 10 min

By James Durkin, JDCS · Updated 6 August 2026

What it costs, honestly.

This is the lesson where a lot of private AI marketing quietly stops being honest. So let's do the arithmetic in the open, including the parts that argue against buying anything. If you take one number away, make it this one: below roughly 30 people, running AI on your own hardware generally does not pay for itself on cost alone. That doesn't kill the idea. It changes what you'd be buying it for.

The hardware, at Australian prices

Real prices from Australian retailers, as at August 2026, so you can see the actual shape of the spend rather than a US list price converted in someone's head.

  • Mac Studio M4 Max, 36GB: $3,499. A sensible one-person or small-team machine.
  • Mac Studio M3 Ultra, 96GB: $6,999. Enough unified memory to hold a genuinely capable model.
  • ASUS Ascent GX10, 128GB: $6,249. A compact desktop AI box.
  • NVIDIA RTX 5090, 32GB: about $6,499. Goes into a workstation you build around it.
  • NVIDIA RTX PRO 6000 Blackwell, 96GB: $18,499. The serious end for a small office serving a team.

Timing matters more than usual right now. Memory prices have spiked: DRAM rose by up to 98% in the first quarter of 2026, with further rises expected. The RTX 5090 in Australia sits roughly 61% above its $4,039 launch price. High-memory Mac Studios have been effectively unavailable. If your situation lets you defer six to twelve months, deferring is a legitimate strategy rather than procrastination, and renting capacity in the meantime is a reasonable bridge.

One thing you can stop worrying about: power. A Mac Studio M4 Max running continuously costs roughly $273 a year in electricity at a 30 cents per kilowatt-hour planning rate. Even the heavier machines land in the hundreds to low thousands. Against several thousand dollars a year of API or per-seat spend, the power bill is noise. People fixate on it because it's the cost they can picture, and it is almost never the one that decides anything.

The seat count where it turns over

The comparison that matters runs at three team sizes, against both a per-seat SaaS subscription and pay-as-you-go API usage, with the on-premise figure including the hardware amortised and a realistic allowance for administration.

  • 10 people. SaaS around $5,400 a year, API around $4,435, on-premise around $7,700. On-premise loses, and not narrowly.
  • 30 people. SaaS around $16,200, API around $13,306, on-premise around $14,050. Roughly a tie, and what tips it is how many hours of administration the box actually consumes.
  • 100 people. SaaS around $54,000, API around $44,352, on-premise around $28,757. On-premise wins clearly, with payback in something like 1.6 to 2.6 years.

So if a supplier is pitching a ten-seat business a server on the promise of savings, the arithmetic isn't there and you should say so. There are three arguments that do hold at ten seats, and they have nothing to do with cost per seat: confidentiality that the cloud cannot give you at any price, unmetered use, and a bill that stays where you put it. Buy it for those reasons if they apply. Don't let anyone dress them up as a saving.

There's a subtler point about unmetered use that the spreadsheet misses. Teams self-censor under per-token billing. People skip the second draft and the exploratory question because they can see the meter running. When the meter is off, usage often climbs several times over, and the value of that extra use can exceed the arithmetic in either direction. It's real, it's hard to quantify, and it deserves a sentence rather than a business case.

The costs everyone forgets

Hardware is the easy number. These are the ones that turn a good decision into a bad one when they get left out of the sums.

  • Maintenance: 5 to 10 hours a month. This is the number consultants omit and clients discover. Model updates, patching, driver and runtime CVEs, the occasional upgrade that quietly changes output quality. At consulting rates that can exceed the hardware amortisation. Either it's a retainer or it's a named person in-house, and either way it belongs in the proposal in writing.
  • A UPS, somewhere between $400 and $2,500 depending on what you're protecting.
  • Circuit capacity. A 1,400 watt draw will trip a 10 amp office circuit. Worth knowing before delivery day, not after.
  • Backups of the right things. Model weights are re-downloadable and don't need protecting. Your fine-tunes, embeddings, vector databases and prompt library are not re-downloadable. Those are the asset, and they're the ones people forget to back up.
  • Noise and heat. Anything drawing over 400 watts continuously does not belong where people take client calls. At 1,400 watts you will beat the air conditioning in an Australian summer.
  • Who fixes it at 2am. Never let the box be the only path to working AI. A cloud fallback keeps the business running on the bad day, and it barely dents the savings.
  • Idle time. Hardware costs the same whether it's busy or not. A box at 8% utilisation still depreciates at full speed, while a cloud bill goes to zero when nobody is using it.

The other side of the ledger: what the cloud has already done

None of the above means cloud pricing is a safe harbour. The useful thing here is that nobody has to predict a price rise in order to plan for one. The record is already published, and it points in both directions. All the API prices below are in US dollars, so an Australian buyer carries currency movement on top of every one of them.

  • OpenAI's flagship API price went up. From US$1.25 per million input tokens and US$10 per million output in August 2025, to US$5 and US$30 by April 2026. Four times the input price in under a year.
  • Anthropic published the date its mid-tier gets dearer. Claude Sonnet 5's introductory US$2 and US$10 rate runs through 31 August 2026, and the standard US$3 and US$15 applies from 1 September 2026. At least that one is on the calendar. Most are not.
  • A price rise that isn't a price rise. Anthropic's own pricing page notes that Claude 4.7 and later models use a newer tokenizer which "produces approximately 30% more tokens for the same text". The rate per token did not move. The bill did.
  • Seats move too, and it happened here. In January 2025 Microsoft bundled Copilot into consumer Microsoft 365 and raised Australian prices: Personal from $109 to $159 a year, a 45% rise, and Family from $139 to $179, a 30% rise. The ACCC brought proceedings alleging Microsoft misled users by not disclosing a third option, the Classic plans, which kept the existing features without Copilot at the old price. Microsoft acknowledged poor communication and offered affected subscribers an eight-week refund window.
  • Prices fall as well, and pretending otherwise is dishonest. Anthropic's top tier dropped from US$15 and US$75 to US$5 and US$25. Cached input runs at 10% of the standard rate at both OpenAI and Anthropic, and batch processing is half price across the major providers.

The pattern worth taking away is not that vendors are villains. It's that the price you signed up on is a commercial decision made under competitive pressure, and competitive pressure is not a contract. That cuts both ways: your bill can fall without you doing anything, and it can rise the same way. Owning the hardware doesn't make AI cheaper for a small team. It makes one line of the budget stop moving, and it takes the decision about that line back off someone else's roadmap.

The bottom line: at 10 seats on-premise loses on cost, at around 30 it's a tie decided by how much administration actually happens, and at 100 it wins clearly. Australian hardware runs from $3,499 for a Mac Studio M4 Max to $18,499 for an RTX PRO 6000, with memory prices spiked and deferral a legitimate option. Power is noise; the 5 to 10 hours a month of maintenance is not. On the other side, OpenAI's flagship input price went up fourfold between August 2025 and April 2026, Anthropic's mid-tier rises on 1 September 2026, a change in how text is counted added about 30% more billable tokens for the same document, and the ACCC has taken Microsoft to court over an Australian price rise with Copilot bundled in. Nobody has to predict a price rise to plan for one. Next up: the hybrid setup most businesses should actually land on.
Quick check

A few quick questions to lock it in. No marks recorded, just for you.

Q1.At roughly what team size does running AI on your own hardware stop losing on cost?

At 10 seats on-premise loses clearly. At 100 it wins clearly, with payback in about 1.6 to 2.6 years. Below 30, buy it for confidentiality and unmetered use rather than savings.

Q2.Which running cost do people most often leave out of the sums?

Power is noise, roughly $273 a year for a Mac Studio M4 Max. Maintenance at consulting rates can exceed the hardware amortisation, so name it in the proposal.

Q3.Anthropic's pricing page notes its newer tokenizer produces approximately how many more tokens for the same text?

The rate per token didn't move. The bill did. It's the clearest example of a cost rise that never appeared as a price rise.

Pick up anywhere

Save your progress

Pop your email in and we'll send you a link to pick up where you left off, on any device. No account needed.

Just for the link to your progress. No spam, and I never share your details.