Every vendor will price the hardware. Almost none will price the three years after it. Here is where the money actually goes on a private AI cloud, and why the number a board is shown at the start is usually about a third of the real one.
A board asks what it would cost to bring AI workloads in-house. Someone goes to a vendor. The vendor produces a bill of materials, which is an honest answer to the question they were asked, and it becomes the number the board carries around for the next six months.
It is the wrong number, and not because anyone lied. The quote covers the first of seven cost components. The other six are where a private cloud actually costs money, and they arrive gradually enough that nobody has to explain them all at once.
Accelerator racks draw many times what a typical enterprise rack does. That is widely known. What is less widely modelled is that the power bill does not scale with the number of racks you install, it scales with what those racks are doing, and training workloads do not behave like inference workloads at all.
A cost model built on nameplate draw will be wrong in both directions across the year. We model it against the actual demand curve, because the difference between a training-heavy estate and an inference-heavy one at the same rack count can be substantial.
Air cooling stops being viable at a density most organisations reach faster than they expect. Once you are into liquid, you are into a facility change rather than an IT purchase, with a different budget, a different approval path and a much longer lead time.
This is the layer that most programmes discover last and can change least quickly. We walk the floor before anything else for exactly this reason. There is no point designing a system the building cannot carry.
East to west traffic inside a training cluster is a fabric problem, and fabric is expensive in a way that surprises people who have only bought enterprise switching. North to south is a transit problem, and if any part of the workload still touches a public cloud service, that egress is a recurring line nobody put in the model.
Operating systems, virtualisation, orchestration, monitoring, backup, and whatever the AI plane needs. Some of it is open source with a support contract attached, which is not the same as free. Some of it is licensed per socket or per core in ways that interact badly with dense accelerator nodes.
This line is rarely enormous but it is almost always missing, and it recurs annually.
Hardware fails. In a rented environment somebody else's problem, in an owned environment yours, and the lead times on accelerator parts have not been reliable for several years. A spares holding is capital sitting on a shelf, and a refresh cycle is a capital event you should be planning from year one rather than discovering in year three.
The largest omission, and the most predictable. Somebody has to monitor it, patch it, plan capacity, and respond when it breaks outside working hours. Doing that properly needs more than one person, because one person cannot be on call permanently.
This is where the shared operations model earns its place. A single operations centre serving several environments spreads that cost in a way a single organisation staffing its own cannot match. It is also the part of the design that has to be settled early, because it determines what the residency boundary has to allow.
Once all seven components are modelled year by year, the total is only half an answer. The number that lets you compare anything to anything is cost per unit of work, which for AI workloads means cost per million tokens or cost per result.
Almost nobody has that number. When they do, it usually came from a vendor estimate rather than a measurement, and vendor estimates on efficiency are not measurements no matter how carefully they are described. We measure it on a representative workload, on hardware matched to the design, and publish the method so a third party can reproduce it.
That is the figure that lets a board decide honestly between building and staying where they are. It is also the figure that no one selling hardware will put in writing, which tells you something about how often it comes out favourably for them.
Take the vendor quote seriously as a hardware quote and stop treating it as a cost model. Ask what the building can carry before you ask what to buy. And decide early who is going to run it at three in the morning, because that decision shapes the residency boundary, the design and about a third of the three year number.
A sovereign cloud assessment gives you a tenderable bill of materials and a measured run-cost model that survives the board.