There is a version of cloud optimisation that is politically easy. You run a report, identify some orphaned snapshots and idle IP addresses, delete them, and declare a small win. Finance is satisfied. The number moves slightly. Everyone moves on.

And then next quarter, the bill is back where it was, because the actual problem was never addressed.

The actual problem, in most cloud estates, is right-sizing. Matching the resources you are running to the work those resources actually do. It sounds obvious. In practice, it is the conversation that organisations consistently avoid, because it requires someone to tell an engineering team that the infrastructure they provisioned is significantly over-specified, and engineering teams do not always respond well to that.

Here is the reality. A workload running at an average CPU utilisation of 15% on a large instance is not a sign of careful capacity planning. It is a sign that the instance was provisioned conservatively, the expected demand never materialised, and nobody has revisited the specification since. This is not unusual. It is the default outcome when there is no active process for reviewing utilisation against cost.

The numbers bear this out. Right-sizing compute alone typically delivers a 20-30% reduction in spend for organisations that do it properly. For a business running a meaningful cloud estate, that is not a rounding error. It is a material number that could fund headcount, product development, or simply improve margins.

So why does it not happen more consistently? A few reasons.

The first is that right-sizing production workloads requires care and testing, and engineering teams are already busy. Downsizing an instance on a customer-facing service without proper staging and rollback planning is a genuine risk. The caution is reasonable. The problem is that reasonable caution often becomes indefinite deferral.

The second is that over-provisioning is comfortable. A larger instance means headroom. It means fewer incidents caused by resource constraints. The cost of that headroom is invisible to the people provisioning it, especially when there is no showback mechanism connecting their infrastructure choices to a budget they care about.

The third is that right-sizing recommendations from cloud providers are not always trustworthy. They tend to focus on CPU utilisation and miss memory constraints entirely. Acting on a provider recommendation without checking both can create a new problem in its place.

None of these is insurmountable. The right approach is to start with utilisation data over a representative period of at least 30 days, check both CPU and memory before drawing any conclusions, and have a direct conversation with the team that owns the workload before making changes. Not as bureaucracy, but because the people closest to the workload will tell you things the metrics will not.

The organisations doing this well have made it a quarterly practice rather than a one-off exercise. The first pass is the hard one. After that, you are maintaining a baseline rather than excavating one.

The savings are there. They have been there for a while. The question is whether the conversation will happen this quarter or next year.