Your AWS Bill Is Too High. Here's Where the Money Usually Goes

A practical, no-hype walkthrough of where AWS bills actually leak — oversized instances, NAT Gateways, idle dev environments, forgotten storage — and how to reduce AWS costs without breaking your stack.

By

UpNext Software — AI, ML & Python Engineering

Your AWS Bill Is Too High. Here's Where the Money Usually Goes

Almost every AWS cost conversation I've had starts the same way: someone forwards me a bill, adds one line of context — "this keeps going up and nobody knows why" — and waits.

And nearly every time, the money is hiding in the same five or six places. Not in exotic services. Not in some clever architectural flaw. In boring, avoidable stuff that nobody owns.

So this post is the walkthrough I usually give. If you want to reduce AWS costs without rewriting your application or hiring a FinOps team, start here.

First, read the bill properly

Most people look at the AWS console summary, see "EC2 — $4,200", and stop. That number tells you nothing actionable. EC2 as a line item bundles compute, EBS volumes, snapshots, Elastic IPs, NAT Gateways and data transfer. They have completely different fixes.

Open Cost Explorer and do this instead:

  • Set granularity to daily, not monthly. Spikes and step-changes are far more informative than totals.
  • Group by Usage Type, not by Service. This is the single most useful change you can make. Suddenly "EC2" becomes "NatGateway-Bytes", "BoxUsage:m5.2xlarge", "EBS:VolumeUsage.gp2" — and you can act.
  • Look at the last 6 months, not the last 30 days. You're looking for the month things changed.

Ninety percent of the time, 80% of the waste sits in five or six usage types. Find those first. Ignore the long tail — chasing $12/month line items is how cost projects die.

Where the money usually goes

1. Instances that are two or three sizes too big

The most common pattern I see: someone provisioned for a launch that never got that traffic, or copied production sizing into staging, and nobody revisited it. CPU sits at 4%. Memory sits at 20%.

Check CloudWatch for average and p95 CPU over the last month. If an instance never crosses 20% CPU and isn't memory-bound, it's a candidate for downsizing. Each halving of instance size halves that line item — this is the highest-leverage change available to most teams, and it takes an afternoon.

One caution: CloudWatch doesn't report memory by default. If you downsize based on CPU alone and the workload is actually memory-hungry, you'll find out the hard way. Install the CloudWatch agent or check with your APM tool first.

2. NAT Gateways (the quiet killer)

This is the one that surprises people most. A NAT Gateway costs roughly $0.045 per hour just to exist — about $32–35/month each — plus a per-GB data processing charge in the same ballpark. Run one per availability zone across three environments and the base cost alone is meaningful before a single byte moves.

The real damage is the data processing. If your private-subnet services pull container images from a public registry, hit S3 over the internet, or ship logs to a third party, every gigabyte is billed twice — once for processing, once for transfer.

Fixes, in order of effort:

  • Add VPC Gateway Endpoints for S3 and DynamoDB. They're free. If your app talks to S3 from a private subnet, this alone can cut a noticeable chunk of NAT traffic.
  • Consolidate NAT Gateways in non-production. Multi-AZ NAT redundancy for a staging environment is paying for an insurance policy nobody needs.
  • Cache container images in ECR rather than pulling from Docker Hub on every deploy.

3. Dev and staging running 24/7

Your team works maybe 45 hours a week. Your development environment runs 168. That means roughly 70% of its cost is pure waste.

A simple scheduler — EventBridge plus a Lambda, or AWS Instance Scheduler — that stops non-production EC2 and RDS instances at 8pm and starts them at 9am on weekdays is one of the best returns on a day of engineering you will ever get. Just make sure someone in a different timezone isn't relying on it, which is a real problem when your team is in Surat and your client is in Texas. Build the schedule around actual working hours, not office hours.

4. Storage nobody remembers creating

Terminate an EC2 instance and, depending on how it was configured, the EBS volume can survive it. Those orphans bill forever. So do:

  • Unattached EBS volumes — check for volumes in the "available" state.
  • Ancient snapshots — automated backups with no lifecycle policy accumulate for years.
  • gp2 volumes — gp3 is generally cheaper and faster. Migrating is a live operation on modern volumes. This is close to free money.
  • S3 with no lifecycle rules — logs, exports and old uploads sitting in Standard forever. Enable S3 Intelligent-Tiering or write lifecycle rules to move objects to Infrequent Access or Glacier.

5. Data transfer out

Data leaving AWS to the internet is billed per GB, and cross-AZ traffic between your own services is billed too. If you're serving images, video or large API responses directly from EC2 or S3, putting CloudFront in front usually reduces the bill, because CloudFront's egress rates are lower and it absorbs repeat requests.

Also worth checking: chatty microservices spread across availability zones. Every internal call crossing an AZ boundary has a cost. It's small per call and large per million.

6. CloudWatch Logs and observability

Log groups default to never expire. I have seen teams paying for four years of debug logs from a service that was decommissioned two years ago. Set retention on every log group — 14 or 30 days is plenty for most applications, with anything you actually need for compliance archived to S3.

Also audit your log volume. Debug-level logging left on in production is an expensive habit.

7. The managed-service convenience premium

Managed services are usually worth paying for. But some are priced for scale you don't have. A small OpenSearch cluster, an idle MSK cluster, a Redis instance used for one cache key, an RDS instance for a database with 200 rows — each of these might be replaceable with something dramatically cheaper.

Be honest here. If a managed service saves your team a day a month of operational pain, it's earning its keep. If it exists because someone added it in a sprint two years ago, kill it.

Commitments: only after you've cleaned up

Savings Plans and Reserved Instances can cut compute costs substantially — often in the 20–40% range for a one-year commitment, more for three years. But do the rightsizing first. Committing to instance sizes you're about to shrink locks in your waste for a year.

Two other levers worth knowing:

  • Graviton (ARM) instances are typically meaningfully cheaper than equivalent x86 for the same performance. If you're running containerised Python, Node or Java, the migration is often just a rebuild of your images. We've moved workloads across with less drama than expected.
  • Spot instances are dramatically cheaper but can be reclaimed with two minutes' notice. Perfect for CI runners, batch jobs, ML training and stateless workers. Not for your primary database.

When your AWS bill is actually fine

I'll say the unpopular thing: sometimes the answer is to leave it alone.

If you're spending $600 a month on AWS and paying engineers well, spending two weeks to save $150/month is a bad trade. That's a real cost with a payback period measured in years, and the engineering attention would have been worth more somewhere else.

My rough rule: below about $1,000–1,500/month, do the one-afternoon cleanup — retention policies, orphaned volumes, dev schedulers, gp3 — and then get back to building product. Deep cost engineering starts paying for itself when the bill is a few thousand dollars a month or growing faster than revenue.

Also: don't optimise a bill that's about to change shape. If you're mid-migration or about to launch something that triples traffic, wait. You'll be optimising for a system that no longer exists.

Make it stay fixed

Cleanups regress. Within six months the same waste reappears unless something is watching. The lightweight version of governance:

  • AWS Budgets with alerts at a few thresholds, going to a channel people actually read.
  • Cost anomaly detection enabled — it's free and catches the "someone left a GPU instance running" class of problem within a day or two.
  • Mandatory tags for environment and service, so next quarter's review takes an hour instead of a week.
  • A 15-minute monthly look at the daily cost graph. That's it. Step-changes are obvious when you check regularly.

We run this discipline on our own infrastructure too, including the environments behind Orbis Lead CRM — partly to keep our margins honest, and partly because you learn a lot more about cost when it's your own money.

If you want a second pair of eyes

We do cloud cost reviews as part of our DevOps work: read the bill properly, rank the fixes by saving-per-hour-of-effort, and tell you plainly which ones aren't worth doing. Often the report ends with "do these four things, ignore the rest." You can see the kind of systems we work on in our work, and if AI or data workloads are driving your spend, that overlaps heavily with our AI & ML development practice.

If your bill is climbing and nobody on the team can explain why, get in touch — send over a Cost Explorer screenshot and we'll tell you what we'd look at first, no strings attached.