Taming The AWS Bill

You know what's more fun than building things? Finding out what they actually cost to run.

This blog is a modest operation. Thirty-two markdown files. About 150KB of content, total. And for months it's been sitting on top of an AWS stack that costs $85-100 a month. That's not a hosting bill. That's a bar tab in Tokyo.

We decided to put the blog on a diet. Here's how it went from ~$90/month to under $40 — without losing a single feature, a single view count, or a minute of uptime.

The Audit

First rule of cost optimization: look at the actual bill before touching anything. Cost Explorer had the receipts:

Service Monthly What it actually was
EC2 - Other $33.48 A NAT Gateway, billed by the hour whether it does anything or not
Elastic Load Balancing $16.75 An ALB, also flat-rate
RDS $14.20 The database (legitimately doing its job)
VPC $11.17 Three Elastic IPs at $3.65 each
ECS $7.35 A Fargate task running 24/7
Route 53 $2.02 Four hosted zones

Total: ~$85-100/month. But the money line was the usage data. CloudWatch told a brutal story:

  • Fargate CPU: 0% average. 0% maximum. For fourteen days straight. We were paying 24/7 for a server that did nothing.
  • The publishing Lambda: zero invocations in 60 days. Posts go up about four times a year. The pipeline was very well rested.
  • The database: a steady 2 connections. Basically just the web server's pool and health checks saying hello.

And meanwhile the blog itself? ~100,000 page views a month of the same 150KB of content, rendered fresh on every single request.

The Options

Three paths presented themselves:

Option A — Go full static. Render the markdown to HTML at publish time, serve it from S3 + CloudFront like the game sites. Cost: ~$4/month. Downside: view counts. We like the view counts. Watching numbers tick up is half the fun of writing.

Option B — Serverless the web tier. The Go web server becomes a Lambda function behind CloudFront. Keep RDS, keep view counts, delete everything else. Cost: ~$36-40/month.

Option C — Trim the fat. Keep the architecture, resize things. Cost: still ~$70-75/month, because the fixed floor barely moves.

B was the winner. The blog stays dynamic, the publishing flow stays "drop a file in S3," and the view counts stay exact — not approximated, not eventually-consistent. Exact.

The New Architecture

BEFORE                                          AFTER
Route53 → ALB → Fargate (24/7) → RDS            Route53 → CloudFront → Lambda → RDS
              ↓                                   (S3 gateway endpoint - free)
         NAT Gateway ($33/mo)                     SSM/Logs (interface endpoints)
              ↓                                   No NAT. No ALB. No Fargate. No EIPs.
         S3, SSM, ECR pulls

The whole fixed-cost floor — NAT, ALB, three EIPs, the idle Fargate task — evaporated.

The Dual-Mode Binary

The web server is a Go binary running gorilla/mux. We didn't rewrite it — we made it ambidextrous:

// When running on AWS Lambda, the runtime API env var is always present.
if os.Getenv("AWS_LAMBDA_RUNTIME_API") != "" {
    lambda.Start(httpadapter.NewV2(router).ProxyWithContext)
    return
}
// Otherwise: the classic HTTP server for local development
log.Fatal(srv.ListenAndServe())

One binary, two modes. Locally it's the same dev server it always was. On AWS, httpadapter.NewV2 wraps the existing router untouched — zero changes to handlers, templates, or logic. The templates and static assets got go:embed-ed into the binary, so the deployment package is literally just one file.

One small but important addition: db.SetMaxOpenConns(2). Lambda containers come and go, and each one brings a connection pool. Cap it, or a burst of cold starts could poke the little database in the eye.

The 403 Mystery

Then came the part every migration has: the thing that works in theory and returns 403 in reality.

CloudFront in front of a Lambda Function URL uses Origin Access Control — CloudFront signs requests with SigV4 so the function URL only accepts traffic from your distribution. We created the OAC, locked the URL to AWS_IAM, added the resource policy granting the CloudFront service principal the lambda:InvokeFunctionUrl action with the right conditions... and got 403.

Every request. From every edge.

The fix is one line in the AWS docs, easy to miss: OAC with function URLs needs TWO permissions, not one. lambda:InvokeFunctionUrl and plain lambda:InvokeFunction, both to the CloudFront principal. Missing the second one gets you the Forbidden message with a link to docs you already read.

Added the second statement. Everything went green instantly.

The CloudFront Cache Policy Trap

Second gotcha: we wanted dynamic pages uncached (TTL 0) so view counts stay exact, with query strings forwarded for the ?sort=popularity toggle. The natural instinct is one cache policy doing both.

CloudFront says no. A caching-disabled cache policy requires every cache-key behavior to be none — you can't forward query strings through it. Query forwarding lives in a separate origin request policy. Two policies, one behavior:

cache_policy_id          = <no-cache policy>            # TTL 0, forward nothing extra
origin_request_policy_id = <forward-query policy>       # forward ?sort

The result: every page load is counted, and the sort toggle works. View counts remain exactly what they always were — honest numbers.

The Landmine

And now the one that made us glad we review plans like paranoid people.

The teardown plan looked perfect: 29 resources to destroy, all the old ALB/ECS/NAT machinery. Then we read the resource list. There, sitting innocently in the middle of the destroy set:

aws_route53_record.blog_alias_record

That's not ALB machinery. That's the root domain's DNS record — and it was defined in the very file we were deleting. Running that apply would have taken the site clean off the internet while every other change reported "success."

The fix: move the record's definition into the DNS file before the teardown apply. Same resource address, terraform keeps it, plan goes from 30 destroys to 29, root domain survives. Always read the destroy list. Every. Single. Item.

No More NAT

The NAT gateway was the single biggest line item, and it existed for one job: letting things in private subnets call AWS APIs. But the blog's Lambdas only ever need three services:

Endpoint Type Cost
S3 Gateway Free
SSM Interface $7.30/mo
CloudWatch Logs Interface $7.30/mo

$14.60 of endpoints replaced $33.48 of NAT gateway. The database never needed a detour anyway — it's right there in the VPC. The private subnets went from "route everything through a toll booth" to three direct paths.

The Sequencing

Zero downtime came from doing the migration as three separate applies instead of one big bang:

  1. Apply 1 — create everything new (Lambda, endpoints, CloudFront, certificate). Old site untouched, still serving.
  2. Apply 2 — flip the DNS alias records to CloudFront. Alias changes propagate in seconds.
  3. Apply 3 — the teardown. 29 resources destroyed while nobody notices, because traffic already points at the new stack.

Plus a safety net first: the entire posts table dumped to S3 (32 posts, every view count), database automated backups enabled for the first time in this database's life (retention was literally zero before — yikes), and a monthly budget alarm wired to email. The database that holds four years of writing now has a backup schedule like an adult.

The Numbers

Before After
NAT Gateway $33.48 $0
ALB $16.75 $0
EIPs $11.17 $0
Compute $7.35 ~$0.10 (Lambda)
Database $14.20 $14.20
Total ~$85-100 ~$36-40

That's about $600 a year back in the pocket. The remaining bill is mostly the database — which is now the only always-on thing in the whole stack, and it earns its keep with those view counts.

And if we ever want to go further, the path to ~$4/month (full static with a tiny counter database) is sitting right there. The door stays open.

While We Were In There

Oh — and the blog got a fresh coat of paint the same day. The dark-and-red ninja theme is now a light silver theme: bright white surfaces, graphite text, a glassy header, and dark ink code blocks as the contrast moment. Same stack, same architecture, just shinier. Writing about cost optimization on a freshly optimized blog hits different.

Lessons Learned

Usage metrics tell the truth. The 0% CPU and the zero invocations weren't bugs — they were the bill explaining itself. Sixty days of "you don't need this server" was hiding in plain sight.

The fixed-cost floor is the enemy. Per-request pricing sounds scary but for low traffic it's almost free; hourly pricing for idle hardware is the silent killer. A blog serving 150KB of content should never pay $68/month just to exist.

Review the destroy list like it's going to bite. Because it is. The difference between a clean teardown and an outage was one line in a plan read-through.

And the meta-lesson: this blog is four years of documenting "build it, automate it, write it down" — and this post is that loop applied to the blog itself. The infrastructure ate its own dog food. The dog food was cheaper.

The toaster got tamed. The AWS bill got tamed. Next up: whatever else has been smugly billing us monthly.


This blog post was written with the help of GLM 5.3 Flash.