A $169-a-month VM beat serverless and Kubernetes for our stack
We priced Fargate, Cloud Run, Container Apps and managed Kubernetes against one $169/month VM, then kept the VM and built blue/green deploys on it.
In July 2026 I priced what it would cost to move Certant's production stack off the single machine it runs on. That machine costs $169 a month at list price (12 threads, 64 GB of RAM). The serverless container platforms came in at $782 to $2,790 a month in Sydney on on-demand rates, and managed Kubernetes landed at 1.2 to 3.9 times the machine. We stayed put.
Staying put is only defensible if the one machine does the things you would otherwise pay a platform for, and the main one is releasing without downtime. Below is the cost study, then how we release on one host: a blue/green deploy that runs two slots side by side with a 24-hour rollback, plus a 20-second heartbeat we added after a production deploy looked hung when it was only pulling a very large image.
What the stack looks like when you measure it
On 5 July 2026 we measured production live, on the v2.3.21 slot. The app is 15 containers (among them the API, the retrieval service, a set of Celery workers, two beat schedulers, the admin UI, the chat front end and an AutoGluon ML worker) plus a shared infrastructure stack of PostgreSQL with pgvector and Apache AGE, Redis, LiteLLM and Celery Flower.
The whole thing uses about 27 GiB of resident memory, and the CPU sits near idle between ingests. The biggest single process is the retrieval service at 8.26 GiB, most of that an in-process semantic-highlight model, followed by PostgreSQL at 4.90 GiB. Everything else is around a gigabyte or well under.
The configured limits are much higher. Add up what each container is allowed to use and you get about 22 vCPU and 84 GB, on a host with 12 threads and 62 GiB usable. The OCR worker has a 20 GB limit sized for a local parser fallback that production does not use (production OCR goes through an API). The ML worker reserves 4 CPUs and 10 GB for AutoGluon jobs that only fire when someone runs an ML node or a forecast.
On one machine that overcommit costs nothing, because the containers are mostly idle and share the same cores.
Serverless bills the allocation, all day
Serverless container platforms charge for what you allocate, 24 hours a day, whether it is used or not. So every container needs an explicit size, and the sensible sizes (measured memory plus headroom, honouring the limits that exist) add back up to the 22 vCPU and 84 GB that the machine was quietly overcommitting.
I priced two scenarios. "As configured" is a lift-and-shift of all 19 containers. "Optimised" merges the two beat schedulers into one, folds two small services into their neighbours, makes the ML worker scale to zero and cuts the OCR worker down to 2 vCPU and 8 GB, which brings it to 15.5 vCPU and 60 GB.
Monthly compute cost in USD, Sydney region, unit prices checked on 5 July 2026 against the AWS Pricing API, the Azure Retail Prices API and the Cloud Run pricing page. Egress, logs and backups are extra.
| Platform | As configured | Optimised |
|---|---|---|
| AWS Fargate (x86, on-demand) | ~$1,115 | ~$782 |
| Google Cloud Run (plus Cloud SQL and Memorystore) | ~$2,310 | ~$1,535 |
| Azure Container Apps (plus managed PostgreSQL and Redis) | ~$2,790 | ~$2,290 |
| The VM we run on | $169 | $169 |
Against the $169 baseline that is 5 to 16 times the cost.
Fargate sells fixed CPU and memory tiers, and a 20 GB container forces the 4-vCPU tier whether you need the cores or not. Cloud Run's instance-based mode charges for whole vCPUs, so every always-on service costs at least one vCPU ($56.76 a month in Sydney before a byte of RAM), and anything over 8 GiB of memory needs at least 4 vCPUs. Container Apps locks CPU to memory at 1:2 up to 4 vCPU and 8 GiB, which penalises memory-heavy services, and its cheap idle rate only applies below 0.01 vCPU and under 1 KB/s of network. A Celery worker long-polling Redis never gets that quiet, so it bills at the active rate around the clock.
Container count itself turned out to be expensive on the per-service-floor platforms. The three container merges were worth about $24 a month on Fargate but about $183 on Cloud Run. Neither Cloud Run nor Container Apps gives you block storage either, so PostgreSQL and Redis become Cloud SQL and Memorystore (or their Azure equivalents), which is already in the numbers above.
The cheapest credible serverless option I found was the optimised stack on Fargate ARM with a one-year Savings Plan, at about $493 a month in Sydney. That needs multi-architecture builds of every image, including PyTorch for the highlight model, and a validation pass on ARM.
The highlight model would also run slower. It does about 0.9 s per chunk on the production machine's cores and about 3.6 s per chunk on our weaker ARM staging host. Cloud vCPUs are a lottery of older cores, so query latency on any of these platforms would land somewhere between the two, and I would have had to budget more CPU for the retrieval service before committing.
Serverless only wins on cost if most of the workers become event-driven and scale to zero. That is an architecture change, and it fights Celery's long-poll model.
Managed Kubernetes costs 1.2 to 3.9 times the VM
Kubernetes changes the maths because the bill is for nodes. Pod requests can sit near measured usage with the burst headroom in limits, which is the same overcommit the single machine gets for free. The 19 containers pack onto three 4 vCPU, 16 GB nodes with room to spare.
Monthly cost for three 4 vCPU, 16 GB nodes plus control plane, about 100 GB of persistent volume and a load balancer, USD, Sydney region, from the same 5 July price check.
| Managed Kubernetes | On-demand | With a one-year commitment |
|---|---|---|
| AWS EKS (m7i.xlarge x3) | ~$653 | ~$441 (Graviton plus Savings Plan) |
| Google GKE Standard, regional (e2-standard-4 x3) | ~$522 | ~$368 |
| Azure AKS (D4as_v5 x3, Standard control plane) | ~$578 | ~$429 (D4s_v5 reserved) |
| GKE Autopilot (bills pod requests) | ~$871 | not priced |
| A budget provider's managed Kubernetes, free control plane | ~$210 to ~$276 | not priced |
That puts the hyperscalers at 1.7 to 3.9 times the VM and the budget provider at 1.2 to 1.6 times. Autopilot prices like serverless because it bills requests, which rather defeats the point for a workload like ours.
At 1.2 times the cost, the budget option was close enough to take seriously. It buys node-loss tolerance, rolling deploys and horizontal scaling, none of which one machine can offer. I decided against it because of the engineering it would take.
Kubernetes would mean a permanent second packaging target. We would maintain Helm charts alongside deploy.sh, and deploy.sh isn't going anywhere, because customers who run Certant on their own hardware need it (the air-gapped audit covers what those installs look like). The shared document storage volume is bind-mounted by eight or more containers, including one holding a SQLite database, so it would need ReadWriteMany storage, and the budget provider's volumes only attach to one node at a time. PostgreSQL would need an operator such as CloudNativePG. The blue/green flow below would have to be rewritten as rolling deploys with an 1,800-second termination grace period.
Even three nodes only buys partial resilience. Losing one leaves 32 GB for a working set of about 29 GB, which you would survive, just about. True N+1 means a fourth node at another $66 to $184 a month depending on the cloud.
So we stayed on the one machine. If scaling or resilience pressure comes back, the next step we have already researched is the budget provider's managed Kubernetes, ahead of the hyperscalers or serverless.
Two slots on one machine
The deploy runs the old and new versions side by side on the same host. Each version is its own Docker Compose project in its own slot (A or B), with its own host ports, and nginx points at whichever slot is live. PostgreSQL, Redis and Flower live in a separate infrastructure project that both slots share and neither deploy touches.
A release goes through these steps, with typical timings from our operator guide.
- Work out the current and new slot, print the plan, and in production stop for an explicit "yes" (about 30 s).
- Make sure the infrastructure project is up (a no-op on re-runs).
- Check out the tagged release and bring up the new version on the idle slot (5 to 10 minutes with cached images, 25 to 40 minutes cold).
- Health-check the new version's HTTP endpoints and every worker container. Any failure tears the new version down and aborts before nginx is touched (3 to 5 minutes).
- Drain the singletons on the old version: the graph-merge worker and both beat schedulers, using Celery's
cancel_consumerand polling until idle. - Re-template the nginx site for the new slot's ports, run
nginx -tand reload. That is the cutover, and it takes about 5 seconds. - Send SIGTERM to the old workers with an 1,800-second grace period so in-flight tasks finish, bring the old project down, and record the new active version and slot.
Getting that to run without anyone hand-editing the host took a validation cycle on our demo and staging machines in May 2026, and four things came out of it.
Both versions originally joined one shared Docker network, so Docker's DNS round-robined service names between blue and green during the overlap, and worker callbacks landed on the wrong version's API about half the time. Now each version gets its own default network for its internal traffic, and only the services that talk to PostgreSQL and Redis also join a shared infrastructure network.
The 1,800-second grace period caused a different problem. Compose passes it through to each container's stop timeout, and plain docker stop honours it, so re-running a half-finished deploy would sit for up to 30 minutes per worker before it even started. The cleanup step now passes -t 10, and the long grace is kept for the cutover, where it belongs.
On a first deploy, the retrieval service's health check allowed 30 seconds to start, but downloading a spaCy model took about 60 on the demo machine, so eight dependent containers stayed stuck in Created. The start period is now 180 seconds.
Capacity matters most if you copy this. On a 2-vCPU demo VM with 48 GB of RAM, running two versions meant 26 application containers at once, and load average went past 120. Health checks timed out and both versions were marked unhealthy. The fail-safe held, though. The new version never passed its health check, so nginx was never re-templated, the state files never changed, and the old version was still answering with HTTP 200 once the new version's containers were killed. A 4-vCPU, 23 GB staging host then completed its first blue/green deploy with peak memory at 22 per cent, leaving about 18 GB free for an overlap. CPU was the bottleneck (memory headroom mattered less than we expected), and the production sizing we wrote down afterwards was at least 8 vCPUs and 32 GB so the overlap has room.
We left one overlap in place. During the roughly 30 seconds when both versions' beat schedulers are up, scheduled tasks can fire twice. That is only safe for idempotent jobs, so the assumption has to be documented.
The 24-hour rollback
When the old version comes down, its containers and network go but its images and checkout stay on disk for 24 hours. Rolling back inside that window skips the full deploy: it flips nginx and the state files back to the old slot. After 24 hours, a rollback is a normal forward deploy of the older tag.
The production deploy of v2.3.0 on 24 June 2026 shows what that looks like in practice. It went through the normal flow, cut over to slot B with 15 of 15 containers healthy, and brought the previous v1.14.21 slot down. The images for v1.14.21 stayed, so the rollback window was intact. We also changed garbage collection to keep the current version plus the two most recently stopped ones regardless of age, so a quiet stretch between releases can't age the rollback target out from under us.
During the same deploy the disk was at 91 per cent, with 475 GB of it Docker build cache that nothing was pruning, which is why disk usage is one of the fields in the heartbeat below. Pruning the cache took it to 37 per cent.
A deploy that looked hung
The app phase of a production deploy runs as an asynchronous job with a 5,400-second budget, polled every 30 seconds. During the v2.3.0 deploy the terminal showed nothing but poll lines for minutes on end, which looked like a hang.
The job was pulling the AutoGluon ML worker image at about 130 Mbps, and every image in that release was big and fresh: the retrieval service image was 11.8 GB and the shared worker image 9.56 GB. From the terminal there was no way to tell that apart from a wedged deploy.
So the deploy wrapper now starts a background heartbeat alongside the poll. Every 20 seconds it connects to the target host and prints one line to stderr:
deploy.sh running MM:SS | pulling: <image> | app containers healthy N/15 | disk X%
That shows how long the deploy has been running, which image it is pulling, how many app containers are healthy and how full the disk is. It only runs on deploys that carry a version, an environment variable turns it off, and a trap kills it on any exit so it never outlives the deploy. We checked it with a live probe against production before relying on it.
If you are weighing the same move, measure resident memory and CPU against the limits you have configured before you price anything. If the gap is as wide as ours, the per-allocation platforms will bill you for it every hour of the month.



