gagarinWriting

Assembly required

Clouds sell parts. Turning twenty of them into somewhere you can deploy is a second job nobody pays for, and what you build is a worse PaaS than the one you could have rented.

A cloud hands you IP ranges, security groups, service accounts, IAM policies, load balancer listeners and virtual machines. None of those is a place to put an application. You get to one by assembling them, and the assembly is work: a few days at the start of a project, an hour or two a week after that, for as long as the project lives. It bills nothing. It ships no feature. Each team that does it builds a private platform-as-a-service out of the same parts, and the one they finish is as good as their best engineer's spare attention, which is not very good. The other option is to pay an infrastructure salary for the same artifact.

The not-so-fun puzzle

Count the objects in the smallest useful deployment. One container serving HTTP, one Postgres behind it, on a domain, with TLS.

On AWS: a VPC, two subnets in two availability zones because the load balancer refuses to start with one, a route table, an internet gateway, a NAT gateway so the container can reach your package registry, two security groups, a task definition, an IAM task role, a second IAM role for the thing that launches the task, a target group, a listener, an ACM certificate with its validation record, a Route 53 record, an RDS subnet group, a parameter group, a secret, and a log group. Twenty resources, give or take your tolerance for defaults. Your application is one line inside one of them. Google Cloud and Azure sell the same box with different part names.

None of the twenty is difficult on its own. The difficulty is that all twenty have to agree, and when two of them disagree the system reports a timeout. You forgot the NAT gateway, so image pulls hang for three minutes and then fail with a registry error that says nothing about routing. The security group allows 5432 from a CIDR that looks right and covers the wrong subnet. The health check path returns 302 and the target group calls it unhealthy, so the load balancer drains the only task you have and the deploy rolls back clean, with no failure to read.

Nowhere in the documentation does anyone promise that a web app with a database is easy to run. They promise a page for each primitive, and they deliver that. The convenience was never offered. We supplied it ourselves and then forgot that we had.

Your time or your money

Google and Stripe employ people whose entire job is the twenty objects. They have platform teams, internal deploy tools and an on-call rotation to go with them, and that is a sane way to spend money at their size.

Below that size you choose between two bills. Hire an infrastructure engineer, who costs about what a senior backend developer costs and writes none of your product. Or give the work to a backend developer, who does it in the gaps between features.

Take the second option and look at what the repo collects. A Terraform module copied from a blog post. A deploy.sh with a set -e at the top and four aws calls under it. A GitHub Actions workflow that works, then a second one for staging that drifts from the first within a month. A naming convention somebody documented in Notion in March. Six weeks later you own a private PaaS with one customer, no roadmap, and a maintainer who would rather be writing the product.

Yet another cloud wrapper

Say the expertise is in the building and the wrapper is good. I have read several of these now, at companies with no connection to each other, and they contain the same verbs. Build an image and push it. Set environment variables for an environment. Deploy a version. Declare that this service may reach that database. Print status. Roll back.

Six verbs, written from scratch at each of those companies, because the problem has one shape and everybody finds it. Writing them again re-derives something the industry settled years ago. No customer pays for your deploy subcommand.

And you keep it. The provider releases a major version that renames an attribute you set in nine places. The node image you pinned goes end-of-life. The engineer who understood the module takes another job, and the module becomes a thing the team edits by copying the block above and changing a string. Internal tools get the maintenance that nothing external demands, which is none.

Who's holding the bag?

Most firms with twelve engineers do not want an infrastructure salary on the books, so the twenty objects land on developers who have never administered a network. From there it goes one of a few ways.

They learn, on project time. Four days reading VPC documentation is four days of a developer's salary spent arriving at the line where the application work starts. They will be competent at it in a year and will have shipped less in the meantime.

They guess, and some guesses cost money. A security group open to 0.0.0.0/0 because that made the timeout stop. A database whose backups have never been restored, so nobody knows whether they work. A secret passed as a build argument and printed into a public CI log. None of these announces itself on the day it is made.

They hand it to an agent. An agent writes plausible Terraform in a minute, and plausible is the output that hurts. The plan applies, the service answers 200, and the bucket holding customer uploads is readable by anyone who knows the name. The agent is working from the same twenty-part box you are, with the same four ways to change one thing and no single source of truth among them, so it guesses in the places you would have guessed.

Then the bill arrives, or the data does not stay where you put it, and somebody signs the incident report. The agent is not on call. Your developer, who was never an infrastructure engineer and was handed the work anyway, holds the bag.

There's a better way

Raise the primitives to the things you say out loud when you describe the system. A project. Services inside it. Managed resources. Which service may reach which. A domain, for the ones the public should see.

gg resource add saas/pg postgres
gg ship         saas/api:8080 --deps pg
gg ship         saas/web:3000 --deps api
gg domain   add saas/web
# => Live at https://web-3cnciet6.apps.gagarin.cloud/

No VPC, because you were never going to express a product decision as a CIDR block. No security groups: --deps pg opens the one route and hands api the credentials for pg in the same call, so the route and the password arrive together or neither of them does. Every service stays private until gg domain add, which puts the public internet behind a command you have to type rather than behind a default you have to remember.

What this does not buy you: containers still crash-loop, and migrations still fail to apply. One class of failure leaves, the class where twenty parts have to agree and the system tells you only that something, somewhere, timed out.

And when the project outgrows us, gg eject writes out the Kubernetes you would have had, which you then own, knobs and all. The reasoning behind all of this is in why coherence, and the objects themselves are in the model.

We are on call for the platform. You are on call for your code. The line sits where you can see it, which is more than the twenty objects ever offered.