Kubernetes is a payroll decision
For a team of a few engineers, the cost of Kubernetes is not compute. It is a second product with its own release train, and the question is who gets paged for it.
Every few months someone asks me whether their three-person company should move to Kubernetes, and every few months I give an answer that has nothing to do with their workload. The workload is almost never the constraint. The constraint is a person. Kubernetes is a second product your company now runs, with its own release schedule, its own upgrade deadlines and its own ways of failing at two in the morning, and the honest question is whether there is someone on the payroll whose job it is to answer that page.
A second release train
Here is the part of Kubernetes that does not appear on the architecture diagram. The project ships a minor version roughly three times a year, and each minor version receives patches for about fourteen months: a year of support and a two-month window in which to upgrade off it. That policy has held since 1.19, and endoflife.date keeps the dates for every version.
Read the chart as a calendar rather than as a feature list. If you stay on a supported version, you upgrade the cluster at least every fourteen months, and in practice every eight or so, because nobody wants to run the last two months of a window. Each upgrade is a control plane change, a node image change, and a walk through the deprecation notes to find the API your ingress controller still uses. Managed offerings do some of that walking for you and none of the reading.
The managed offerings also enforce the calendar rather than relax it. A version that leaves upstream support leaves the provider's support soon after, and most providers will move a cluster off an unsupported version on their own schedule if you have not moved it on yours. So the fourteen months is not a suggestion with a managed control plane. It is the longest you can go without someone on your team doing an upgrade, and the upgrade is the same work whoever hosts the API server: read what changed, test what you run against it, and find out what broke on the day.
Your application has its own release train, the one you chose, with deploys whenever you like. The cluster's train runs on the project's schedule, whether or not this quarter is a good time. For a large team the two trains have different drivers. For a small one they have the same driver, and the collisions land on that person's calendar with no regard for what the product needed that week.
The second-pager rule
So the rule I use is about staffing, not about scale. Adopt Kubernetes when there is a second person who can be paged for the cluster, for etcd, upgrades, the network plugin, the ingress and the certificates, independently of the person paged for the application. Until that second pager exists, a container runtime on one or two machines is the more reliable system, not the less ambitious one.
The reason is what happens during an incident when one person holds both pagers. A cluster problem is, by construction, also an application outage, because the application runs on the cluster. The single engineer is now debugging a node that stopped scheduling pods while the product is down, and there is nobody free to look at the product, to answer the customer, or to notice that the real cause was a bad deploy an hour earlier. Two pagers held by one person are not redundancy. They are one person with two ways to be woken.
Compare the alternative. SocialSure's platform ships as Docker containers through a CI/CD pipeline onto a small number of machines, with structured logs, metrics and alerting around it, and that arrangement has never once needed a person who was not also the person who wrote the code. When that setup breaks, the failure is one of a handful of things, each of them visible from a shell, each of them fixable by a person who understands the application, because the runtime is doing almost nothing the application did not ask for. There is no second product to page for. The boring setup has a shorter list of ways to be wrong, and on a small team the length of that list is the reliability.
What people think the decision is about
The usual argument for Kubernetes on a small team is a list of capabilities: rolling deploys, health checks, restarts, secrets, autoscaling, service discovery. Every item on the list is real and every item is available without the cluster. A process manager restarts a crashed container. A reverse proxy does health checks and a rolling swap between two containers. Secrets come from the environment or a vault. Autoscaling, for a product whose traffic a small team can describe from memory, is one spare machine. The list is a list of things Kubernetes does well at a scale where doing them by hand would take a team; it is not a list of things only Kubernetes can do.
The other usual argument is that you will need it later, so you should learn it now. That one is half right. Learning it is cheap and worth doing. Running it in production before you can staff it is the expensive part, because the learning happens during incidents, and incidents on a small team are the product being down.
The cell people get wrong
The bottom-right cell is the interesting one: the workload genuinely needs orchestration, many services, real scaling, and there is still only one person. The tempting answer is a managed control plane, and it helps, because the cloud provider takes etcd and the control plane upgrades off your list. It does not take the node upgrades, the deprecated APIs, the network plugin, the ingress, the certificates, or the reading. A managed cluster removes the first pager's hardest hour and leaves the rest of the night.
The honest options in that cell are two. Hire the second person before the cluster, not after it, so that the cluster arrives with someone whose job it is. Or rent the whole platform, the kind where you push a container and never see a node, and accept the price and the lock-in as the cost of not staffing an operations role. Both are payroll decisions. Neither is an architecture decision, which is why the architecture diagram never settles the argument.
When the answer flips
It flips on the day the second pager is real, and not before. Not when traffic doubles, not when a client asks for a compliance checklist that mentions orchestration, not when a new hire arrives who used it at their last job. The person has to exist, be on the rota, and be able to fix the cluster without the application engineer in the room. On that day Kubernetes stops being a second product with one owner and becomes what it was designed to be, a shared platform with a team behind it, and the fourteen-month windows become a schedule rather than a threat.
Until then, the most reliable thing a small team can run is the thing it fully understands, on the fewest machines that will hold it, with a pager that only rings for one product. That is not a lack of ambition. It is arithmetic about who is awake.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS