11 July 2026 · 5 min read

Kubernetes is a payroll decision

For a team of a few engineers, the cost of Kubernetes is not compute. It is a second product with its own release train, and the question is who gets paged for it.

Every few months someone asks me whether their three-person company should move to Kubernetes, and every few months I give an answer that has nothing to do with their workload. The workload is almost never the constraint. The constraint is a person. Kubernetes is a second product your company now runs, with its own release schedule, its own upgrade deadlines and its own ways of failing at two in the morning, and the honest question is whether there is someone on the payroll whose job it is to answer that page.

A second release train

Here is the part of Kubernetes that does not appear on the architecture diagram. The project ships a minor version roughly three times a year, and each minor version receives patches for about fourteen months: a year of support and a two-month window in which to upgrade off it. That policy has held since 1.19, and endoflife.date keeps the dates for every version.

Support windows of five Kubernetes minor versions A Gantt chart from January 2024 to January 2027. Version 1.30 is supported from April 2024 to June 2025, 1.31 from August 2024 to October 2025, 1.32 from December 2024 to February 2026, 1.33 from April 2025 to June 2026, and 1.34 from August 2025 to October 2026. Each bar is about fourteen months long and a new one starts every four months. Fourteen months each, a new one every four Patch support windows, Kubernetes 1.30 to 1.34 Jan 2024 Jan 2025 Jan 2026 Jan 2027 1.30 1.31 1.32 1.33 1.34
Source: release and end-of-life dates from endoflife.date, following the policy on kubernetes.io/releases.

Read the chart as a calendar rather than as a feature list. If you stay on a supported version, you upgrade the cluster at least every fourteen months, and in practice every eight or so, because nobody wants to run the last two months of a window. Each upgrade is a control plane change, a node image change, and a walk through the deprecation notes to find the API your ingress controller still uses. Managed offerings do some of that walking for you and none of the reading.

The managed offerings also enforce the calendar rather than relax it. A version that leaves upstream support leaves the provider's support soon after, and most providers will move a cluster off an unsupported version on their own schedule if you have not moved it on yours. So the fourteen months is not a suggestion with a managed control plane. It is the longest you can go without someone on your team doing an upgrade, and the upgrade is the same work whoever hosts the API server: read what changed, test what you run against it, and find out what broke on the day.

Your application has its own release train, the one you chose, with deploys whenever you like. The cluster's train runs on the project's schedule, whether or not this quarter is a good time. For a large team the two trains have different drivers. For a small one they have the same driver, and the collisions land on that person's calendar with no regard for what the product needed that week.

Two release trains on one calendar Two horizontal lanes across a year. The upper lane shows the application's deploys as frequent small marks chosen by the team. The lower lane shows the cluster's forced upgrades as three larger marks at fixed points. Two of the cluster marks land in the same weeks as product launches, shown as highlighted collisions. Whose calendar is it One year of an application's deploys and its cluster's forced upgrades upgrade collision App deploys launch launch Cluster upgrades collision collision
Illustrative: the cluster's dates are set by the project's cadence, the launches by the product; on a small team both land on one person.

The second-pager rule

So the rule I use is about staffing, not about scale. Adopt Kubernetes when there is a second person who can be paged for the cluster, for etcd, upgrades, the network plugin, the ingress and the certificates, independently of the person paged for the application. Until that second pager exists, a container runtime on one or two machines is the more reliable system, not the less ambitious one.

The reason is what happens during an incident when one person holds both pagers. A cluster problem is, by construction, also an application outage, because the application runs on the cluster. The single engineer is now debugging a node that stopped scheduling pods while the product is down, and there is nobody free to look at the product, to answer the customer, or to notice that the real cause was a bad deploy an hour earlier. Two pagers held by one person are not redundancy. They are one person with two ways to be woken.

Compare the alternative. SocialSure's platform ships as Docker containers through a CI/CD pipeline onto a small number of machines, with structured logs, metrics and alerting around it, and that arrangement has never once needed a person who was not also the person who wrote the code. When that setup breaks, the failure is one of a handful of things, each of them visible from a shell, each of them fixable by a person who understands the application, because the runtime is doing almost nothing the application did not ask for. There is no second product to page for. The boring setup has a shorter list of ways to be wrong, and on a small team the length of that list is the reliability.

What people think the decision is about

The usual argument for Kubernetes on a small team is a list of capabilities: rolling deploys, health checks, restarts, secrets, autoscaling, service discovery. Every item on the list is real and every item is available without the cluster. A process manager restarts a crashed container. A reverse proxy does health checks and a rolling swap between two containers. Secrets come from the environment or a vault. Autoscaling, for a product whose traffic a small team can describe from memory, is one spare machine. The list is a list of things Kubernetes does well at a scale where doing them by hand would take a team; it is not a list of things only Kubernetes can do.

The other usual argument is that you will need it later, so you should learn it now. That one is half right. Learning it is cheap and worth doing. Running it in production before you can staff it is the expensive part, because the learning happens during incidents, and incidents on a small team are the product being down.

The decision as two questions A two by two grid. Horizontal axis: does the workload need orchestration, no on the left and yes on the right. Vertical axis: is there a second person who can be paged for the cluster, no at the bottom and yes at the top. Bottom left: containers on a box. Bottom right: managed Kubernetes only if someone is hired first, or a platform that runs it for you. Top left: still a box, the second person has better things to do. Top right: Kubernetes. Still a box or two the second person has better work Kubernetes two trains, two drivers Containers on a box most small products Hire first, or rent the platform a cluster nobody can page for Does the workload need orchestration: no to yes Is there a second pager: no to yes
Illustrative: the second-pager rule as a grid; only one cell is a cluster you run yourself.

The cell people get wrong

The bottom-right cell is the interesting one: the workload genuinely needs orchestration, many services, real scaling, and there is still only one person. The tempting answer is a managed control plane, and it helps, because the cloud provider takes etcd and the control plane upgrades off your list. It does not take the node upgrades, the deprecated APIs, the network plugin, the ingress, the certificates, or the reading. A managed cluster removes the first pager's hardest hour and leaves the rest of the night.

The honest options in that cell are two. Hire the second person before the cluster, not after it, so that the cluster arrives with someone whose job it is. Or rent the whole platform, the kind where you push a container and never see a node, and accept the price and the lock-in as the cost of not staffing an operations role. Both are payroll decisions. Neither is an architecture decision, which is why the architecture diagram never settles the argument.

When the answer flips

It flips on the day the second pager is real, and not before. Not when traffic doubles, not when a client asks for a compliance checklist that mentions orchestration, not when a new hire arrives who used it at their last job. The person has to exist, be on the rota, and be able to fix the cluster without the application engineer in the room. On that day Kubernetes stops being a second product with one owner and becomes what it was designed to be, a shared platform with a team behind it, and the fourteen-month windows become a schedule rather than a threat.

Until then, the most reliable thing a small team can run is the thing it fully understands, on the fewest machines that will hold it, with a pager that only rings for one product. That is not a lack of ambition. It is arithmetic about who is awake.

KubernetesSmall TeamsOperations
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS