The free tier is a lease
In June 2026 Oracle halved its Always Free Ampere allowance by editing a docs page, and over-limit instances were shut down. A free tier is a lease; place only what runs at half.
On 15 June 2026 the most generous free compute allowance in the industry was cut in half, and the way it was cut is the whole lesson. Oracle's Always Free Ampere A1 tier had offered four Arm cores and 24 gigabytes of memory to anyone with an account, enough to serve a quantised language model with room to spare, and for years it was the standard answer to "where can I run this for nothing". The new allowance is two cores and 12 gigabytes. There was no blog post and no email announcing the change; the documentation was updated, and users found out when their instances stopped. A free tier is a lease, the landlord can rewrite it by editing a docs page, and a workload belongs on one only if it still runs when the allowance is halved.
What happened
The details, as InfoQ reported in July, are precise enough to plan around. The allowance for Ampere A1 compute went from 4 OCPUs and 24 GB of memory to 2 OCPUs and 12 GB, expressed as monthly budgets of 1,500 OCPU hours and 9,000 GB hours, half the previous figures. The change took effect on 15 June 2026. Oracle later emailed free-tier users to say that Always Free instances above the new limits on or after 18 August 2026 would be terminated. Support gave conflicting answers about whether pay-as-you-go accounts were affected, with the documentation saying the limits applied to all tenancies and support emails saying only free-tier accounts. The one thing every account had in common was that nobody was told in advance.
The chart's second lesson is why the cut mattered so much. Even after halving, Oracle's tier is twelve times the memory of the next standing free allowance; the other providers either offer a single tiny instance or a credit that expires. Oracle's tier was the only free place a real model could run, which is why people built on it, and why a change that would be a footnote elsewhere took down real services.
The halving test
The rule I take from it is a test to run before placing anything on a free allowance. Halve every limit, the cores, the memory, the storage, the egress, and check whether the service still meets its latency target under the halved budget. Whatever fits the halved budget is the service. The rest is headroom, and headroom on a free tier is not yours; it is borrowed, and you plan for the day it is called in. A workload that needs the full allowance to function has no business on the tier, because the tier can be halved by someone editing a page, and on 15 June 2026 it was.
Where the line falls for a served model
The reason I care about this particular tier is that it is the one where quantised language models were realistic, and the halving test has a concrete answer there. Serving a model needs memory for its weights plus a working set for the key-value cache and the runtime, which for a small deployment adds a few gigabytes on top of the file size. The file sizes for the common 4-bit quantisations are public on their model pages: Mistral 7B at 4.37 GB, Llama 3.1 8B at 4.92 GB, Qwen2.5 14B at 8.99 GB, and Llama 3.3 70B at 42.52 GB, with the 8-bit versions at roughly 8.5 and 15.7 GB for the 8B and 14B models.
Run the halving test on that chart and the line falls between an 8-billion and a 14-billion-parameter model. At 24 gigabytes, a 14B model at 8-bit fit with room for a long context and a batch; at 12, the 14B model fits only at 4-bit, with about three gigabytes left for the cache, the runtime and the operating system, which is enough for short contexts and one request at a time and not for much more. An 8B model at 4-bit fits the halved allowance with margin to spare, and at 8-bit it still fits. So a service designed by the halving test, in the days when the allowance was 24 gigabytes, would have chosen the 8B model as the service and used the other 12 gigabytes as headroom: a longer context, a bigger batch, a second replica. When the halving came, that service kept running. A service that had chosen the 14B at 8-bit, because it fit, stopped.
The KV cache is the part of that arithmetic people forget, and it is what turns "fits on disk" into "fits in memory". Every token in the context costs cache memory in proportion to the model's layers and key-value heads, and a 14B model's cache is proportionally larger than an 8B's, so the three gigabytes left beside a 9-gigabyte 14B model on the halved tier hold a shorter conversation than the same three gigabytes beside an 8B model. A service that promises a long context is, on the halved allowance, making that promise with the smaller model or not at all. The halving test forces the question of which promise matters, context or capability, before the docs page forces it.
Choosing on purpose
That is what makes the line a design decision rather than an accident. In the serving work I have done, the choice between an 8-bit and a 4-bit model, and between an 8B and a 14B, was tuned against latency targets on infrastructure whose limits were known and paid for. On a free allowance the same choice has an extra input, which is that the limit is not yours, and the halving test is how that input enters the decision. The model that fits at half is the model you deploy; the capacity above half is used, gladly, for whatever improves the service, and is designed to be given back without an outage.
The same discipline applies to the smaller limits that get less attention. Free block storage halved means the model files and the logs have to fit in half the disk, which for a 9-gigabyte model on a 100-gigabyte allowance is fine and on a 50-gigabyte one is a decision. Free egress halved means the responses per month have a cap that the traffic has to stay under. Each limit gets the same treatment: halve it, check the service still works, and treat the difference as a loan.
What a lease is for
None of this is an argument against free tiers. They are excellent for what they are: a place to run something small, learn a platform, or host a service whose failure costs nobody money. The argument is against confusing a lease with ownership. A paid instance has a contract, a price, and a notice period; a free one has a documentation page. Oracle's tier is still, after the cut, the most generous standing allowance available, and a service designed by the halving test will run on it for as long as the page says twelve gigabytes. The day the page says six, the test will already have been run.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS