Adam Hagestedt. ← all essays
Monthly essay

Personal AI is in Uber time

Cheap personal stickers train the habit. Production hosting is where the subsidy ends.

Continues the argument in Intelligence spend is the new cloud spend: that piece was about pooling intelligence spend and bring-your-own-inference. This one is the operator sequel, the cliff when you leave the subsidized lane.

Here is the claim, up front. Personal AI is in Uber time. The sticker price on free and cheap consumer tiers is not the cost of serving real usage. It is a land-grab price, trained to make “just ask the model” feel free. The bill shows up when a team leaves that lane and has to host, provision, or self-run the same class of work.

I buy and operate intelligence for a company north of 3,000 people. I also use the personal products like everyone else. The gap between those two surfaces is the story. One feels like a $0 or $20 habit. The other looks like cloud spend with a new name.

What “Uber time” means here

Remember the ride-share years when a cross-town trip cost less than the fuel and labor underneath it. Capital bought habit. Fares corrected later. Public writing in 2026 already uses that metaphor for tokens, and for good reason. The pattern is not “AI is fake.” The pattern is “the price you feel is not the cost you will own.”

Two moves in August made the split visible.

1/ OpenAI put a current-generation model, GPT-5.6 Luna, in front of Free and Go ChatGPT users and removed the cap on text chats (tools still limited). Habit first. Meter later.

2/ In the same news cycle, reporting on DeepSeek described peak-hour output price increases on the order of 4x to 5x, plus peak and off-peak tiers. The cheap era was a strategy. Strategies expire.

Neither move proves inference suddenly got expensive in a physics sense. Per-token costs have fallen hard over the last year and a half. What they prove is different: labs choose who absorbs the remaining cost, for how long, and on which surface. Free personal chat, loss-leader API rates, and enterprise contracts are different subsidy instruments. They do not move together.

The three cliffs operators actually hit

The recruiter-friendly demo dies in three places. Same product story. Different invoice.

1/ Personal sticker → production meter. A prototype that lived on a free consumer tier or a free API project meets billing. On Gemini’s API surface, enabling billing on a project removes the free tier for that project. Every call meters from the first token. That is not a bug. It is the moment testing economics end.

2/ On-demand demo → provisioned capacity. Bedrock Model Units and Azure provisioned throughput units stop billing like a toy and start billing like reserved capacity. You pay for the floor whether traffic shows up. Public FinOps writeups keep landing the same rule of thumb: provisioned only beats on-demand once you sustain high utilization, often in the 60% to 70%+ range, on a flat-enough load shape. Spiky internal tools fail that test. Customer-facing steady load can pass it. Either way, the cost shape changed.

3/ “We’ll just self-host the open model” → always-on GPUs. Open weights are not free inference. They swap a per-token bill for a per-hour GPU bill. Idle nights and weekends do not negotiate with Azure.

That third cliff is the one people still romanticize. So put a number on it.

The ~$70k/month number, with the assumptions attached

Public Azure list pricing for Standard_ND96isr_H100_v5 (eight NVIDIA H100 GPUs) in the cheapest major US regions sits at about $98.32 per hour on Linux pay-as-you-go. At a 730-hour month, that is roughly $72k. Call it ~$70k/month as a ballpark.

I am treating that as an estimate with assumptions, not a Finance-confirmed quote and not a private SKU special.

Assumptions:

1/ One always-on node. No multi-node HA, no spare for failover.

2/ Region near East US / West US 2 class list prices. Other regions run higher.

3/ Pay-as-you-go Linux. A one-year reservation can cut on the order of ~35%. A three-year cut can land near ~55%. Spot is cheaper and not a production serving plan.

4/ Compute only. Egress, disks, model storage, orchestration layers, and people time sit on top.

5/ Model class: a GLM-class open stack that published serving guides still place on roughly eight H100-class GPUs at FP8 for serious context, not a 9B laptop demo.

What that number is good for: showing the cut when a team leaves subsidized personal or API stickers and owns production hosting on a familiar enterprise Azure path. What it is not: proof that every open model costs $70k, or that specialist GPU clouds cannot undercut hyperscaler list. They often can. The cliff is still real. The shape of the bill changes from “tokens I forgot to count” to “capacity I have to keep warm.”

Continuity with intelligence spend

In July I argued that intelligence spend is starting to look like cloud spend: commits, pools, marketplace gravity, and bring-your-own-inference as an RFP line. That essay was about how enterprises will buy.

This one is about what breaks before the pool is mature. Teams still prove value on subsidized surfaces. Then a pilot wants latency guarantees, data boundaries, or predictable capacity. Then someone says self-host. Then finance sees a line that looks like a small HPC budget for one model family.

BYOI does not remove the cliff. It moves who pays it. If your product only works because you are reselling subsidized tokens, you are on the same ground I warned about in July. If your internal agent only works because nobody is metering the personal tier it was born on, you are there too.

What gets cut when the subsidy thins

Show the cut. Do not invent traits.

1/ Unlimited-feeling chat dies first. Free tiers keep guardrails and tool caps. Power users get steered into paid seats.

2/ Flat “AI for everyone” budgets die second. When enterprise moves from soft seats to token-aware billing, finance can finally see which workflows burn. That is healthy. It also kills vanity rollouts.

3/ Naive self-host savings die third. If utilization is low, the open model on an idle eight-GPU node loses to a hosted open-model API. The middle lane exists for a reason.

The honest counterpoint: capability per dollar is still getting better. Builders who treat today’s price as a permanent floor still get hurt. Both can be true.

What buyers should do now

Treat personal AI stickers as training wheels, not a cost model.

1/ Pick one production workflow and meter it end to end for two weeks on the path you would actually run (API on-demand, provisioned, or self-host). Cost per task beats vibes.

2/ Run the same workflow at 2x and 3x token or capacity cost before you promise headcount or SLA outcomes. If the business case snaps, the business case was the subsidy.

3/ Ask the BYOI question in every agentic buy. Route what you can onto capacity you already own. Pool the rest under one owner, the same way you already own cloud.

What vendors should do now

Stop selling demos that only pencil on someone else’s loss leader.

1/ Price the workflow outcome, not the hidden tokens.

2/ Accept the buyer’s model and account where you can.

3/ Publish a production cost shape: on-demand, provisioned, and self-host or VPC options, with the cliffs named. Buyers will find them anyway.

What I’d do Monday

1/ Prove it on one team. Choose one workflow with a real user and a real failure mode. Kill the personal-tier prototype path for that workflow. Put it on a metered project with a hard budget cap.

2/ Write the cliff in one page. For that workflow, show today’s bill, the provisioned option, and a one-node self-host estimate with the assumptions above. Same quality bar. Three numbers.

3/ Expand on pull, not slideware. Only widen access when a second team asks with a use case and a budget owner. No company-wide “unlimited” banner.

4/ Operationalize so it survives you. Name a single owner for intelligence run rate. Add cost-per-task to the weekly ops review. Document the failover model (which cheaper model or cached path catches overflow).

5/ Rehearse the exit. If the current provider raised prices 3x next month, which workloads move to hosted open weights, which stay frontier, and which get deleted? Write that list while the sticker is still cheap.

Personal AI can stay magical on the phone. Production AI has to survive the invoice. Uber time ends for every market that finds product-market fit. The operators who win are the ones who meter before the habit hardens.