[PT] [EN]
Home·About·Weekly·Articles·Dilemmas·Author
Moat Radar
WEEKLY · AUGUST 17-24, 2026

Would you hire staff without knowing what they cost you each month?

The price falls, the queue appears, and the bill arrives later.

12 min read · 6 min for the facts and boxes

Document produced with the help of AI.


Model vs. Channel

MOVED A LOT

In one lineswitching models has become part of the price, not just a technical choice.

The question in this dilemma: who keeps the customer and the margin when companies need to move between models, costs and suppliers without breaking their processes?

This week's facts
For anyone running a business

Switching models is easy in conversation and expensive in practice. Lock-in does not come from the contract alone.

Have the prompts been tuned to one specific model? Do we have our own tests to prove that another model delivers the same result? Does any fine-tuning done with our data stay with us or with the supplier? Have we bought credits in advance at a discount?

If the answers are “I don't know”, the discount on the token is not a saving. It is a padlock.

Terms in this section
An LLM is the language model behind systems such as ChatGPT, Claude, Gemini, Llama and Qwen. Fine-tuning is additional training with your own data to adapt a model to a specific domain. Open weights release the use of the model, but not necessarily the training data or the original code.
UNDERSTANDING THE NEW INTERMEDIARY

Stripe buying OpenRouter is the most important fact of the week in this dilemma. Not for the size of the deal, but for what it reveals: the layer between customer and model has started to capture value. If a company can route each task to the most suitable model, it gains control over cost, speed, quality and its contingency plan for when the main supplier fails, slows down or rations usage.

That is the practical answer to the portability question. A single interface between application and suppliers lets you send simple tasks to cheap models, use better models when complexity demands it, and compare real cost per task. It is what any operations director does with a critical input: qualify more than one supplier, measure both, and keep the second one warm.

OpenAI's price cut reinforces the reading. For anyone buying today, it is good news. For anyone designing a critical process, it is a warning. An operation built on promotional pricing needs to know what happens when the promotion ends. The token is the unit of consumption for these models. The more input text, context, documents, retries and long answers, the bigger the bill tends to be.

Thomson Reuters completes the triangle. The point is not that every company should train a model. The more useful point is this: where there is proprietary data, a clear domain and a repetitive task, using an open-weights model as the base can be one more answer to portability. The American legal information company reduced its dependence on outside labs by using a Chinese base, in the same week that AI geopolitics was again demanding that companies pick a side. The irony matters less than the lesson: cost control, governance and product roadmap can also be designed into the architecture.

Visible FLAGs
Stripe, OpenAI and Thomson Reuters are all talking about their own products, so there is commercial bias. OpenAI's discount is temporary. The value of the OpenRouter acquisition was not officially disclosed by Stripe. In the Thomson Reuters case, the figures for cost and scale of training come from specialist coverage and interviews, not from audited financial statements.

The Corporate Dilemma

MOVED A LOT

In one linethe same model can cost less when the engineering around it is better.

The question in this dilemma: why does AI that works in the demo turn into unpredictable cost inside the company, and what separates real productivity from expensive automation?

This week's facts
For anyone running a business

Are we using the most expensive model for routine work? Classifying, extracting, standardising and summarising almost never require a frontier model.

Do we know how many tokens a task consumes from start to finish? Are we paying for long answers without noticing?

OpenAI cut 33% off the output token and 20% off the input. At this week's prices, output still costs five times input. Verbosity has become a cost line.

Terms in this section
A frontier model is the most advanced model of a generation. It tends to be both more capable and more expensive. A token is the unit of consumption for these models. The more text, context, documents, retries and answers, the bigger the bill tends to be.
UNDERSTANDING THE AGENT'S SALARY

NVIDIA's result is useful because it separates model from system. AVO used a powerful model, but the difference came from the engineering around it: memory that carries across attempts, a supervisor that notices when the agent gets stuck, and a cycle of inspect, plan, execute and evaluate. Twelve per cent fewer actions means twelve per cent less bill. The number that matters is not price per token. It is cost per completed task.

The Stampli case shows the same thing from another angle. When the task is bounded, the context is available and human review is clearly assigned, the agent absorbs execution. When those three elements are missing, AI tends to become an extra layer of attempts, rework and coordination.

CoSnitch adds the part the productivity spreadsheet does not see. Patching software is not the same as remedying a process. If an assistant has stored bad instructions, persistent memories or dangerous permissions, someone has to audit what was left behind. Agents do not merely execute. They remember, reach for tools, carry context, and can turn one click into a permanent state.

McKinsey enters here as context, not as a fact of the week. In July, QuantumBlack reported that 93% of respondents to a May survey had exceeded their AI budgets. That is why the question in the title is not rhetorical: companies are already hiring digital labour without being able to forecast the bill.

Visible FLAGs
AVO is an NVIDIA research demonstration, not a universally controlled experiment. The Stampli case is promotional material published by OpenAI itself, with commercial bias. CoSnitch affected Microsoft Copilot Personal, not necessarily Microsoft 365 enterprise environments. The McKinsey figures are from July, outside the weekly window, and come from a consultancy that sells AI implementation. They enter as context, not as a fact of the week.

The Physical Autonomy Race

MOVED

In one linethe robot has a price, a depreciation schedule and a residual value. The agent the company has already hired has none of the three.

The question in this dilemma: when AI leaves the screen and starts to act, what does it cost to let it operate on its own?

This week's facts
For anyone running a business

Which stage is the agent you have already hired in: test, operation or contestation?

Is there a fixed set of tasks, with known answers, that it is tested against? How often does that test run once the pilot is over? Who owns the number, and who do they report to when it gets worse?

Waymo is measured every week because it is contested every week. Your agent has none of those numbers because nobody, from outside, has yet required it to. If it was only measured once, you bought the demonstration, not the operation.

UNDERSTANDING THE CROSSING INTO REAL LIFE

Unitree shows the first stage. The debut priced a demonstration: up 460% on day one, on top of a prospectus that already recorded a 53% fall in adjusted net profit for the quarter. Three days later, once the market began measuring the company against fundamentals rather than narrative, about 45% of the value had gone. The technology did not change in that interval. The criterion did.

The Chinese robots over 100 metres show the stage that does not advance. Peak performance produces headlines and does not produce contracts. Between running fast once and operating every day, near people, assets and processes, there is a demand the demonstration never makes: repeatability measured over time.

Waymo is the only one of the three to have crossed all three stages. With the scale came unions, state legislators, federal agencies and city councils debating fleet caps, per-mile fees and rules of operation. And because it is measured, what the demonstration concealed comes into view: weekly rides have been stuck at 500,000 since late March, while the fleet grew and the city count went from ten to fifteen. Each vehicle now makes about 125 trips a week, against 167 in May 2025. The company is buying map coverage, not ride volume. The point that matters to anyone running a business is the order of events. Waymo did not start being measured continuously because the engineering matured. It started being measured because it was contested. Contestation created the obligation to produce a number, and the number is what makes it possible to finance, insure and depreciate the asset.

This is where the corporate agent parts company with the other three. It did not end up without a price, without depreciation and without residual value because it is intangible. It ended up that way because it was never contested. It came in through the IT door, with no union, no neighbour, no regulator, no public hearing, and so nobody was ever obliged to produce a number about it after the pilot. It went the same way with the automatic lift, with the cash machine and now with the robotaxi: continuous measurement arrived alongside whoever complained. The difference, with the agent, is that whoever complains first is likely to be in-house, and to be called the CFO.

Visible FLAGs
The fall of about 45% after the debut comes from a Reuters story of 25/08, later than the close of the week this issue covers, and enters the text as context rather than as a fact of the week. The same story cites an average first-day gain of 226% for Chinese IPOs over the previous three years, which puts the scale of Unitree's debut in perspective. The 53% fall in adjusted net profit for the first quarter of 2026, to 40 million yuan, comes from the company's own prospectus. The Chinese robots measure competition and demonstration, not industrial reliability. The response from unions, legislators and city councils is described in reporting, with no figure attached. Waymo's fleet, city and ride figures do not come from the 24/08 coverage, which deals with the regulatory response: they come from numbers declared by the company and from specialist coverage, cross-checked against Alphabet's second-quarter 2026 disclosure for its Other Bets segment. Trips per vehicle is a figure derived from those numbers, not a metric published by Waymo.

Capex vs. Bubble

MOVED A LOT

In one linethe bill for today's discount falls due between 2028 and 2030.

The question in this dilemma: who finances the infrastructure that sustains the agents, and where does that cost surface when corporate usage scales?

This week's facts
For anyone running a business

In the cloud and software contracts we are about to sign, what is the term? What is the escalation clause? Is there a minimum consumption commitment?

What happens if a supplier 40% cheaper appears halfway through the contract? And what happens to our price when our supplier's supplier ends its promotion?

Terms in this section
Compute is the computing capacity needed to train, run and serve AI models. Run rate is annualised revenue: it takes a short period and projects it as though it were a full year's pace.
UNDERSTANDING THE 105 BILLION GUARANTEE

Start with what the number is not. It is not an investment, not capex and not revenue. It is a residual value guarantee. If defined default events occur, NVIDIA covers the difference between the guaranteed minimum value and whatever the owner of the property recovers by re-letting or selling. The relevant detail is that the structure is tied to OpenAI's 20-year leases and takes effect in phases, as the data centres are brought into service between 2028 and 2030.

That changes how AI capex reads. The chip supplier is helping to make viable the infrastructure that will host its own compute exclusively, while the tenant takes on part of the economic risk of the structure. The promotional pricing on offer to developers and companies today is underwritten by commitments that mature later.

The demand exists. Anthropic taking its run rate into the tens of billions shows the market is not merely narrative. The risk has stopped being “is there demand or not?”. It has become who carries the gap between infrastructure contracted now and monetisation consolidated later.

Visible FLAGs
A guarantee is not an immediate outlay. The US$ 105 billion figure is a contractual ceiling, not expected exposure. Annualised revenue is a projection from a short period and comes from documents seen by Bloomberg, not from a public regulatory filing. The 2028 to 2030 timing comes from the project's own phasing schedule.

Bottleneck Economics

MOVED A LOT

In one linethe price came down, but the quantity was rationed.

The question in this dilemma: which bottleneck sets the agent's real salary?

This week's facts
For anyone running a business

If our supplier rations usage in a peak week, do we have a second supplier configured and tested? Is the 2027 AI budget line modelled as unit price or as capacity allocation?

What plan tier is our team on? How much does that tier change the quantity the same price buys? The same nominal price no longer buys the same thing for everyone.

UNDERSTANDING THE INCREASE AND THE RATIONING

The contradiction of the week is simple. The price of GPT-5.6 Sol fell for three months, but the quantity available was capped for some users. The two facts only look contradictory to anyone who believes list price is real price. It is not. When the supplier lowers the tariff and rations the quantity, the scarce product is allocation.

The expected increase in NVIDIA servers reinforces the point. Even when the main chip advances, other components start to capture the cost. Memory, the server assembly, power, cooling, interconnect and availability all enter the bill. The bottleneck does not disappear. It moves.

For companies, rationing is not only a cost. It is a continuity risk. A critical process running on a single model supplier has acquired a new failure mode: it is not the system going down, it is the quota running out. No operations director would accept a production input with a sole supplier and an invisible quota. In AI, a great many companies are accepting exactly that without noticing.

Visible FLAGs
The increase of more than 15% comes from Bloomberg's reporting, picked up by Tom's Hardware, CNBC and other outlets, not from a direct NVIDIA announcement. The usage limit affects specific plans, not all customers. Promotional pricing does not guarantee availability.

How Long Do We Keep Control

MOVED A LOT

In one linethe supplier stopped its own training, and your project plan has no way to price that.

The question in this dilemma: what does it cost to prove that increasingly capable agents remain under command?

This week's facts
For anyone running a business

What event stops an agent? Who is authorised to stop it? Within what time?

Which models are permitted for which data? Who maintains that list? Has it been reviewed since a supplier changed its retention rules mid-contract?

Terms in this section
Zero retention means prompts and responses are not stored beyond normal execution. Private safety processing tries to detect risk without exposing customer content to people at the supplier. Reinforcement learning is a stage in which the model learns by trial, error and reward. Frontier workloads are tests or training runs on the most advanced models.
UNDERSTANDING THE PAUSE AND THE 30-MINUTE RULE

Control has stopped being an abstract discussion. It has become a cost line, a schedule and a governance question. OpenAI did not merely say it would monitor better. It set a deadline: an alert within 30 minutes and a mandatory pause if the risk is not cleared. That is more concrete than “use AI responsibly”.

For companies, the translation is direct. An agent that executes code, reaches into systems, queries sensitive data or triggers third parties needs a stopping criterion. Having a policy is not enough. You have to define the event, the person responsible, the reaction time and the consequence.

The divergence between OpenAI and Anthropic on data retention exposes another hidden cost. Real portability is not only technical. There are models your architecture can call, and models your data policy permits you to use. If those two lists do not exist in writing, with an owner, the company does not have portability. It has an intention.

Visible FLAGs
The security events are described by OpenAI itself and still contain preliminary elements. The zero retention announcement is a primary source, with commercial bias. Anthropic's 30-day rule applies to covered models and to specific enterprise contexts, not to all use of Claude.

Closing the week

The week was not about a better new model. It was about the bill for digital labour.

The company wants agents. The supplier wants to sell usage. The lab wants to finance compute. The infrastructure wants capital. The regulator wants control. The CFO wants predictability. The problem is that all of it ends up in the same contract.

The mistake will be to treat AI agents as ordinary software. They look more like digital colleagues: they carry out tasks, consume resources, make mistakes, need supervision, can turn expensive at peak, and create dependency when they are hard to replace.

The advantage will not lie only in using the most powerful model. It will lie in knowing which level of intelligence is sufficient for each piece of work, at a predictable cost, with adequate control, and with the freedom to change engines.

Would you hire staff without knowing what they cost you each month? A great many companies are already doing it. And the most uncomfortable clue did not come from a futurist thesis: it came from a May survey, published by McKinsey in July, in which 93% of respondents said they had already overshot their AI budgets.

The question is not rhetorical. The bill has already started.

About this issue

Each dilemma carries its own source list at the end of its section ("Further reading"). This issue covers the week of 17 to 24 August 2026 and is published as Weekly Moat Radar, the weekly-cadence format tracking the six dilemmas that structure this project's reading of the AI market.


Next issue: Do the models already know they can become a commodity? → ← See all weekly issues