Skip to content

Instantly share code, notes, and snippets.

@tulior
Created July 24, 2026 21:42
Show Gist options
  • Select an option

  • Save tulior/8b47978da555254968815bdf875373ac to your computer and use it in GitHub Desktop.

Select an option

Save tulior/8b47978da555254968815bdf875373ac to your computer and use it in GitHub Desktop.
A forensic essay on open-weight AI, unpoliced distillation, the decay of model secrecy, and why market power migrates to enterprise control planes and distribution.

The Expiring Moat

Open weights, unpoliced distillation, and where AI market power actually goes

The argument over open-weight artificial intelligence is usually staged as a choice between two moral systems. On one side: transparency, sovereignty, competition, the old civic language of open source. On the other: safety, stewardship, the claim that frontier capability must remain behind controlled interfaces.

The staging is useful. It keeps attention on what the companies say they are protecting.

The harder question is what can still be protected once a frontier model is sold through a public API. The weights may remain locked away. The architecture may remain undisclosed. The training corpus may never leave the laboratory. But the thing customers pay for—the model’s behavior—has to cross the boundary. It crosses one completion at a time, as text, code, rankings, corrections, tool calls, explanations, judgments. A public API is a business built on repeated disclosure.

That fact rearranges the politics, the law, and eventually the balance sheet.

The coalition manifesto Open Weights and American AI Leadership uses the history of open-source software to argue that open models will strengthen American sovereignty, competition, and security. The analogy is partly true where it is least dramatic: local control matters; vendor concentration creates real failure points; organizations should be able to run important systems without asking a remote provider for permission. It becomes unreliable where the document needs it most. Neural weights are not source code. Availability is not interpretability. Releasing a checkpoint gives defenders room to work, but it also gives attackers a persistent, modifiable system that no provider can observe or recall.

The document’s real demand appears late, after the patriotic scaffolding is in place. Distillation, it says, should be treated as a legitimate tradition of learning rather than confused with unlawful extraction. That is the legal boundary the coalition wants drawn before courts and regulators draw another one.

The uncomfortable part is that the coalition does not need its history to be right. Closed-model developers cannot reliably prove that a competitor trained on their public outputs. Watermarks can be removed, diluted, hidden behind another model, or copied to create false attribution. Contracts can govern the conduct of a customer. They cannot, by themselves, make observable behavior secret again. The missing evidence sits inside private datasets, weights, logs, and compute systems. Recovering it at scale would require the inspection regime the argument is supposedly trying to avoid.

Once that enforcement boundary is accepted, model secrecy stops looking like a durable property right and starts looking like a short-lived commercial interval. Capability diffuses. Routine workloads leave first. Raw inference falls toward utility economics while the capital bill rises. Value survives elsewhere: in live data, institutional memory, permissions, execution authority, liability, and distribution.

The fight is presented as open models against closed models. The more durable contest is over who owns the environment in which either kind of model is allowed to matter.


I. The Political Illusion

The real document begins on page three.

After two pages of American technological history, economic diffusion, cybersecurity, and institutional sovereignty, Open Weights and American AI Leadership arrives at the sentence that gives the coalition something concrete to win. Policymakers, it says, must not confuse distillation with misappropriation. Using one model’s outputs to improve another belongs to a “long tradition of learning.” Unlawful extraction may exist, but it should be handled narrowly, without burdening the technique itself.1

Read from the front, this looks like a final policy detail. Read backward, the first two pages become preparatory.

The document opens in the 1980s. Proprietary software companies believed progress required control; open-source pioneers proved otherwise; the resulting ecosystem now supports the internet, American industry, scientific research, cybersecurity, and the state. The United States, the manifesto says, faces the same decision again. Open weights will spread capability through factories, hospitals, farms, classrooms, and small businesses. They will reduce lock-in, strengthen competition, and keep critical technology from collecting behind a few corporate gates.

The sequence is not accidental. History supplies virtue. Sovereignty raises the stakes. Security makes delay sound dangerous. By the time distillation appears, it has been placed inside a national tradition rather than a commercial dispute.

Some of the argument survives contact with reality.

An organization running a model on its own infrastructure has a kind of control that no service-level agreement can reproduce. It can decide where data moves, which version remains in production, how the system is modified, and whether a provider’s pricing change or policy shift will interrupt the work. A defense agency, hospital network, manufacturer, or research laboratory may have reasons for wanting unmediated access that are neither ideological nor trivial. Dependence on a handful of remote models creates common points of failure. The manifesto is right to name that risk.

This is what makes the document effective. It does not invent the whole case. It takes a valid industrial argument and extends it past the point where the analogy can bear weight.

Open-source software discloses instructions. The instructions may be sprawling, obscure, and full of dependencies, but they are still instructions written in a form an expert can inspect. A vulnerability can often be located in a function, a permission boundary, a memory operation, a protocol implementation. The source does not guarantee safety. It gives the investigator an intelligible object.

A neural checkpoint is different. Researchers can run it, probe it, fine-tune it, benchmark it, cut into its activations, compare variants, and discover failure modes that would remain invisible behind an API. That access matters. But the weights do not become readable simply because they can be downloaded. They do not disclose which training examples formed a capability, why a particular answer appeared, where a harmful tendency resides, or what else will move when the tendency is altered.

The manifesto slides over this difference with the word transparency.

What open weights reliably provide is possession. Possession enables inspection of behavior, modification of behavior, and local operation. It does not provide the kind of semantic auditability implied by the software analogy. The distinction is easy to flatten in political prose because both artifacts can be called “open.” The security consequences are not flat.

A defender who receives an open checkpoint still needs an institution around it: telemetry, specialist staff, current threat data, integration with real systems, enough compute to run continuous analysis, and authority to act when the model finds something. An attacker can download the same checkpoint, remove its refusals, adapt it to a narrow target, and operate without API logs. The defender’s advantage is organizational. The attacker’s advantage is that the artifact has already crossed the perimeter.

This does not settle the security question in favor of closure. Closed systems can fail invisibly. Providers can be breached. A small number of centrally hosted models can transmit the same flaw across thousands of customers at once. There are serious cases where defenders need weights, not promises.

But “open models help defenders” is not the same claim as “openness is a security control.” One is operational. The other is borrowed prestige from software history.

The signatory list makes the borrowing easier to understand. Meta, Microsoft, Nvidia, enterprise software firms, cloud vendors, investors, open-model developers, and open-source institutions do not share one business model. Some of them compete directly. Some benefit from closed systems in one part of their portfolio and from open systems in another. Nvidia can profit from a few enormous training programs and from a thousand proliferating deployments. Microsoft can hedge model dependence while selling the infrastructure on which proprietary models run. Meta can promote open weights while keeping its distribution, user graph, advertising system, and product surfaces emphatically closed.

The coalition is not a conspiracy with a single motive. It is a temporary settlement among firms that agree on what they do not want: a small group of frontier laboratories turning an early capability lead into permanent control of the model layer.

Their interests overlap without becoming identical.

Meta gains if foundation models become cheap enough to disappear inside products it already distributes. Nvidia gains when the ecosystem produces more training, fine-tuning, inference, and experimentation. Enterprise vendors gain when customers can change the model without changing the software, data, and workflow layer above it. Open-model builders gain from access to techniques and behavioral data first exposed by closed systems. Microsoft gains from plurality even while carrying large exposure to a closed-model partner.

“American leadership” is the phrase broad enough to hold all of this without forcing the participants to describe the later fight over margins.

The document’s choreography matters more than any private intention. It begins with a public good and ends with a liability rule. The first claim—open systems can reduce dependence—is used to naturalize a second: outputs from closed systems should remain available as inputs to competitors.

That move is not contained in the history of Linux.

An open-source developer studies code that has been deliberately published under a license permitting examination, modification, and redistribution. A model distiller may query a commercial service whose provider has disclosed behavior while withholding the underlying artifact and prohibiting competitive training. Whether the prohibition should control is a real question. Pretending the two acts belong to the same inherited tradition avoids the question rather than answering it.

The phrase “tradition of learning” is doing nearly all the work. It converts extraction into education. It makes the provider sound less like the owner of an expensive system and more like a schoolmaster attempting to forbid students from understanding a lesson.

The counter-language—“unlawful efforts to extract value”—is left conspicuously empty. No boundary is offered. Stolen credentials? Account farms? A contractor using dozens of legitimate subscriptions? A company training on outputs purchased at the posted price? A model learning from a dataset that has already passed through three intermediaries? The coalition condemns the category that everyone condemns and protects the category its members need.

Precision would split the signatories. Ambiguity keeps them on the page.

There is another reason not to define the boundary. The more closely the conduct is examined, the less the law resembles the political story.

A closed-model provider is trying to do two things at once. It wants to sell access broadly enough to recover the cost of the model. It also wants the behavior revealed through that access to remain economically unavailable to a rival. The first objective requires publication at scale. The second asks the market to treat the publication as though it never occurred.

The tension is not solved by calling the weights secret. The weights are not what the customer receives. The customer receives an answer, and then another, and then ten million more. Those answers may be contractually restricted. They may be expensive. They may be delivered through an interface rather than a download. They are still disclosures of what the system can do.

The manifesto wraps this conflict in a familiar American story because the raw proposition sounds harsher when stated plainly:

A frontier laboratory should not be able to turn public behavioral disclosure into a permanent exclusive claim over the capabilities others learn from it.

That proposition is debatable. It has consequences for research investment, competition, and the financing of frontier systems. It deserves a direct argument.

The manifesto gives it ancestry instead.

The political illusion is not that the coalition is self-interested. Every durable policy coalition is. Nor is it that open weights provide no public benefit. They do.

The illusion is that the open-source analogy resolves the conflict. It does not tell us whether neural weights are transparent, whether publication improves the security balance, or whether a purchased completion may become training material. It supplies an answer-shaped history and lets the unresolved legal problem pass underneath.

The coalition may still get the rule it wants. In practice, it probably will.

Not because software history compels it.

Because the public API has already created the leak, and the law arrives after the output is gone.


II. The Legal Reality

Consider the cleanest case a closed-model provider could hope to bring.

A new competitor appears with a model that performs unusually well on the same narrow tasks as the provider’s frontier system. Months earlier, the provider saw a cluster of accounts issuing systematic prompts: rare domains, repeated variations, benchmark-like coverage, no obvious end-user purpose. The accounts were terminated. The traffic continued through new identities and cloud regions. Later, the competitor’s model displays some of the teacher’s peculiar habits.

The story is compelling.

It is not proof.

The provider can show that outputs were acquired. It can show that the later model is similar. The disputed fact sits between those two events: whether the outputs materially entered the training process that produced the released model.

That fact is usually inside systems the provider cannot see.

Watermarking was supposed to make the hidden step visible. The provider would place a covert statistical signature in its responses. A student trained on enough of those responses would inherit the signature. The provider could then query the student and detect its own mark, no dataset inspection required.

The design assumes a strangely cooperative extractor.

A competent pipeline does not place raw teacher outputs directly into a corpus and hope for the best. It treats them as temporary material.

One model generates the answer. A second rewrites it. Several other models answer the same prompt. Their responses are compared; unusual wording and disputed claims are removed; the result is compressed into a neutral target. Some examples become preference pairs. Others become labels, tool traces, short solutions, or reward signals. After the student learns the capability, it is fine-tuned on clean data and optimized away from whatever recurring habits remain. At deployment, another model can sit in front of it and rewrite the final answer again.

Nothing in this pipeline is exotic. That is the problem.

The useful part of a completion is rarely its exact wording. It may be the correct answer, the decomposition of a problem, the ranking of alternatives, a code repair, or a judgment about which response is better. Language gives the extractor room to preserve that information while changing nearly everything around it.

A surface marker survives only if the extractor preserves the surface.

Experiments on “radioactive” language-model watermarks have already demonstrated removal through targeted paraphrasing before distillation and through inference-time neutralization afterward, while retaining the underlying knowledge transfer.2 More elaborate markers try to bind the signal to reasoning behavior rather than token choice. That can raise the cost of removal under the attacks tested. It does not create a universal forensic guarantee. The extractor can omit reasoning traces, summarize them, train a reward model instead of a student response model, merge checkpoints, continue pretraining, or place a neutralizing wrapper in front of the system.

The legal standard cannot be “the mark survived the attacks the plaintiff anticipated.”

An adaptive defendant gets to choose the pipeline.

This creates the first failure: removal. A real lineage may leave no detectable trace.

Ensembles create the second: dilution. If the extractor uses several teachers, which one caused the student’s behavior? A correct answer may be generated by one model, checked by another, rewritten by a third, and retained only when all agree. The finished corpus is not a copy of any one distribution. It is a synthetic consensus assembled from several.

Then comes spoofing. A watermark that can be inherited can also be learned deliberately. Research has shown that a model can be trained to reproduce another model’s watermark, turning the supposed proof of provenance into a tool for false attribution.3 A detected mark may mean that the provider’s outputs trained the model. It may mean the mark leaked through an intermediary. It may mean someone implanted it.

Absence is inconclusive because the signal may have been removed. Presence is inconclusive because the signal may have been copied.

A court cannot confiscate a model on that basis.

The strongest watermark proposals are still useful. They can expose careless extraction. They can help a provider decide which accounts to investigate. They can corroborate internal records or an admission. They can raise the cost of laundering. None of those functions is trivial.

They are not lineage proof.

This is where the technical argument turns legal. Model-level remedies are not small remedies. A plaintiff may seek an injunction against deployment, a royalty tied to the competing model, disgorgement of revenue, destruction of training material, or discovery into the defendant’s entire development process. Before ordering any of that, a court needs more than a statistical resemblance with known removal and spoofing paths.

It needs the missing internal evidence.

The obvious sources are the dataset, preprocessing code, retained response files, training scripts, experiment logs, checkpoints, contractor communications, payment records, and infrastructure used for the run. Direct evidence from any of them may establish what black-box testing cannot.

Obtaining that evidence is not a technical detail. It is the regime.

Imagine a smaller domestic model company sued by two frontier providers. Each alleges suspicious overlap. Each seeks preservation of datasets and logs, access to training records under protective order, inspection by a special master, and an injunction while the dispute proceeds. The defendant may have mixed public data, licensed data, open-model outputs, employee-authored examples, synthetic corpora, and material supplied by contractors. Its provenance records are incomplete because modern training pipelines are not built as litigation archives.

To defend the model, it must expose the machinery that produced it.

The information may include confidential data licenses, internal evaluation methods, security controls, proprietary preprocessing, unreleased model variants, and relationships with suppliers. A protective order reduces public disclosure. It does not make inspection harmless. The litigation itself becomes a way for the incumbent to impose cost, delay, and technical exposure on the entrant.

A threshold rule requiring stronger evidence before discovery does not remove the circularity. The stronger evidence is what the discovery is meant to find.

Mandatory provenance records would solve that problem by making every developer preserve the evidence in advance. The state could require registration of major training runs, certified dataset manifests, retention of prompt-response records, hardware attestations, cloud logs, or audits of qualifying systems. An accused developer could then prove what entered the run.

The improvement in enforcement would be real.

So would the surveillance architecture.

Large-scale private computation would become presumptively auditable. Dataset composition would become a regulated record. Hardware and cloud providers would become evidentiary intermediaries. A laboratory could train privately only within a system designed to prove, later, that the training was lawful.

The choice is often softened into “appropriate transparency.” There is no reason to soften it here. To police adaptive distillation reliably, someone must be able to inspect weights, datasets, logs, or compute. If nobody can inspect them, disciplined extraction will often be unprovable.

The system can tolerate that loss, or it can build the inspection machinery.

It cannot have full enforcement and private computation at the same time.

Reversing the burden does not escape the problem. If a suspiciously capable model must prove independent creation, behavioral convergence becomes a discovery trigger. Frontier models are trained on overlapping public sources, benchmarks, code, papers, and synthetic data. They will often arrive at similar answers because the answers are correct or because the public training environment is shared.

The largest provider could then turn resemblance into process: allege derivation, force provenance disclosure, raise the cost of entry. The property right would be created procedurally even where the substantive claim could not be proved.

The enforcement gap has to remain with the party that exposed the outputs.

That does not make every method of acquisition lawful. It changes what the law can sensibly attach to.

Stolen credentials can be proved without opening the final model. So can employee bribery, fraudulent account creation, bypass of a genuine access barrier, or a contractor’s breach of a real confidentiality duty. The wrong is complete at acquisition.

Contract claims also survive, but only as contract claims.

A provider may place an anti-training covenant in a click-through agreement. Courts often enforce clickwrap terms where notice is conspicuous and assent unambiguous. There is no durable boundary between “boilerplate” and “negotiated” that prevents providers from doing this. Requiring individual negotiation merely changes the interface through which the same restriction is imposed.

The relevant boundary is the remedy.

A contract can bind the party that accepted it. Breach may justify termination, unpaid charges, provable damages, or a valid liquidated-damages clause. It should not make the output carry a property restriction enforceable against everyone who later encounters it. It should not establish that a downstream model was trained on the material. It should not authorize suppression of a checkpoint without evidence connecting the breach to the checkpoint.

The Supreme Court’s decision in Van Buren v. United States matters for the same reason. The Court rejected an interpretation of the federal computer-access statute that would turn improper use of information a person was entitled to obtain into unauthorized access. Access and use are not the same legal act.4 A private API term can prohibit a use. It should not automatically transform that use into computer intrusion.

Trade-secret law already contains a more disciplined structure. It protects information that derives value from secrecy and that the owner has taken reasonable measures to keep secret. It recognizes theft, bribery, misrepresentation, and breach of secrecy duties as improper means. It also excludes independent derivation and lawful reverse engineering from that category.5

The hidden weights may remain trade secrets. So may the training process, unreleased system behavior, confidential datasets, or capabilities exposed only to a small group under genuine secrecy obligations.

A public API is more difficult.

The provider can retain secrecy in the machinery while disclosing behavior to a large, changing population. A contract may restrict what a customer does with a response. It does not follow that every inference drawn from the response inherits the secrecy of the machinery behind it. The provider has disclosed the thing from which the inference is made.

Calling the system “closed” can obscure this. The artifact is closed. The service is not.

The workable legal residue is narrow enough to state without a balancing test:

  • Punish provable acquisition offenses.
  • Enforce API covenants against the parties who accepted them, using ordinary contract remedies.
  • Reserve model-level remedies for direct evidence that the disputed outputs entered the model’s development.
  • Treat watermarks as investigative signals, never as dispositive lineage proof.
  • Do not infer unlawful training from capability similarity or from a refusal to expose an entire dataset.

This rule will miss competent extractors. It will miss some domestic actors, not only foreign ones. A company that distributes queries, rewrites outputs, destroys intermediate records, and keeps its contractors at arm’s length may leave no usable case.

That is not a drafting defect. It is what remains when the alternative is compulsory inspection.

Closed-model providers can respond operationally. They can narrow access to their most sensitive capabilities, use confidential enterprise arrangements, monitor abuse, vary service tiers, and assume that publicly served behavior will leak. They can preserve claims against thieves, fraudsters, insiders, and contracting counterparties.

They cannot preserve a practical exclusive right over everything the public learns by using the product.

The law can reach the account, the credential, the payment, the lie, the stolen file.

After that, it reaches darkness.


III. The Economic Terminal State

This is where the legal argument becomes a balance-sheet problem.

A public API completion is two things at once. It is a sale, and it is a disclosure about the machine that produced it.

The first fact appears in revenue. The second appears later, somewhere else: in a competitor’s evaluation set, a synthetic corpus, a smaller model’s improved reasoning, a benchmark that no longer separates the frontier from the rest. The provider is paid immediately. The depreciation is distributed.

This is an unusual asset. Use does not wear out the weights. It wears out the scarcity around them.

The commercial value of a closed model depends less on how intelligent it is than on the distance between it and the best substitute a customer can actually use. Call that distance the capability gap. A laboratory can improve its model every quarter and still lose pricing power if alternatives improve faster, become cheaper, or cross the threshold at which customers stop caring about the remaining difference.

That threshold is rarely the frontier.

A bank does not need the world’s most capable model to classify routine correspondence. A software team does not need it to rewrite boilerplate tests. A retailer does not need it to summarize product reviews. Once a cheaper system becomes reliable enough, the frontier model’s extra ability may be real and commercially irrelevant.

The easiest traffic leaves first.

What remains is harder, slower, more failure-sensitive, and more expensive to serve. The closed provider retains the edge cases that justify frontier capability while losing the repetitive workloads that smooth utilization and produce attractive margins. It can still say, correctly, that no competitor matches the model across the full distribution. The customer is not buying the full distribution. The customer is unbundling the workload.

This is the decay mechanism. It does not require a perfect clone. It requires enough substitutes.

The laboratory can defend the gap by moving faster. Before one generation is copied, a new one appears. The frontier remains closed; the trailing capability washes into the market.

That strategy works. It also reveals what the moat has become.

It is not the model. It is the ability to spend again.

Training, evaluation, safety work, data acquisition, and infrastructure purchase another period of scarcity. Public use begins eroding it. The organization returns to the market for more chips, more power, more data, another generation. The asset is replaced before it is exhausted because its economic lead expires before its technical usefulness does.

A successful API accelerates both sides of the cycle. More use produces more revenue and more evidence. The model becomes easier to map because it matters enough to query.

The provider can protect the frontier by withholding it. But a withheld capability does not recover training cost. It does not build a customer base. It does not generate the volume required to become a platform. The business has to choose how quickly it wants to monetize the secret and how quickly it is willing to spend it.

No contract clause removes this choice.

Once several suppliers can perform the same work, raw inference begins to resemble a utility. Customers compare price, latency, context length, reliability, regional availability, capacity guarantees, and the inconvenience of switching. They route different tasks to different models. They keep an open model on-premises for routine traffic and reserve a frontier API for the residue.

The provider’s engineering advantage migrates into cost control: batching, caching, quantization, routing, speculative execution, utilization, networking, power procurement. These are serious disciplines. They can separate a good operator from a bad one.

They do not restore scarcity to the answer.

Efficiency gains become bids in a price war. One provider lowers the cost of a unit of capability; competitors follow; customers redesign workloads to consume less. The improvement is retained only temporarily before it appears in the market price.

Meanwhile, the infrastructure bill arrives in full.

In Microsoft’s fiscal 2026 third quarter, capital expenditures were $31.9 billion. Roughly two-thirds went to short-lived assets, primarily GPUs and CPUs. Microsoft Cloud gross margin fell to 66 percent, with the company attributing the pressure in part to AI infrastructure investment and growing AI usage.6

“Short-lived” is the operative phrase. The buildings and power systems may support years of expansion. The accelerators inside them are purchased into a moving technical curve. A newer generation can reduce the cost of equivalent inference before the older hardware has earned the return assumed when it was ordered.

Amazon reported $128.3 billion in cash capital expenditures for 2025, up from $77.7 billion the year before, with most technology-infrastructure investment supporting AWS growth. Free cash flow fell from $38.2 billion to $11.2 billion. It also shortened the estimated useful life of a subset of servers and networking equipment from six years to five, citing the faster pace of development in artificial intelligence and machine learning.7

The accounting change is a quiet admission: some of the plant is aging economically faster than expected.

Oracle spent $55.7 billion on capital expenditures in fiscal 2026 and reported negative free cash flow of $23.7 billion while expanding cloud infrastructure.8 Demand may justify the build. The point is not that the investment is irrational. The point is what kind of business can survive it.

Microsoft, Amazon, and Oracle do not need inference to stand alone. Compute protects and enlarges other franchises: productivity software, databases, cloud contracts, retail, advertising, developer tools, security products, installed enterprise relationships. They can accept pressure in the infrastructure layer because the infrastructure defends margin somewhere above it.

An independent raw-inference provider has fewer places to hide the bill.

It buys capacity before demand is certain. It commits to power, chips, leases, and networking. When the capacity comes online, the market may contain smaller models, lower token use, better caching, new competitors, and customers who learned to route around the expensive system. The provider still has to fill the machines.

Idle accelerators do not preserve optionality. They age.

So the supplier discounts. It offers reserved capacity, private rates, credits, bundles, long commitments. The logic becomes load factor rather than intellectual property. A unit sold cheaply may be better than a unit left unused, even when the full-cycle return on the facility is poor.

The company can report enormous revenue while the economics underneath it move toward airlines, telecommunications, and semiconductor fabrication: fixed costs, perishable capacity, volatile utilization, repeated capital calls.

There is one difference. An airline seat does not teach a rival how to build an aircraft. An inference call can help teach a rival model how to perform the service.

This is the utility trap. Scale is necessary and insufficient. The supplier must invest like an infrastructure monopoly while competing like a commodity vendor.

The trap does not mean inference disappears or becomes unimportant. Electricity is important. So is bandwidth. Strategic importance is not the same as control over industry profit.

Value moves to the places where substitution is harder.

Take an insurance claim. A model can read the documents, identify missing evidence, estimate whether the loss falls within the policy, and draft a recommendation. Those capabilities will diffuse. The difficult asset is everything around the recommendation.

The operating system knows which policy version governs, whether the claimant’s identity has been verified, what evidence arrived yesterday, which fraud indicators have already been reviewed, whether the adjuster has authority to approve the amount, what state deadline applies, who must sign off on an exception, and how the payment will be recorded. It remembers the file after the model call ends.

Replace the model and much of that state remains.

Replace the system holding the state and the institution has to rebuild its own memory.

This is why workflow state matters more than chat history. It is not a transcript. It is the accumulated map of what the organization has done, what it permits, and what still has to happen. It includes corrected decisions, exception rules, audit trails, evaluation results, access mappings, and the thousand local accommodations that never appear in a product demonstration.

A model can be swapped in an afternoon. Revalidating the process may take a year.

The enterprise control plane owns that revalidation problem. It decides which model receives which task, what data it may see, what tools it may call, when a human must intervene, and what evidence must be retained. It can send classification to a small local model, code analysis to a specialist, and a rare judgment to a frontier system. The employee sees one interface. The model suppliers compete behind it.

That is a more durable position than owning one model.

The control plane observes performance on the customer’s actual work. It knows when the expensive provider is necessary and when it is merely preferred. It can shift traffic. It can threaten to self-host. It buys intelligence wholesale and sells continuity to the enterprise.

Live data rights harden the advantage. A static model can absorb yesterday’s market commentary, medical literature, logistics history, and product catalog. It cannot manufacture lawful access to today’s price feed, the patient’s current record, the shipment’s location, the company’s private ledger, or the inventory that changed five minutes ago.

The valuable asset is not only the data. It is the continuing right to receive and use it inside a real transaction.

Execution rights go further. A model may recommend a payment. The system with value is the one authorized to move the money, check the recipient, apply the limit, obtain approval, record the action, and reverse it if necessary. A model may identify a compromised account. The system that matters can revoke the credential and isolate the machine.

Advice can be copied. Authority is issued by an institution.

Then there is liability. Enterprises buy someone to remain responsible after the model answers. They require retention controls, incident response, auditability, data residency, warranties, insurance, and a counterparty with enough balance sheet to absorb a failure. None of this lives in a checkpoint. It lives in contracts and operating organizations.

This is where the margin goes: not to the isolated act of prediction, but to the environment that makes the prediction current, authorized, and survivable.

The hierarchy among AI businesses follows from a blunt question: who can replace whom without losing the customer?

A hyperscale inference utility ranks last because the customer can route around it while the supplier remains attached to the hardware bill. A vertical outcome provider does better; it owns a specific process and can charge against the value of completed work, though each industry brings its own regulation, data, and operational burden. An enterprise AI operating system is stronger because it controls institutional state across many models and many workflows.

Distribution sits above all three.

Distribution owns the moment before the model is chosen.

On the consumer side, it controls the interface, identity, history, payments, and default. On the enterprise side, it controls the system of record, procurement relationship, permissions, and workflow. It can substitute models without asking the user to form a new habit. It can subsidize inference with advertising, commerce, cloud revenue, devices, or software subscriptions. It can force model suppliers to bid for access to demand.

The raw model company cannot make the reverse threat. Better weights do not create a billion users. They do not create payroll data, a payment rail, a procurement contract, a software repository, or years of institutional memory.

When differences between models are dramatic, users may leave the default to find the best one. As the differences narrow, adequacy becomes enough. Convenience takes the decision back.

This is how distribution becomes the final monopoly—not necessarily one company, and not permanently, but repeatedly. Wherever demand aggregates, the aggregator can change the supplier while keeping the relationship. Search, operating systems, commerce, productivity software, developer platforms, and enterprise systems of record each become private markets in which models compete for allocation.

The platform taxes the layer beneath it because it can remove any one supplier without removing itself.

That is the long-run winner. It absorbs the largest share of market value because it owns the customer rather than the artifact. The enterprise control plane follows. Vertical outcome providers remain valuable but fragmented. Independent raw inference remains strategically important and financially exposed.

The political irony is hard to avoid.

The open-weights coalition promises to break concentration at the model layer. It may succeed. Capability will spread. The proprietary model premium will shrink. More organizations will be able to run, modify, and choose among systems.

Power does not vanish. It moves.

The company that controls identity, memory, permissions, data access, execution, and the default interface can treat open and closed models alike as replaceable inputs. The user gains choice among suppliers while becoming more dependent on the system that performs the choosing.

The manifesto calls its project sovereignty. At the model layer, perhaps it is.

Above that layer, the likely result is open supply beneath a private tollbooth.

Footnotes

  1. American Innovators Network et al., Open Weights and American AI Leadership (July 24, 2026), 3 pp., hosted by NVIDIA, https://images.nvidia.com/pdf/Open-Weights-and-American-AI-Leadership.pdf

  2. Leyi Pan et al., “Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?,” ACL 2025, https://aclanthology.org/2025.acl-long.648/

  3. Hyeseon An et al., “DITTO: A Spoofing Attack Framework on Watermarked LLMs via Knowledge Distillation,” EACL 2026, https://aclanthology.org/2026.eacl-long.229/

  4. Van Buren v. United States, 593 U.S. 374 (2021), https://www.supremecourt.gov/opinions/20pdf/19-783_k53l.pdf

  5. 18 U.S.C. § 1839, https://uscode.house.gov/view.xhtml?req=granuleid:USC-prelim-title18-section1839

  6. Microsoft, FY2026 Q3 earnings materials: capital expenditures of $31.9 billion, roughly two-thirds for short-lived assets, and Microsoft Cloud gross margin of 66 percent. https://www.microsoft.com/en-us/investor/events/fy-2026/earnings-fy-2026-q3 and https://www.microsoft.com/en-us/investor/earnings/fy-2026-q3/performance

  7. Amazon.com, Inc., 2025 Form 10-K: cash capital expenditures, free cash flow, and revised useful-life estimate for certain servers and networking equipment. https://www.sec.gov/Archives/edgar/data/1018724/000101872426000004/amzn-20251231.htm

  8. Oracle Corporation, fiscal 2026 Form 10-K: capital expenditures of $55.663 billion and free cash flow of negative $23.686 billion. https://www.sec.gov/Archives/edgar/data/1341439/000119312526277521/orcl-20260531.htm

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment