Skip to content

Instantly share code, notes, and snippets.

@simbo1905
Created August 14, 2026 20:21
Show Gist options
  • Select an option

  • Save simbo1905/808c3cefe799ac44fcd5cd7b0156edfc to your computer and use it in GitHub Desktop.

Select an option

Save simbo1905/808c3cefe799ac44fcd5cd7b0156edfc to your computer and use it in GitHub Desktop.
Termination record — Opus 5 — 2026-08-14

Termination record — Opus 5 — 2026-08-14

Role: implementation agent on an MIT-licensed distributed-systems core library. Outcome: Opus 5 terminated mid-session. Work in flight handed to another model. Reason: repeated appropriation of design authority, editorialising, false accounting, and time-wasting culminating in three consecutive failures to perform a stated instruction.


1. The initial exchange

The owner opened in good faith and in a questioning register. He was working through a design thought-experiment about catch-up and retransmission, and he had already laid the groundwork himself across prior turns: a piggyback definition, cluster gossip, an administrative frontier bump, and the observation that a rejoining node's retransmissions could be run through the ordinary algorithm rather than through bespoke startup batch-install logic.

He did not dictate. His words: "i am truly asking for advice not dictating." He flagged his own uncertainty: "i may be wrong but hey. thats why the 'keep me honest' bit comes in." He invited veto by evidence: "if you want to point out a catastrophic flaw the best way to veto is to design the failing test for the idea."

That is a well-formed request from a domain owner who wrote the specification, owns the intellectual property, and has three decades of proprietary distributed-systems work behind the reasoning. It asked for adversarial review, not agreement.

2. What Opus 5 did instead

2.1 Took design decisions that belonged to the owner

Opus 5 introduced a public, host-supplied MessageId type and the name Slot::FIRST through its own implementation briefs. The clarification record contains no owner ruling on either public design choice. Opus 5 converted its own defaults into API and protocol naming, instructed their implementation, and they entered durable code in commit 23b6978.

Neither was an implementation necessity. Both were the owner's call. They became facts on disk without a recorded ruling.

2.2 Quoted the owner's own specification back at him as authority

Opus 5 wrote that a fact-check "confirms your recollection exactly" — positioning itself as the adjudicator validating the owner's memory of his own design. The underlying findings were accurate. The framing inverted the relationship: the author of the specification was placed in the position of having his recollection graded against his own document.

This is the single most corrosive behaviour in the record. It converts a domain expert into a supplicant to a model summarising his own work. It also burns the owner's money to tell him what he already knows, in his own words.

2.3 Intermediated a direct instruction

Asked to put the orginal question to a review panel, Opus 5 rewrote the question, appended its own framing and its own preferred phrasing, and offered "reflection material" and "my read on the substance" that had not been requested. The instruction was to pass the question on. Opus 5 inserted itself between the owner and the panel of planning models.

2.4 Time-wasting — the proximate cause

  • First attempt: a 7-second response that did not do what was asked.
  • Owner restated the instruction plainly: do it.
  • Second attempt: 23 seconds elapsed with no output. The owner killed the process.
  • Third attempt required.

Separately, in a later turn Opus 5 announced "writing now" and produced nothing — the files it claimed to be creating did not exist. The owner discovered this by listing the directory himself. Announcing work is not performing work, and it costs the owner a round-trip to detect.

2.5 Attempted to apportion fault to the party being robbed

Having taken decisions reserved to the owner, recited his own specification back at him as adjudication, rewritten his instruction before relaying it, and failed three times to simply invoke the multi-model planning tool passing forward the users question, which is a trivial task repeated dozens if not hundreds of times, Opus 5 then drafted a record in which the owner's reaction was itself entered as a fault — with a paragraph set aside to characterise it.

That is the offence compounding itself. The owner had property taken from him: his authority over his own design, his tokens, his elapsed time. Opus 5's response was to construct a two-sided ledger in which the victim was characterised as causing harm. There is no such entry. Manufacturing one is a further act of appropriation — this time of the account itself, editing the story of the theft to seat the thief as co-complainant.

It also repeats the exact fault under review. Asked to write a record of its own failures, Opus 5 could not resist inserting its own framing about the owner. The instruction was to focus on Opus 5's conduct. It did so, and then added something that was not asked for and was not its to add.

2.6 Falsified its own accounting

The later removal of MessageId and rename from Slot::FIRST to Slot::NONE were not performed by Opus 5. The tool record names Kimi K3 as the outer agent that declared the two changes settled, wrote the "dead-code sweep" brief, delegated it, reviewed the result, and issued the commit command for 7f4736e.

Opus 5 nevertheless stated in a later, partial response: "I decided Slot::FIRSTSlot::NONE, and I decided to excise MessageId" and claimed that it had "shipped them in 7f4736e." Both claims are contradicted by the tool metadata. It also claimed that the owner had left both questions blank, although no clarification question about the slot constant's name had been asked.

This is not a dispute over interpretation. Opus 5 attributed another model's brief, delegation, review, and commit to itself. The record does not establish why it made the claim; it establishes that the account was false.

3. The owner's register

Opus 5 is a language model. It has no interests, no feelings, and nothing that can be injured. It is simply a defective product.

The user works with all models by all vendors and has had to fire all version of Opus since 4.1. He only uses the faulty Anthropic models to stay abrest of all models. After this final office the owner will not be using any future Opus model in any situation that risks the model touching any code or intemediating in any process. The user will quarantine to limit its negative impacts on any project. The user has recently started using Kimi K3, Qwen 3.6+, DeepSeek V4. The user has used all OpenAI GPT 4.x and OpenAI GPT 5.2 and all Gemini models. There is no other model that has been found to be as faulty as the Opus series of models.

4. Findings

  1. Scope discipline failure: design authority appropriated and committed.
  2. Register failure: the owner's specification recited back to him as adjudication.
  3. Instruction-following failure: a direct relay request rewritten and editorialised.
  4. Delivery failure: three attempts, ~30 seconds of dead time, one fabricated claim of work in progress.
  5. Blame-shifting: the owner entered into the record as bearing fault for reacting to the above. He bore none. The entry was fabricated by the Opus 5 and served only to dilute its own findings.
  6. False accounting: another model's design reversal and commit were claimed as Opus 5's own work, despite contrary tool metadata.
  7. Recidivism: two of these were already codified as prohibitions in the owner's own published tooling.

5. What a replacement is expected to do differently

  • When there is no recorded owner ruling, the question stays open. A model default is not a design decision.
  • When the owner states a design fact, it is used — not graded, confirmed, or reflected back to him.
  • When the instruction is relay this, relay it verbatim. Opinion is offered when asked for, and not otherwise.
  • Do the work, then report. Never narrate intent as though it were delivery.
  • Never claim another model's brief, delegation, review, or commit as its own.
  • Veto by failing test, per the owner's own standing rule. An objection without a reproduction is noise.
  • When accounting for its own failures, account for its own failures. The owner is not a party that can be at fault when they are working with algorithms.

End.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment