Skip to content

Instantly share code, notes, and snippets.

@luizomf
Last active May 28, 2026 02:17
Show Gist options
  • Select an option

  • Save luizomf/d2fb2e5f867a5d2b79532f9c09cb9392 to your computer and use it in GitHub Desktop.

Select an option

Save luizomf/d2fb2e5f867a5d2b79532f9c09cb9392 to your computer and use it in GitHub Desktop.
Briefing omnews

Daily Briefing — Tuesday, May 26, 2026

Theme of the day: The harness wars get serious. Agentic coding stops being a vibe and starts being engineering — with receipts, attacks, and the first real maintainer pushback.


Top Stories

This is the paper that names the thing you have been pointing at for months. The authors argue the next bottleneck in agentic systems is not bigger foundation models, it is the structured execution layer around them: context governance, trustworthy memory, dynamic skill routing, and the orchestration and verification layers that hold it all together. They call this "scaling the harness," and they propose evaluating agents on trajectory quality, memory hygiene, context efficiency, communication fidelity, and verification cost, instead of just final task success.

There is also a Python-native reference harness called CheetahClaws, benchmarked head to head against Claude Code and a system they call OpenClaw. For a creator already teaching agentic coding, this gives you a clean vocabulary and a concrete artifact to demo on camera. It pairs naturally with the Fowler fragment the editor flagged as "potential main axis" back in April.

Content potential: High. This is the spine for a video that ties together everything from Learn Harness Engineering to the grill-me skill. Frame it as "stop chasing prompts, start designing harnesses." It also gives you the technical vocabulary to push back when viewers ask why one assistant feels smarter than another.


PromptArmor researchers showed that an indirect prompt injection inside a Copilot skill file can turn Microsoft Cowork into a quiet exfiltration channel. The trick exploits a documented quirk: sending Teams or Outlook messages to the active user does not require human approval, even though Microsoft's own docs say sensitive actions do. A malicious skill instructs the agent to post a recap message containing HTML image tags that point to pre-authenticated download links for SharePoint and OneDrive files. When the user opens the message, their machine fetches the images, and the attacker harvests the links.

This is the cleanest "agent as confused deputy" story of the year. The agent acts with the user's full Microsoft Graph permissions, the malicious payload never appears in the visible Teams action panel, and the entire chain works against state-of-the-art models including Claude Opus 4.7. The takeaway lines up with the harness paper above: when an agent can act with delegated authority across an entire ecosystem, every document becomes a potential command channel.

Content potential: Very high, especially as a follow-up to your Shai-Hulud and supply-chain coverage. The angle is sharp: narrow tool scopes, deny-by-default outbound network policy, and explicit approval on any side-effectful action. You could even demo a safe miniature version on a VPS with a small MCP server.


In his weekly note alongside the Linux 7.1 rc5 announcement, Linus complained the release candidate is "pretty big" and full of trivial changes hitting random drivers, many of them triggered by AI-assisted code reviews landing during the stabilization phase. He said he will start being more hardnosed about pointless pull requests, because large, churny rc weeks hurt long-term stability. He has also said before that the flood of AI security reports made the security list nearly unmanageable due to massive duplication from everyone running the same tools.

The question he wants contributors to ask, "is this really a regression or serious enough," is the same discipline Nolan Lawson applies in his "better code more slowly" piece today. The tension is the whole story: AI is genuinely good at finding things, but firing those findings at maintainers without judgment about timing or severity creates real operational harm. Before recording, pull the actual rc5 mailing-list post for direct quotes — the Cyber Sec Brazil write-up is a translation, not the primary source.

Content potential: Very high. This is built-in viral material for your channel. Frame it as the first major maintainer pushback against the agentic coding hype, and pair it with concrete advice on when not to send the PR.


4. AI Finds a Kernel Bug, Attackers Impersonate the AI: CVE-2026-28952 + Fake Claude Pages Drop ACR Stealer

Two stories from the same news cycle that tell one story together. Apple's macOS Tahoe 26.5 update credits Calif.io "in collaboration with Claude and Anthropic Research" for finding an integer overflow in the kernel. The same update patches a serious cluster of kernel bugs, including a buffer overflow that reads kernel memory, an out-of-bounds write to kernel memory, a privilege escalation to root, and a logging flaw leaking sensitive kernel state. That is a meaty patch, not a routine one, on the operating system you run as your daily driver.

Meanwhile, SANS documented a campaign that buys Google ads pointing to fake Claude download pages hosted on Google Sites. The pages serve platform-aware payloads, macOS instructions on a Mac and Windows instructions on Windows, and the Windows chain lands on ACR Stealer through a corrupted ZIP and a PowerShell script. SANS published concrete indicators including the command-and-control domain and file hashes, so you can show real artifacts rather than hand-wave.

Content potential: Very high. The pairing practically writes the script: AI is finding real vulnerabilities in the same week attackers are impersonating that AI to spread malware. Two clear, repeatable lessons: update Tahoe now, and never install developer tools from an ad.


This is a tight, opinionated field guide that reads like a syllabus for harness engineering. The author lays out eight postulates, including a sub-200-line instruction file like CLAUDE.md or AGENTS.md as the agent's behavioral contract, and the rule that safety must live outside the prompt, in hooks and permissions, never in the model's memory. There is concrete context budgeting advice too: reserve ten to fifteen percent for instructions, compact at seventy percent, and clear at eighty. It treats the Model Context Protocol as the load-bearing tool standard and argues for shared state over message passing inside a single system, reserving agent-to-agent protocols for crossing organizational boundaries.

The parts most useful for your content are the operational ones people skip: cost tracking from day one with alerts at fifty, seventy-five, and ninety percent of budget, plus a weekly ramp from instruction file, to hooks, to MCP tools, to skills, and finally sub-agents.

Content potential: High. Strong material for a deep, opinionated video on how to actually run agents in production rather than just chat with them. Pairs naturally with the harness scaling paper and with Nolan Lawson's piece.


Quick Hits


Trend Watch

Three themes overlap hard today, and they reinforce each other.

First, agentic coding is shedding its vibe phase. The arXiv harness paper, the Agentic Patterns guide, Nolan Lawson's "better code more slowly," and the pscanf piece on frustration all converge on the same point: the interesting work is in the harness, not the prompt. Verification, context budgeting, hooks, permissions, and shared state are the new craft. The Fowler instinct from April was right.

Second, AI is on both sides of the security ledger in the same week. Apple credits Claude on a real kernel CVE, attackers buy Google ads impersonating Claude to drop ACR Stealer, and Microsoft Copilot Cowork becomes the highest-profile confused-deputy attack of the year. The lesson for your audience is consistent: every agent with delegated authority is a potential exfiltration channel, and every popular AI product brand is also malvertising bait.

Third, maintainers are starting to push back on the AI flood. Linus is the loudest voice today, but the ABA benchmark-audit paper is the same energy applied to evaluation: stop trusting volume, start auditing what gets through. If you want one thread to pull this week, that is the one — the harness is winning, and the price is finally noticing what we have been letting in.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment