Theme of the day: The harness wars get serious. Agentic coding stops being a vibe and starts being engineering — with receipts, attacks, and the first real maintainer pushback.
This is the paper that names the thing you have been pointing at for months. The authors argue the next bottleneck in agentic systems is not bigger foundation models, it is the structured execution layer around them: context governance, trustworthy memory, dynamic skill routing, and the orchestration and verification layers that hold it all together. They call this "scaling the harness," and they propose evaluating agents on trajectory quality, memory hygiene, context efficiency, communication fidelity, and verification cost, instead of just final task success.
There is also a Python-native reference harness called CheetahClaws, benchmarked head to head against Claude Code and a system they call OpenClaw. For a creator already teaching agentic coding, this gives you a clean vocabulary and a concrete artifact to demo on camera. It pairs naturally with the Fowler fragment the editor flagged as "potential main axis" back in April.
Content potential: High. This is the spine for a video that ties together everything from Learn Harness Engineering to the grill-me skill. Frame it as "stop chasing prompts, start designing harnesses." It also gives you the technical vocabulary to push back when viewers ask why one assistant feels smarter than another.
PromptArmor researchers showed that an indirect prompt injection inside a Copilot skill file can turn Microsoft Cowork into a quiet exfiltration channel. The trick exploits a documented quirk: sending Teams or Outlook messages to the active user does not require human approval, even though Microsoft's own docs say sensitive actions do. A malicious skill instructs the agent to post a recap message containing HTML image tags that point to pre-authenticated download links for SharePoint and OneDrive files. When the user opens the message, their machine fetches the images, and the attacker harvests the links.
This is the cleanest "agent as confused deputy" story of the year. The agent acts with the user's full Microsoft Graph permissions, the malicious payload never appears in the visible Teams action panel, and the entire chain works against state-of-the-art models including Claude Opus 4.7. The takeaway lines up with the harness paper above: when an agent can act with delegated authority across an entire ecosystem, every document becomes a potential command channel.
Content potential: Very high, especially as a follow-up to your Shai-Hulud and supply-chain coverage. The angle is sharp: narrow tool scopes, deny-by-default outbound network policy, and explicit approval on any side-effectful action. You could even demo a safe miniature version on a VPS with a small MCP server.
In his weekly note alongside the Linux 7.1 rc5 announcement, Linus complained the release candidate is "pretty big" and full of trivial changes hitting random drivers, many of them triggered by AI-assisted code reviews landing during the stabilization phase. He said he will start being more hardnosed about pointless pull requests, because large, churny rc weeks hurt long-term stability. He has also said before that the flood of AI security reports made the security list nearly unmanageable due to massive duplication from everyone running the same tools.
The question he wants contributors to ask, "is this really a regression or serious enough," is the same discipline Nolan Lawson applies in his "better code more slowly" piece today. The tension is the whole story: AI is genuinely good at finding things, but firing those findings at maintainers without judgment about timing or severity creates real operational harm. Before recording, pull the actual rc5 mailing-list post for direct quotes — the Cyber Sec Brazil write-up is a translation, not the primary source.
Content potential: Very high. This is built-in viral material for your channel. Frame it as the first major maintainer pushback against the agentic coding hype, and pair it with concrete advice on when not to send the PR.
4. AI Finds a Kernel Bug, Attackers Impersonate the AI: CVE-2026-28952 + Fake Claude Pages Drop ACR Stealer
Two stories from the same news cycle that tell one story together. Apple's macOS Tahoe 26.5 update credits Calif.io "in collaboration with Claude and Anthropic Research" for finding an integer overflow in the kernel. The same update patches a serious cluster of kernel bugs, including a buffer overflow that reads kernel memory, an out-of-bounds write to kernel memory, a privilege escalation to root, and a logging flaw leaking sensitive kernel state. That is a meaty patch, not a routine one, on the operating system you run as your daily driver.
Meanwhile, SANS documented a campaign that buys Google ads pointing to fake Claude download pages hosted on Google Sites. The pages serve platform-aware payloads, macOS instructions on a Mac and Windows instructions on Windows, and the Windows chain lands on ACR Stealer through a corrupted ZIP and a PowerShell script. SANS published concrete indicators including the command-and-control domain and file hashes, so you can show real artifacts rather than hand-wave.
Content potential: Very high. The pairing practically writes the script: AI is finding real vulnerabilities in the same week attackers are impersonating that AI to spread malware. Two clear, repeatable lessons: update Tahoe now, and never install developer tools from an ad.
This is a tight, opinionated field guide that reads like a syllabus for harness engineering. The author lays out eight postulates, including a sub-200-line instruction file like CLAUDE.md or AGENTS.md as the agent's behavioral contract, and the rule that safety must live outside the prompt, in hooks and permissions, never in the model's memory. There is concrete context budgeting advice too: reserve ten to fifteen percent for instructions, compact at seventy percent, and clear at eighty. It treats the Model Context Protocol as the load-bearing tool standard and argues for shared state over message passing inside a single system, reserving agent-to-agent protocols for crossing organizational boundaries.
The parts most useful for your content are the operational ones people skip: cost tracking from day one with alerts at fifty, seventy-five, and ninety percent of budget, plus a weekly ramp from instruction file, to hooks, to MCP tools, to skills, and finally sub-agents.
Content potential: High. Strong material for a deep, opinionated video on how to actually run agents in production rather than just chat with them. Pairs naturally with the harness scaling paper and with Nolan Lawson's piece.
- Automated Benchmark Auditing for AI Agents — Agentic auditor across 168 benchmarks finds critical issues in over twenty-five percent of tasks; fixing them shifts model rankings by nearly ten percent.
- Trustworthy Software Project Generation with an ITP — An agent generates a fully verified RISC-V CPU interpreter in about thirty minutes using Rocq, where the same pipeline with Dafny does not complete.
- How Agentic AI Coding Assistants Become the Attacker's Shell — Formal threat model for why every Claude Code or Codex install is a potential remote shell when it eats untrusted artifacts. Article body was empty, verify before going deep.
- Every Frontier AI Is INTJ — Five hundred ninety-seven of six hundred personality-test runs across Claude, GPT, Gemini, Grok, GLM, and MiniMax come back INTJ. The personality is the product.
- Talk Python: Great Docs — New Python documentation tool from Posit, ships Markdown plus llms.txt files for both humans and agents. Static output, drops cleanly on any VPS.
- Como construí a arquitetura de um SaaS de WhatsApp usando RAG — Brazilian dev walks through a Node-based RAG architecture with WebSocket human handoff. Clean Portuguese case study you can riff on.
- Como separei classificação de reescrita para reduzir alucinação de LLM — Two-stage Haiku-plus-Sonnet pattern that splits "what can I inject" from "how to inject." Solid LangGraph-adjacent pattern.
- Using AI to Write Better Code More Slowly — Nolan Lawson uses a Claude skill that runs Claude, Codex, and Cursor Bugbot in parallel and only fixes the criticals. Harness engineering for code review.
- KnowledgeDeliver LMS Zero-Day Dropping Godzilla and Cobalt Strike — Hard-coded ASP.NET machineKey in a vendor template breaks every install. Sister story to your Laravel-Lang piece.
- CosyEdit2: Speech-Editing-Oriented RL Unlocks Better Zero-Shot TTS — Training for stricter speech-edit consistency improves zero-shot TTS as a side effect. Worth tracking alongside Kokoro and OmniVoice.
- Shard: 10x KV Cache Compression — Drop-in HuggingFace cache compresses Llama 3.1 8B's KV memory roughly ten times with no measurable loss. Perfect for your M4 Max local-LLM rig.
- Per-Interpreter GIL in mod_wsgi 6.0.0 — Graham Dumpleton ships WSGIPerInterpreterGIL, the most concrete proof yet that post-GIL Python is real for production web workloads.
- Agentic Patterns Field Guide — Already covered above as a Top Story.
- Motorola Phones Hijacking the Amazon App for Affiliate Codes — Smart Feed rewrites Amazon links through a redirect chain that injects an influencer's affiliate code. Trust failure on a nineteen-hundred-dollar device.
- Anubis OSS 3.6 with Direct Ollama Model Downloads — Local LLM benchmarking app for Apple Silicon, now installable via Homebrew Cask.
- All Your GUCs in a Row: client_min_messages — Christophe Pettus untangles the Postgres parameter everyone misuses. Five-minute Postgres short.
- Performance of Rust Language (slides) — Honest, evergreen breakdown of where Rust pays a safety tax against C++ and where idiomatic code closes the gap.
- ScreenLeak: Near-Frontier PII Removal at 9 ms CPU — Tiny ONNX models for stripping personal info from computer-use data; even accurate detectors still leak between thirty-six and eighty percent of the time when summarizing screens.
- How Shamir's Secret Sharing Works — Crisp evergreen explainer with a real-world Ente Legacy Kit twist. Information-theoretic, not just computational, security.
- The User Is Visibly Frustrated — Argues that the frustration of coding agents is a UX failure, not a capability failure. Strong companion to your harness coverage.
- Why the Smart Home Bubble Popped — Post-mortem on a decade of IoT, with Home Assistant plus agentic AI as the escape route. Older piece, evergreen framing.
- Sway 1.12 with HDR10 on Vulkan — Wayland color and dynamic range finally catching up. Good data point for the Linux-as-creator-desktop conversation.
- Built LeetCode for Linux (tmpfs.tech) — Real disposable Linux environments, Glicko2 rankings. Made for a hands-on shell short.
- CelerLog: Fast Log Parsing via Dynamic Routing — Routes only semantically tricky log lines to an LLM, the rest go through a fast statistical processor. Eighty to ninety-four percent fewer tokens. Article body was empty, verify before going deep.
- Local Repo/Pkg Caching for Homelabs — Reader prompt for a tutorial on multi-distro mirrors over Tailscale or WireGuard.
- Secvant: Browser File Encryption with WebCrypto — Local-first encryption demo. Honest teaching moment about why client-side crypto only works when the page itself is trustworthy.
Three themes overlap hard today, and they reinforce each other.
First, agentic coding is shedding its vibe phase. The arXiv harness paper, the Agentic Patterns guide, Nolan Lawson's "better code more slowly," and the pscanf piece on frustration all converge on the same point: the interesting work is in the harness, not the prompt. Verification, context budgeting, hooks, permissions, and shared state are the new craft. The Fowler instinct from April was right.
Second, AI is on both sides of the security ledger in the same week. Apple credits Claude on a real kernel CVE, attackers buy Google ads impersonating Claude to drop ACR Stealer, and Microsoft Copilot Cowork becomes the highest-profile confused-deputy attack of the year. The lesson for your audience is consistent: every agent with delegated authority is a potential exfiltration channel, and every popular AI product brand is also malvertising bait.
Third, maintainers are starting to push back on the AI flood. Linus is the loudest voice today, but the ABA benchmark-audit paper is the same energy applied to evaluation: stop trusting volume, start auditing what gets through. If you want one thread to pull this week, that is the one — the harness is winning, and the price is finally noticing what we have been letting in.