Skip to content

Instantly share code, notes, and snippets.

View chunhualiao's full-sized avatar
🎯
Focusing

Chunhua Liao chunhualiao

🎯
Focusing
View GitHub Profile

Grok Heavy

生物学能否教我们如何核查主张?

主张核查看起来像当代媒体问题。它比语言更古老。凡是复制信息的生命系统,凡是必须判断一个信号是食物、威胁还是噪声的生命系统,都已经在会让新闻编辑室脸红的约束下运行过事实核查流水线:能量有限、传感器嘈杂、对手会进化,而且没有中央主编。

第一性原理的表述很简单。主张是关于世界的一条消息。核查是估计这条消息更像忠实报告,还是更像突变、猜测或攻击的过程。生物学已经为基因组、为身体、也为必须依据不完整报告行动的动物群体做过这件事。有用的问题不是自然是否“在事实核查”,而是哪些机制可以迁移,哪些隐喻掩盖的比揭示的更多,以及信息科学里哪些已经有实验证据。

本文从这些机制重建问题。它追溯这一类比如何在冷战时期进入心理学,分子生物学如何把保真变成分层工程,以及人工免疫系统与多智能体检测器的当前工作实际表明了什么。它也说明类比在何处失败。自身免疫不是俏皮警告,而是检测器分不清损伤与差异时必须付出的代价。

in math, there is LEAN . can we have similar things for natural language claims? it would be wonderful if we can have LEAN-like thing for a claim to easily check it.

Opus 5.5 Max:

A "Lean for Natural-Language Claims": What Exists, Where It Breaks, and How to Build One (State of the Art, September 2026)

Partially: the checking of inference from stated premises can already be made Lean-like and is deployed commercially, but the two other things a claim needs, true premises and a faithful formalization, cannot be kernel-checked, so the best achievable artifact is a "claim certificate" that proves validity, records evidence with provenance and calibrated uncertainty, and exposes the formalization for independent audit.

TL;DR

Sihao Huang (黄思豪) is a technologist and policy specialist whose career has moved unusually fast from rockets and quantum hardware to White House AI policy and, as of September 2026, the compute strategy of a frontier lab.

He is now Head of Frontier Compute Strategy at Anthropic, working with co-founder and Chief Compute Officer Tom Brown on infrastructure expansion, industry coalitions, and planning for rapid AI progress.

Brown welcomed him by noting that Huang “loves technology: he’s worked on rocket engines and collects chips/wafers as a hobby,” and calling him “one of the best people I know for thinking about where AI’s headed.”

That pairing is the point of the hire. Anthropic’s bottleneck is no longer only chips and data centers. It is also power, supply chains, export rules, and government relationships. Huang’s résumé sits at that intersection.

So, in the domain of online search, what's the state of the art to ensure the assurance of search results? Like how dependable the information obtained is? And if the search query involves multiple steps or multiple hops, and what is the SOTA to ensure it can really find the information people want.

Grok Heavy

I’ll look up current research and industry practice on search reliability and multi-hop retrieval so the answer reflects the actual state of the art.I have enough on citation failure rates; next I’ll fill in traditional ranking systems and the latest multi-hop / deep-research benchmarks.Assurance in online search now means two different things, and they are not solved by the same machinery.

The first is document-level quality: did the engine surface pages that are relevant, authoritative, and not spam? The second is claim-level dependability: if the system answers rather than just ranks, is each statement supported by the sources it cites, and can a multi-step query actually be completed without lo

So, in the domain of online search, what's the state of the art to ensure the assurance of search results? Like how dependable the information obtained is? And if the search query involves multiple steps or multiple hops, and what is the SOTA to ensure it can really find the information people want.

The state of the art is moving from “retrieve relevant pages and cite them” toward “maintain and verify an evidence-backed research process.” But two very different assurance problems must be separated:

Information dependability: Are the sources credible, and do they actually support the answer?

Search adequacy: Did the system find enough of the right information to satisfy the user’s real question—including all necessary steps, constraints, and potentially contradictory evidence?

As of September 11, 2026, the first problem has relatively concrete implementations, such as claim-level grounding checks. The second is less settled: recent benchmarks explicitly target missing answers, premature stopping, and fail

So what's the state of the art to verify a claim using online search , including all its supportive evidences recursively?

Grok Heavy

I’ll look up current claim-verification systems that search the web and recursively check supporting evidence, not just single-shot fact checks.Recursive claim verification is now a search-and-argument problem, not a “Google the sentence and trust the top hit” problem. The state of the art treats a claim as a tree: split it, search for each leaf, then treat every supporting document as a new claim until the evidence bottoms out in primary sources, independent corroboration, or an explicit “unverifiable.”

That last part is the part most products skip.

What “verify a claim with all supportive evidence, recursively” actually is

find original letter, then critically comment on it using facts and reasoning:

Fields Medal Winners Warn AI Disrupts Math's Core Values Last updated 1 hour ago The joint statement argues that AI companies treat tough math problems as quick benchmarks, often announcing solutions without proper writeups, citations, or isolation of new methods. Signatories like Artur Avila and Manjul Bhargava stress that mathematics thrives on shared ideas through talks and refinement, not 'mass production' that raises attribution and plagiarism issues. They call on the math community, AI firms, and society to address this misalignment, amid debates where critics see protectionism and supporters value human integration of results.

The original

The letter is "A Severe Misalignment of AI in Mathematics," published 11 September 2026. Two canonical locations:

there are already tons of AI agent building systems. what is the big deal of the following on? investigate and generate a detailed comparables reports:

OpenAI Launches Agents API for Easy Cloud Agent Building Last updated 43 minutes ago OpenAI unveiled its Agents API in public beta on September 10, a managed service powered by the same Codex harness behind ChatGPT Work. It takes care of orchestration, long sessions, context, and tools, letting builders focus on unique features and get started in under a minute with their own tools and sandboxes from partners like Modal and Vercel. Early users see it as a game-changer, like AWS for agents, shifting focus to specialized workflows in areas like insurance or security while OpenAI manages the infrastructure.

The “tons of agent systems already exist” observation is correct. The September 10, 2026 Agents API is not another LangGraph. It is OpenAI selling the hosted Codex harness: the production loop that already runs Codex and ChatGPT Work, as a managed cloud

Two Kinds of Editing

What genome engineering and model weight surgery are teaching each other — and what they're both about to run into

A long read. Roughly 30 minutes.


I. Four hours and ten years

@chunhualiao
chunhualiao / AI surpasses Humans.md
Created July 30, 2026 20:50
AI surpasses Humans

Grok

漫长的较量:人工智能如何在各项任务中超越人类——以及未来二十年可能发生什么

七十多年来,人工智能一直以人类表现为衡量标准。这些比较最初只是思想实验和实验室好奇,后来变成公开表演,再变成悄无的技术里程碑,最终演变成感知、语言、策略和科学领域一系列 cascading 的静默革命。模式始终一致:一个曾经被认为需要独特人类判断力的领域,往往比专家预测的更快被专门系统攻克,然后前沿继续向前推进。

本文追溯机器在特定、明确界定的任务上跨越人类水平表现的历史节点。内容仅基于已确立的历史事件、已发表的基准测试和同行评审或广泛报道的结果。随后转向近期未来——大约未来十年和二十年——依据截至2026年中期可见的能力轨迹,以及跟踪这些发展的专家小组中位数预测。

奠基:从迷宫到完美信息游戏