Skip to content

Instantly share code, notes, and snippets.

@IgorWarzocha
Created July 16, 2026 20:11
Show Gist options
  • Select an option

  • Save IgorWarzocha/d4410d801374323674fee33da34e4ce0 to your computer and use it in GitHub Desktop.

Select an option

Save IgorWarzocha/d4410d801374323674fee33da34e4ce0 to your computer and use it in GitHub Desktop.
Pi explorer benchmark: GPT-5.6 Luna vs Terra across reasoning levels

Pi explorer benchmark

Three repository-exploration tasks, run sequentially with normal Pi skills, extensions, prompt templates, and project instructions enabled. Each run used an ephemeral session and the existing explorer system prompt.

  1. trace a deferred custom tool through its loader, runtime bridge, and subprocess
  2. trace Agent Pages creation from its external boundary to browser-visible state
  3. trace the principal Hermes Agent CLI/model/tool/final-response path

The first comparison covered tasks 1 and 2:

Model / effort Task 1 Task 2 Average Tool calls Output tokens Reasoning tokens
Luna Low 144.8s 57.2s 101.0s 13 4,131 362
Luna High 169.4s 114.1s 141.7s 21 12,946 5,862
Terra Low 76.4s 71.9s 74.1s 17 6,265 1,154
Terra High 129.8s 115.3s 122.6s 15 11,252 3,682

Quality notes

  • Luna Low: good concise TypeScript tracing, but materially missed the vendored Rust/V8 host source on task 1 and called it unavailable.
  • Luna High: corrected that miss, but was much slower and more verbose.
  • Terra Low: complete runtime trace on task 1 and the strongest task 2 answer, including the important distinction that Vite's module registry makes a new page visible independently of index refresh.
  • Terra High: accurate and complete, but did not add a material result over Terra Low.

Verdict

Terra Low is the best explorer default in this sample. It was the fastest aggregate configuration, complete on both tasks, and avoided High's reasoning/output expansion. Luna Low is the cheaper lightweight option, but its task-1 miss and highly variable latency make it a weaker default.

Task 3 follow-up: broad Hermes runtime trace

Hermes' repo-local semantic-grep extension stalled before agent_start when Pi launched from the repository root. The valid comparison therefore launched from $HOME, kept global resources enabled, explicitly appended Hermes' AGENTS.md, and anchored every tool call with the absolute repository path. The two invalid startup attempts are retained as results/*.invalid-hermes-cwd.*.

Model / effort Wall time Tool calls Input Cache read Output Reasoning
Luna Low 121.4s 23 190,897 1,917,440 4,404 530
Luna High 356.1s 35 326,310 5,188,608 11,096 2,870
Terra Low 124.6s 18 194,224 1,613,568 4,639 568
Terra High 148.4s 20 194,915 1,906,176 6,161 1,067

Task 3 narrows the Low-speed result to a tie: Luna Low was 3.1 seconds faster, while Terra Low reached a stronger, equally complete result with five fewer tool calls. Luna High over-explored into a compaction and took nearly six minutes. Terra High remained controlled but did not materially improve Terra Low's answer.

Across all three tasks, Terra Low remains the strongest fixed explorer default: it won the first two tasks decisively, tied Luna Low on the broad third task, and was consistently complete. Luna Low remains a credible shallow/default budget choice when nominal usage cost matters more than occasional misses.

XHigh / Max follow-up

Model / effort Task 1 Task 2 Task 3 Average Tool calls Output Reasoning
Luna XHigh 352.4s 231.8s 347.7s 310.6s 113 33,612 17,301
Luna Max 268.1s 131.9s 357.3s 252.4s 96 31,559 15,293
Terra XHigh 180.9s 188.5s 248.0s 205.8s 59 26,964 12,645
Terra Max 218.6s 158.9s 273.8s 217.1s 73 28,949 13,101

All runs completed within the ten-minute cap. XHigh and Max produced complete answers but no material improvement over Terra Low; they mainly inspected dependency/runtime details beyond the requested principal path. Luna XHigh was the worst overall explorer configuration, making 113 tool calls. Max happened to beat XHigh for Luna while XHigh beat Max for Terra, showing that stochastic tool-plan shape dominated the nominal effort ordering in this small sample.

Terra Low remains the overall winner at 90.9 seconds average. Terra XHigh, the best extreme-effort configuration, was 2.26× slower and used 7.34× as many reasoning tokens.

@ValadaresX

Copy link
Copy Markdown

There aren't many comments here, but probably a lot of people have already seen this article you wrote. It's very good and I've already sent the link to several people. Keep it up. You write well, very useful information. Thank you.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment