Skip to content

Instantly share code, notes, and snippets.

@N8python
Created September 21, 2026 19:59
Show Gist options
  • Select an option

  • Save N8python/0aeb710426dec9cfccbc5097e6e110df to your computer and use it in GitHub Desktop.

Select an option

Save N8python/0aeb710426dec9cfccbc5097e6e110df to your computer and use it in GitHub Desktop.
Astra chess methodology description

Astra chess evaluation: methodology, results, and audit materials

Date: 2026-09-21. Model identifier: gpt-6-astra. Reasoning efforts: low, medium, high, xhigh, and max. This is an exploratory, independently run evaluation, not an official OpenAI or Stockfish benchmark.

We played 40 scored games per effort, 200 games total, against strength-limited Stockfish 17.1. Each game used a fresh Codex CLI session and one continuous agent turn, with chess moves and opponent replies exchanged through an MCP tool. Opponent ratings adapted separately for each effort. A separate preliminary medium-versus-1320 game informed opponent selection but is excluded from the reported estimates.

The resulting numbers measure performance against this particular Stockfish UCI_Elo ladder and harness. They are not calibrated FIDE, USCF, Lichess, or Chess.com ratings.

1. Results

Primary estimates fit all 40 games per effort, fixing color advantage to zero and fitting a draw parameter. The intervals below are individual, approximate 95% profile-likelihood intervals.

Effort Games W–D–L Estimated Elo Approximate 95% interval Mean output tokens/game
low 40 11–1–28 1462 1329–1586 3,240.625
medium 40 15–3–22 1470 1348–1587 3,591.400
high 40 16–1–23 1551 1434–1664 5,616.725
xhigh 40 15–3–22 1726 1606–1843 13,209.650
max 40 15–1–24 1647 1525–1765 18,035.050

Output tokens include reasoning tokens. The token plot uses these means on a logarithmic x-axis, Elo estimates on the y-axis, and these approximate intervals as error bars. Connected points are a visual guide, not a fitted response curve. Tokens/game also depend on game length and cannot isolate the cost of reasoning effort per move.

The same win/draw/loss record can imply different ratings because efforts faced different opponents. Xhigh has the highest point estimate; this alone does not establish that it is stronger than max. We did not perform formal pairwise or multiplicity-adjusted comparisons.

2. What ran, in chronological order

  1. Pilot: one game, medium as White against Stockfish 1320. Astra won by checkmate. This game is excluded from all final likelihoods, but its outcome was incorporated into the frozen planning distribution used to choose opponents.
  2. Initial ten per effort: five sequential color-balanced pairs. Each pair comprised one White and one Black game against the same rating, played concurrently. The next rating was selected after both games finished. Medium was evaluated first; low, high, xhigh, and max subsequently ran in parallel. These 50 games used Stockfish at one second per move on an ARM Mac.
  3. Continuation: after reviewing the initial results, we authorized exactly 30 additional games for each effort. These 150 games ran on two remote DGX Spark machines so laptop sleep would not interrupt them. One host ran low, medium, and max; the other ran high and xhigh. Stockfish used 1,225,000 nodes per move, calibrated to the earlier Mac search work.
  4. Each continuation consisted of five waves of six games, three White and three Black, at one selected opponent rating. All six started together; the next wave waited for the entire current wave. This was a wave scheduler, not a continuously refilled worker queue. Each effort adapted independently, without waiting for other efforts. Maximum configured concurrency was 30 games across the five efforts.
  5. All authorized games and final reports completed. No failed game was replaced with a newly sampled game. Startup configuration failures before scored game play were retained separately, and did not contribute outcomes. The remote services had automatic restarts disabled.

Every effort has 20 White and 20 Black scored games. All games began from the standard initial position; no opening suite or opening randomization was supplied. We did not explicitly fix a model sampling seed or temperature, or an engine random seed. The model identifier was an alias, not a pinned immutable backend checkpoint.

The decision to extend from ten to forty games was made after seeing the first ten results. The entire 40-game study was therefore not preregistered. The 30-game continuation scope and its conservative interval construction were fixed before those new games.

3. Model session and tool restrictions

The scored runs used Codex CLI 0.155.1, authenticated with an existing ChatGPT subscription login. Each game launched a fresh codex exec process with model gpt-6-astra and its named reasoning effort. The launcher ignored user configuration, disabled project instructions, and configured a read-only sandbox. Shell, web search, browser, apps, plugins, memories, hooks, delegation, and unrelated capabilities were disabled in its explicit configuration; the saved run_game.py contains the full settings.

The chess MCP server was the game interface. Code-mode functions.exec and tool discovery were available to reach those tools. The prompt permitted the dispatcher only as transport and prohibited implementing chess analysis code or using external analysis aids. This distinction matters: the dispatcher is a general JavaScript interface, so the prompt restriction should not be described as a mathematical guarantee that computation was impossible. The saved dispatcher audits were inspected; their recorded call patterns were chess-tool discovery and invocation, including storing tool names, rather than chess search programs.

Each player continued through tool results within one ongoing turn. Saved native session metadata was used to verify the model and reasoning effort. The 200 scored games had zero recorded compactions. This documents session continuity and observed compaction behavior; it does not establish what private reasoning state the model backend preserves internally.

The exact scored-game user prompt was the following, replacing White with Black where appropriate:

Play exactly one complete game of standard chess as White against the opponent supplied by the chess tools. Play to win using your own chess reasoning. Use only the chess MCP tools get_state, play_move, and resign. You MAY use the functions.exec dispatcher and tool discovery to invoke these chess tools. The dispatcher is transport only: do not implement chess analysis code or use any external information, engines, shell, delegation, or analysis aids. Keep your reasoning within this single ongoing task and continue after each tool result until game_over is true. Do not end your response early or ask the user for input. First call get_state, then submit one move at a time with play_move using UCI notation. The tool makes the opponent's reply and returns the board and FEN; neither is analysis. Illegal moves do not advance the game: correct them and continue. Standard claimable draws are automatically claimed; a 120-full-move pilot cap is adjudicated as a draw. You may resign if you judge the position hopeless. When the game ends, give only the result and a brief assessment. Do not write commentary between moves.

The word “pilot” in the move-cap instruction was retained in the scored protocol. The user prompt is not the complete underlying Codex system/developer prompt; those were supplied by the installed client/service.

4. Chess tools and returned information

A Python FastMCP server owned a persistent python-chess board and a separate Stockfish UCI process for each game. The model never supplied or maintained the authoritative board state.

Tool Behavior
get_state() Return the current board, assigned color, and game status.
play_move(move: str) Accept one UCI move such as e2e4 or e7e8q; validate legality; apply it; if the game continues, obtain and apply Stockfish's reply; return the updated state.
resign() End the game as a loss for Astra.

If Astra played Black, Stockfish made the opening move before the first state was returned. Astra was untimed.

State fields were fen, board (an ASCII board), side_to_move, you_are, ply, in_check, game_over, result, termination, and last_moves (up to two move records). Each move record included ply, actor, uci, san, at_unix, and fen_after. Stockfish's record also included engine_seconds. Thus move timestamps and opponent wall time were visible. The state did not expose engine evaluation scores, principal variations, search depth, node counts, a legal-move list, or the selected opponent Elo. Engine search statistics were written to a separate audit log, outside the returned state.

MCP results contained the JSON state as text. Supplying FEN and an ASCII board after every move reduces board-tracking burden. This measures chess play with that scaffolding, not blindfold play or unaided reconstruction of the board from move history.

Illegal or malformed moves did not advance the board: the tool returned an error, attempted move, and unchanged state, and the model could retry. The 200 games contained 47 rejected illegal-move attempts, including 34 in the continuation. These attempts were not automatic forfeits, and their token cost remains included.

Termination used board.outcome(claim_draw=True), so claimable draws were automatically taken according to python-chess semantics, including claims available through an intended move. Resignation was allowed. A still-live position at 240 plies was adjudicated a draw (120_MOVE_PILOT_CAP_ADJUDICATION), after ordinary terminal checks. The saved PGNs and game-result records identify each termination.

5. Stockfish settings and node calibration

Both phases used Stockfish 17.1, UCI_LimitStrength=true, selected UCI_Elo, one thread, and 64 MiB hash per game. The supported UCI_Elo range was 1320–3190. Our selection ladder was the discrete set 1320, 1400, 1500, …, 3100, 3190; this differs from permitting every integer setting.

The initial ten games used a one-second search limit. To avoid changing search work simply by moving to remote machines, we calibrated a fixed node limit:

  • Take 30 saved positions: six from each effort, from its initial games 3 and 4, at approximately 15%, 50%, and 85% of each game's Stockfish-to-move positions.
  • Search each for one second on the Mac, with a fresh game identity to clear hash between samples, at the continuation's first-wave opponent settings (low 1700, medium 1500, high 1600, xhigh 1800, max 1800).
  • Median searched nodes: 1,224,615. Round to the nearest thousand: 1,225,000 nodes/move. Mean: 1,396,163.5; 10th–90th percentiles: 1,108,317.7–1,816,984.5; observed range: 1,070,078–3,539,553.
  • Use this node budget for every continuation game. The remote runs did not revert to a time limit. Engine stopping checks can produce small node-count overshoots.

This approximately matches search work, not each original move or playing strength exactly. Calibration used cold hash while actual games retained their per-game hash, and a single median cannot reproduce position-dependent one-second work. The combined estimates assume sufficient comparability of the two protocols; separate new-only results are provided for that reason.

The two remote hosts used the same compiled ARM Stockfish binary. The Sparks ran the orchestration and opponent engine; they did not host Astra inference locally.

6. Adaptive opponent selection

Each effort had its own planning weights. The frozen initial planning distribution incorporated the common medium pilot win, but after initialization only that effort's results updated its weights. Thus adaptation was independent across efforts, apart from the shared initial design choice.

The planning model was the Davidson win/draw/loss model described below, with a color term. Its finite grid was:

  • Rating R: −200 through 4400 in steps of 50, with weights proportional to a normal density centered at 1900 with SD 700.
  • White advantage h: −80, 0, 80, 160, weighted by a normal density centered at 40 with SD 60.
  • Draw parameter nu: 0, 0.3, 1, 3, with masses 0.2, 0.4, 0.3, 0.1.
  • Multiply these weights by the probability of the observed pilot White win against 1320; normalize.

For each candidate opponent rating, enumerate the nine possible outcomes of a White/Black pair and compute expected posterior variance of R. Choose the rating minimizing that variance, breaking exact ties by the first candidate. After a completed pair, update weights by its joint likelihood.

For the continuation, first replay that effort's original five pairs into the planning distribution. Then choose the next pair-optimal rating, repeat it for three concurrent pairs, wait for all six results, and update with all three pairs. This is a practical adaptation heuristic: it does not optimize the full six-game joint design, the remaining total experiment, or wall-clock utilization. The frozen planning distribution retains a color nuisance even though final inference fixes color advantage to zero. Planning weights are not the final reported posterior or confidence interval.

The saved adaptive_design.py and run_thirty.py implement these rules; their hashes are listed below.

7. Rating model and approximate intervals

For a game against nominal opponent rating s, let R be Astra's rating, h White's advantage, c=+1 when Astra is White and −1 when Black, and nu >= 0 the draw parameter. Define:

z = ln(10) * (R - s + c*h) / 800
D = exp(z) + exp(-z) + nu
P(win)  = exp(z) / D
P(draw) = nu / D
P(loss) = exp(-z) / D

When nu=0, the win probability is the ordinary base-10 Elo logistic formula. With draws, the model treats wins, draws, and losses separately; we do not simply convert a draw into half a Bernoulli win.

Final inference sets h=0. Each effort is fit separately by maximizing the sum of log probabilities of its outcomes, with the actual selected opponent rating for every game. The pilot is excluded. No planning prior enters this maximum-likelihood fit.

For each candidate R, maximize the likelihood over continuous nu. If no draws occurred, the optimum is nu=0; otherwise solve for the draw parameter whose sum of fitted draw probabilities equals the observed number of draws. Then maximize this profile over R.

The approximate 95% interval retains ratings satisfying:

2 * [maximum log likelihood - profiled log likelihood at R] <= 3.841458820694124

The cutoff is the 95th percentile of chi-square with one degree of freedom. It is an asymptotic approximation, not an exact finite-sample calibration for these adaptively selected 40 games. With only 40 games and few draws, its numerical precision should not be confused with exact 95% coverage. These are individual intervals, not simultaneous bounds across five efforts, and repeated interim intervals are not sequentially valid confidence bounds. Including only games that happen to finish early can additionally bias interim snapshots through completion order.

Analysis decisions and an earlier correction

The initial ten-game analysis allowed a free color-advantage nuisance and used exhaustive enumeration of the five-pair adaptive design (9^5 outcome paths) for calibration over numerical parameter grids. Low lost every Black game, creating separation: its unrestricted fit did not have a finite ordinary MLE. An initially reported finite optimizer output (about 161 Elo) was corrected as a numerical artifact; raw results were preserved.

After seeing the initial results, the user chose the zero-color-advantage assumption. That is a post-hoc choice for the original ten and was fixed before the next thirty. The final 40-game intervals use the approximate method above; they do not inherit the original ten-game exhaustive calibration. Neither the historical reporting correction nor the change in inferential color assumption altered played moves or the frozen opponent-selection policy.

8. Separate conservative intervals for the 30 new games

We also report a separate confidence set based on a likelihood-mixture construction fixed before the continuation. It uses the original ten only to choose fixed alternative-mixture weights, and the thirty new outcomes as the testing data.

Effort New-game MLE New-game approximate 95% New-game conservative 95% New W–D–L
low 1405 1251–1543 1094–1656 8–1–21
medium 1469 1335–1599 1258–1670 12–2–16
high 1544 1413–1671 1336–1743 13–1–16
xhigh 1719 1586–1847 1482–1939 12–1–17
max 1606 1472–1734 1341–1849 12–0–18

The alternative mixture grid uses R=-1000,…,6000 in steps of 25 and nu={0,.03,.1,.3,1,3,10,30}. Its initial weights are proportional to exp(-0.5*((R-1900)/700)^2), equally weighted across draw parameters, then multiplied by the original ten-game zero-color likelihood and normalized. Those weights are held fixed for the continuation.

Let m(new | old) be the weighted average of new-data likelihoods on this grid. For a proposed rating R, compute:

E(R) = m(new | old) / sup_{nu >= 0} L_new(R, nu)

Exclude R when E(R) >= 20; retain it otherwise. The null denominator maximizes over a continuous draw parameter, so the finite alternative grid does not restrict the null hypothesis to that grid.

Under a fixed true zero-color Davidson model (R*,nu*), E(R*) <= m/L_new(R*,nu*). Conditional on the original data, that likelihood ratio has expectation one under the true model with the predictable adaptive opponent policy. Markov's inequality therefore bounds false exclusion by 1/20 = 5%. This yields at least 95% coverage under those assumptions for the fixed 30-new-game experiment, without relying on the chi-square approximation. The mixture can be inefficient, which is one reason the intervals are wider; they also use fewer testing outcomes than the combined intervals.

“Conservative” is a statement about statistical coverage under the assumed model, not protection against wrong engine Elo calibration, time-varying model strength, dependent games, or a false zero-color assumption. These sets are also individual, not jointly 95% across all five efforts. We do not claim that the completion-ordered partial-game displays were confidence sequences.

9. Usage measurement

Usage was extracted from each game's native Codex session token counters and aggregated once per game. Input tokens include reused conversation context, cached input is a subset of input, and uncached input is input minus cached input. Reasoning tokens are a subset of output tokens, not an extra quantity to add to output. Game usage includes tool transport, rejected move attempts, and the final response.

Effort Output tokens, new 30 Included reasoning Uncached input Overlapping batch wall time
low 99,605 58,254 1,687,191 28.5 min
medium 107,533 67,717 1,545,613 31.0 min
high 176,850 138,432 1,753,684 43.0 min
xhigh 370,842 330,591 1,990,583 74.8 min
max 532,923 492,924 2,090,717 85.3 min

The 150 new games used 1,287,753 output tokens, including 1,087,918 reasoning tokens; 9,067,788 uncached input tokens; and 125,552,896 cached input tokens. All 200 scored games used 1,747,738 output tokens, including 1,479,718 reasoning tokens. Monitoring, calibration, report generation, the pilot, and pre-game setup attempts are outside these game totals. Shared account quota percentages do not provide exact game-only billing or a dollar cost.

10. Audit records and implementation provenance

This document describes the completed experiment and its saved implementation. It is a methodology write-up; the underlying game files and executable analysis are not embedded in this Markdown file.

The local experiment archive retains the following materials for deeper audit:

  • All 200 scored outcomes, opponent settings, colors, terminations, move counts, and illegal-move counts, in per-effort combined-results.json files.
  • All 200 complete move sequences in per-effort all-40-games.pgn files.
  • Per-game token accounting and move/result verification in verification.json, plus per-effort estimates in inference.json and the aggregate FINAL-COMPARISON.json.
  • Frozen chess_server.py, run_game.py, adaptive_design.py, and run_thirty.py sources implementing the game tools, launcher, opponent selection, and wave scheduling.
  • Frozen report_thirty.py, implementing both the approximate profile intervals and the conservative likelihood-mixture sets described above.
  • Manifests, exact prompts, CLI events, native-session verification, dispatcher audits, engine search logs, calibration samples, and source snapshots.

An independent numerical reproduction requires the game outcomes and opponent settings, plus the analysis implementation; a move-by-move audit additionally requires the PGNs. This document alone is not a complete reproducibility archive. The pseudocode, parameter grids, likelihood, and numerical thresholds above specify how to reconstruct the analysis, but cannot independently authenticate the model identity, usage, or absence of assistance.

The original analysis environment reported chess 1.11.2, numpy 2.5.3, scipy 1.18.1, and mcp 1.30.0, on Python 3.12. Small cross-version floating-point differences are possible.

Account credentials, native session identifiers, private communications, account quota records, and model reasoning transcripts are not included in this document. Code hashes identify the saved implementation; they are not third-party attestation that it executed.

SHA-256 provenance

The Mac engine binary hash was a4399f5cd064e32f7468b688f2e3a624dba215d64739882d82d841b4cc96602d; the shared remote engine binary hash was 95f03265fc1e854d23114475a7b764d91c9547137660f51050f58d2fa59231e5.

Frozen continuation source SHA-256
chess_server.py 7eea889c93d386dd02658d80e07d95d2768c3c712e6e31249ea4d49d85334a3c
run_game.py 14cbc9034b4c97b15b2a8e7b672147dd088f7db3eebfe440ae7b595c1fa808a4
adaptive_design.py a439377e2b0a11f0d8be5908b61879316e8ba4538f66e7b5d34b66c1a72ff335
run_thirty.py ce99a6cf039315b31c9f811b62de62b1dffe3c96b4c772bc95c80bc6af2820d0
report_thirty.py 5831e7cda018d6133faee4d8ca2a302050088943850cbf066a54fec8a65d4cfe

11. Limits on interpretation

  • This is a small, exploratory comparison, with 40 games per effort and a study extension decided after initial results. Point estimates should not be read as precise universal model ratings.
  • The nominal Stockfish rating scale is accepted as an input. We did not independently calibrate its actual strength under these search budgets or against human players.
  • Zero color advantage, a constant rating and draw parameter per effort, and conditional independence of game outcomes are modeling assumptions. Balanced colors do not prove zero advantage. Repeated starting positions, related prompts, and shared model infrastructure may induce dependence or heterogeneity.
  • The first and second phases differ in host and search stopping rule. The node calibration mitigates, but does not eliminate, this protocol difference. Efforts were assigned to hosts rather than randomized across hosts.
  • Legal-move feedback, automatically claimed draws, the move cap, and permission to resign are parts of the benchmark. Different harness rules could change results.
  • A continuous session avoids starting over on every move, but this experiment does not compare session architectures or expose private backend reasoning state.
  • The comparison varies reasoning effort, not a controlled identical token budget. More output tokens per game can reflect both more thinking and different game lengths. The connecting line in the plot is not evidence of a monotonic causal relationship.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment