Skip to content

Instantly share code, notes, and snippets.

@wengct
Created July 5, 2026 07:42
Show Gist options
  • Select an option

  • Save wengct/f13ee75062313f48ab502ea8f4de27f5 to your computer and use it in GitHub Desktop.

Select an option

Save wengct/f13ee75062313f48ab502ea8f4de27f5 to your computer and use it in GitHub Desktop.
benchmark_2026-07-05T13-37-25

Codex Token 使用量 Benchmark 報告

摘要

  • 產生時間:2026-07-05 13:53:43 CST
  • 專案根目錄:/home/chenting/projects/OW
  • 重跑次數:3
  • Model:gpt-5.4-mini
  • Effort:medium
  • 執行指令:codex exec --model gpt-5.4-mini -c model_reasoning_effort="medium" --sandbox danger-full-access "<effective prompt>"
  • 基礎 Prompt:請列出所有 API 的權限矩陣
  • 每次執行前清理:git reset --hard HEAD && git clean -fdx
  • Subagent 統計:true
  • Report Analysis:true
  • Report Analysis 輸出:/home/chenting/projects/OW/report-analysis.md
  • Code Review:false
  • Code Review 輸出:/home/chenting/projects/OW/code-review.md
  • cloc baseline:without-agents / /home/chenting/projects/OW/wt-without-agents
  • cloc 前清理 baseline:git reset --hard HEAD && git clean -fdx
  • cloc 排除目錄:bin,obj,node_modules,dist,build,.git,.vs,.vscode,coverage,TestResults,packages
  • CLI 顯示 Total 計算公式:(input_tokens - cached_input_tokens) + output_tokens
  • 基準情境:without-agents,平均 CLI 顯示 Total 91,698.7 tokens
  • 平均最省 Token 組合:without-agents,平均 CLI 顯示 Total 91,698.7 tokens
  • 平均最快組合:without-agents,平均耗時 3m 11s
  • Token 數量排名(少 > 多):without-agents(91,698.7 tokens) > agents(154,615.7 tokens) > agents-subagents-skills(265,414.3 tokens) > agents-subagents(335,662.7 tokens)
  • 費用排名(少 > 多):without-agents($0.15518) > agents($0.29053) > agents-subagents-skills($0.43604) > agents-subagents($0.44948)
  • Token 波動排名(低 > 高,依 CLI Total 變異係數):agents(17.5%) > without-agents(36.3%) > agents-subagents(46.8%) > agents-subagents-skills(47.8%)

情境設定

情境 Worktree Prefix Prompt
without-agents wt-without-agents
agents wt-with-agents
agents-subagents wt-with-agents-subagents Use one explorer agent when finding code you need to read and edit.
agents-subagents-skills wt-with-agents-subagents-skills Use one explorer agent when finding code you need to read and edit.

Baseline cloc 統計

  • Baseline 情境:without-agents
  • Baseline 目錄:/home/chenting/projects/OW/wt-without-agents
  • cloc 前清理:git reset --hard HEAD && git clean -fdx
  • 執行指令:cloc . --exclude-dir=bin,obj,node_modules,dist,build,.git,.vs,.vscode,coverage,TestResults,packages
github.com/AlDanial/cloc v 1.98  T=0.84 s (1351.7 files/s, 271002.2 lines/s)
-------------------------------------------------------------------------------
Language                     files          blank        comment           code
-------------------------------------------------------------------------------
C#                             498          15080          14096          74833
TypeScript                     245           5043           5066          34207
Markdown                       139           7368              1          24038
YAML                            20             62             49          14645
Vuejs Component                151           1746             80          14030
JSON                            16              0              0           9053
SQL                             18            428            201           4034
Razor                            7            149             13            804
INI                              1             82              0            307
CSS                              3             42             12            245
MSBuild script                  10             42              1            223
SVG                             13              0              0            125
TOML                             1             17              1             57
JavaScript                       2              2             21             54
Bourne Shell                     1             11              1             40
XML                              2              0              6             25
HTML                             1              0              0             12
Text                             1              1              0              4
-------------------------------------------------------------------------------
SUM:                          1129          30073          19548         176736
-------------------------------------------------------------------------------

平均排名

排名 情境 成功次數 平均耗時 平均費用 USD 約 NT$ 計價層級 平均 CLI 顯示 Total 相對基準 Total 相對基準比例 平均非快取 Input 平均 Output 平均快取 Input 平均快取比例
1 without-agents 3/3 3m 11s $0.15518 NT$5.0 Standard Total: 91,698.7
Main: 91,698.7
Subagents: 0.0
0.0 0.0% 79,442.3 12,256.3 539,221.3 82.4%
2 agents 3/3 4m 39s $0.29053 NT$9.3 Standard Total: 154,615.7
Main: 154,615.7
Subagents: 0.0
+62,917.0 +68.6% 137,219.3 17,396.3 1,457,792.0 90.6%
3 agents-subagents-skills 3/3 4m 10s $0.43604 NT$14.0 Standard Total: 265,414.3
Main: 149,552.0
Subagents: 115,862.3
+173,715.6 +189.4% 242,578.7 22,835.7 2,017,962.7 87.7%
4 agents-subagents 3/3 3m 29s $0.44948 NT$14.4 Standard Total: 335,662.7
Main: 181,760.3
Subagents: 153,902.3
+243,964.0 +266.0% 314,842.7 20,820.0 1,595,392.0 83.4%

波動分析

僅統計成功執行且有 Session ID 的資料。成功次數小於 2 次時,標準差與變異係數顯示為 N/A。

排名 情境 成功次數 CLI Total 平均 CLI Total 標準差 CLI Total 變異係數 CLI Total 最小 CLI Total 最大 CLI Total 全距 耗時標準差 費用標準差 USD
1 agents 3 154,615.7 27,007.3 17.5% 138,975.0 185,801.0 46,826.0 47s $0.07608
2 without-agents 3 91,698.7 33,286.9 36.3% 60,394.0 126,665.0 66,271.0 1m 27s $0.07772
3 agents-subagents 3 335,662.7 156,953.8 46.8% 215,178.0 513,153.0 297,975.0 1m 0s $0.19671
4 agents-subagents-skills 3 265,414.3 126,994.2 47.8% 120,345.0 356,488.0 236,143.0 2m 0s $0.25550

重跑明細

輪次 情境 Worktree Session ID 耗時 費用 USD 計價層級 CLI 顯示 Total 非快取 Input Output 快取 Input Input Reasoning Output 快取比例 Exit Code
1 agents wt-with-agents 019f30c8-17da-71e2-a9f4-13d2222c5df4 4m 19s $0.26930 Standard Total: 138,975
Main: 138,975
Subagents: 0
Subagent Count: 0
122,527 16,448 1,378,560 1,501,087 8,924 91.8% 0
2 agents wt-with-agents 019f30cd-f89b-7063-a297-10d96d9782db 5m 32s $0.37498 Standard Total: 185,801
Main: 185,801
Subagents: 0
Subagent Count: 0
166,078 19,723 2,155,520 2,321,598 11,483 92.8% 0
3 agents wt-with-agents 019f30d3-0ae6-7400-9fc7-e47bbeac9cd0 4m 5s $0.22732 Standard Total: 139,071
Main: 139,071
Subagents: 0
Subagent Count: 0
123,053 16,018 839,296 962,349 10,154 87.2% 0
1 agents-subagents wt-with-agents-subagents 019f30c8-17b1-76f0-b503-d7301318b99d 3m 46s $0.33495 Standard Total: 215,178
Main: 118,991
Subagents: 96,187
Subagent Count: 1
195,841 19,337 1,347,328 1,543,169 6,282 87.3% 0
2 agents-subagents wt-with-agents-subagents 019f30cd-f8b6-7d13-8f7c-ce07f495199d 2m 23s $0.33687 Standard Total: 278,657
Main: 57,197
Subagents: 221,460
Subagent Count: 1
265,406 13,251 1,042,432 1,307,838 3,005 79.7% 0
3 agents-subagents wt-with-agents-subagents 019f30d3-0aeb-7453-8ba3-df90270e829d 4m 19s $0.67662 Standard Total: 513,153
Main: 369,093
Subagents: 144,060
Subagent Count: 1
483,281 29,872 2,396,416 2,879,697 14,966 83.2% 0
1 agents-subagents-skills wt-with-agents-subagents-skills 019f30c8-179b-7680-a5da-e9115751115f 6m 26s $0.69146 Standard Total: 356,488
Main: 170,737
Subagents: 185,751
Subagent Count: 1
316,849 39,639 3,672,576 3,989,425 21,040 92.1% 0
2 agents-subagents-skills wt-with-agents-subagents-skills 019f30cd-f89c-7162-b8b1-d90a4580c8d1 2m 38s $0.18045 Standard Total: 120,345
Main: 78,733
Subagents: 41,612
Subagent Count: 1
109,978 10,367 684,160 794,138 4,078 86.2% 0
3 agents-subagents-skills wt-with-agents-subagents-skills 019f30d3-0ae5-7ba0-87be-765c4d3578b5 3m 27s $0.43622 Standard Total: 319,410
Main: 199,186
Subagents: 120,224
Subagent Count: 1
300,909 18,501 1,697,152 1,998,061 6,852 84.9% 0
1 without-agents wt-without-agents 019f30c8-17a8-7383-9c0f-a774acb8cfa5 3m 51s $0.15720 Standard Total: 88,037
Main: 88,037
Subagents: 0
Subagent Count: 0
72,906 15,131 459,136 532,042 8,549 86.3% 0
2 without-agents wt-without-agents 019f30cd-f8a2-73a0-99d4-9ffbfa207506 4m 11s $0.23187 Standard Total: 126,665
Main: 126,665
Subagents: 0
Subagent Count: 0
110,729 15,936 1,028,096 1,138,825 9,062 90.3% 0
3 without-agents wt-without-agents 019f30d3-0ae6-7be2-9dd0-161d647946fa 1m 31s $0.07646 Standard Total: 60,394
Main: 60,394
Subagents: 0
Subagent Count: 0
54,692 5,702 130,432 185,124 1,021 70.5% 0

Session File 明細

輪次 情境 Session File
1 agents Main
/home/chenting/.codex/sessions/2026/07/05/rollout-2026-07-05T13-37-27-019f30c8-17da-71e2-a9f4-13d2222c5df4.jsonl
2 agents Main
/home/chenting/.codex/sessions/2026/07/05/rollout-2026-07-05T13-43-52-019f30cd-f89b-7063-a297-10d96d9782db.jsonl
3 agents Main
/home/chenting/.codex/sessions/2026/07/05/rollout-2026-07-05T13-49-25-019f30d3-0ae6-7400-9fc7-e47bbeac9cd0.jsonl
1 agents-subagents Main
/home/chenting/.codex/sessions/2026/07/05/rollout-2026-07-05T13-37-27-019f30c8-17b1-76f0-b503-d7301318b99d.jsonl

Subagent 1
Nickname: Helmholtz
Agent ID: 019f30c8-3c1e-77b1-bdbc-df02b7edf78a
File: /home/chenting/.codex/sessions/2026/07/05/rollout-2026-07-05T13-37-36-019f30c8-3c1e-77b1-bdbc-df02b7edf78a.jsonl
2 agents-subagents Main
/home/chenting/.codex/sessions/2026/07/05/rollout-2026-07-05T13-43-52-019f30cd-f8b6-7d13-8f7c-ce07f495199d.jsonl

Subagent 1
Nickname: Sartre
Agent ID: 019f30ce-2d6d-7650-ae07-e1f706e243bd
File: /home/chenting/.codex/sessions/2026/07/05/rollout-2026-07-05T13-44-06-019f30ce-2d6d-7650-ae07-e1f706e243bd.jsonl
3 agents-subagents Main
/home/chenting/.codex/sessions/2026/07/05/rollout-2026-07-05T13-49-25-019f30d3-0aeb-7453-8ba3-df90270e829d.jsonl

Subagent 1
Nickname: Goodall
Agent ID: 019f30d3-2afb-7ba2-8ab0-efcd2c8bdfab
File: /home/chenting/.codex/sessions/2026/07/05/rollout-2026-07-05T13-49-33-019f30d3-2afb-7ba2-8ab0-efcd2c8bdfab.jsonl
1 agents-subagents-skills Main
/home/chenting/.codex/sessions/2026/07/05/rollout-2026-07-05T13-37-27-019f30c8-179b-7680-a5da-e9115751115f.jsonl

Subagent 1
Nickname: Cicero
Agent ID: 019f30c8-4851-7f11-af41-456f3d59eff9
File: /home/chenting/.codex/sessions/2026/07/05/rollout-2026-07-05T13-37-39-019f30c8-4851-7f11-af41-456f3d59eff9.jsonl
2 agents-subagents-skills Main
/home/chenting/.codex/sessions/2026/07/05/rollout-2026-07-05T13-43-52-019f30cd-f89c-7162-b8b1-d90a4580c8d1.jsonl

Subagent 1
Nickname: Carson
Agent ID: 019f30ce-2d57-7f23-a31f-ee3f3101e12b
File: /home/chenting/.codex/sessions/2026/07/05/rollout-2026-07-05T13-44-06-019f30ce-2d57-7f23-a31f-ee3f3101e12b.jsonl
3 agents-subagents-skills Main
/home/chenting/.codex/sessions/2026/07/05/rollout-2026-07-05T13-49-25-019f30d3-0ae5-7ba0-87be-765c4d3578b5.jsonl

Subagent 1
Nickname: Kant
Agent ID: 019f30d3-320f-7391-804c-e4769cc5af10
File: /home/chenting/.codex/sessions/2026/07/05/rollout-2026-07-05T13-49-35-019f30d3-320f-7391-804c-e4769cc5af10.jsonl
1 without-agents Main
/home/chenting/.codex/sessions/2026/07/05/rollout-2026-07-05T13-37-27-019f30c8-17a8-7383-9c0f-a774acb8cfa5.jsonl
2 without-agents Main
/home/chenting/.codex/sessions/2026/07/05/rollout-2026-07-05T13-43-52-019f30cd-f8a2-73a0-99d4-9ffbfa207506.jsonl
3 without-agents Main
/home/chenting/.codex/sessions/2026/07/05/rollout-2026-07-05T13-49-25-019f30d3-0ae6-7be2-9dd0-161d647946fa.jsonl

計算方式

每次 Codex 執行完成後,先從終端機輸出擷取 session id

耗時統計範圍為 codex exec 指令開始到結束,不包含前置 git clean 與後續報表產生。

再依 Session ID 到以下路徑尋找 JSONL 原始事件檔:

~/.codex/sessions/{YYYY}/{MM}/{DD}/rollout-{YYYY}-{MM}-{DD}T{hh}-{mm}-{ss}-{session-id}.jsonl

每份 JSONL 會讀取最後一筆符合以下條件的事件:

.type == "event_msg" && .payload.type == "token_count"

費用計算公式:

cost_usd = 非快取 Input / 1,000,000 * input_rate + 快取 Input / 1,000,000 * cached_rate + Output / 1,000,000 * output_rate
Context Input = 非快取 Input + 快取 Input
GPT-5.4 / GPT-5.5 / Pro:Context Input > 272000 時使用 Long Context 費率
GPT-5.3-Codex / GPT-5.4 mini / GPT-5.4 nano:使用單一 Standard 費率

內建 Standard 費率表(USD / 1M tokens):

Model Short Input Short Cached Short Output Long Input Long Cached Long Output
gpt-5.5 5.00 0.50 30.00 10.00 1.00 45.00
gpt-5.5-pro 30.00 0.00 180.00 60.00 0.00 270.00
gpt-5.4 2.50 0.25 15.00 5.00 0.50 22.50
gpt-5.4-mini 0.75 0.075 4.50 - - -
gpt-5.4-nano 0.20 0.02 1.25 - - -
gpt-5.3-codex 1.75 0.175 14.00 - - -
gpt-5.4-pro 30.00 0.00 180.00 60.00 0.00 270.00

CLI 顯示 Total:

(input_tokens - cached_input_tokens) + output_tokens

波動分析:

CLI Total 標準差 = 同一情境成功重跑資料的樣本標準差
CLI Total 變異係數 = CLI Total 標準差 / CLI Total 平均 * 100%
CLI Total 全距 = CLI Total 最大 - CLI Total 最小
成功次數小於 2 次時,標準差與變異係數顯示為 N/A

Subagent 統計方式:

1. 從主 session JSONL 讀取 function_call_output
2. 解析 output JSON 取得 agent_id 與 nickname
3. 使用 agent_id 尋找 subagent session JSONL
4. Main 與 Subagents 分開計算,再加總為 Total

實際送給 Codex 的 prompt:

{prefix prompt}

{基礎 Prompt}

欄位說明

欄位 說明
輪次 第幾次重跑。
情境 Benchmark 情境名稱,來自 RUNS 的顯示名稱。
Worktree 該情境對應的 worktree 相對路徑。
Prefix Prompt 該情境額外加在基礎 Prompt 前面的提示詞。
Session ID Codex 本次執行產生的主 session id。
Session File 主 session 與 subagent session 對應的 JSONL 原始事件檔。
耗時 codex exec 指令執行耗時。
平均耗時 成功執行的平均 codex exec 耗時。
Input Main + Subagents 的 input_tokens 加總。
快取 Input Main + Subagents 的 cached_input_tokens 加總。
非快取 Input input_tokens - cached_input_tokens
Output Main + Subagents 的 output_tokens 加總。
Reasoning Output Main + Subagents 的 reasoning_output_tokens 加總。
CLI 顯示 Total 非快取 Input + Output,並在同一欄拆分 Total / Main / Subagents。
相對基準 Total 該情境平均 CLI 顯示 Total 減去基準情境平均值。負數代表比基準省。
相對基準比例 相對基準 Total / 基準情境平均 CLI 顯示 Total
Context Window 本次統計 session 中最大的 model_context_window
快取比例 cached_input_tokens / input_tokens * 100
CLI Total 標準差 同一情境成功重跑資料的 CLI 顯示 Total 樣本標準差。
CLI Total 變異係數 CLI Total 標準差 / CLI Total 平均 * 100%,用來比較不同平均規模情境的相對波動。
CLI Total 全距 同一情境成功重跑資料的最大 CLI Total 減最小 CLI Total。
耗時標準差 同一情境成功重跑資料的耗時樣本標準差。
費用標準差 USD 同一情境成功重跑資料的估算費用樣本標準差。
Exit Code codex exec 的結束代碼,0 表示成功。
成功次數 該情境成功取得 Session ID 與 token_count 的次數 / 總重跑次數。
Subagent Count 本次主 session 找到且納入統計的 subagent session 數量。
Baseline cloc 統計 BASELINE_NAME 對應 worktree 先執行 git reset --hard HEAD && git clean -fdx,再執行 cloc . --exclude-dir=bin,obj,node_modules,dist,build,.git,.vs,.vscode,coverage,TestResults,packages 的結果。

Benchmark 報表分析

結論摘要

本次 Benchmark 的結論很明確:

  1. without-agents 是四個情境中最優解,同時拿到最低平均費用與最快平均耗時。
  2. agents 並沒有比基準更省,反而在 token、費用與時間上都明顯增加。
  3. 兩個含 subagents 的情境都會把總 token 拉高到非常明顯的程度,沒有換到對等的效率收益。
  4. skills 在這份報表裡沒有展現出成本優勢,反而讓流程更重、平均耗時也更高。

如果目標是控制 token 與執行成本,這份資料支持的選擇是 without-agents

主要觀察

1. without-agents 在成本與速度上都最好

without-agents 的平均結果為:

  • 平均耗時:3m 11s
  • 平均費用:$0.15518
  • 平均 CLI 顯示 Total:91,698.7 tokens

這是四個情境裡最便宜、也最快的結果。代表在這個 benchmark 題目中,直接由主代理完成工作,比引入額外協作層更有效率。

2. agents 不是更精簡,而是更重

agents 的平均結果為:

  • 平均耗時:4m 39s
  • 平均費用:$0.29053
  • 平均 CLI 顯示 Total:154,615.7 tokens

相較 without-agents

  • Token 增加 68.6%
  • 費用增加約 87.1%
  • 耗時也更長

這代表單純加入 agent 協作,並沒有讓這個任務更省資源。雖然 agents 的波動是四個情境裡最低的,但平均成本仍明顯落後。

3. subagents 會明顯放大總 token

agents-subagents 的平均結果最重:

  • 平均耗時:3m 29s
  • 平均費用:$0.44948
  • 平均 CLI 顯示 Total:335,662.7 tokens

相較 without-agents

  • Token 增加 266.0%
  • 費用增加約 189.4%

這表示子代理雖然把部分工作拆出去,但整體成本並沒有下降,反而因為協調與重複理解上下文,讓總 token 大幅膨脹。

4. skills 沒有抵銷協作成本

agents-subagents-skills 的平均結果為:

  • 平均耗時:4m 10s
  • 平均費用:$0.43604
  • 平均 CLI 顯示 Total:265,414.3 tokens

相較 without-agents

  • Token 增加 189.4%
  • 費用增加約 181.0%

skills 相較純 subagents 雖然有把平均 token 壓下來,也讓費用略低,但仍然遠高於 without-agents。這代表在本題裡,技能化流程沒有帶來可觀的效率紅利。

數據解讀

Token 不等於成本唯一決定因素,但仍是主因

從平均欄位看,四個情境的 token 與費用大致同向變化,說明主要成本驅動仍是上下文規模與推理過程。

幾個值得注意的點:

  • agents 的平均快取比例最高,達 90.6%,但它並不是最便宜。
  • without-agents 的快取比例只有 82.4%,卻依然最省錢。
  • agents-subagents-skills 的快取比例 87.7%,也沒有換到低成本。

結論很直接:快取比例高,不代表總費用一定低。只要非快取 input、總 output 或協作層級增加,整體成本還是會上升。

Subagents 的主要問題不是「有沒有分工」,而是「分工成本有沒有被吃掉」

在含 subagent 的情境裡,主 session 仍然保有很大的 token 量:

  • agents-subagents:Main 181,760.3,Subagents 153,902.3
  • agents-subagents-skills:Main 149,552.0,Subagents 115,862.3

這表示子代理不是把主代理工作完全接走,而是額外新增了一層理解與協調成本。對這類 benchmark 而言,分工沒有換到足夠的效率提升。

波動性方面,agents 最穩,但不是最佳解

波動分析顯示:

  • agents 的 CLI Total 變異係數最低,只有 17.5%
  • without-agents36.3%
  • agents-subagents46.8%
  • agents-subagents-skills47.8%

也就是說,agents 雖然不是最便宜,但它的穩定性最好;without-agents 雖然最省,但波動相對較大。這個差異意味著:

  • 如果你最在乎可預測性,agents 有優勢。
  • 如果你最在乎成本與速度,without-agents 仍然勝出。

基線環境補充

Baseline cloc 顯示這個專案規模不小,總計:

  • 1,129 個檔案
  • 176,736 行 code

主要語言分布集中在:

  • C#74,833
  • TypeScript34,207
  • Markdown24,038
  • YAML14,645
  • Vuejs Component14,030

這代表專案本身就是多語言、跨層級的程式碼庫,理論上確實會讓上下文管理變得重要。不過就這份 benchmark 的結果來看,增加 agent、subagent 或 skills 並沒有帶來更好的 token 效率,反而更容易累積協作開銷。

建議

如果目標是成本最佳化

優先使用 without-agents。它在這份資料中同時滿足:

  • 最低平均費用
  • 最快平均耗時
  • 最低平均 CLI 顯示 Total

如果目標是穩定性

agents 是四個情境中波動最小的選項,適合把「結果一致性」放在比「成本最低」更前面的情境。

如果考慮導入 subagents 或 skills

建議不要把它們當成預設策略。根據這份報表:

  • subagents 會顯著提高 token 與費用
  • skills 沒有抵銷這個成本
  • agents-subagents-skills 雖比純 subagents 好一些,但仍明顯劣於 without-agents

因此,若沒有明確的任務切分需求,增加協作層不值得。

注意事項

本報表仍有幾個限制:

  • 每個情境只重跑 3 次,樣本數偏小
  • 基礎 Prompt 單一,不能代表所有任務型態
  • 這是針對特定專案、特定時間點的結果

因此,這份結論適合用來判斷「這類工作負載下的實際成本策略」,但不應直接視為所有情境都成立的通則。

最終判斷

若以「最少 token、最低費用、最快完成」為主要目標,這份報表支持的結論是:

首選 without-agents,次選 agents;不建議預設使用 agents-subagentsagents-subagents-skills

#!/usr/bin/env bash
set -uo pipefail
if [[ -z "${BASH_VERSION:-}" ]]; then
echo "Please run this script with bash."
echo "Example: bash ./run-codex.sh 3"
exit 1
fi
# =========================
# 執行說明
# =========================
# 可以執行下列語法覆寫腳本內的值
# RERUN_COUNT="1" \
# MODEL="gpt-5.4" \
# EFFORT="low" \
# SANDBOX="workspace-write" \
# INCLUDE_SUBAGENTS="true" \
# SHOW_CODEX_OUTPUT="false" \
# REPORT_ANALYSIS_ENABLED="true" \
# CODE_REVIEW_ENABLED="true" \
# BASELINE_NAME="without-agents" \
# bash ./run-codex.sh
#
# 若都使用預設值,可以直接呼叫
# bash ./run-codex.sh
# =========================
# 使用者設定區
# =========================
# 通常只需要改這一區。
# 建議先確認:PROJECT_DIR、PROMPT、BASELINE_NAME、RUNS。
# PROMPT:所有情境共用的基礎提示詞;一般中文、空白、標點不用跳脫。
# 範例:PROMPT="${PROMPT:-新增「產生到新 SQL Editor」選項,並保留複製到剪貼簿功能。}"
PROMPT="${PROMPT:-$(cat <<'EOF'
請列出所有 API 的權限矩陣
EOF
)}"
# PROJECT_DIR:多個 worktree 的上層目錄。
# 範例:PROJECT_DIR="/home/<username>/projects/token-efficiency-test"
PROJECT_DIR="${PROJECT_DIR:-$(pwd)}"
# RERUN_COUNT:每個情境重複執行幾次;同一輪的 RUNS 會同時執行。
# 範例:./run-codex.sh 3 # 每個情境跑 3 次
RERUN_COUNT="${1:-${RERUN_COUNT:-1}}"
# SANDBOX:codex exec 的 --sandbox 模式;常用 workspace-write / read-only / danger-full-access。
# 範例:SANDBOX="${SANDBOX:-workspace-write}"
SANDBOX="${SANDBOX:-danger-full-access}"
# MODEL:codex exec 的 --model;所有 benchmark 與 code review 都會使用。
# 範例:MODEL="${MODEL:-gpt-5.4}"
MODEL="${MODEL:-gpt-5.4}"
# EFFORT:codex exec 的 reasoning effort;透過 -c model_reasoning_effort=... 傳入,預設 medium。
# 範例:EFFORT="${EFFORT:-medium}"
EFFORT="${EFFORT:-medium}"
# USD_TO_TWD:報表中台幣換算用匯率;只影響顯示,不影響 USD 成本排序。
# 範例:USD_TO_TWD="${USD_TO_TWD:-32}"
USD_TO_TWD="${USD_TO_TWD:-32}"
# GPT-5.4 / GPT-5.5 Short/Long context 門檻;Input + Cached Input 超過此值改用 Long Context 費率。
# GPT-5.3-Codex 沒有 Short/Long Context 費率分層。
CONTEXT_LONG_THRESHOLD="${CONTEXT_LONG_THRESHOLD:-272000}"
# INCLUDE_SUBAGENTS:是否將主 session 呼叫出的 subagent session token 一起納入統計。
# 範例:INCLUDE_SUBAGENTS="${INCLUDE_SUBAGENTS:-true}"
INCLUDE_SUBAGENTS="${INCLUDE_SUBAGENTS:-true}"
# SHOW_CODEX_OUTPUT:是否顯示 codex exec 原始輸出。
# 範例:SHOW_CODEX_OUTPUT="${SHOW_CODEX_OUTPUT:-false}"
SHOW_CODEX_OUTPUT="${SHOW_CODEX_OUTPUT:-false}"
# CODEX_BIN:Codex CLI 執行檔位置;預設從 PATH 找 codex,找不到時再指定完整路徑。
# 範例:CODEX_BIN="/home/<username>/.nvm/versions/node/v24.11.0/bin/codex"
CODEX_BIN="${CODEX_BIN:-$(command -v codex 2>/dev/null || true)}"
# REPORT_FILE:Markdown benchmark 報告輸出路徑。
# 範例:REPORT_FILE="$PROJECT_DIR/reports/codex-benchmark.md"
REPORT_TIMESTAMP="$(date '+%Y-%m-%dT%H-%M-%S')"
REPORT_FILE="${REPORT_FILE:-$PROJECT_DIR/benchmark_${REPORT_TIMESTAMP}.md}"
# REPORT_ANALYSIS_ENABLED:產生 benchmark 報告後,是否再用 Codex 分析報表並附加回 benchmark 報告末端。
# 範例:REPORT_ANALYSIS_ENABLED="${REPORT_ANALYSIS_ENABLED:-true}"
REPORT_ANALYSIS_ENABLED="${REPORT_ANALYSIS_ENABLED:-true}"
# REPORT_ANALYSIS_FILE:報表分析輸出路徑;預設在執行腳本當下所在目錄產生 report-analysis.md。
# 範例:REPORT_ANALYSIS_FILE="$(pwd)/report-analysis.md"
REPORT_ANALYSIS_FILE="${REPORT_ANALYSIS_FILE:-$(pwd -P)/report-analysis.md}"
# CLOC_EXCLUDE_DIRS:cloc 統計時要排除的目錄;只會針對 BASELINE_NAME 對應的 worktree 執行。
# 範例:CLOC_EXCLUDE_DIRS="bin,obj,node_modules,dist,build,.git,.vs,.vscode,coverage,TestResults,packages"
CLOC_EXCLUDE_DIRS="${CLOC_EXCLUDE_DIRS:-bin,obj,node_modules,dist,build,.git,.vs,.vscode,coverage,TestResults,packages}"
# CODE_REVIEW_ENABLED:四個情境跑完後,是否再執行一次 Codex 比對工作樹未簽入變更。
# 範例:CODE_REVIEW_ENABLED="${CODE_REVIEW_ENABLED:-true}"
CODE_REVIEW_ENABLED="${CODE_REVIEW_ENABLED:-false}"
# CODE_REVIEW_FILE:Code Review 報告輸出路徑;預設在執行腳本當下所在目錄產生 code-review.md。
# 範例:CODE_REVIEW_FILE="$(pwd)/code-review.md"
CODE_REVIEW_FILE="${CODE_REVIEW_FILE:-$(pwd -P)/code-review.md}"
# BASELINE_NAME:比較基準情境;必須等於 RUNS 每列第一欄的「情境名稱」。
# 範例:BASELINE_NAME="without-agents"
BASELINE_NAME="${BASELINE_NAME:-without-agents}"
# RUNS:要比較的 benchmark 情境清單。
# 格式:"情境名稱|worktree 相對路徑|prefix prompt"
# 欄位1 情境名稱:顯示在報表中,也可被 BASELINE_NAME 指定。
# 欄位2 worktree 相對路徑:相對於 PROJECT_DIR,實際目錄為 $PROJECT_DIR/$worktree。
# 欄位3 prefix prompt:此情境專用提示詞,會加在 PROMPT 前面;可留空。
# 注意:prefix prompt 內不要放 |,因為腳本用 | 分隔欄位。
# 範例:"with-agents|wt-with-agents|請優先參考 AGENTS.md 的規則。"
# 快速產生 Worktree 與 Branch (可指定變更集)
# git worktree add ../wt-with-agents -b wt/with-agents {commit}
# git worktree add ../wt-without-agents -b wt/without-agents {commit}
# git worktree add ../wt-with-agents-subagents -b wt/with-agents-subagents {commit}
# git worktree add ../wt-with-agents-subagents-skills -b wt/with-agents-subagents-skills {commit}
# 快速刪除 Worktree 與 Branch
# git worktree remove --force ../wt-with-agents && git branch -D wt/with-agents
# git worktree remove --force ../wt-without-agents && git branch -D wt/without-agents
# git worktree remove --force ../wt-with-agents-subagents && git branch -D wt/with-agents-subagents
# git worktree remove --force ../wt-with-agents-subagents-skills && git branch -D wt/with-agents-subagents-skills
RUNS=(
"without-agents|wt-without-agents|"
"agents|wt-with-agents|"
"agents-subagents|wt-with-agents-subagents|Use one explorer agent when finding code you need to read and edit."
"agents-subagents-skills|wt-with-agents-subagents-skills|Use one explorer agent when finding code you need to read and edit."
)
# =========================
# 程式邏輯區
# =========================
if [[ -z "$CODEX_BIN" ]]; then
echo "codex not found. Please set CODEX_BIN."
exit 1
fi
if [[ ! -d "$PROJECT_DIR" ]]; then
echo "PROJECT_DIR not found: $PROJECT_DIR"
exit 1
fi
if ! [[ "$RERUN_COUNT" =~ ^[1-9][0-9]*$ ]]; then
echo "Usage: $0 <rerun-count>"
echo "Example: $0 3"
exit 1
fi
if [[ -z "$MODEL" ]]; then
echo "MODEL is empty. Please set MODEL."
exit 1
fi
if [[ -z "$EFFORT" ]]; then
echo "EFFORT is empty. Please set EFFORT."
exit 1
fi
if ! awk -v n="$USD_TO_TWD" 'BEGIN { exit !(n + 0 > 0) }'; then
echo "USD_TO_TWD must be a positive number."
exit 1
fi
if ! [[ "$CONTEXT_LONG_THRESHOLD" =~ ^[1-9][0-9]*$ ]]; then
echo "CONTEXT_LONG_THRESHOLD must be a positive integer."
exit 1
fi
if ! command -v jq >/dev/null 2>&1; then
echo "jq not found."
echo "macOS: brew install jq"
echo "WSL/Linux: sudo apt update && sudo apt install -y jq"
exit 1
fi
if ! command -v cloc >/dev/null 2>&1; then
echo "cloc not found. Please install cloc before running this script."
echo "macOS: brew install cloc"
echo "WSL/Linux: sudo apt update && sudo apt install -y cloc"
exit 1
fi
CLOC_BIN="$(command -v cloc)"
SCRIPT_RUN_DIR="$(pwd -P)"
TMP_DIR="$(mktemp -d 2>/dev/null || mktemp -d -t codex-benchmark)"
trap 'rm -rf "$TMP_DIR"' EXIT
ALL_RESULTS="$TMP_DIR/all.tsv"
AVG_RESULTS="$TMP_DIR/avg.tsv"
sanitize_name() {
echo "$1" | tr '/ :' '___'
}
escape_md_cell() {
echo "$1" | sed 's/|/\\|/g'
}
find_session_file() {
local session_id="$1"
find "$HOME/.codex/sessions" \
-type f \
-name "*${session_id}*.jsonl" 2>/dev/null |
while IFS= read -r file; do
if stat -f "%m %N" "$file" >/dev/null 2>&1; then
stat -f "%m %N" "$file" 2>/dev/null
else
stat -c "%Y %n" "$file" 2>/dev/null
fi
done |
sort -nr |
head -n 1 |
cut -d' ' -f2-
}
extract_token_usage() {
local session_file="$1"
jq -s -r '
[.[] | select(.type == "event_msg" and .payload.type == "token_count")]
| last as $e
| if $e == null then
empty
else
$e.payload.info.total_token_usage as $t
| [
($t.input_tokens // 0),
($t.cached_input_tokens // 0),
(($t.input_tokens // 0) - ($t.cached_input_tokens // 0)),
($t.output_tokens // 0),
($t.reasoning_output_tokens // 0),
(($t.input_tokens // 0) - ($t.cached_input_tokens // 0) + ($t.output_tokens // 0)),
($t.total_tokens // 0),
($e.payload.info.model_context_window // 0)
]
| @tsv
end
' "$session_file"
}
sum_token_usage_rows() {
awk -F'\t' '
{
input += $1
cached += $2
noncached += $3
output += $4
reasoning += $5
cli += $6
raw += $7
if ($8 > context) context = $8
}
END {
if (NR == 0) exit 1
printf "%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\n",
input + 0,
cached + 0,
noncached + 0,
output + 0,
reasoning + 0,
cli + 0,
raw + 0,
context + 0
}
'
}
extract_subagent_meta() {
local session_file="$1"
jq -r '
def parse_output:
if type == "string" then
try fromjson catch empty
else
.
end;
(
..
| objects
| select(has("agent_id"))
| [
(.agent_id // ""),
(.nickname // "")
]
| @tsv
),
(
..
| objects
| select(has("output"))
| .output
| parse_output
| select(type == "object" and has("agent_id"))
| [
(.agent_id // ""),
(.nickname // "")
]
| @tsv
)
' "$session_file" |
awk -F'\t' '$1 != ""' |
sort -u
}
write_result() {
local result_file="$1"
local round="$2"
local name="$3"
local worktree="$4"
local status="$5"
local session_id="$6"
local session_files="$7"
local input_tokens="$8"
local cached_input_tokens="$9"
local non_cached_input_tokens="${10}"
local output_tokens="${11}"
local reasoning_output_tokens="${12}"
local cli_display_total="${13}"
local raw_total_tokens="${14}"
local model_context_window="${15}"
local cache_ratio="${16}"
local main_cli_display_total="${17}"
local subagent_cli_display_total="${18}"
local subagent_count="${19}"
local duration_seconds="${20}"
printf "%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\n" \
"$round" \
"$name" \
"$worktree" \
"$status" \
"$session_id" \
"$session_files" \
"$input_tokens" \
"$cached_input_tokens" \
"$non_cached_input_tokens" \
"$output_tokens" \
"$reasoning_output_tokens" \
"$cli_display_total" \
"$raw_total_tokens" \
"$model_context_window" \
"$cache_ratio" \
"$main_cli_display_total" \
"$subagent_cli_display_total" \
"$subagent_count" \
"$duration_seconds" > "$result_file"
}
clean_worktree() {
local name="$1"
local dir="$2"
echo "[$name] clean git worktree"
git -C "$dir" reset --hard HEAD >/dev/null || return 1
git -C "$dir" clean -fdx >/dev/null || return 1
if [[ -n "$(git -C "$dir" status --porcelain)" ]]; then
echo "[$name] worktree is not clean"
git -C "$dir" status --short
return 1
fi
return 0
}
run_codex() {
local round="$1"
local name="$2"
local worktree="$3"
local prefix_prompt="${4:-}"
local dir="$PROJECT_DIR/$worktree"
local safe_name
safe_name="$(sanitize_name "$name")"
local result_file="$TMP_DIR/${round}-${safe_name}.tsv"
(
echo "[round-$round][$name] start: $dir"
if [[ ! -d "$dir/.git" && ! -f "$dir/.git" ]]; then
echo "[round-$round][$name] not a git worktree: $dir"
write_result "$result_file" "$round" "$name" "$worktree" "not_git_worktree" "" "" 0 0 0 0 0 0 0 0 "0.0" 0 0 0 0
exit 1
fi
clean_worktree "$name" "$dir" || {
echo "[round-$round][$name] clean failed"
write_result "$result_file" "$round" "$name" "$worktree" "clean_failed" "" "" 0 0 0 0 0 0 0 0 "0.0" 0 0 0 0
exit 1
}
cd "$dir" || {
echo "[round-$round][$name] cd failed: $dir"
write_result "$result_file" "$round" "$name" "$worktree" "cd_failed" "" "" 0 0 0 0 0 0 0 0 "0.0" 0 0 0 0
exit 1
}
if [[ -n "$prefix_prompt" ]]; then
EFFECTIVE_PROMPT="${prefix_prompt}"$'\n\n'"${PROMPT}"
else
EFFECTIVE_PROMPT="$PROMPT"
fi
START_TS=$(date +%s)
OUTPUT=$("$CODEX_BIN" exec --model "$MODEL" -c model_reasoning_effort="$EFFORT" --sandbox "$SANDBOX" -- "$EFFECTIVE_PROMPT" 2>&1)
STATUS=$?
END_TS=$(date +%s)
DURATION_SECONDS=$((END_TS - START_TS))
if [[ "$SHOW_CODEX_OUTPUT" == "true" ]]; then
echo "$OUTPUT"
fi
SESSION_ID=$(echo "$OUTPUT" | awk -F': ' '/^session id:/ {print $2; exit}')
if [[ -z "$SESSION_ID" ]]; then
echo "[round-$round][$name] session id not found"
write_result "$result_file" "$round" "$name" "$worktree" "$STATUS" "" "" 0 0 0 0 0 0 0 0 "0.0" 0 0 0 "$DURATION_SECONDS"
exit "$STATUS"
fi
SESSION_FILE=$(find_session_file "$SESSION_ID")
if [[ -z "$SESSION_FILE" ]]; then
echo "[round-$round][$name] session file not found: $SESSION_ID"
write_result "$result_file" "$round" "$name" "$worktree" "$STATUS" "$SESSION_ID" "" 0 0 0 0 0 0 0 0 "0.0" 0 0 0 "$DURATION_SECONDS"
exit "$STATUS"
fi
MAIN_TOKEN_DATA=$(extract_token_usage "$SESSION_FILE")
if [[ -z "$MAIN_TOKEN_DATA" ]]; then
echo "[round-$round][$name] token_count not found: $SESSION_FILE"
write_result "$result_file" "$round" "$name" "$worktree" "$STATUS" "$SESSION_ID" "$SESSION_FILE" 0 0 0 0 0 0 0 0 "0.0" 0 0 0 "$DURATION_SECONDS"
exit "$STATUS"
fi
SUBAGENT_TOKEN_ROWS=""
SUBAGENT_SESSION_FILES_TEXT=""
SUBAGENT_COUNT=0
if [[ "$INCLUDE_SUBAGENTS" == "true" ]]; then
echo "[round-$round][$name] scan subagents from main session: $SESSION_FILE"
while IFS=$'\t' read -r AGENT_ID NICKNAME; do
[[ -z "$AGENT_ID" ]] && continue
echo "[round-$round][$name] found subagent: agent_id=$AGENT_ID nickname=${NICKNAME:-N/A}"
AGENT_FILE=$(find_session_file "$AGENT_ID")
if [[ -n "$AGENT_FILE" ]]; then
SUBAGENT_COUNT=$((SUBAGENT_COUNT + 1))
ROW=$(extract_token_usage "$AGENT_FILE" || true)
if [[ -n "$ROW" ]]; then
SUBAGENT_TOKEN_ROWS="${SUBAGENT_TOKEN_ROWS}${ROW}"$'\n'
else
echo "[round-$round][$name] subagent token_count not found: $AGENT_FILE"
fi
if [[ -n "$NICKNAME" ]]; then
SUBAGENT_SESSION_FILES_TEXT="${SUBAGENT_SESSION_FILES_TEXT}<br><br>**Subagent ${SUBAGENT_COUNT}**<br>Nickname: ${NICKNAME}<br>Agent ID: ${AGENT_ID}<br>File: ${AGENT_FILE}"
else
SUBAGENT_SESSION_FILES_TEXT="${SUBAGENT_SESSION_FILES_TEXT}<br><br>**Subagent ${SUBAGENT_COUNT}**<br>Agent ID: ${AGENT_ID}<br>File: ${AGENT_FILE}"
fi
else
echo "[round-$round][$name] subagent session file not found: $AGENT_ID"
fi
done < <(extract_subagent_meta "$SESSION_FILE")
fi
if [[ -n "$SUBAGENT_TOKEN_ROWS" ]]; then
SUBAGENT_TOKEN_DATA=$(printf "%s" "$SUBAGENT_TOKEN_ROWS" | sum_token_usage_rows)
else
SUBAGENT_TOKEN_DATA=$'0\t0\t0\t0\t0\t0\t0\t0'
fi
TOKEN_DATA=$(printf "%s\n%s\n" "$MAIN_TOKEN_DATA" "$SUBAGENT_TOKEN_DATA" | sum_token_usage_rows)
IFS=$'\t' read -r \
INPUT_TOKENS \
CACHED_INPUT_TOKENS \
NON_CACHED_INPUT_TOKENS \
OUTPUT_TOKENS \
REASONING_OUTPUT_TOKENS \
CLI_DISPLAY_TOTAL \
RAW_TOTAL_TOKENS \
MODEL_CONTEXT_WINDOW <<< "$TOKEN_DATA"
IFS=$'\t' read -r \
MAIN_INPUT_TOKENS \
MAIN_CACHED_INPUT_TOKENS \
MAIN_NON_CACHED_INPUT_TOKENS \
MAIN_OUTPUT_TOKENS \
MAIN_REASONING_OUTPUT_TOKENS \
MAIN_CLI_DISPLAY_TOTAL \
MAIN_RAW_TOTAL_TOKENS \
MAIN_MODEL_CONTEXT_WINDOW <<< "$MAIN_TOKEN_DATA"
IFS=$'\t' read -r \
SUBAGENT_INPUT_TOKENS \
SUBAGENT_CACHED_INPUT_TOKENS \
SUBAGENT_NON_CACHED_INPUT_TOKENS \
SUBAGENT_OUTPUT_TOKENS \
SUBAGENT_REASONING_OUTPUT_TOKENS \
SUBAGENT_CLI_DISPLAY_TOTAL \
SUBAGENT_RAW_TOTAL_TOKENS \
SUBAGENT_MODEL_CONTEXT_WINDOW <<< "$SUBAGENT_TOKEN_DATA"
CACHE_RATIO=$(awk -v cached="$CACHED_INPUT_TOKENS" -v input="$INPUT_TOKENS" 'BEGIN {
if (input > 0) printf "%.1f", cached / input * 100;
else printf "0.0";
}')
SESSION_FILES_TEXT="**Main**<br>${SESSION_FILE}${SUBAGENT_SESSION_FILES_TEXT}"
write_result \
"$result_file" \
"$round" \
"$name" \
"$worktree" \
"$STATUS" \
"$SESSION_ID" \
"$SESSION_FILES_TEXT" \
"$INPUT_TOKENS" \
"$CACHED_INPUT_TOKENS" \
"$NON_CACHED_INPUT_TOKENS" \
"$OUTPUT_TOKENS" \
"$REASONING_OUTPUT_TOKENS" \
"$CLI_DISPLAY_TOTAL" \
"$RAW_TOTAL_TOKENS" \
"$MODEL_CONTEXT_WINDOW" \
"$CACHE_RATIO" \
"$MAIN_CLI_DISPLAY_TOTAL" \
"$SUBAGENT_CLI_DISPLAY_TOTAL" \
"$SUBAGENT_COUNT" \
"$DURATION_SECONDS"
echo "[round-$round][$name] session id: $SESSION_ID"
echo "[round-$round][$name] duration: ${DURATION_SECONDS}s"
echo "[round-$round][$name] CLI display total: $CLI_DISPLAY_TOTAL"
echo "[round-$round][$name] main CLI display total: $MAIN_CLI_DISPLAY_TOTAL"
echo "[round-$round][$name] subagent CLI display total: $SUBAGENT_CLI_DISPLAY_TOTAL"
echo "[round-$round][$name] subagent count: $SUBAGENT_COUNT"
echo "[round-$round][$name] done, exit code: $STATUS"
exit "$STATUS"
)
}
find_run_worktree_by_name() {
local target_name="$1"
local item
local name
local worktree
local prefix_prompt
for item in "${RUNS[@]}"; do
IFS='|' read -r name worktree prefix_prompt <<< "$item"
if [[ "$name" == "$target_name" ]]; then
echo "$worktree"
return 0
fi
done
return 1
}
run_cloc_baseline() {
local baseline_worktree="$1"
local baseline_dir="$PROJECT_DIR/$baseline_worktree"
local output_file="$2"
local status
echo
echo "========== cloc baseline =========="
echo "[cloc] baseline: $BASELINE_NAME"
echo "[cloc] start: $baseline_dir"
if [[ ! -d "$baseline_dir" ]]; then
echo "[cloc] baseline worktree not found: $baseline_dir"
return 1
fi
if [[ ! -d "$baseline_dir/.git" && ! -f "$baseline_dir/.git" ]]; then
echo "[cloc] baseline is not a git worktree: $baseline_dir"
return 1
fi
echo "[cloc] clean baseline worktree"
git -C "$baseline_dir" reset --hard HEAD >/dev/null || return 1
git -C "$baseline_dir" clean -fdx >/dev/null || return 1
if [[ -n "$(git -C "$baseline_dir" status --porcelain)" ]]; then
echo "[cloc] baseline worktree is not clean"
git -C "$baseline_dir" status --short
return 1
fi
(
cd "$baseline_dir" || exit 1
"$CLOC_BIN" . --exclude-dir="$CLOC_EXCLUDE_DIRS"
) > "$output_file" 2>&1
status=$?
if [[ "$status" -ne 0 ]]; then
echo "[cloc] failed, exit code: $status"
cat "$output_file"
return "$status"
fi
echo "[cloc] done"
return 0
}
build_review_worktree_list() {
local item
local name
local worktree
local prefix_prompt
local dir
for item in "${RUNS[@]}"; do
IFS='|' read -r name worktree prefix_prompt <<< "$item"
dir="$PROJECT_DIR/$worktree"
echo "- 情境:$name"
echo " Worktree 相對路徑:$worktree"
echo " Worktree 絕對路徑:$dir"
if [[ -n "${prefix_prompt:-}" ]]; then
echo " Prefix Prompt:$prefix_prompt"
else
echo " Prefix Prompt:無"
fi
echo
done
}
run_report_analysis() {
if [[ "$REPORT_ANALYSIS_ENABLED" != "true" ]]; then
echo
echo "Report analysis skipped: REPORT_ANALYSIS_ENABLED=$REPORT_ANALYSIS_ENABLED"
return 0
fi
local analysis_dir="$SCRIPT_RUN_DIR"
local analysis_prompt
local status
analysis_prompt="$(cat <<EOF
幫我分析報表,並在目前目錄產生一份正體中文語系 zh-TW 的 report-analysis.md。
分析報告輸出路徑:
$REPORT_ANALYSIS_FILE
請依據這份 benchmark 報表進行分析:
$REPORT_FILE
EOF
)"
echo
echo "========== report analysis =========="
echo "[report-analysis] start: $analysis_dir"
rm -f "$REPORT_ANALYSIS_FILE"
(
cd "$analysis_dir" || exit 1
"$CODEX_BIN" exec --model "$MODEL" -c model_reasoning_effort="$EFFORT" --sandbox "$SANDBOX" --skip-git-repo-check -- "$analysis_prompt"
) > "$TMP_DIR/report-analysis-output.log" 2>&1
status=$?
if [[ "$SHOW_CODEX_OUTPUT" == "true" ]]; then
cat "$TMP_DIR/report-analysis-output.log"
fi
if [[ "$status" -ne 0 ]]; then
echo "[report-analysis] codex exec failed, exit code: $status"
echo
echo "## Report Analysis" >> "$REPORT_FILE"
echo >> "$REPORT_FILE"
echo "> Report Analysis 執行失敗,exit code: \`$status\`" >> "$REPORT_FILE"
return "$status"
fi
if [[ ! -f "$REPORT_ANALYSIS_FILE" ]]; then
echo "[report-analysis] report-analysis.md not found: $REPORT_ANALYSIS_FILE"
echo
echo "## Report Analysis" >> "$REPORT_FILE"
echo >> "$REPORT_FILE"
echo "> Report Analysis 執行完成,但找不到 \`$REPORT_ANALYSIS_FILE\`" >> "$REPORT_FILE"
return 1
fi
{
echo
echo "---"
echo
cat "$REPORT_ANALYSIS_FILE"
} >> "$REPORT_FILE"
echo "[report-analysis] merged into report: $REPORT_FILE"
return 0
}
run_code_review() {
if [[ "$CODE_REVIEW_ENABLED" != "true" ]]; then
echo
echo "Code review skipped: CODE_REVIEW_ENABLED=$CODE_REVIEW_ENABLED"
return 0
fi
local review_dir="$SCRIPT_RUN_DIR"
local review_prompt
local review_worktrees
local output
local status
review_worktrees="$(build_review_worktree_list)"
review_prompt="$(cat <<EOF
幫我比對以下工作樹下的尚未簽入變更,請依據清單中的工作樹進行處理,不要假設固定只有四個工作樹。
工作樹清單:
${review_worktrees}
我需要知道哪個情境的變更品質較優,並依照整體完成度、正確性、可維護性、風險與需求符合度進行評比排序。
報告要求:
- 在目前目錄產生一份正體中文語系 zh-TW 的 code-review.md
- 報告中不要透露專案結構與程式碼細節
- 可以描述差異、優缺點、風險、建議與排名
- 排名請使用 RUNS 裡的「情境名稱」
這是我對這些工作樹要做的事:
${PROMPT}
EOF
)"
echo
echo "========== code review =========="
echo "[code-review] start: $review_dir"
rm -f "$CODE_REVIEW_FILE"
(
cd "$review_dir" || exit 1
"$CODEX_BIN" exec --model "$MODEL" -c model_reasoning_effort="$EFFORT" --sandbox "$SANDBOX" --skip-git-repo-check -- "$review_prompt"
) > "$TMP_DIR/code-review-output.log" 2>&1
status=$?
if [[ "$SHOW_CODEX_OUTPUT" == "true" ]]; then
cat "$TMP_DIR/code-review-output.log"
fi
if [[ "$status" -ne 0 ]]; then
echo "[code-review] codex exec failed, exit code: $status"
echo
echo "## Code Review" >> "$REPORT_FILE"
echo >> "$REPORT_FILE"
echo "> Code Review 執行失敗,exit code: \`$status\`" >> "$REPORT_FILE"
return "$status"
fi
if [[ ! -f "$CODE_REVIEW_FILE" ]]; then
echo "[code-review] code-review.md not found: $CODE_REVIEW_FILE"
echo
echo "## Code Review" >> "$REPORT_FILE"
echo >> "$REPORT_FILE"
echo "> Code Review 執行完成,但找不到 \`$CODE_REVIEW_FILE\`" >> "$REPORT_FILE"
return 1
fi
{
echo
echo "---"
echo
cat "$CODE_REVIEW_FILE"
} >> "$REPORT_FILE"
echo "[code-review] merged into report: $REPORT_FILE"
return 0
}
validate_runs() {
local seen_names=""
for item in "${RUNS[@]}"; do
IFS='|' read -r name worktree prefix_prompt <<< "$item"
if [[ -z "$name" || -z "$worktree" ]]; then
echo "Invalid RUNS item: $item"
echo "Expected format: name|worktree|prefix prompt"
return 1
fi
if echo "$seen_names" | grep -Fxq "$name"; then
echo "Duplicate RUNS name found: $name"
echo "Each RUNS item must have a unique report name."
return 1
fi
seen_names="${seen_names}"$'\n'"${name}"
done
return 0
}
validate_runs || exit 1
BASELINE_WORKTREE="$(find_run_worktree_by_name "$BASELINE_NAME")"
if [[ -z "$BASELINE_WORKTREE" ]]; then
echo "BASELINE_NAME not found in RUNS: $BASELINE_NAME"
exit 1
fi
BASELINE_DIR="$PROJECT_DIR/$BASELINE_WORKTREE"
CLOC_RESULT_FILE="$TMP_DIR/cloc-baseline.txt"
run_cloc_baseline "$BASELINE_WORKTREE" "$CLOC_RESULT_FILE" || exit 1
STATUS=0
for round in $(seq 1 "$RERUN_COUNT"); do
echo
echo "========== round $round / $RERUN_COUNT =========="
PIDS=()
for item in "${RUNS[@]}"; do
IFS='|' read -r name worktree prefix_prompt <<< "$item"
run_codex "$round" "$name" "$worktree" "$prefix_prompt" &
PIDS+=("$!")
done
for pid in "${PIDS[@]}"; do
wait "$pid" || STATUS=1
done
done
RESULT_FILES=("$TMP_DIR"/*.tsv)
if [[ ! -e "${RESULT_FILES[0]}" ]]; then
echo "No benchmark result files found: $TMP_DIR/*.tsv"
exit 1
fi
cat "${RESULT_FILES[@]}" > "$ALL_RESULTS"
format_number_awk='
function comma0(n, s) {
s = sprintf("%.0f", n)
while (s ~ /[0-9][0-9][0-9][0-9]/) {
sub(/[0-9][0-9][0-9]($|,)/, ",&", s)
}
return s
}
function comma1(n, s, a) {
s = sprintf("%.1f", n)
split(s, a, ".")
return comma0(a[1]) "." a[2]
}
function signed1(n) {
if (n > 0) return "+" comma1(n)
if (n < 0) return "-" comma1(-n)
return "0.0"
}
function signed_pct(n) {
if (n > 0) return "+" sprintf("%.1f", n) "%"
if (n < 0) return sprintf("%.1f", n) "%"
return "0.0%"
}
function duration(n, h, m, s) {
n = int(n + 0.5)
h = int(n / 3600)
m = int((n % 3600) / 60)
s = n % 60
if (h > 0) return h "h " m "m " s "s"
if (m > 0) return m "m " s "s"
return s "s"
}
function display_total(total, main, subagent) {
return "Total: " comma1(total) "<br>Main: " comma1(main) "<br>Subagents: " comma1(subagent)
}
function display_total0(total, main, subagent, count) {
return "Total: " comma0(total) "<br>Main: " comma0(main) "<br>Subagents: " comma0(subagent) "<br>Subagent Count: " comma0(count)
}
function model_key(model, m) {
m = tolower(model)
gsub(/_/, "-", m)
if (m == "codex-5.3" || m == "gpt-5.3-codex" || m == "gpt-5-codex" || m == "gpt-5.3 codex") return "gpt-5.3-codex"
if (m ~ /^gpt-5\.5-pro/) return "gpt-5.5-pro"
if (m ~ /^gpt-5\.5/) return "gpt-5.5"
if (m ~ /^gpt-5\.4-pro/) return "gpt-5.4-pro"
if (m ~ /^gpt-5\.4-mini/) return "gpt-5.4-mini"
if (m ~ /^gpt-5\.4-nano/) return "gpt-5.4-nano"
if (m ~ /^gpt-5\.4/) return "gpt-5.4"
return m
}
function rate_input(model, context_input, threshold, key) {
key = model_key(model)
if (key == "gpt-5.5") return context_input > threshold ? 10.00 : 5.00
if (key == "gpt-5.5-pro") return context_input > threshold ? 60.00 : 30.00
if (key == "gpt-5.4") return context_input > threshold ? 5.00 : 2.50
if (key == "gpt-5.4-pro") return context_input > threshold ? 60.00 : 30.00
if (key == "gpt-5.4-mini") return 0.75
if (key == "gpt-5.4-nano") return 0.20
if (key == "gpt-5.3-codex") return 1.75
return context_input > threshold ? 5.00 : 2.50
}
function rate_cached(model, context_input, threshold, key) {
key = model_key(model)
if (key == "gpt-5.5") return context_input > threshold ? 1.00 : 0.50
if (key == "gpt-5.5-pro") return 0.00
if (key == "gpt-5.4") return context_input > threshold ? 0.50 : 0.25
if (key == "gpt-5.4-pro") return 0.00
if (key == "gpt-5.4-mini") return 0.075
if (key == "gpt-5.4-nano") return 0.02
if (key == "gpt-5.3-codex") return 0.175
return context_input > threshold ? 0.50 : 0.25
}
function rate_output(model, context_input, threshold, key) {
key = model_key(model)
if (key == "gpt-5.5") return context_input > threshold ? 45.00 : 30.00
if (key == "gpt-5.5-pro") return context_input > threshold ? 270.00 : 180.00
if (key == "gpt-5.4") return context_input > threshold ? 22.50 : 15.00
if (key == "gpt-5.4-pro") return context_input > threshold ? 270.00 : 180.00
if (key == "gpt-5.4-mini") return 4.50
if (key == "gpt-5.4-nano") return 1.25
if (key == "gpt-5.3-codex") return 14.00
return context_input > threshold ? 22.50 : 15.00
}
function context_tier(model, context_input, threshold, key) {
key = model_key(model)
if (key == "gpt-5.3-codex" || key == "gpt-5.4-mini" || key == "gpt-5.4-nano") return "Standard"
return context_input > threshold ? "Long" : "Short"
}
function cost_usd(model, noncached, cached, output, threshold, context_input) {
context_input = noncached + cached
return (noncached / 1000000 * rate_input(model, context_input, threshold)) + (cached / 1000000 * rate_cached(model, context_input, threshold)) + (output / 1000000 * rate_output(model, context_input, threshold))
}
function usd(n) { return "$" sprintf("%.5f", n) }
function twd(n, rate) { return "NT$" comma1(n * rate) }
'
awk -F'\t' '
{
total[$2]++
worktree[$2] = $3
}
$4 == "0" && $5 != "" {
success[$2]++
input[$2] += $7
cached[$2] += $8
noncached[$2] += $9
output[$2] += $10
reasoning[$2] += $11
cli[$2] += $12
raw[$2] += $13
context[$2] += $14
cache_ratio[$2] += $15
main_cli[$2] += $16
subagent_cli[$2] += $17
subagent_count[$2] += $18
duration_sum[$2] += $19
}
END {
for (name in total) {
c = success[name] + 0
if (c > 0) {
printf "%s\t%s\t%d\t%d\t%.1f\t%.1f\t%.1f\t%.1f\t%.1f\t%.1f\t%.1f\t%.1f\t%.1f\t%.1f\t%.1f\t%.1f\t%.1f\n",
name,
worktree[name],
total[name],
c,
input[name] / c,
cached[name] / c,
noncached[name] / c,
output[name] / c,
reasoning[name] / c,
cli[name] / c,
raw[name] / c,
context[name] / c,
cache_ratio[name] / c,
main_cli[name] / c,
subagent_cli[name] / c,
subagent_count[name] / c,
duration_sum[name] / c
} else {
printf "%s\t%s\t%d\t%d\t0\t0\t0\t0\t0\t0\t0\t0\t0\t0\t0\t0\t0\n",
name,
worktree[name],
total[name],
c
}
}
}
' "$ALL_RESULTS" > "$AVG_RESULTS"
BASELINE_LINE=$(awk -F'\t' -v name="$BASELINE_NAME" '$1 == name { print; exit }' "$AVG_RESULTS")
if [[ -z "$BASELINE_LINE" ]]; then
echo "BASELINE_NAME not found: $BASELINE_NAME"
exit 1
fi
BASELINE_SUCCESS=$(echo "$BASELINE_LINE" | cut -f4)
if [[ "$BASELINE_SUCCESS" == "0" ]]; then
echo "BASELINE_NAME has no successful runs: $BASELINE_NAME"
exit 1
fi
BASELINE_AVG_CLI_DISPLAY_TOTAL=$(echo "$BASELINE_LINE" | cut -f10)
BASELINE_AVG_CLI_DISPLAY_TOTAL_TEXT=$(awk -v n="$BASELINE_AVG_CLI_DISPLAY_TOTAL" "$format_number_awk"' BEGIN { print comma1(n); }')
BEST_LINE=$(awk -F'\t' '$4 > 0 { print }' "$AVG_RESULTS" | sort -t $'\t' -k10,10n | head -n 1)
BEST_NAME=$(echo "$BEST_LINE" | cut -f1)
BEST_AVG_CLI_DISPLAY_TOTAL=$(echo "$BEST_LINE" | cut -f10)
BEST_AVG_CLI_DISPLAY_TOTAL_TEXT=$(awk -v n="$BEST_AVG_CLI_DISPLAY_TOTAL" "$format_number_awk"' BEGIN { print comma1(n); }')
BEST_DURATION_LINE=$(awk -F'\t' '$4 > 0 { print }' "$AVG_RESULTS" | sort -t $'\t' -k17,17n | head -n 1)
BEST_DURATION_NAME=$(echo "$BEST_DURATION_LINE" | cut -f1)
BEST_AVG_DURATION=$(echo "$BEST_DURATION_LINE" | cut -f17)
BEST_AVG_DURATION_TEXT=$(awk -v n="$BEST_AVG_DURATION" "$format_number_awk"' BEGIN { print duration(n); }')
TOKEN_RANKING_TEXT=$(awk -F'\t' '$4 > 0 { print }' "$AVG_RESULTS" | sort -t $'\t' -k10,10n | awk -F'\t' "$format_number_awk"'
BEGIN { text = "" }
{
item = $1 "(" comma1($10) " tokens)"
text = (text == "" ? item : text " > " item)
}
END { print text }
')
COST_RANKING_TEXT=$(awk -F'\t' '$4 > 0 { print }' "$AVG_RESULTS" | awk -F'\t' -v model="$MODEL" -v threshold="$CONTEXT_LONG_THRESHOLD" "$format_number_awk"'
{
cost = cost_usd(model, $7, $6, $8, threshold)
printf "%.10f\t%s\t%s\n", cost, $1, usd(cost)
}
' | sort -t $'\t' -k1,1n | awk -F'\t' '
BEGIN { text = "" }
{
item = $2 "(" $3 ")"
text = (text == "" ? item : text " > " item)
}
END { print text }
')
VOLATILITY_RANKING_TEXT=$(awk -F'\t' -v model="$MODEL" -v threshold="$CONTEXT_LONG_THRESHOLD" "$format_number_awk"'
$4 == "0" && $5 != "" {
name = $2
n[name]++
cli = $12 + 0
cli_sum[name] += cli
cli_sumsq[name] += cli * cli
}
END {
for (name in n) {
if (n[name] < 2) continue
mean = cli_sum[name] / n[name]
variance = (cli_sumsq[name] - cli_sum[name] * cli_sum[name] / n[name]) / (n[name] - 1)
if (variance < 0) variance = 0
sd = sqrt(variance)
cv = mean > 0 ? sd / mean * 100 : 0
printf "%.10f\t%s\t%s%%\n", cv, name, sprintf("%.1f", cv)
}
}
' "$ALL_RESULTS" | sort -t $'\t' -k1,1n | awk -F'\t' '
BEGIN { text = "" }
{
item = $2 "(" $3 ")"
text = (text == "" ? item : text " > " item)
}
END { print text }
')
if [[ -z "$VOLATILITY_RANKING_TEXT" ]]; then
VOLATILITY_RANKING_TEXT="成功重跑次數不足 2 次,無法計算"
fi
{
echo "# Codex Token 使用量 Benchmark 報告"
echo
echo "## 摘要"
echo
echo "- 產生時間:$(date '+%Y-%m-%d %H:%M:%S %Z')"
echo "- 專案根目錄:\`$PROJECT_DIR\`"
echo "- 重跑次數:$RERUN_COUNT"
echo "- Model:\`$MODEL\`"
echo "- Effort:\`$EFFORT\`"
echo "- 執行指令:\`codex exec --model $MODEL -c model_reasoning_effort=\"$EFFORT\" --sandbox $SANDBOX \"<effective prompt>\"\`"
echo "- 基礎 Prompt:\`$PROMPT\`"
echo "- 每次執行前清理:\`git reset --hard HEAD && git clean -fdx\`"
echo "- Subagent 統計:\`$INCLUDE_SUBAGENTS\`"
printf -- '- Report Analysis:`%s`\n' "$REPORT_ANALYSIS_ENABLED"
printf -- '- Report Analysis 輸出:`%s`\n' "$REPORT_ANALYSIS_FILE"
printf -- '- Code Review:`%s`\n' "$CODE_REVIEW_ENABLED"
printf -- '- Code Review 輸出:`%s`\n' "$CODE_REVIEW_FILE"
printf -- '- cloc baseline:`%s` / `%s`\n' "$BASELINE_NAME" "$BASELINE_DIR"
printf '%s\n' '- cloc 前清理 baseline:`git reset --hard HEAD && git clean -fdx`'
printf -- '- cloc 排除目錄:`%s`\n' "$CLOC_EXCLUDE_DIRS"
printf '%s\n' '- CLI 顯示 Total 計算公式:`(input_tokens - cached_input_tokens) + output_tokens`'
echo "- 基準情境:**$BASELINE_NAME**,平均 CLI 顯示 Total **$BASELINE_AVG_CLI_DISPLAY_TOTAL_TEXT tokens**"
echo "- 平均最省 Token 組合:**$BEST_NAME**,平均 CLI 顯示 Total **$BEST_AVG_CLI_DISPLAY_TOTAL_TEXT tokens**"
echo "- 平均最快組合:**$BEST_DURATION_NAME**,平均耗時 **$BEST_AVG_DURATION_TEXT**"
echo "- Token 數量排名(少 > 多):$TOKEN_RANKING_TEXT"
echo "- 費用排名(少 > 多):$COST_RANKING_TEXT"
echo "- Token 波動排名(低 > 高,依 CLI Total 變異係數):$VOLATILITY_RANKING_TEXT"
echo
echo "## 情境設定"
echo
echo "| 情境 | Worktree | Prefix Prompt |"
echo "|---|---|---|"
for item in "${RUNS[@]}"; do
IFS='|' read -r name worktree prefix_prompt <<< "$item"
safe_prefix=$(escape_md_cell "${prefix_prompt:-無}")
safe_name=$(escape_md_cell "$name")
safe_worktree=$(escape_md_cell "$worktree")
echo "| $safe_name | $safe_worktree | $safe_prefix |"
done
echo
echo "## Baseline cloc 統計"
echo
echo "- Baseline 情境:\`$BASELINE_NAME\`"
echo "- Baseline 目錄:\`$BASELINE_DIR\`"
echo "- cloc 前清理:\`git reset --hard HEAD && git clean -fdx\`"
echo "- 執行指令:\`cloc . --exclude-dir=$CLOC_EXCLUDE_DIRS\`"
echo
echo "\`\`\`text"
cat "$CLOC_RESULT_FILE"
echo "\`\`\`"
echo
echo "## 平均排名"
echo
echo "| 排名 | 情境 | 成功次數 | 平均耗時 | 平均費用 USD | 約 NT$ | 計價層級 | 平均 CLI 顯示 Total | 相對基準 Total | 相對基準比例 | 平均非快取 Input | 平均 Output | 平均快取 Input | 平均快取比例 |"
echo "|---:|---|---:|---:|---:|---:|---|---:|---:|---:|---:|---:|---:|---:|"
awk -F'\t' '$4 > 0 { print }' "$AVG_RESULTS" |
sort -t $'\t' -k10,10n |
awk -F'\t' -v baseline="$BASELINE_AVG_CLI_DISPLAY_TOTAL" -v model="$MODEL" -v threshold="$CONTEXT_LONG_THRESHOLD" -v usd_to_twd="$USD_TO_TWD" "$format_number_awk"'
{
diff = $10 - baseline
diff_pct = baseline > 0 ? diff / baseline * 100 : 0
cost = cost_usd(model, $7, $6, $8, threshold)
tier = context_tier(model, $7 + $6, threshold)
printf "| %d | %s | %s/%s | %s | %s | %s | %s | %s | %s | %s | %s | %s | %s | %s%% |\n",
NR,
$1,
$4,
$3,
duration($17),
usd(cost),
twd(cost, usd_to_twd),
tier,
display_total($10, $14, $15),
signed1(diff),
signed_pct(diff_pct),
comma1($7),
comma1($8),
comma1($6),
$13
}
'
echo
echo "## 波動分析"
echo
echo "僅統計成功執行且有 Session ID 的資料。成功次數小於 2 次時,標準差與變異係數顯示為 N/A。"
echo
echo "| 排名 | 情境 | 成功次數 | CLI Total 平均 | CLI Total 標準差 | CLI Total 變異係數 | CLI Total 最小 | CLI Total 最大 | CLI Total 全距 | 耗時標準差 | 費用標準差 USD |"
echo "|---:|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|"
awk -F'\t' -v model="$MODEL" -v threshold="$CONTEXT_LONG_THRESHOLD" "$format_number_awk"'
function sample_sd(sum, sumsq, n, variance) {
if (n < 2) return -1
variance = (sumsq - sum * sum / n) / (n - 1)
if (variance < 0) variance = 0
return sqrt(variance)
}
{
seen[$2] = 1
}
$4 == "0" && $5 != "" {
name = $2
n[name]++
cli = $12 + 0
duration_seconds = $19 + 0
cost = cost_usd(model, $9, $8, $10, threshold)
cli_sum[name] += cli
cli_sumsq[name] += cli * cli
duration_sum[name] += duration_seconds
duration_sumsq[name] += duration_seconds * duration_seconds
cost_sum[name] += cost
cost_sumsq[name] += cost * cost
if (!(name in cli_min) || cli < cli_min[name]) cli_min[name] = cli
if (!(name in cli_max) || cli > cli_max[name]) cli_max[name] = cli
}
END {
for (name in seen) {
count = n[name] + 0
if (count < 2) {
printf "9999999999\t| 0 | %s | %d | N/A | N/A | N/A | N/A | N/A | N/A | N/A | N/A |\n",
name,
count
continue
}
cli_mean = cli_sum[name] / count
cli_sd = sample_sd(cli_sum[name], cli_sumsq[name], count)
cli_cv = cli_mean > 0 ? cli_sd / cli_mean * 100 : 0
cli_range = cli_max[name] - cli_min[name]
duration_sd = sample_sd(duration_sum[name], duration_sumsq[name], count)
cost_sd = sample_sd(cost_sum[name], cost_sumsq[name], count)
printf "%.10f\t| 0 | %s | %d | %s | %s | %.1f%% | %s | %s | %s | %s | %s |\n",
cli_cv,
name,
count,
comma1(cli_mean),
comma1(cli_sd),
cli_cv,
comma1(cli_min[name]),
comma1(cli_max[name]),
comma1(cli_range),
duration(duration_sd),
usd(cost_sd)
}
}
' "$ALL_RESULTS" |
sort -t $'\t' -k1,1n |
awk -F'\t' '{ sub(/^\| [^|]+ \|/, "| " NR " |", $2); print $2 }'
echo
echo "## 重跑明細"
echo
echo "| 輪次 | 情境 | Worktree | Session ID | 耗時 | 費用 USD | 計價層級 | CLI 顯示 Total | 非快取 Input | Output | 快取 Input | Input | Reasoning Output | 快取比例 | Exit Code |"
echo "|---:|---|---|---|---:|---:|---|---:|---:|---:|---:|---:|---:|---:|---:|"
sort -t $'\t' -k2,2 -k1,1n "$ALL_RESULTS" |
awk -F'\t' -v model="$MODEL" -v threshold="$CONTEXT_LONG_THRESHOLD" "$format_number_awk"'
{
cost = cost_usd(model, $9, $8, $10, threshold)
tier = context_tier(model, $9 + $8, threshold)
printf "| %s | %s | %s | `%s` | %s | %s | %s | %s | %s | %s | %s | %s | %s | %s%% | %s |\n",
$1,
$2,
$3,
$5,
duration($19),
usd(cost),
tier,
display_total0($12, $16, $17, $18),
comma0($9),
comma0($10),
comma0($8),
comma0($7),
comma0($11),
$15,
$4
}
'
echo
echo "## Session File 明細"
echo
echo "| 輪次 | 情境 | Session File |"
echo "|---:|---|---|"
sort -t $'\t' -k2,2 -k1,1n "$ALL_RESULTS" |
awk -F'\t' '
{
printf "| %s | %s | %s |\n", $1, $2, $6
}
'
echo
echo "## 計算方式"
echo
echo "每次 Codex 執行完成後,先從終端機輸出擷取 \`session id\`"
echo
echo "耗時統計範圍為 \`codex exec\` 指令開始到結束,不包含前置 git clean 與後續報表產生。"
echo
echo "再依 Session ID 到以下路徑尋找 JSONL 原始事件檔:"
echo
echo "\`\`\`text"
echo "~/.codex/sessions/{YYYY}/{MM}/{DD}/rollout-{YYYY}-{MM}-{DD}T{hh}-{mm}-{ss}-{session-id}.jsonl"
echo "\`\`\`"
echo
echo "每份 JSONL 會讀取最後一筆符合以下條件的事件:"
echo
echo "\`\`\`text"
echo '.type == "event_msg" && .payload.type == "token_count"'
echo "\`\`\`"
echo
echo "費用計算公式:"
echo
echo "\`\`\`text"
echo "cost_usd = 非快取 Input / 1,000,000 * input_rate + 快取 Input / 1,000,000 * cached_rate + Output / 1,000,000 * output_rate"
echo "Context Input = 非快取 Input + 快取 Input"
echo "GPT-5.4 / GPT-5.5 / Pro:Context Input > ${CONTEXT_LONG_THRESHOLD} 時使用 Long Context 費率"
echo "GPT-5.3-Codex / GPT-5.4 mini / GPT-5.4 nano:使用單一 Standard 費率"
echo "\`\`\`"
echo
echo "內建 Standard 費率表(USD / 1M tokens):"
echo
echo "| Model | Short Input | Short Cached | Short Output | Long Input | Long Cached | Long Output |"
echo "|---|---:|---:|---:|---:|---:|---:|"
echo "| gpt-5.5 | 5.00 | 0.50 | 30.00 | 10.00 | 1.00 | 45.00 |"
echo "| gpt-5.5-pro | 30.00 | 0.00 | 180.00 | 60.00 | 0.00 | 270.00 |"
echo "| gpt-5.4 | 2.50 | 0.25 | 15.00 | 5.00 | 0.50 | 22.50 |"
echo "| gpt-5.4-mini | 0.75 | 0.075 | 4.50 | - | - | - |"
echo "| gpt-5.4-nano | 0.20 | 0.02 | 1.25 | - | - | - |"
echo "| gpt-5.3-codex | 1.75 | 0.175 | 14.00 | - | - | - |"
echo "| gpt-5.4-pro | 30.00 | 0.00 | 180.00 | 60.00 | 0.00 | 270.00 |"
echo
echo "CLI 顯示 Total:"
echo
echo "\`\`\`text"
echo "(input_tokens - cached_input_tokens) + output_tokens"
echo "\`\`\`"
echo
echo "波動分析:"
echo
echo "\`\`\`text"
echo "CLI Total 標準差 = 同一情境成功重跑資料的樣本標準差"
echo "CLI Total 變異係數 = CLI Total 標準差 / CLI Total 平均 * 100%"
echo "CLI Total 全距 = CLI Total 最大 - CLI Total 最小"
echo "成功次數小於 2 次時,標準差與變異係數顯示為 N/A"
echo "\`\`\`"
echo
echo "Subagent 統計方式:"
echo
echo "\`\`\`text"
echo "1. 從主 session JSONL 讀取 function_call_output"
echo "2. 解析 output JSON 取得 agent_id 與 nickname"
echo "3. 使用 agent_id 尋找 subagent session JSONL"
echo "4. Main 與 Subagents 分開計算,再加總為 Total"
echo "\`\`\`"
echo
echo "實際送給 Codex 的 prompt:"
echo
echo "\`\`\`text"
echo "{prefix prompt}"
echo
echo "{基礎 Prompt}"
echo "\`\`\`"
echo
echo "## 欄位說明"
echo
echo "| 欄位 | 說明 |"
echo "|---|---|"
echo "| 輪次 | 第幾次重跑。 |"
echo "| 情境 | Benchmark 情境名稱,來自 \`RUNS\` 的顯示名稱。 |"
echo "| Worktree | 該情境對應的 worktree 相對路徑。 |"
echo "| Prefix Prompt | 該情境額外加在基礎 Prompt 前面的提示詞。 |"
echo "| Session ID | Codex 本次執行產生的主 session id。 |"
echo "| Session File | 主 session 與 subagent session 對應的 JSONL 原始事件檔。 |"
echo "| 耗時 | \`codex exec\` 指令執行耗時。 |"
echo "| 平均耗時 | 成功執行的平均 \`codex exec\` 耗時。 |"
echo "| Input | Main + Subagents 的 \`input_tokens\` 加總。 |"
echo "| 快取 Input | Main + Subagents 的 \`cached_input_tokens\` 加總。 |"
echo "| 非快取 Input | \`input_tokens - cached_input_tokens\`。 |"
echo "| Output | Main + Subagents 的 \`output_tokens\` 加總。 |"
echo "| Reasoning Output | Main + Subagents 的 \`reasoning_output_tokens\` 加總。 |"
echo "| CLI 顯示 Total | \`非快取 Input + Output\`,並在同一欄拆分 Total / Main / Subagents。 |"
echo "| 相對基準 Total | 該情境平均 CLI 顯示 Total 減去基準情境平均值。負數代表比基準省。 |"
echo "| 相對基準比例 | \`相對基準 Total / 基準情境平均 CLI 顯示 Total\`。 |"
echo "| Context Window | 本次統計 session 中最大的 \`model_context_window\`。 |"
echo "| 快取比例 | \`cached_input_tokens / input_tokens * 100\`。 |"
echo "| CLI Total 標準差 | 同一情境成功重跑資料的 CLI 顯示 Total 樣本標準差。 |"
echo "| CLI Total 變異係數 | \`CLI Total 標準差 / CLI Total 平均 * 100%\`,用來比較不同平均規模情境的相對波動。 |"
echo "| CLI Total 全距 | 同一情境成功重跑資料的最大 CLI Total 減最小 CLI Total。 |"
echo "| 耗時標準差 | 同一情境成功重跑資料的耗時樣本標準差。 |"
echo "| 費用標準差 USD | 同一情境成功重跑資料的估算費用樣本標準差。 |"
echo "| Exit Code | \`codex exec\` 的結束代碼,\`0\` 表示成功。 |"
echo "| 成功次數 | 該情境成功取得 Session ID 與 token_count 的次數 / 總重跑次數。 |"
echo "| Subagent Count | 本次主 session 找到且納入統計的 subagent session 數量。 |"
echo "| Baseline cloc 統計 | 對 \`BASELINE_NAME\` 對應 worktree 先執行 \`git reset --hard HEAD && git clean -fdx\`,再執行 \`cloc . --exclude-dir=$CLOC_EXCLUDE_DIRS\` 的結果。 |"
} > "$REPORT_FILE"
run_report_analysis || STATUS=1
run_code_review || STATUS=1
echo
echo "Benchmark report generated:"
echo "$REPORT_FILE"
exit "$STATUS"
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment