Benchmark: ci-23600665687 / baseline-1
Branch: brian/trie-updates-data-size-metric
Date: 2025-06-24
Workload: 500 blocks, ~1 GGas each, ~12.7K txs each (5,048,389 total)
The InstrumentedStateProvider wraps every MDBX read during EVM
execution. These are the reads that hit PlainAccountState,
PlainStorageState, and Bytecodes tables.
| Read Type | p50 | p90 | p99 | p99.9 | max |
|---|---|---|---|---|---|
| Account | 4.0 µs | 4.0 µs | 4.0 µs | 4.0 µs | 2,803 µs |
| Storage | 7.4 µs | 7.4 µs | 7.4 µs | 7.4 µs | 5,226 µs |
| Code | 2.0 µs | 3.3 µs | 6.5 µs | 372 µs | 1,418 µs |
Account and storage reads are extremely consistent (p50 = p99) with rare outliers at p99.9+. Code reads are fastest (most are cached) but show more variance. The max latencies (~3–5 ms) correspond to MDBX page faults hitting cold storage.
| Read Type | Total Ops | Duration (s) | Avg Latency (µs) | Sustained IOPS |
|---|---|---|---|---|
| Account | 7,383,075 | 42.7 | 5.8 | 172,906 |
| Storage | 17,284,633 | 128.9 | 7.5 | 134,093 |
| Code | 1,328,312 | 4.2 | 3.2 | 316,265 |
| Total | 25,996,020 | 175.8 | 6.8 | 147,873 |
| Read Type | Reads/TX | Time/TX (µs) |
|---|---|---|
| Account | 1.46 | 8.5 |
| Storage | 3.42 | 25.5 |
| Code | 0.26 | 0.8 |
| Total | 5.15 | 34.8 |
Each transaction costs 34.8 µs of state read time, dominated by storage reads (73% of read time) despite account reads being nearly as frequent.
| p50 | p90 | p99 | max | |
|---|---|---|---|---|
| Account read time | 75.7 ms | 100.9 ms | 126.6 ms | 137.8 ms |
| Storage read time | 101.6 ms | 148.7 ms | 267.7 ms | 334.7 ms |
| Code read time | 5.3 ms | 8.1 ms | 11.2 ms | 12.4 ms |
| Total read time | 182.6 ms | 257.7 ms | 405.5 ms | 484.9 ms |
| Execution time | 684 ms | 896 ms | 911 ms | 986 ms |
| Reads as % of execution | 27% | 29% | 45% | 49% |
State reads consume 27–45% of block execution time. The p99 spike to 45% is driven by storage read variance (267ms vs 101ms at p50).
The proof workers (30 storage + 30 account) run pipelined behind EVM
execution. Their cursor reads hit the AccountsTrie, StoragesTrie,
HashedAccounts, and HashedStorages MDBX tables.
| Cursor | Table | Avg Latency (µs/op) |
|---|---|---|
| Trie seek | Account trie | 6.6 |
| Trie seek | Storage tries | 5.6 |
| Hashed (all ops) | Hashed accounts | 1.6 |
| Hashed (all ops) | Hashed storage | 1.6 |
Trie seeks are ~4× slower than hashed cursor ops because trie B-tree nodes are larger and the seek path is deeper.
These are histogram quantiles recorded per worker at shutdown — one observation per worker per block.
Trie cursor seeks:
| Worker Type | p50 | p90 | p99 | max |
|---|---|---|---|---|
| Storage worker | 964 | 3,813 | 4,369 | 4,923 |
| Account worker | 0 (idle) | 4,428 | 5,111 | 5,763 |
Hashed cursor ops (seek + next + is_empty):
| Worker Type | p50 | p90 | p99 | max |
|---|---|---|---|---|
| Storage worker | 4,192 | 12,417 | 14,608 | 16,441 |
| Account worker | 0 (idle) | 8,434 | 9,798 | 11,020 |
Account workers show p50 = 0 because not all workers receive work every block — the load is concentrated on fewer workers.
Per-worker-per-block cursor duration:
| Cursor | Worker Type | p50 | p90 | p99 | max |
|---|---|---|---|---|---|
| Trie | Storage | 4.7 ms | 10.9 ms | 15.3 ms | 19.4 ms |
| Trie | Account | 0 ms | 15.4 ms | 19.8 ms | 30.2 ms |
| Hashed | Storage | 5.7 ms | 9.6 ms | 13.1 ms | 16.9 ms |
| Hashed | Account | 0 ms | 10.1 ms | 13.1 ms | 19.5 ms |
| Cursor | Total Ops | Duration (s) | IOPS (per active worker) |
|---|---|---|---|
| Trie (state) | 44,173,283 | 290.3 | 152,167 |
| Trie (storage) | 64,542,314 | 358.5 | 180,062 |
| Hashed (state) | 82,572,378 | 134.2 | 615,294 |
| Hashed (storage) | 189,047,450 | 297.1 | 636,141 |
| Total | 380,335,425 | 1,080.1 | — |
| Cursor | Ops/TX | Aggregate µs/TX | Wall-clock µs/TX (÷60 workers) |
|---|---|---|---|
| Trie (all) | 21.5 | 128.5 | 2.1 |
| Hashed (all) | 53.8 | 85.4 | 1.4 |
| Total | 75.3 | 213.9 | 3.6 |
Each transaction generates 75 cursor reads across the proof workers, costing 214 µs of aggregate worker time. With 60 workers in parallel, the wall-clock contribution is only 3.6 µs/tx.
| Source | Reads/TX | µs/TX | Runs On | Bottleneck? |
|---|---|---|---|---|
| EVM state reads | 5.15 | 34.8 | Main execution thread | ✅ Yes — serial |
| Proof cursor reads | 75.3 | 3.6 (wall) | 60 parallel workers | No — parallelized |
| Total | 80.5 | 38.4 | — | — |
EVM state reads are the bottleneck because they run serially on the execution thread. Proof worker reads are 14.7× more ops per TX but effectively free due to parallelism.
EVM state reads are serial and directly gate execution throughput. At 34.8 µs/tx, state reads alone limit throughput to:
1 / 34.8 µs = ~28,700 TPS theoretical read ceiling
| TPS | Read Time/s | % of 1 Core | Read IOPS Required | Sustainable? |
|---|---|---|---|---|
| 1,000 | 34.8 ms | 3.5% | 5,150 | ✅ |
| 5,000 | 174 ms | 17.4% | 25,750 | ✅ |
| 10,000 | 348 ms | 34.8% | 51,500 | |
| 15,000 | 522 ms | 52.2% | 77,250 | |
| 20,000 | 696 ms | 69.6% | 103,000 | 🔴 Little room for compute |
| 25,000 | 870 ms | 87.0% | 128,750 | 🔴 Near read ceiling |
| 28,700 | 1,000 ms | 100% | 147,873 | 🔴 Theoretical limit |
| 50,000+ | >1s | >100% | >257K | 🔴 Exceeds single-core |
At the measured 147K sustained IOPS, the read path saturates a single core at ~28.7K TPS. This is the read-only ceiling — EVM compute consumes the remainder of execution time, so the practical limit is lower.
Proof cursor reads scale with TPS but are parallelized across workers. Each TX generates 75.3 cursor ops costing 213.9 µs of aggregate worker time. The question is whether the worker pool can keep up.
| TPS | Aggregate Read Time/s | Workers Needed (100% util) | With 60 Workers |
|---|---|---|---|
| 1,000 | 213.9 ms | 0.2 | ✅ 0.4% utilized |
| 5,000 | 1.07 s | 1.1 | ✅ 1.8% utilized |
| 10,000 | 2.14 s | 2.1 | ✅ 3.6% utilized |
| 50,000 | 10.7 s | 10.7 | ✅ 17.8% utilized |
| 100,000 | 21.4 s | 21.4 | ✅ 35.7% utilized |
| 200,000 | 42.8 s | 42.8 | |
| 280,000 | 59.9 s | 59.9 | 🔴 ~100% utilized |
Proof workers are not the bottleneck. With 60 workers, the cursor read pool can sustain ~280K TPS before saturating. Scaling beyond that just requires more workers.
| Component | Read Ceiling (TPS) | Limiting Factor |
|---|---|---|
| EVM state reads | ~28,700 | Single-threaded MDBX reads at 148K IOPS |
| Proof cursor reads | ~280,000 | 60-worker parallelism |
| Practical limit | ~15,000 | Reads + EVM compute must share one core |
The EVM state read path is the binding constraint. At 15K TPS, reads consume ~52% of the execution core, leaving the other ~48% for actual EVM computation — already tight.
- State caching — the CachedStateProvider already helps; a larger or smarter cache reduces cold reads. The p50 = p99 latency uniformity suggests most reads are already cache-warm at 4–7 µs; the tail (max ~5 ms) indicates occasional page faults
- Parallel EVM execution — running transactions on multiple cores parallelizes reads, effectively multiplying the read ceiling by core count
- Faster storage — the 5.8–7.5 µs average latency is dominated by MDBX B-tree traversal, not disk I/O; an in-memory index or hash table could cut this to <1 µs
- Read-ahead / prefetching — prewarming already does this for known access lists; better prediction reduces cold misses
- Reducing reads per TX — access list preloading or stateless execution could cut the 5.15 reads/tx ratio
| Resource | Per-TX Cost | Ceiling (TPS) | Bottleneck Type |
|---|---|---|---|
| EVM state reads | 34.8 µs | ~28,700 | Single-threaded |
| Trie+hashed writes | 38.8 µs | ~25,800 | Sequential MDBX inserts |
| Proof cursor reads | 3.6 µs (wall) | ~280,000 | Parallelized |
Reads and writes have nearly identical per-TX costs (~35–39 µs) and nearly identical ceilings (~26–29K TPS). Neither dominates — they're balanced bottlenecks. Optimizing only one shifts the wall to the other.
At the current benchmark workload (~1,000 TPS sustained):
- Reads use 3.5% of capacity
- Writes use 3.9% of capacity
- Both scale linearly and hit their respective walls together around 20–25K TPS
| Metric | Scope | What It Measures |
|---|---|---|
sync.state_provider.account_fetch_latency |
per-read | MDBX PlainAccountState read latency |
sync.state_provider.storage_fetch_latency |
per-read | MDBX PlainStorageState read latency |
sync.state_provider.code_fetch_latency |
per-read | MDBX Bytecodes read latency |
sync.state_provider.total_{account,storage,code}_fetch_latency |
per-block | Total read time per block |
trie.cursor.operations{type,operation} |
per-worker-block | Trie cursor op count per worker per block |
trie.cursor.overall_duration{type} |
per-worker-block | Trie cursor total time per worker per block |
trie.hashed_cursor.operations{type,operation} |
per-worker-block | Hashed cursor op count per worker per block |
trie.hashed_cursor.overall_duration{type} |
per-worker-block | Hashed cursor total time per worker per block |