Skip to content

Instantly share code, notes, and snippets.

@mediocregopher
Last active March 26, 2026 16:12
Show Gist options
  • Select an option

  • Save mediocregopher/5b9a6e3e0d907a07e541d9653fda8396 to your computer and use it in GitHub Desktop.

Select an option

Save mediocregopher/5b9a6e3e0d907a07e541d9653fda8396 to your computer and use it in GitHub Desktop.

Reth Trie & Hashed State Write Analysis

Benchmark: ci-23600665687 / baseline-1
Branch: brian/trie-updates-data-size-metric
Date: 2025-06-24
Workload: 500 blocks, ~1 GGas each, ~10–15K txs each


Block Workload Characteristics

p50 p90 p99 max
Gas used 1.015 GGas 1.034 GGas 1.047 GGas 1.050 GGas
Transactions 12,755 14,055 14,520 15,621
Accounts touched 17,036 18,492 19,940 20,827
Storage slots updated 13,563 14,477 15,821 17,240
Bytecodes updated 40 51 75 76

Totals across 500 blocks: 509.9 GGas, 5,048,389 txs.


1. Per-Block Write Sizes

Trie Updates (TrieUpdatesSorted serialized bytes)

p50 p90 p99 max min
Bytes 37.7 MB 42.1 MB 43.1 MB 44.5 MB 33.8 MB
Entries 91,373 101,793 103,829 108,023 80,999
Bytes/entry 412 413 415 412 417

Hashed State (HashedPostStateSorted serialized bytes)

p50 p90 p99 max min
Bytes 2.62 MB 2.70 MB 2.90 MB 2.94 MB 2.33 MB
Entries 30,739 31,898 33,400 34,626 26,444
Bytes/entry 85 85 87 85 88

Combined

p50 p90 p99
Total bytes/block 40.3 MB 44.8 MB 46.0 MB

2. Per-GGas Write Rates

Normalized by dividing per-block write bytes by per-block gas used.

Trie Updates per GGas

p50 p90 p99
Bytes/GGas 37.1 MB 40.7 MB 41.2 MB
Entries/GGas 90,023 98,445 99,167

Hashed State per GGas

p50 p90 p99
Bytes/GGas 2.58 MB 2.61 MB 2.77 MB
Entries/GGas 30,285 30,849 31,904

Combined per GGas

p50 p90 p99
Total bytes/GGas 39.7 MB 43.3 MB 44.0 MB

At today's 36M gas limit (~0.036 GGas), a full mainnet block would write roughly 1.4 MB of trie+hashed state data. At a hypothetical 1 GGas target, expect ~40 MB/block.


3. Per-Transaction Write Rates

Normalized by dividing per-block write bytes by per-block tx count.

Trie Updates per Transaction

p50 p90 p99
Bytes/tx 2,955 2,996 2,969
Entries/tx 7.2 7.2 7.1

Hashed State per Transaction

p50 p90 p99
Bytes/tx 205 192 200
Entries/tx 2.4 2.3 2.3

Combined per Transaction

p50 p90 p99
Total bytes/tx 3,160 3,188 3,169

Each transaction produces roughly 3.2 KB of trie+hashed state write data on average, with trie updates (branch node changes) accounting for 93% of the volume.


4. Actual DB Write Throughput (save_blocks batches)

Persistence batches ~3 blocks together (500 blocks / 166 batches).

Per-Batch Sizes

p50 p90 p99 max
Trie bytes/batch 75.6 MB 84.7 MB 84.7 MB 86.0 MB
Hashed bytes/batch 5.8 MB 6.3 MB 6.3 MB 6.5 MB
Total bytes/batch 81.4 MB 91.0 MB 91.0 MB 92.5 MB

Per-Batch Write Duration

p50 p90 p99 max
Trie write time 624 ms 667 ms 667 ms 722 ms
Hashed write time 378 ms 405 ms 405 ms 410 ms
Total write time 1,002 ms 1,072 ms 1,072 ms 1,132 ms
Total persistence 1,762 ms

Sustained Write Throughput

Data Total Written Total Duration Throughput
Trie updates 12.11 GB 110.9 s 112 MB/s
Hashed state 0.96 GB 59.5 s 16.5 MB/s

Trie writes are 12.6× larger than hashed state writes but only 1.9× slower, because MDBX trie table writes have sequential key patterns. The hashed state tables (HashedAccounts + HashedStorages) use keccak256 hashes as keys, producing a random insertion pattern that is ~7× slower per byte.

Write Amplification (batch vs per-block)

Data Per-Block Total (sum) Actual Written (sum) Ratio
Trie updates 17.37 GB 12.11 GB 0.70×
Hashed state 1.20 GB 0.96 GB 0.80×

Batching reduces actual writes by 20–30% since intermediate trie nodes from earlier blocks in the batch are superseded by later updates.


5. Write Budget as % of Pipeline Time

Phase Duration (s) % of wall clock (790s)
Trie update writes 110.9 14.0%
Hashed state writes 59.5 7.5%
Other persistence (headers, bodies, receipts, tx lookups) 122.0 15.4%
Total persistence 292.4 37.0%
Block execution 441.5 55.9%
Other 56.1 7.1%

Trie + hashed state writes consume 21.6% of wall-clock time and 58.3% of persistence time.


6. Scaling Implications by TPS

Write volume scales linearly with transaction count. Each transaction produces a remarkably stable ~3.2 KB of trie + hashed state write data (2,955 bytes trie + 205 bytes hashed). This makes sense: each transaction touches a fixed set of accounts and storage slots, and trie node mutations are proportional to state changes, not gas burned.

Sustained Write Data Rate

TPS Trie Data Rate Hashed Data Rate Combined
100 0.3 MB/s 0.02 MB/s 0.3 MB/s
1,000 2.8 MB/s 0.2 MB/s 3.0 MB/s
5,000 14.1 MB/s 1.0 MB/s 15.1 MB/s
10,000 28.2 MB/s 2.0 MB/s 30.1 MB/s
50,000 140.8 MB/s 9.8 MB/s 150.6 MB/s
100,000 281.7 MB/s 19.5 MB/s 301.2 MB/s
200,000 563.3 MB/s 39.1 MB/s 602.4 MB/s

Measured Backend Throughput

Backend Measured Throughput Bottleneck
Trie updates (MDBX) 112 MB/s Sequential B-tree inserts
Hashed state (MDBX) 16.5 MB/s Random-key inserts (keccak256 hashes)
MDBX commit (fsync) ~10 GB/s effective ~8 ms/MB of dirty data

Where the Write Wall Hits

The storage backend can sustain a maximum write rate before falling behind. Per transaction, the write time cost is:

  • Trie: 2,955 bytes ÷ 112 MB/s = 26.4 µs/tx
  • Hashed: 205 bytes ÷ 16.5 MB/s = 12.4 µs/tx
  • Combined: 38.8 µs/tx

This gives a theoretical ceiling of ~25,800 TPS if the storage backend did nothing but write trie + hashed state data. In practice, fsync overhead and other persistence work reduce this.

TPS Write Utilization Sustainable?
1,000 3.9% ✅ Comfortable — this benchmark
5,000 19.4% ✅ Headroom for execution + fsync
10,000 38.8% ⚠️ Writes consume ~40% of backend capacity
15,000 58.2% ⚠️ Tight — leaves ~40% for fsync + other writes
20,000 77.6% 🔴 Likely unsustainable with fsync overhead
25,000 97.0% 🔴 At theoretical write ceiling
50,000+ >100% 🔴 Exceeds backend throughput

What Dominates: Trie vs Hashed State

Trie Hashed
% of write bytes 93% 7%
% of write time 68% 32%

Trie updates dominate on volume (93%) but hashed state writes are disproportionately slow per byte (7× lower throughput due to random key insertion), so hashed state consumes 32% of write wall time despite being only 7% of the data.

Overcoming the Write Wall

To sustain >25K TPS, the write path needs fundamental changes:

  • Faster storage backend — io_uring, direct I/O, or replacing MDBX
  • Write batching across blocks — already saves 20–30%; larger batches amortize fsync and reduce trie update volume as intermediate nodes are superseded
  • Pipelining writes behind execution — current architecture already does this, so write throughput only needs to match sustained TPS, not burst
  • Reducing hashed state write volume — the 7× throughput penalty on random-key tables means even small reductions in hashed state writes have outsized impact on total write time

Metric Reference

Metric Scope What It Measures
sync.block_validation.trie_updates_sorted_data_bytes per-block Serialized trie update bytes computed at validation
sync.block_validation.trie_updates_sorted_size per-block Trie update entry count
sync.block_validation.hashed_state_sorted_data_bytes per-block Serialized hashed state bytes computed at validation
sync.block_validation.hashed_post_state_size per-block Hashed state entry count
storage.providers.database.save_blocks.write_trie_updates per-batch Trie write duration in MDBX
storage.providers.database.save_blocks.write_trie_updates_data_bytes per-batch Trie bytes actually written to MDBX
storage.providers.database.save_blocks.write_hashed_state per-batch Hashed state write duration in MDBX
storage.providers.database.save_blocks.write_hashed_state_data_bytes per-batch Hashed state bytes actually written to MDBX
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment