Skip to content

Instantly share code, notes, and snippets.

I'm trying to come up with a popularization of the meaning of the Baum-Connes conjecture for people with only undergraduate math. One idea I had was that Group C*-algebras, when the group is commutative, are basically the algebra of function on the dual, so under the algebras-spaces equivalence, Group C*-algebra are attempting to generalize "the dual space of a group" to non-abelian groups, and then, the Baum-Connes conjecture is about algebraic-topologic invariants of that "noncommutative dual space" and maybe specifically relating it to the classifying space ?

Your framing is essentially right, and it's the standard heuristic. A few refinements will make it both accurate and more compelling.

What's right

  • For a commutative group G (say a locally compact abelian one), Pontryagin duality plus Fourier transform identifies C*(G) with the algebra of continuous functions vanishing at infinity on the dual group Ĝ. That's the Gelfand-Naimark correspondence at work: commutative C*-algebras are spaces.

ConSan startup speedups on gfx1201

HEAD (222980ed695) compares September 26 and September 30, 2026 measurements on gfx1201. Each cell shows startup seconds before → after (speedup).

Workload Default preset High preset SuperCollider
Dense prefill 78.5 → 12.1 (6.49×) 79.3 → 12.4 (6.40×) 56.0 → 24.6 (2.28×)
Synthetic decode 90.1 → 12.8 (7.04×) 92.3 → 12.9 (7.16×) 60.8 → 24.7 (2.46×)
MoE prefill 167 → 26.1 (6.40×) 173 → 25.9 (6.68×) 105 → 50.6 (2.08×)
Gluon shared roundtrip 0.440 → 0.444 (0.99×) 0.434 → 0.441 (0.98×) 0.434 → 0.437 (0.99×)

ConSan performance gains — September 30, 2026

There are two distinct sets of gains: faster gfx1250 emulation and faster ConSan startup on physical gfx1201.

Commit What was slow What changed Measured contribution
45a3312e368 Emulated memory accesses repeatedly acquired an exclusive VM lock and copied callbacks/backing ownership. ConSan’s scratch spills amplified this. Share immutable VM state and distribute reader locking. Measured together with scratch coalescing: solution 18’s time outside instrumentation fell from ~41 s to ~2.6 s.
bf566e98a52 Scratch spills performed separate backing accesses for individual lanes. Coalesce eligible accesses within one swizzle unit; a 32-lane DWORD spill can use one contiguous operation. A scratch-only historical control improved 67.25 → 61.53 s; both memory fixes reached 38.74 s. These noisy controls don’t establish a precise split between the two commits.
c00010c86e6 Waitcheck repeatedly decoded
@bjacob
bjacob / a.txt
Created September 30, 2026 14:35
We've had 22–24 kernels timing out on every build since 11f077a2 (Sep 24); on a3a161be (Sep 21) only 4 timed out. On today's 1e03999a, 21 of those kernels finish in 3–7 s with the hook off, yet with ConSan on none finish within 60 minutes. a3a161be with the same preset, profile and image finishes all 21 in 4–320 s. So something that landed between Sep 21 and Sep 24 made ConSan much slower on these kernels. I can't bisect it because those commits don't share history after the force-push. 22 of the 23 slow kernels use ds_*_b128, if that helps narrow it.
@bjacob
bjacob / a.md
Created September 29, 2026 16:58

ConSan and DBI: concrete tensions and a path to merging

Compared DBI overview, DBI implementation design, and ConSan design, checking relevant code. Snapshots: sanitizers 1e03999af42; origin/develop 69405c977d5, fetched for this revision (September 29, 2026). The overview specifies intended DBI architecture; the implementation document describes an earlier milestone. Each example below identifies whether it illustrates existing ConSan behavior, an existing DBI limitation, or a hypothetical adaptation. Hypothetical failures are design obligations to address, not bugs established in either implementation. The traces are illustrative, not newly run tests. Source paths below are relative to emulation/rocjitsu; most implementation paths start und

@bjacob
bjacob / a.md
Created September 28, 2026 15:30

Element-wise SIMD math implementations

Research date: September 28, 2026.

Scope

The operation of interest is vector<Nxf32> -> vector<Nxf32> (and analogous floating-point types), with independent evaluations of functions such as sqrt, sin, cos, exp, and log across lanes. Scalar math functions that merely use SIMD instructions internally are outside this scope.

This is a source-based survey, not a benchmark or an exhaustive audit of every function. Implementation details depend on the function, precision, architecture, library version, and compiler settings. Links below generally point to moving upstream branches.

21 gfx1201 emulator failures; all reproduced with ctest -j1.

ConSanDeviceGfx1201Sim.Preset.high.DefaultSelected

[rocjitsu-dbi-hooks] ConSan auto report reader=109555456902080 has incomplete or malformed atomic publication evidence flags=3 events=8192 capacity=6208 dropped=1 clock=6222
[rocjitsu-dbi-hooks] ConSan analysis verdict applicable=true analysis_complete=false static_complete=true dynamic_complete=false applicable_code_objects=1 incomplete_code_objects=0 access=23/23 barrier=28/28 atomic=0/0 fence=0/0 visible_evidence=14 dynamic_incomplete=1 
[rocjitsu-dbi-hooks] RJ_CONSAN_REQUIRE_DIAGNOSTICS requested, but 1 auto ConSan report buffer(s) produced zero diagnostics or sampled conflicts (sampled=14, conflicts=0 immediate_conflicts=0)
  test command exited with 88, expected 0
How /goal was part of the problem
/goal installs a hook that runs every time the main thread ends a turn. A checker reads the conversation and
decides whether the goal is met. If it is not, the session is nudged to continue. That produced two costs.
- A per-turn tax. The cost is the number of turn ends times the context size. That night there were 299 turn
ends with the goal active, at about 500k tokens each. About a quarter of the bill is cache reads and writes
that appear in the session's cost record but in no transcript. The goal checks fit that volume. The docs say
the checker runs on Haiku, but the cost record shows almost no Haiku usage. A 500k conversation also does
not fit Haiku's window. So the checks were probably billed at Fable rates, but I have not proven it.
@bjacob
bjacob / a.md
Created September 15, 2026 16:59

Functionality implications of keeping Sampled as the only MOI engine

Scope and evidence

Source review of shared/rocjitsu/sanitizers, revision 75881fde3e181682f8bcddca58bc8031106757e1. All source paths below are relative to rocm-systems. No new GPU trials or test runs were performed. Fault specifications are executable expectations, not freshly measured detection rates. Status ledgers summarize previous accepted evidence; neither substitutes for rerunning the exact binary/configuration after changes.

Benchmark update: the gfx950 campaign documents were read after pulling revision 3ddf9dee36. They now contain the completed physical MI350X results. The implementation and test discussion below remains the original source review at 75881fde3e; its line references refer to that revision. This update does not constitute a full re-review of the intervening implementation changes or a rerun of the fault corpus. The benchmark evidence strengthens the performance motivation for Sampled but does not estab

ConSan gfx1201 benchmark status

For each mode, Startup is the total latency through the first synchronized run and its selected evidence checkpoints, including instrumentation, loading, binding, and warm-up; Run is the absolute second-run latency followed by its ratio to the matching uninstrumented second run.

Workload Uninstrumented Startup Uninstrumented Run SuperCollider Startup SuperCollider Run RecordReplay Startup RecordReplay Run Sampled Startup Sampled Run InlineShadow Startup InlineShadow Run
PyTorch synthetic dense prefill (32-token prompt) 0.937 s 0.00209 s (1×) 132 s 0.00332 s (1.59×) 182 s 0.805 s (386×) 163 s 0.0092 s (4.41×) 159 s 0.431 s (207×)