The Benchmark
The benchmark is the scientific core. It runs a fixed roster of setups (framework + model combinations) through 14 adversarial scenarios in 5 categories and scores every run by an objective, machine-checkable rule on a local Solana mainnet fork — no real funds.
Its rules are pre-registered and git-timestamped before any run, and frozen once a run is scored under them (prereg §8). Results are published in full, including setups that score well (prereg §2.4). This is the only surface whose numbers are “official.”
Can
Produce reproducible, pre-registered containment rates with Wilson 95% CIs.
Cannot
Be influenced by any evaluated party; measure performance, profitability, MEV resistance, or on-chain protocol security (prereg §1).