The problem
Foundation MLIPs and property models are exploding — but industrial users cannot tell which model is trustworthy for their material class. Vendor benchmarks are marketing; academic benchmarks use public test sets with leakage, measure physics metrics instead of engineering KPIs, and are saturating. No independent testing lab exists at the application level.
Our approach
- Private, rotating test sets — no leakage, annual rotation
- High-level references — DFT, CCSD(T) where needed, curated experimental data
- Blind protocol — unlabeled structures out, predictions back; weights never leave the vendor, labels never leave us
- Application scorecards per material class — accuracy, failure rate, out-of-distribution transfer, uncertainty quality, cost, reproducibility
First vertical: liquid electrolytes
Battery electrolytes are our starting point — the biggest budget in atomistic ML, and no independent benchmark exists.
- Ionic conductivity
- Viscosity
- Li⁺ transference number
- Diffusion coefficients
- Dielectric constant & density
Roadmap
- Now: standardized MD protocol, 40–60 electrolyte systems with experimental references
- Next: evaluate open models (UMA, MACE, Orb, MatterSim, SevenNet, …)
- Then: public leaderboard, preprint, and extensions to solid-state electrolytes, polymers, CO₂ sorbents
For model teams & researchers
We evaluate open models for free — the leaderboard will never be paywalled. If you build a foundation model or have electrolyte reference data, we would like to talk. This is a research collaboration, not a service offering.