A single weighted score across cloud and open models, built from a small set of public benchmarks. Move the slider to see how the ranking shifts depending on whether you value raw capability or task accuracy more.
Demo data: the scores below are illustrative placeholders (Jan–Jul 2026), not verified live benchmark results. Swap in your own sourced numbers in the JSON block before using this for anything real.
Weighting
Performance33%
Accuracy33%
Cost efficiency34%
Weights are normalized so they always sum to 100%, shown above each slider. Performance ≈ Arena preference + reasoning speed. Accuracy ≈ MMLU, GPQA, coding & multimodal correctness. Cost efficiency ≈ inverse of blended $/M tokens.