MANIFESTO
·
·
4 min read
Proof Over Promises: Why We Publish Our Benchmarks
Every AI company publishes numbers. Most of them are selected, rounded, and framed until they say whatever marketing needed. We think that's backwards: benchmarks should be evidence, not advertising. So here's our policy, in writing.
What we publish
- Full breakdowns, not just headlines. Our FORGE-100 result isn't "65.2%" — it's five categories with base scores, gains, task counts, and grading rubric, all on one page.
- The unflattering parts too. Constrained design went from 4% to 21%. That's still a low absolute number, and we published it anyway — because a 5x lift on the hardest tasks tells you more than a polished average.
- Methodology before marketing. Model, effort level, task count, judge, rubric — all stated up front, before any claim.
What we won't do
- Cherry-pick the one category that looks best and crop the rest.
- Compare our best configuration against a competitor's weakest.
- Hide custom task sets behind vague names without describing them.
Why this matters for you
You're being asked to trust an intelligence layer with real work. Trust shouldn't come from a landing page — it should come from numbers you can interrogate. Every claim on this site links to the data behind it. If you find a flaw in our evaluation, tell us: we'd rather fix the benchmark than defend the score.
Make your AI think like a pro.
Connect any model and get expert reasoning. Free tier available.