BENCHMARK RESULTS
·
DeepSeek V4 Flash + Kohnex: +22.5 Points on FORGE-100
Same model, same max effort — 42.7% to 65.2% across 100 custom expert tasks, with the full category breakdown.
Read more →Proof over promises: benchmarks, reasoning, and real results.
Same model, same max effort — 42.7% to 65.2% across 100 custom expert tasks, with the full category breakdown.
Read more →Our evaluation policy in writing — full breakdowns, unflattering numbers included, methodology before marketing.
Read more →A 2.5B model beating giants on math and code confirms it — reasoning beats raw scale.
Read more →