← All posts
BENCHMARK RESULTS · · 6 min read

DeepSeek V4 Flash + Kohnex: +22.5 Points on FORGE-100

We built FORGE-100 because public benchmarks stopped answering the question we actually care about: does this make a model reason better on expert work? So we wrote 100 custom tasks across five expert categories, ran DeepSeek V4 Flash at maximum reasoning effort twice — once alone, once with Kohnex Brain injected — and graded every answer on a 0/1/2 partial-credit rubric with the same model as judge.

The headline number

Model only: 42.7% → Model + Kohnex: 65.2% (+22.5 points). Same model, same max effort. The only difference was Kohnex.

Where the gains came from

How we kept it honest

Identical prompts, identical reasoning effort, one variable. Every response was scored on completeness, correctness, depth of analysis, actionable recommendations, edge-case coverage, production-readiness, and engineering judgment. This is an internal evaluation on a custom task set — not a public benchmark, not independently audited — and we publish the full breakdown on our benchmarks page so you can inspect every category instead of trusting a headline.

That last part is the point. Proof over promises.

Make your AI think like a pro.

Connect any model and get expert reasoning. Free tier available.