Benchmark Results
Benchmarks
Expert reasoning evaluations — same model, same max effort. Only difference: Kohnex.
Overall Score
DeepSeek V4 Flash · MAX reasoning effort · 100 tasks
Same model, same reasoning effort — only difference: Kohnex
By Category
MAX vs MAX + Kohnex · 20 tasks each
Methodology
How these numbers were produced
MODEL
DeepSeek V4 Flash at maximum reasoning effort on both arms — baseline and Kohnex-assisted alike.
TASK SET
100 custom expert tasks — 20 in each of 5 categories spanning diagnosis, debugging, review, synthesis, and design.
GRADING
Each response scored by the same model acting as judge, on a 0/1/2 partial-credit rubric across 7 quality dimensions.
FAIR COMPARISON
Identical prompts, identical effort. The single variable between the two arms: Kohnex Brain.
Scores come from an internal evaluation on a custom task set — not a public benchmark, and not yet independently audited.
NEXT UP
More benchmarks on the way
New model evaluations are in progress. They will appear here in this index as soon as they ship.
Disclaimer: Internal benchmark scores, not independently audited. FORGE-100 is a custom task set authored by Kohnex (42.7% vs 65.2% on deepseek-v4-flash at max effort). Results reflect this specific evaluation.