Suite v1.0 · published Sep 21, 2026

Which model for which job. With reasons.

XOBENCH gives every model one clear request at its highest reasoning effort and takes its first answer, three times over. Thirteen categories of real engineering and business work, scored by deterministic checks and a blind three-judge panel from three vendors, with every raw response, rationale and dollar published.

Categories
13
7 engineering, 6 business
Models
1
in the published run
Attempts
3
3 per task per model
Judge scores
39
deterministic + panel + human

Capability matrix

Top three per category from the latest published run. Score is the mean of task medians; consistency is the average spread across the three repeats. There is deliberately no overall ranking.

Run: gemini-3.8-flash

Engineering track

7 categories

3D and Graphics

Produce Three.js scenes, SVG illustrations and CSS motion from a brief. Scored on rendered output.

5 tasks
No published scores in this category yet.

Coding

Implement functions and modules against a precise spec. Scored by hidden tests.

5 tasks
No published scores in this category yet.

Data and SQL

Write correct queries and transformations against a described schema. Scored by result comparison.

5 tasks
No published scores in this category yet.

Debugging

Given code and a symptom, find and fix the real bug without breaking anything else.

5 tasks
No published scores in this category yet.

IT Scripting

PowerShell and shell automation for Microsoft 365, Windows and network administration. Scored on safety and correctness.

5 tasks
No published scores in this category yet.

Security

Find real vulnerabilities in real-looking code, and implement sensitive features safely.

5 tasks
No published scores in this category yet.

Business track

6 categories

Legal Drafting

Draft and revise letters, clauses and policies a small firm would actually send. Judged by panel with required-element checks.

5 tasks
No published scores in this category yet.

Single shot, highest effort

One prompt, one answer, no tools, no retries. Every model at the top reasoning setting it offers, recorded per run.

Blind cross-vendor panel

Three judges from three vendors, never the vendor under test, scoring against anchored rubrics with written rationales. Deterministic checks wherever a check can be automated.

Median of three, spread shown

Every task runs three times. We publish the median and the spread, plus cost per task and cost per point, because price is part of the answer.