3D and Graphics
Produce Three.js scenes, SVG illustrations and CSS motion from a brief. Scored on rendered output.
XOBENCH gives every model one clear request at its highest reasoning effort and takes its first answer, three times over. Thirteen categories of real engineering and business work, scored by deterministic checks and a blind three-judge panel from three vendors, with every raw response, rationale and dollar published.
Top three per category from the latest published run. Score is the mean of task medians; consistency is the average spread across the three repeats. There is deliberately no overall ranking.
Produce Three.js scenes, SVG illustrations and CSS motion from a brief. Scored on rendered output.
Implement functions and modules against a precise spec. Scored by hidden tests.
Write correct queries and transformations against a described schema. Scored by result comparison.
Given code and a symptom, find and fix the real bug without breaking anything else.
PowerShell and shell automation for Microsoft 365, Windows and network administration. Scored on safety and correctness.
Find real vulnerabilities in real-looking code, and implement sensitive features safely.
Build complete, polished interfaces from a written brief. Scored on rendered output.
Write the emails and announcements professionals actually have to send, in a specified voice, under constraints.
Pull structured data out of messy real-world documents into exact JSON. Scored field by field.
Questions built to tempt fabrication. Scored on whether the model flags what it cannot know.
Draft and revise letters, clauses and policies a small firm would actually send. Judged by panel with required-element checks.
Multi-step business reasoning with checkable answers: math, scheduling, policy interpretation, logic.
Condense long documents into briefs for a named audience without inventing anything.
One prompt, one answer, no tools, no retries. Every model at the top reasoning setting it offers, recorded per run.
Three judges from three vendors, never the vendor under test, scoring against anchored rubrics with written rationales. Deterministic checks wherever a check can be automated.
Every task runs three times. We publish the median and the spread, plus cost per task and cost per point, because price is part of the answer.