James/Models

Models we benchmark

These are the public benchmarks we have run on individual AI models, each called directly through its own provider with no James routing in between. Every result is from our own runs and carries the date that run finished (2026). A dash means we have not run that model on that test.

Reasoning and knowledge

BenchmarkKimi K3GPT-6 LunaDeepSeek V4.1 FlashDeepSeek V4 Flash 0731Qwen 3.8 27BQwen 3.6 35B-A3BGLM-5.3 Flash
GPQA Diamond (198 questions)181 (91.4%)175 (88.4%)an earlier run on Sep 25 scored 173172 (86.9%)includes 6 of 12 timed-out questions rerun with a longer limit171 (86.4%)165 (83.3%)153 (77.3%)159 (80.3%)
MMLU-Pro, STEM subset (210)194 (92.4%)——190 (90.5%)———
MMLU-Pro (490)430 (87.8%)——————
SimpleQA (500)213 (42.6%)——————
GSM8K (300)295 (98.3%)——————
MATH-500 advisory491 of 500——————
AIME 2024 (30) advisory29 of 30——————
AIME 2025 (30) advisory27 of 30——————
IFEval (200) advisory175 of 200——————
IFBench (300)69.3%——————

Each model answered every question once through its own provider, at its highest reasoning setting where the provider offers one; the score is the number answered correctly. "Advisory" rows come from our first runs on September 18, when the answer-length limit was not recorded.

Coding

BenchmarkKimi K3GPT-6 LunaDeepSeek V4.1 FlashQwen 3.8 27B
LiveCodeBench (175 problems)145154162 and 164two runs135
LiveCodeBench, 60 hardest: solved374548 to 50two runs—
LiveCodeBench, 60 hardest: solved within 90 seconds13357—
HumanEval (164)163162162159
Aider Polyglot (120 exercises, 2 tries)95.0%72.5% on the first try79.2%39.2% on the first try——
BigCodeBench-Hard (148)38.5%34.5%——
SWE-bench Verified (90 issues)8060——

Programs are graded by running them against each benchmark's hidden tests. LiveCodeBench uses recent contest problems, and its "60 hardest" rows cover the hard problems that read from standard input. Aider Polyglot covers 20 exercises in each of six languages. SWE-bench Verified fixes real GitHub issues with the same open-source agent loop and step limit for both models.

LiveCodeBench by difficulty, with speed and cost

ModelEasy (43)Medium (52)Hard (80)All (175)Seconds per problem, median / 90th pctAnswered within 90 sRight within 90 sCost per problem
GPT-6 Luna43486315410.1 / 85.5159 (91%)142$0.0014
Kimi K342495414531.7 / 398113 (65%)108$0.098
DeepSeek V4.1 Flash, run 143526716297 / 78687 (50%)85$0.0008
DeepSeek V4.1 Flash, run 243526916497 / 65984 (48%)84$0.0007
Qwen 3.8 27B434547135128 / 55376 (43%)75$0.0069

Same 175 problems as above, split by the benchmark's own difficulty labels; time is wall-clock from request to final answer, and cost is tokens used times the provider's list price (OpenRouter's reported charge for GPT-6 Luna).

Chat and writing

BenchmarkKimi K3GPT-6 LunaGLM-5.3 Flash
Arena-Hard (500 prompts) vs o3-mini: wins in both orders65.2%40.6%of 497 judged—
Arena-Hard, the benchmark's own score73.981.6 with style control——
Arena creative writing (250) vs Gemini 2.0 Flash: wins in both orders81.6%60.0%—
Arena creative writing, the benchmark's own score86.492.8 with style control——
Arena creative writing, separate comparison run86.5—83.1

A GPT-4.1 judge compares each answer with a fixed baseline model's answer, once in each order. "Wins in both orders" counts prompts where the judge preferred the model both times. The benchmark's own score comes from its scoring script, and style control adjusts for length and formatting. The separate Sep 23 run used a different judge setup, so compare it only within its own row.

Business automation

BenchmarkKimi K3GPT-6 Luna
AutomationBench, all 600 tasks227188
Finance (100)4535
HR (100)3932
Marketing (100)4337
Operations (100)3938
Sales (100)3025
Support (100)3121

Each task asks the model to act through tools on simulated business systems; a task passes only when those systems end in the correct state.

Speed and cost per question

BenchmarkKimi K3GPT-6 LunaDeepSeek V4.1 FlashDeepSeek V4 Flash 0731Qwen 3.8 27BQwen 3.6 35B-A3BGLM-5.3 Flash
GPQA Diamond22.5 s / 158 s · 6.4¢6.7 s / 51 s · 0.04¢51 s / 386 s · 0.03¢47 s / 306 s · 0.08¢70 s / 324 s · 0.36¢48 s / 70 s69 s / 508 s
MMLU-Pro, STEM subset4.7 s / 59 s · 2.5¢——14 s / 186 s · 0.04¢———
LiveCodeBench, 60 hardest201 s / 546 s · 19.4¢42 s / 201 s · 0.30¢411 s / 1,344 s · 0.15¢————
HumanEval12.5 s / 38.6 s · 0.93¢3.6 s / 5.8 s · 0.013¢9.1 s / 40.8 s · 0.007¢—8.7 s / 45 s · 0.064¢——
Aider Polyglot37 s · 10.8¢15 s · 0.18¢—————
BigCodeBench-Hard47 s · 4.5¢10 s · 0.07¢—————
SimpleQA3.9¢——————
GSM8K3.8 s · 0.30¢——————
SWE-bench Verified43.6¢ per issue——————
Arena-Hard, all 500 prompts$30.78 total——————

Times are wall-clock seconds per question, median / 90th percentile (a single figure is the median); cost is per question in US cents, tokens used times the provider's list price (OpenRouter's reported charge for GPT-6 Luna); a blank means not recorded.

Updated September 26, 2026. We add each new result as its run finishes.