Reasoning and knowledge
| Benchmark | Kimi K3 | GPT-6 Luna | DeepSeek V4.1 Flash | DeepSeek V4 Flash 0731 | Qwen 3.8 27B | Qwen 3.6 35B-A3B | GLM-5.3 Flash |
|---|---|---|---|---|---|---|---|
| GPQA Diamond (198 questions) | 181 (91.4%) | 175 (88.4%)an earlier run on Sep 25 scored 173 | 172 (86.9%)includes 6 of 12 timed-out questions rerun with a longer limit | 171 (86.4%) | 165 (83.3%) | 153 (77.3%) | 159 (80.3%) |
| MMLU-Pro, STEM subset (210) | 194 (92.4%) | — | — | 190 (90.5%) | — | — | — |
| MMLU-Pro (490) | 430 (87.8%) | — | — | — | — | — | — |
| SimpleQA (500) | 213 (42.6%) | — | — | — | — | — | — |
| GSM8K (300) | 295 (98.3%) | — | — | — | — | — | — |
| MATH-500 advisory | 491 of 500 | — | — | — | — | — | — |
| AIME 2024 (30) advisory | 29 of 30 | — | — | — | — | — | — |
| AIME 2025 (30) advisory | 27 of 30 | — | — | — | — | — | — |
| IFEval (200) advisory | 175 of 200 | — | — | — | — | — | — |
| IFBench (300) | 69.3% | — | — | — | — | — | — |
Each model answered every question once through its own provider, at its highest reasoning setting where the provider offers one; the score is the number answered correctly. "Advisory" rows come from our first runs on September 18, when the answer-length limit was not recorded.
Coding
| Benchmark | Kimi K3 | GPT-6 Luna | DeepSeek V4.1 Flash | Qwen 3.8 27B |
|---|---|---|---|---|
| LiveCodeBench (175 problems) | 145 | 154 | 162 and 164two runs | 135 |
| LiveCodeBench, 60 hardest: solved | 37 | 45 | 48 to 50two runs | — |
| LiveCodeBench, 60 hardest: solved within 90 seconds | 13 | 35 | 7 | — |
| HumanEval (164) | 163 | 162 | 162 | 159 |
| Aider Polyglot (120 exercises, 2 tries) | 95.0%72.5% on the first try | 79.2%39.2% on the first try | — | — |
| BigCodeBench-Hard (148) | 38.5% | 34.5% | — | — |
| SWE-bench Verified (90 issues) | 80 | 60 | — | — |
Programs are graded by running them against each benchmark's hidden tests. LiveCodeBench uses recent contest problems, and its "60 hardest" rows cover the hard problems that read from standard input. Aider Polyglot covers 20 exercises in each of six languages. SWE-bench Verified fixes real GitHub issues with the same open-source agent loop and step limit for both models.
LiveCodeBench by difficulty, with speed and cost
| Model | Easy (43) | Medium (52) | Hard (80) | All (175) | Seconds per problem, median / 90th pct | Answered within 90 s | Right within 90 s | Cost per problem |
|---|---|---|---|---|---|---|---|---|
| GPT-6 Luna | 43 | 48 | 63 | 154 | 10.1 / 85.5 | 159 (91%) | 142 | $0.0014 |
| Kimi K3 | 42 | 49 | 54 | 145 | 31.7 / 398 | 113 (65%) | 108 | $0.098 |
| DeepSeek V4.1 Flash, run 1 | 43 | 52 | 67 | 162 | 97 / 786 | 87 (50%) | 85 | $0.0008 |
| DeepSeek V4.1 Flash, run 2 | 43 | 52 | 69 | 164 | 97 / 659 | 84 (48%) | 84 | $0.0007 |
| Qwen 3.8 27B | 43 | 45 | 47 | 135 | 128 / 553 | 76 (43%) | 75 | $0.0069 |
Same 175 problems as above, split by the benchmark's own difficulty labels; time is wall-clock from request to final answer, and cost is tokens used times the provider's list price (OpenRouter's reported charge for GPT-6 Luna).
Chat and writing
| Benchmark | Kimi K3 | GPT-6 Luna | GLM-5.3 Flash |
|---|---|---|---|
| Arena-Hard (500 prompts) vs o3-mini: wins in both orders | 65.2% | 40.6%of 497 judged | — |
| Arena-Hard, the benchmark's own score | 73.981.6 with style control | — | — |
| Arena creative writing (250) vs Gemini 2.0 Flash: wins in both orders | 81.6% | 60.0% | — |
| Arena creative writing, the benchmark's own score | 86.492.8 with style control | — | — |
| Arena creative writing, separate comparison run | 86.5 | — | 83.1 |
A GPT-4.1 judge compares each answer with a fixed baseline model's answer, once in each order. "Wins in both orders" counts prompts where the judge preferred the model both times. The benchmark's own score comes from its scoring script, and style control adjusts for length and formatting. The separate Sep 23 run used a different judge setup, so compare it only within its own row.
Business automation
| Benchmark | Kimi K3 | GPT-6 Luna |
|---|---|---|
| AutomationBench, all 600 tasks | 227 | 188 |
| Finance (100) | 45 | 35 |
| HR (100) | 39 | 32 |
| Marketing (100) | 43 | 37 |
| Operations (100) | 39 | 38 |
| Sales (100) | 30 | 25 |
| Support (100) | 31 | 21 |
Each task asks the model to act through tools on simulated business systems; a task passes only when those systems end in the correct state.
Speed and cost per question
| Benchmark | Kimi K3 | GPT-6 Luna | DeepSeek V4.1 Flash | DeepSeek V4 Flash 0731 | Qwen 3.8 27B | Qwen 3.6 35B-A3B | GLM-5.3 Flash |
|---|---|---|---|---|---|---|---|
| GPQA Diamond | 22.5 s / 158 s · 6.4¢ | 6.7 s / 51 s · 0.04¢ | 51 s / 386 s · 0.03¢ | 47 s / 306 s · 0.08¢ | 70 s / 324 s · 0.36¢ | 48 s / 70 s | 69 s / 508 s |
| MMLU-Pro, STEM subset | 4.7 s / 59 s · 2.5¢ | — | — | 14 s / 186 s · 0.04¢ | — | — | — |
| LiveCodeBench, 60 hardest | 201 s / 546 s · 19.4¢ | 42 s / 201 s · 0.30¢ | 411 s / 1,344 s · 0.15¢ | — | — | — | — |
| HumanEval | 12.5 s / 38.6 s · 0.93¢ | 3.6 s / 5.8 s · 0.013¢ | 9.1 s / 40.8 s · 0.007¢ | — | 8.7 s / 45 s · 0.064¢ | — | — |
| Aider Polyglot | 37 s · 10.8¢ | 15 s · 0.18¢ | — | — | — | — | — |
| BigCodeBench-Hard | 47 s · 4.5¢ | 10 s · 0.07¢ | — | — | — | — | — |
| SimpleQA | 3.9¢ | — | — | — | — | — | — |
| GSM8K | 3.8 s · 0.30¢ | — | — | — | — | — | — |
| SWE-bench Verified | 43.6¢ per issue | — | — | — | — | — | — |
| Arena-Hard, all 500 prompts | $30.78 total | — | — | — | — | — | — |
Times are wall-clock seconds per question, median / 90th percentile (a single figure is the median); cost is per question in US cents, tokens used times the provider's list price (OpenRouter's reported charge for GPT-6 Luna); a blank means not recorded.
Updated September 26, 2026. We add each new result as its run finishes.