Coder Role Eval
Runs models through Roomote's full coding workflow on real repository, with a fixed model for other roles, then evaluates the delivered PR against the task's acceptance criteria.
Score: Quality of the final PR produced.
| # | ||||
|---|---|---|---|---|
| 1 | GPT 5.6 LunaMEDBalance Pick | 96% | 780s | $0.21 |
| 2 | Grok 4.5MEDBalance Pick | 96% | 997s | $3.28 |
| 3 | Claude Sonnet 5MEDBalance Pick | 96% | 1,300s | $4.80 |
| 4 | Kimi K3MEDBalance Pick | 96% | 1,530s | $5.36 |
| 5 | Claude Fable 5MEDBalance Pick | 96% | 1,431s | $23.96 |
| 6 | Claude Opus 5MEDBalance Pick | 96% | 1,907s | $16.45 |
| 7 | GPT 5.6 TerraMEDBalance Pick | 90% | 1,028s | $2.07 |
| 8 | GLM 5.2MEDBalance Pick | 90% | 2,470s | $2.19 |
| 9 | GPT 5.6 SolMEDBalance Pick | 90% | 1,513s | $10.80 |
| 10 | Qwen3.8 MaxMEDBalance Pick | 89% | 3,127s | $6.93 |
| 11 | Minimax M3MEDBalance Pick | 86% | 2,169s | $2.22 |
| 12 | Kimi K2.7 CodeMEDBalance Pick | 85% | 3,878s | $5.49 |
| 13 | Deepseek V4 Flash 0731MEDBalance Pick | 81% | 1,653s | $0.24 |
| 14 | Deepseek V4 ProMEDBalance Pick | 79% | 1,933s | $0.44 |
| 15 | Gemini 3.6 FlashMEDBalance Pick | 74% | 1,270s | $8.20 |
| 16 | Claude Haiku 4.5MEDBalance Pick | 36% | 891s | $1.97 |
Based on 6 test cases.

