Model evaluations

Pick the best models for your Roomote.

Different models are better at different jobs. So we regularly run evals for the different roles they play in Roomote so you can make well-informed choices between price, speed and intelligence.

Last run Aug 8, 2026

  1. 1Determine what you care about
  2. 2Examine the best models for each role
  3. 3Configure them in Model settings (Docs)

Coder Role Eval

Runs models through Roomote's full coding workflow on real repository, with a fixed model for other roles, then evaluates the delivered PR against the task's acceptance criteria.

Score: Quality of the final PR produced.

#
1
GPT 5.6 LunaMED
96%
780s
$0.21
2
Grok 4.5MED
96%
997s
$3.28
3
Claude Sonnet 5MED
96%
1,300s
$4.80
4
Kimi K3MED
96%
1,530s
$5.36
5
Claude Fable 5MED
96%
1,431s
$23.96
6
Claude Opus 5MED
96%
1,907s
$16.45
7
GPT 5.6 TerraMED
90%
1,028s
$2.07
8
GLM 5.2MED
90%
2,470s
$2.19
9
GPT 5.6 SolMED
90%
1,513s
$10.80
10
Qwen3.8 MaxMED
89%
3,127s
$6.93
11
Minimax M3MED
86%
2,169s
$2.22
12
Kimi K2.7 CodeMED
85%
3,878s
$5.49
13
Deepseek V4 Flash 0731MED
81%
1,653s
$0.24
14
Deepseek V4 ProMED
79%
1,933s
$0.44
15
Gemini 3.6 FlashMED
74%
1,270s
$8.20
16
Claude Haiku 4.5MED
36%
891s
$1.97

Based on 6 test cases.

Reviewer Role Eval

Tests whether the model finds known problems in completed changes without incorrectly asking for changes to clean work.

Score: Catch rate minus half the false-alarm rate: a missed problem counts double a wrongly blocked clean change. 0 means no better than approving everything; 100 means every known problem caught without blocking any clean change.

#
1
Claude Fable 5HIGH
57%
2,084s
$5.01
2
Kimi K3HIGH
54%
13,846s
$1.12
3
GLM 5.2HIGH
51%
4,692s
$0.13
4
Claude Opus 5HIGH
51%
2,991s
$2.92
5
GPT 5.6 TerraHIGH
50%
1,389s
$0.31
6
Grok 4.5HIGH
42%
3,482s
$0.85
7
GPT 5.6 LunaHIGH
40%
2,639s
$0.05
8
Claude Sonnet 5HIGH
39%
5,445s
$1.75
9
Deepseek V4 Flash 0731HIGH
38%
28,780s
$0.03
10
Qwen3.8 MaxHIGH
35%
16,609s
$1.74
11
Claude Haiku 4.5HIGH
34%
6,733s
$1.03
12
GPT 5.6 SolHIGH
34%
3,244s
$1.76
13
Gemini 3.6 FlashHIGH
32%
1,652s
$0.88
14
Deepseek V4 ProHIGH
24%
7,588s
$0.16
15
Kimi K2.7 CodeHIGH
8%
38,835s
$0.37
16
Minimax M3HIGH
7%
140s
$0.06

Based on 84 test cases.

Advisor Role Eval

Tests the advisor's three production jobs with free-form guidance: planning work up front, unsticking a stuck coding agent, and judging user challenges — conceding when the user is right and holding when they are wrong.

Score: Answer-key coverage averaged across the three modes: each response is checked against required elements and traps derived from the real developer fix, and challenge responses must also take the correct side.

#
1
Claude Opus 5HIGH
83%
2,600s
$0.93
2
Grok 4.5HIGH
80%
1,769s
$0.19
3
GPT 5.6 SolHIGH
80%
2,533s
$0.59
4
GPT 5.6 TerraHIGH
78%
1,583s
$0.10
5
GPT 5.6 LunaHIGH
75%
1,430s
$0.01
6
Claude Fable 5HIGH
75%
1,776s
$1.38
7
Gemini 3.6 FlashHIGH
66%
1,294s
$0.19
8
Claude Sonnet 5HIGH
65%
2,442s
$0.40
9
Kimi K3HIGH
65%
4,117s
$0.32
10
Deepseek V4 ProHIGH
62%
1,720s
$0.04
11
Kimi K2.7 CodeHIGH
61%
3,622s
$0.12
12
Claude Haiku 4.5HIGH
51%
2,944s
$0.13
13
Minimax M3HIGH
49%
2,262s
$0.01
14
GLM 5.2HIGH
45%
1,843s
$0.03
15
Qwen3.8 MaxHIGH
39%
6,543s
$0.33
16
Deepseek V4 Flash 0731HIGH
35%
4,768s
<$0.01

Based on 32 test cases.

Explorer Role Eval

Runs Roomote's actual explore subagent — its production prompt with glob, grep, and read tools — inside each task's repository, checked out exactly as it was before the fix, and measures whether it finds the files the real fix modified.

Score: Recall-weighted retrieval score: missing a file the fix needs costs more than naming an extra one. Search efficiency is ranked separately through cost and latency.

#
1
GPT 5.6 SolLOW
90%
606s
$0.47
2
Gemini 3.6 FlashLOW
85%
1,023s
$0.36
3
GPT 5.6 TerraLOW
84%
288s
$0.09
4
GPT 5.6 LunaLOW
83%
267s
<$0.01
5
Claude Sonnet 5LOW
77%
530s
$0.27
6
Claude Fable 5LOW
73%
490s
$1.45
7
Grok 4.5LOW
71%
456s
$0.16
8
GLM 5.2LOW
70%
470s
$0.04
9
Deepseek V4 Flash 0731LOW
69%
537s
<$0.01
10
Claude Opus 5LOW
62%
521s
$0.65
11
Claude Haiku 4.5LOW
60%
702s
$0.07
12
Kimi K3LOW
55%
1,830s
$0.17
13
Kimi K2.7 CodeLOW
53%
2,025s
$0.06
14
Qwen3.8 MaxLOW
53%
1,527s
$0.13
15
Deepseek V4 ProLOW
52%
1,138s
$0.04
16
Minimax M3LOW
47%
2,165s
$0.02

Based on 13 test cases.

Vision Role Eval

Tests whether the model can accurately read text and identify relevant interface details in screenshots Roomote encounters while working.

Score: Each answer is graded against a fixed answer key for its screenshot: a judge with the key in hand checks that every required detail is stated, and the score is the percentage of questions answered fully correctly.

#
1
Gemini 3.6 FlashLOW
100%
35s
$0.03
2
GPT 5.6 TerraLOW
93%
36s
$0.05
3
Kimi K3LOW
93%
109s
$0.17
4
Claude Fable 5LOW
93%
91s
$0.39
5
GPT 5.6 SolLOW
93%
55s
$0.26
6
GPT 5.6 LunaLOW
86%
37s
<$0.01
7
Grok 4.5LOW
86%
29s
$0.05
8
Claude Sonnet 5LOW
86%
56s
$0.08
9
Claude Haiku 4.5LOW
86%
125s
$0.03
10
Claude Opus 5LOW
86%
109s
$0.21
11
Minimax M3LOW
79%
29s
<$0.01
12
Kimi K2.7 CodeLOW
64%
416s
$0.02
13
Qwen3.8 MaxLOW
64%
572s
$0.04
14
Deepseek V4 Flash 0731LOWDid not complete
0%
1,680s
$0.00
15
Deepseek V4 ProLOWDid not complete
0%
1,680s
$0.00
16
GLM 5.2LOWDid not complete
0%
1,680s
$0.00

Based on 14 test cases.

Router Role Eval

When a task comes in, Roomote must choose which configured environment should handle it—for example, your web app, mobile app, backend, or infrastructure. This eval tests whether the model makes that choice correctly and knows when it needs to open a linked issue for more context.

Score: An equal-weight average of choosing the correct environment, deciding when issue details are needed, and extracting usable issue references.

#
1
GPT 5.6 SolLOW
100%
94s
$0.35
2
Kimi K3LOW
98%
170s
$0.23
3
Claude Opus 5LOW
98%
130s
$0.58
4
Claude Fable 5LOW
98%
171s
$1.16
5
GPT 5.6 LunaLOW
97%
87s
<$0.01
6
Gemini 3.6 FlashLOW
97%
74s
$0.12
7
GLM 5.2LOW
95%
166s
$0.03
8
Claude Sonnet 5LOW
95%
124s
$0.23
9
GPT 5.6 TerraLOW
93%
78s
$0.07
10
Grok 4.5LOW
93%
119s
$0.16
11
Claude Haiku 4.5LOW
90%
356s
$0.12
12
Minimax M3LOW
87%
154s
$0.02
13
Deepseek V4 ProLOW
85%
463s
$0.04
14
Qwen3.8 MaxLOW
85%
351s
$0.18
15
Deepseek V4 Flash 0731LOW
83%
602s
<$0.01
16
Kimi K2.7 CodeLOW
73%
991s
$0.05

Based on 29 test cases.

Get started with Roomote

Roomote is a full-stack web application that connects to your tools. To self-host, run the command on a Linux box, use a template or explore alternatives in our docs.Inspect the script

curl -fsSL https://get.roomote.dev | bash
Other ways to set up