Octavus Benchmarks
Real-world, human-easy / agent-hard computer use: live browsers, native desktop apps, system dialogs, local files, and rendered documents - the interaction layer other benchmarks abstract away.
How to read the cards
The three dimensions
Each dimension on its own, in its native units and ranked best-first - the raw numbers behind the ratings. Speed and cost are per task, averaged over the whole run; cost is the model / provider cost only - never the platform-inclusive figure.
Sample tasks
6 recordings from the runs behind this board. Pick a task to watch exactly what the model did, step by step.
Methodology
How every result on this board is produced, scored, and verified.
About the Benchmark
Each task asks an autonomous agent to complete an ordinary computer task through the interfaces people actually use, then verify the result. Grading is AI-against-a-rubric: a shared Octavus grader worker scores each submission against the task's rubric and answer key, and a deterministic pass rule (criterion weights + a pass threshold, plus any hard gates) is applied by the harness. The published pass rate is over every task that ran; a task whose answer key has not yet been gold-verified against the live target is still counted, but flagged for review. Octavus runs the benchmark as an ordinary Workforce Agents API consumer - no privileged path - and the displayed cost is the underlying model / provider cost.
Leaderboard
Measured by Octavus, ranked by overall rating. Sort any column, filter by provider, effort, or model; every row opens the model's full per-task run.
| # | Model | Verification | ||||
|---|---|---|---|---|---|---|
| 1 | gemini-3.8-flash GoogleHigh | 84Competence 85 · Low Cost 72 · Speed 93 | 85.2%75/88 | $193.27$2.20 / task | 14h 18m9m 45s / task | Verified |
| 2 | gpt-5.6-luna OpenAIHigh | 83Competence 73 · Low Cost 100 · Speed 95 | 72.7%64/88 | $16.45$0.187 / task | 12h 29m8m 31s / task | Verified |
| 3 | claude-opus-5-5 AnthropicHigh | 82Competence 80 · Low Cost 72 · Speed 100 | 79.5%70/88 | $188.75$2.14 / task | 8h 4m5m 30s / task | Verified |
| 4 | claude-opus-4-8 AnthropicHigh | 80Competence 81 · Low Cost 63 · Speed 93 | 80.7%71/88 | $423.78$4.82 / task | 14h 41m10m 1s / task | Verified |
| 5 | grok-4.7 xAIHigh | 77Competence 78 · Low Cost 62 · Speed 88 | 78.4%69/88 | $456.52$5.19 / task | 23h 5m15m 45s / task | Verified |
| 6 | claude-sonnet-5 AnthropicHigh | 76Competence 74 · Low Cost 66 · Speed 92 | 73.9%65/88 | $310.99$3.53 / task | 15h 59m10m 54s / task | Verified |
Cost is the model / provider cost. Octavus runs every model through the same public product surface any customer uses - no privileged path.