Benchmark report

Benchmark Results

Full 200-task Harbor runs over a hard-heavy multilingual coding subset, shown as visual comparisons for pass count, recorded cost, and cost efficiency.

Summary

Near-peer pass counts with lower recorded spend.

Klyrune kept pass counts close to Codex on both tested model families while using materially less recorded budget in these runs.

AutoCodeBench Pass Rate

Score, higher is better
64.0%
Codex GPT-5.5 (xhigh)
61.5%
Klyrune GPT-5.5 (xhigh)
53.0%
Codex DeepSeek V4 FLASH 0731 (max)
52.5%
Klyrune DeepSeek V4 FLASH 0731 (max)

Cost and language detail

Recorded spend and strongest language groups.

Recorded Cost

USD, lower is better
$18.45
Codex GPT-5.5 (xhigh)
$9.68
Klyrune GPT-5.5 (xhigh)
$1.24
Codex DeepSeek V4 FLASH 0731 (max)
$0.69
Klyrune DeepSeek V4 FLASH 0731 (max)

Language Highlights

Passes out of 10
10/10
Ruby GPT-5.5 (xhigh)
9/10
C# GPT-5.5 (xhigh)
9/10
Ruby DeepSeek V4 FLASH 0731 (max)
8/10
Racket DeepSeek V4 FLASH 0731 (max)