Benchmark report
Benchmark Results
Full 200-task Harbor runs over a hard-heavy multilingual coding subset, shown as visual comparisons for pass count, recorded cost, and cost efficiency.
Summary
Near-peer pass counts with lower recorded spend.
Klyrune kept pass counts close to Codex on both tested model families while using materially less recorded budget in these runs.
Cost and language detail
Recorded spend and strongest language groups.
Recorded Cost
USD, lower is betterLanguage Highlights
Passes out of 10Evidence links