§ models · storyline
Kimi K3 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing
Kimi K3 provides 2.8x cost efficiency on pass@4 versus GPT-5.6 Sol, with hybrid routing reaching 85.6% on DeepSWE benchmarks.
We ran 904 DeepSWE rollouts on Kimi K3 and GPT-5.6 Sol. Sol leads pass@1; Kimi K3 wins pass@4 at 2.8x the solves per dollar, and routing between them reaches ~85.6%.
§ sources1 publication · timeline below