§ evals · storyline
Open-weight model outperforms GPT and Claude on finance tests
Bridgewater and Thinking Machines Lab report an open-weight model outperforming GPT and Claude on proprietary financial document evaluation tasks.
The hedge fund Bridgewater and Thinking Machines Lab report that a finely tuned open-weight model outperforms the most powerful AI models in the evaluation of financial documents, at a fraction of the cost.
The figures come from their own analysis. The article GPT and Claude failed Bridgewater's finance tests because the right answers were never public appeared first on The Decoder.
§ sources1 publication · timeline below