A technical comparison of Anthropic's Claude Sonnet 5, Sonnet 4.6, and Opus 4.8, focusing on agentic coding benchmarks, API pricing, and cost-performance metrics.
A comparison of Anthropic's Claude Sonnet 5, Sonnet 4.6, and Opus 4.8, focusing on agentic coding benchmarks, API pricing, and cost-performance trade-offs.
Stanford researchers released TRACE, an open-source, capability-targeted agentic training system designed to improve agent performance by turning recurrent failures into synthetic reinforcement learning environments.
Skyfall AI released MORPHEUS, a persistent enterprise simulation benchmark designed to necessitate continual reinforcement learning in non-stationary environments.
Prime Intellect released Verifiers v1, a tool for decomposing agentic reinforcement learning environments into composable tasksets, harnesses, and runtimes.
A University of Michigan team introduced NeuroVFM, a neuroimaging foundation model trained with Vol-JEPA on over 5 million uncurated clinical MRI and CT volumes.
NVIDIA released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed hybrid MoE LLM designed to significantly increase server throughput while maintaining model accuracy.
NVIDIA released a compressed hybrid MoE LLM, Nemotron-Labs-3-Puzzle-75B-A9B, designed to deliver 2.03x server throughput compared to previous iterations.
Google AI Studio introduced an 'Import from GitHub' feature in Build mode, allowing users to turn existing repositories into editable and deployable applications.
SpaceXAI launched Grok 4.5, a model trained with Cursor specifically for coding and agentic tasks, featuring enhanced token efficiency and competitive pricing.
Ant Group’s Robbyant open-sourced LingBot-Vision, a 1B-parameter, boundary-centric vision foundation model designed to improve dense spatial perception for embodied AI.
OpenAI launched GPT-Live, a full-duplex voice model family for ChatGPT, featuring two versions (GPT-Live-1 and mini) that delegate deep reasoning to GPT-5.5.
NVIDIA released Audex (Nemotron-Labs-Audex-30B-A3B), a unified audio-text LLM capable of understanding and generating general audio while maintaining high text intelligence.
Google AI Studio introduced an 'Import from GitHub' feature in its Build mode, allowing developers to turn existing repositories into editable and deployable applications.
Anthropic has introduced Claude Science Beta, a multi-agent AI workbench designed for reproducible workflows in genomics, proteomics, and cheminformatics.
Interfaze released diffusion-gemma-asr-small, an open-source diffusion-based automatic speech recognition model capable of transcribing six languages using a parallel denoising decoder.
Mistral AI released Leanstral 1.5, an Apache-2.0 licensed Lean 4 code agent model that successfully solved 587 out of 672 problems on the PutnamBench benchmark.
Interfaze released diffusion-gemma-asr-small, an open-source diffusion-based automatic speech recognition (ASR) model capable of transcribing six languages.
WebBrain is a new open-source, local-first AI browser agent for Chrome and Firefox that automates tasks by reading pages, with support for local models like llama.cpp and Ollama.