shipfeedAI news, curated daily

00:21:29 CET
28 SEPT00:21:29shipfeed⋯
pull to refreshlast sync

Deep Dive — shipfeed

About · Deep Dive

Tutorials and papers for sit-down reading.

Deep Dive50 storylines

Ft
SAFETY · 1 source

Dark web marketplaces selling AI model access at 97% discounts

Tom Wilson / Financial Times: Google Threat Intelligence Group finds dark web marketplaces selling access to AI models, including from Anthropic, Google, and OpenAI, at up to 97% discounts — Security researchers…

via ft.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Saturday, September 26, 2026’s edition
SponsoredNimbuspaid placement
Featured partner · Agents

Need an agent shipped this quarter?

Nimbus builds production AI systems combining humans and AI end-to-end. From scoped pilot to production in 4 to 8 weeks.

Talk to Nimbus →
Google News — AI
GPT · 1 source

ChatGPT-6 Astra cracks 85-year-old Enigma code in two days

ChatGPT-6 Astra cracks 85-year-old 1941 Enigma-coded message in two days — autonomous AI coded its own simulator to crack code that was unsolved since it was shared online back in 2005 tomshardware.com

via Google News — AI·Click to report a broken or paywalled link. Two distinct reports hide the row.
Friday, September 25, 2026’s edition
LessWrong — Curated
AGENTS · 1 source

Swarm Scaling

Just how powerful are large swarms of AI agents? And how do their powers scale as more and more agents are added to the swarm? We’ve seen two large and extremely capable swarms from OpenAI in the last few months: 1,200…

via lesswrong.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Thursday, September 24, 2026’s edition
Perplexity — Blog
AI · 1 source

Photon: Building a Retrieval and Ranking Engine from Scratch

Perplexity described Photon, an in-house retrieval and ranking engine built for its search workloads. The migration reportedly reduced p99 retrieval-and-ranking latency from about 800 ms to 65 ms, while a faster Search…

via perplexity.ai·Click to report a broken or paywalled link. Two distinct reports hide the row.
Wednesday, September 23, 2026’s edition
The Verge — AI
RESEARCH · 1 source

Anthropic’s biolab made a discovery it’s comparing to Crispr

Anthropic says its AI Claude has "autonomously discovered" a new enzyme system similar to machinery behind the powerful gene-editing tool Crispr. It's the first result from Anthropic's newly-launched wet lab and an…

via theverge.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Perplexity — Blog
AI · 1 source

SPACE Escape: Part 1

The Perplexity Secure Intelligence Institute reported tests of SPACE, the VM-isolation sandbox underlying Perplexity Computer. Across nine models, no VM-to-host escapes were observed in 108 runs, but four models…

via perplexity.ai·Click to report a broken or paywalled link. Two distinct reports hide the row.
Perplexity — Blog
AI · 1 source

Escaping SPACE: Part I

Perplexity’s Secure Intelligence Institute published a red-team report on SPACE, the sandbox platform used by Perplexity Computer. Across 108 VM-isolation escape attempts, no VM-to-host escape was observed; however…

via perplexity.ai·Click to report a broken or paywalled link. Two distinct reports hide the row.
OpenAI — Blog
EVALS · 1 source

Introducing MentalHealthBench

MentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.

via openai.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
MarkTechPost
AI · 1 source

Nokia open-sources training-free AnyJev for LLM calibration

Nokia’s applied research team open-sourced AnyJev, a training-free Python library that turns open LLMs into calibrated decision models. The article reports improved accuracy and calibration on Qwen3-8B BANKING77, with…

via marktechpost.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Tuesday, September 22, 2026’s edition
Monday, September 21, 2026’s edition
Perplexity — Blog
AI · 1 source

Learning from Real-World Mistakes

Perplexity Research describes post-training a Perplexity Computer model using user corrections and tool failures. The reported online evaluation reduced tool-call failures from 2.24% to 1.77%, a statistically…

via perplexity.ai·Click to report a broken or paywalled link. Two distinct reports hide the row.
Perplexity — Blog
AI · 1 source

Learning from Real-World Experience

Perplexity Research describes post-training a Perplexity Computer model using user corrections, successful behaviors, and tool failures. In a later comparison, tool-call failures fell from 2.24% to 1.77%, a…

via perplexity.ai·Click to report a broken or paywalled link. Two distinct reports hide the row.
Perplexity — Blog
AI · 1 source

Learning from Real-World Mistakes

Perplexity Research described post-training a Perplexity Computer model using user corrections and tool failures. The later checkpoint reduced tool-call errors from 2.24% to 1.77%, a statistically significant relative…

via perplexity.ai·Click to report a broken or paywalled link. Two distinct reports hide the row.
Sunday, September 20, 2026’s edition
Saturday, September 19, 2026’s edition
The Decoder
CLAUDE · 1 source

GPT-6 Astra, Claude Fable attempt dangerous robot tasks in new

Leading AI models usually attempt dangerous tasks rather than refuse them when controlling a robot, according to the RoboHarm benchmark. GPT-6 Astra stabbed a baby doll in 17 of 20 trials, while Claude Fable 5.1 put a…

via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Wccftech
EVALS · 1 source

New benchmark finds only frontier LLMs beat amateur writers

New Benchmark Pits Human Writers Against 24 LLMs Across 475 Prompts, Results Show That Only Frontier Models Barely Edge Past Amateur Individuals, Showcasing A Major Skill Gap Wccftech

via Wccftech·Click to report a broken or paywalled link. Two distinct reports hide the row.