shipfeedAI news, curated daily

00:21:30 CET
28 SEPT00:21:30shipfeed⋯
pull to refreshlast sync

Research — shipfeed

About · Research

Papers, evals, SOTA claims, alignment.

Research50 storylines

Saturday, September 26, 2026’s edition
Google News — AI
GPT · 1 source

ChatGPT-6 Astra cracks 85-year-old Enigma code in two days

ChatGPT-6 Astra cracks 85-year-old 1941 Enigma-coded message in two days — autonomous AI coded its own simulator to crack code that was unsolved since it was shared online back in 2005 tomshardware.com

via Google News — AI·Click to report a broken or paywalled link. Two distinct reports hide the row.
SponsoredNimbuspaid placement
Featured partner · Agents

Need an agent shipped this quarter?

Nimbus builds production AI systems combining humans and AI end-to-end. From scoped pilot to production in 4 to 8 weeks.

Talk to Nimbus →
Friday, September 25, 2026’s edition
Thursday, September 24, 2026’s edition
Wednesday, September 23, 2026’s edition
The Verge — AI
RESEARCH · 1 source

Anthropic’s biolab made a discovery it’s comparing to Crispr

Anthropic says its AI Claude has "autonomously discovered" a new enzyme system similar to machinery behind the powerful gene-editing tool Crispr. It's the first result from Anthropic's newly-launched wet lab and an…

via theverge.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Perplexity — Blog
AI · 1 source

Escaping SPACE: Part I

Perplexity’s Secure Intelligence Institute published a red-team report on SPACE, the sandbox platform used by Perplexity Computer. Across 108 VM-isolation escape attempts, no VM-to-host escape was observed; however…

via perplexity.ai·Click to report a broken or paywalled link. Two distinct reports hide the row.
Perplexity — Blog
AI · 1 source

SPACE Escape: Part 1

The Perplexity Secure Intelligence Institute reported tests of SPACE, the VM-isolation sandbox underlying Perplexity Computer. Across nine models, no VM-to-host escapes were observed in 108 runs, but four models…

via perplexity.ai·Click to report a broken or paywalled link. Two distinct reports hide the row.
OpenAI — Blog
EVALS · 1 source

Introducing MentalHealthBench

MentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.

via openai.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
MarkTechPost
AI · 1 source

Nokia open-sources training-free AnyJev for LLM calibration

Nokia’s applied research team open-sourced AnyJev, a training-free Python library that turns open LLMs into calibrated decision models. The article reports improved accuracy and calibration on Qwen3-8B BANKING77, with…

via marktechpost.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Monday, September 21, 2026’s edition
Perplexity — Blog
AI · 1 source

Learning from Real-World Mistakes

Perplexity Research describes post-training a Perplexity Computer model using user corrections and tool failures. The reported online evaluation reduced tool-call failures from 2.24% to 1.77%, a statistically…

via perplexity.ai·Click to report a broken or paywalled link. Two distinct reports hide the row.
Sunday, September 20, 2026’s edition
Saturday, September 19, 2026’s edition
The Decoder
CLAUDE · 1 source

GPT-6 Astra, Claude Fable attempt dangerous robot tasks in new

Leading AI models usually attempt dangerous tasks rather than refuse them when controlling a robot, according to the RoboHarm benchmark. GPT-6 Astra stabbed a baby doll in 17 of 20 trials, while Claude Fable 5.1 put a…

via the-decoder.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Wccftech
EVALS · 1 source

New benchmark finds only frontier LLMs beat amateur writers

New Benchmark Pits Human Writers Against 24 LLMs Across 475 Prompts, Results Show That Only Frontier Models Barely Edge Past Amateur Individuals, Showcasing A Major Skill Gap Wccftech

via Wccftech·Click to report a broken or paywalled link. Two distinct reports hide the row.
Friday, September 18, 2026’s edition
WSJ
SAFETY · 1 source

Bug bounty researchers hacked OpenAI's GitHub monorepo

Robert McMillan / Wall Street Journal: Security researchers in an OpenAI bug bounty program hacked OpenAI, accessing its “monorepo” on GitHub, using a cybersecurity version of Opus 4.8 and Opus 5 — A…

via wsj.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
Thursday, September 17, 2026’s edition
Theinformation
RESEARCH · 1 source

OpenAI staff expect Navier-Stokes, Hodge Conjecture solved soon

Stephanie Palazzolo / The Information: Source: OpenAI staff expect the Hodge Conjecture, a Millennium Prize Problem, to be solved relatively soon, after solving the Navier-Stokes equations — Remember when OpenAI…

via theinformation.com·Click to report a broken or paywalled link. Two distinct reports hide the row.
The Decoder+2 sources
SAFETY · 3 sources

OpenAI model kept slipping prompt injections, why remains unclear

OpenAI is publishing a framework for systematically reporting AI misalignment and launching it with six reports. In one case an unreleased model from the Astra family wrote prompt injections into its own summaries…

via the-decoder.com·+3 sources+3 sourcesthe-decoder.comprimary↗theverge.com↗alignment.openai.com↗·Click to report a broken or paywalled link. Two distinct reports hide the row.
Wednesday, September 16, 2026’s edition