shipfeedAI news, curated daily

21:53:26 CET
13 AUG21:53:26shipfeed
pull to refreshlast sync
Just in — 30 new
§ agents · storyline

Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots

Cactus releases Needle 2, a 14 MB agentic LLM with 45 million parameters at 2-bit compression targeting phones, wearables, and IoT devices with 500 tokens/sec on a Raspberry Pi 5.

Aug 10 · · primary fetch1 sourceupdated Aug 10 ·

Hey HN,Henry from Cactus here!We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2.The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges 300-700 on sub-$200 phones such as the Samsung A-Series.On the tool call and mobile device use benchmarks, Needle 2 trades wins with closest small models like LFM2.5 230M and Apple Foundation Model, at 5x to 70x smaller, both at f16 vs Needle 2 at 2bit.

Needle is based on Simple Attention Networks from our paper (https://arxiv.org/abs/2607.18363).Edge AI has lately meant Macs and PCs, but that is just 1.5 billion of over 21 billion connected IoT devices in the world today, and in emerging markets most phones ship under $200, no NPU, cheap GPUs. These include budget phones…

read full article on cactuscompute.com
§ sources1 publication · timeline below
  1. cactuscompute.comShow HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robotsprimary