Needle is an 8-29 MB foundation-model binary with a Python library for running AI automation such as tool calls and structured extraction on tiny devices. It beats models 10x its size on mobile tool calls and matches 2-3x bigger models on extraction.
Project overview
By packing tool-calling and extraction performance that beats models 10x its size (tool calls) and matches 2-3x bigger models (extraction) into an 8-29 MB binary, Needle targets phones, wearables, robots, smart homes, cars and microcontrollers where larger general models are impractical.
Project type
AI Agent · Model Runtime · Document Processing
Use cases
Automation
Deployment
Refer to project documentation
License
Apache-2.0
Best for
Developers building AI agent automation that must run on tiny devices such as phones, wearables, robots, smart homes, cars and microcontrollers.
Key capabilities
Given app-exposed functions, Needle picks the right ones and fills every argument from the user's request; multiple requests yield multiple ordered calls; uncovered requests return an empty list rather than a guess.
Limitations and risks
General chat capacity is traded away to optimise for tool calls and extraction on tiny devices; do not adopt it as a general chatbot.
Getting started
Install with a single pip command (pip install cactus-needle), write a @needle.tool function, then run a query via needle.Needle(tools=[...]).run(query).
Evidence and sources
README: The whole model is a single 8-29 MB binary built on our Simple Attention Network, and we trade general chat capacity to beat models 10x its size on mobile tool calls and match 2-3…
README: **Tool calls**: given the functions your app exposes, Needle picks the right ones and fills every argument from what the user said. Ask for two things and you get two calls in ord…
README: pip install cactus-needle
README: Decorate a function: the signature gives the argument types, the docstring is the tool description, and `run()` completes the loop, executing your function and returning its resul…
README: Every turn returns one JSON object with `function_calls`, the model's `reasoning` and a calibrated `confidence`; an off-topic request returns an empty list rather than a guess.