AI · Background · 6 min read
Where inference runs
Some decisions cannot wait for a data centre. The model has to live on the machine, or on a computer close to it.
A delivery robot at a junction has a fraction of a second. A round trip to a distant GPU is a design, not a law of nature, and sometimes it is the wrong design. Latency, a dead link, and the cost of streaming every camera frame all push computation outward — onto the robot, into the building, or onto a machine nearby.
Three places a model can run
- On the robot. Lowest delay, tightest power and heat budget, smallest model.
- In the site: a box in the closet, a factory server, a vehicle gateway. Shared by several machines that already trust the building.
- On a network of other people’s computers. Decentralized compute — phones and spare processors in a network like Acurast, or GPU markets built for rendering and inference — is an attempt to sell that nearby or on-demand thought as a service.
The honest version of “AI meets DePIN” is usually this third place. A robot does not buy a whitepaper. It buys an answer: classify this frame, plan this short motion, transcribe this noise, check this anomaly. Price it per answer, measure the delay, and ask what happens when the helper is offline. If those numbers are missing, the integration is a slide.
Small models are a feature
The models that survive on a machine are often distilled, quantized, and boring to demo next to a frontier chatbot. That is fine. Physical AI is a stack of specialists with a thinner general model above them, not one oracle in the head. When a team says they “run a large model on the robot,” ask about watts, frames per second, and what still has to be escalated.