Boston Dynamics runs robot brains on TPUs; rivals stay on Jetson
TL;DR
- Boston Dynamics offloads its System 2 planner to Google TPUs; the model runs at hundreds of billions to a trillion parameters, too big for a robot.
- For 96 robots, aggregate TCO is $14.97/hr on-device with Jetson Thor, $15.61 on RTX 6000 Pro offload, and $18.63 on B300.
- Sunday Robotics pivoted to on-device inference after home WiFi jitter proved unreliable; its ACT-2 model now reports 99.1% laundry-folding success.
Boston Dynamics runs the planning brain of its robots on Google TPUs, in a datacenter, with a System 2 model estimated at hundreds of billions to roughly a trillion parameters. "A model that size will not fit on a robot, which is why Boston Dynamics splits the stack across the network," the SemiAnalysis newsletter argues in a long piece on where robot inference should live.
The economics look tight only on paper. For a fleet of 96 robots, offloading planning to a B300 aggregates to $18.63 per hour, versus $15.61 on an RTX 6000 Pro and $14.97 running inference on-device with Jetson Thor. But Jetson Thor delivers "only about 1/10th the FLOPs of a GB200 and roughly 1/30th of its memory bandwidth," and the datacenter path lets a B300 sustain 12 robots within a 500 ms chunk deadline where an RTX Pro 6000 handles four.
Utilization warps the picture. Home robots run 1-2 hours a day, only 4-8% of the clock, while a datacenter GPU is assumed at around 90%. At those numbers, offloaded silicon reaches 46% of on-robot TCO per PFLOP for industrial deployments but only 12% for home ones. Real deployments have gone the other way from that math: Agility Robotics keeps inference on Jetson at BMW and its other customer sites specifically to avoid touching site networks, while Boston Dynamics leans on the network because generality wins over predictability in its factories.
The network is the wall.
"Home WiFi turned out to be a mess of dead zones, neighbor interference, badly configured mesh routers, and asymmetric links," the SemiAnalysis piece writes of Sunday Robotics. "Latency was mostly fine. Jitter was the problem." Wireless round trips run 10 to 50 milliseconds; a 100 Hz control loop has a 10 ms deadline. Sunday pivoted to on-device inference and now reports 99.1% laundry-folding success with its ACT-2 model.
Silicon supply complicates the choice. Jetson is a "thin red sliver" of NVIDIA accelerator output, competing for the same TSMC N4 capacity as the datacenter parts, with margins in the mid-60s versus mid-to-high 70s for datacenter accelerators. A million Jetson-class chips in 2030 works out to "only about ten thousand wafers for the year." The piece expects a cascade of approaches, with complex safety decisions staying on the robot regardless of where the planner lives.
Shared on Bluesky by 1 AI expert
Originally reported by newsletter.semianalysis.com
Read the original article →Original headline: Where Does a Robot Think — On-Device vs Datacenter Inference