github.com via Hacker News

Show HN: ESP32-S3 Cluster Runs a 0.4B BitNet 1.58-Bit LLM Across Seven Microcontrollers

Edge AI Open Source Inference ai-research edge-ai

Summary

An open-source project distributes a 0.4B language model across seven ESP32-S3 microcontrollers using 1.58-bit BitNet ternary quantization, with one master node handling tokenization/embedding and six compute nodes running four transformer blocks each. Hidden states pass over SPI daisy-chain; embeddings sit in ~14MB of flash as INT4 while KV caches live in per-node PSRAM.