HeadlinesBriefing favicon HeadlinesBriefing.com

ESP32-S3 7-Node LLM Cluster Runs 0.4B Model

Hacker News •
×

A distributed pipeline inference engine on multiple ESP32S3 running 1.58-bit (Bit Net) Language model. Architecture This project runs a sliced 0.5B LLM across a cluster of 7 ESP32s3. One act as master and others are node. The master node runs the tokenizer and embeding and the other attention layer and MLP ran on the nodes. The master and node communicate through high speed SPI Daisy-Chain.

The master orchestrates input processing and final output generation. It runs a BPE tokenizer and INT4 token embeddings (~14MB in Flash), transmitting hidden state vectors to compute nodes via SPI CH A. Compute nodes handle transformer layers with 1.58-bit ternary quantization for attention and MLP operations, storing KV caches in PSRAM.

The pipeline flows through 6 compute nodes, each processing 4 transformer blocks with RMSNorm, RoPE, and assembly-optimized MAC operations. The final node returns results to the master for final normalization and LM head projection. Greedy sampling produces the next token ID.

The project includes comprehensive tooling: Python scripts for quantization, model packing, and tokenizer serialization, plus ESP-IDF firmware for master and node boards. Documentation covers workflow, flashing procedures, and hardware wiring.