An open-source AI accelerator, developed by AI. open TPU brings the lessons of auto-arch-tournament to AI accelerators. It asks two questions: how far can AI agents go at hardware design, and can they build the chip that runs their own inference?
open TPU is also a learning project. The whole accelerator lives in one small monorepo that you can read end to end: the hardware design (System Verilog), the instruction set, a bit-exact simulator, a kernel language and its compiler, and the host software that drives a real PCIe card.
The design runs ten modern models with their real weights on an Inspur YPCB-00338 card (Xilinx Kintex-7 xc7k480t, two DDR3 channels), and the card produces the same tokens as the simulator, bit for bit. Measured on the card: the first three models on 2026-09-29 with the production image deploy_champ_e698dcd7. Build B decodes LFM2-2.6B, Smol LM3 and Phi-4-mini 8-9% faster than e698dcd7, at 91-94% of the DRAM peak.
Source: Hacker News · Summarized by HeadlinesBriefing