FeSens/openTPU (Apache-2.0) hit the Hacker News front page (250+ points, 300+ comments). Reusing the AI-driven method from auto-arch-tournament (AI-designed RISC-V cores), the author had AI agents iterate on an inference accelerator; one repo holds SystemVerilog RTL, the ISA, a bit-exact simulator, a kernel language/compiler and a profiler. On a Kintex-7 xc7k480t PCIe card it runs Qwen3, Qwen3.5, Gemma 4, Phi-4-mini and others with real weights, token-for-token identical to the simulator; LFM2.5-230M decodes at 85.8 tok/s in 4-bit and Qwen3.5-35B-A3B reaches 3.95 tok/s with host-streamed experts.
Key Takeaways
- ✓Full stack in one repo: SystemVerilog RTL, 8x32-bit-word ISA, bit-exact Python simulator, @ol.jit kernel compiler and Lens profiler, Apache-2.0 (GitHub)
- ✓Measured device decode (4-bit, int8 head): LFM2.5-230M 85.8 tok/s, Qwen3-0.6B 31.3, Qwen3.5-2B 12.09, Gemma 4 E2B 12.14 tok/s
- ✓Uses 82-94% of the 17.1 GB/s DDR3-1066 peak while decoding; 133.33 MHz clock, four-column systolic array
- ✓MoE offload: Qwen3.5-35B-A3B at 3.95 tok/s (153 MB of experts streamed per token, 62% slot hit), LFM2.5-8B-A1B at 10.6 tok/s (offload docs)
- ✓4-bit weights (FP4 with two-level block scales, 4.25 bits/weight) decode 40-45% faster than int8, with per-model perplexity cost documented
Key Decision Metrics at a Glance
Turn your technical choice into a development budget
Compare 40 dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
openTPU (github.com/FeSens/openTPU, Apache-2.0) asks how far AI agents can go at hardware design and whether they can build the chip that runs their own inference. It applies the author's auto-arch-tournament method (previously used for AI-designed RISC-V cores); on Hacker News the author said the accelerator started at a few tokens per second and reached 80+ tok/s on small models through a recursive self-improvement loop. The monorepo contains SystemVerilog RTL, an ISA of 8x32-bit-word instructions, a bit-exact Python simulator that serves as the spec, a Python-embedded kernel language (@ol.jit) and compiler, host tools (otpu-chat, otpu-smi) and the Lens profiler with roofline and per-cycle timelines. The design is intentionally simple: one instruction per cycle across DMA, an int8 systolic matrix unit, an fp32 vector unit and a quantizer, with no caches or hidden scheduling. On an Inspur YPCB-00338 card (Kintex-7 xc7k480t, two DDR3-1066 channels, 17.1 GB/s peak, 133.33 MHz) it runs ten-plus real models token-for-token identical to the simulator: LFM2.5-230M at 85.8 tok/s, Qwen3-0.6B at 31.3, Qwen3.5-2B at 12.09 and Gemma 4 E2B at 12.14 tok/s (4-bit, int8 head, device decode), using 82-94% of DRAM bandwidth. MoE models larger than the card's 4 GiB stream experts from the host: Qwen3.5-35B-A3B runs at 3.95 tok/s. All numbers are self-reported in the README; there is no independent reproduction or like-for-like GPU comparison yet. Everything except the card runs on a laptop via the simulator.
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.