Skip to content
HN On Hacker News ↗

GitHub - FeSens/openTPU: An open-source AI accelerator, developed by AI: RTL, ISA, simulator, compiler and profiler in one repo. Runs Qwen3, LFM2.5 and Qwen3.5 on a Kintex-7 PCIe card.

▲ 345 points • 401 comments • by fsbonetto • 3d ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this text is a mix of AI and human-written content.

77 %

AI likelihood · overall

AI
18% human-written 82% AI-generated
SEGMENTS · HUMAN 2 of 4
SEGMENTS · AI 1 of 4
WORD COUNT 451
PEAK AI % 77% · §3
Analyzed
Oct 6
backend: pangram/v3.3
Segments scanned
4 windows
avg 113 words each
Distribution
18 / 82%
human / AI fraction
Verdict
AI
Pangram v3.3

Article text · 451 words · 4 segments analyzed

Human AI-generated
§1 Mixed · 53%

An open-source AI accelerator, developed by AI. openTPU brings the lessons of auto-arch-tournament to AI accelerators. It asks two questions: how far can AI agents go at hardware design, and can they build the chip that runs their own inference?

§2 Human · 23%

otpu-chat running LFM2.5-230M on the FPGA card (left), with otpu-smi showing the card's utilization and DRAM bandwidth (right). A place to learn openTPU is also a learning project.

§3 AI · 77%

The whole accelerator lives in one small monorepo that you can read end to end: the hardware design (SystemVerilog), the instruction set, a bit-exact simulator, a kernel language and its compiler, and the host software that drives a real PCIe card. If you want to understand how an AI accelerator works, from a matmul in Python down to the wires, this is a good place to start. Results The design runs ten modern models with their real weights on an Inspur YPCB-00338 card (Xilinx Kintex-7 xc7k480t, two DDR3 channels), and the card produces the same tokens as the simulator, bit for bit.

§4 Human · 10%

Model Weights Decode, device Decode, wall Prefill, device DRAM while decoding LFM2.5-230M int8 59.0 tok/s 52.3 tok/s 295.6 tok/s 14.5 GB/s (85% of peak) LFM2.5-230M 4-bit, int8 head 85.8 tok/s 82.1 tok/s 335.4 tok/s 14.1 GB/s (82%) Qwen3-0.6B int8 21.6 tok/s 21.3 tok/s 92.1 tok/s 14.4 GB/s (84%) Qwen3-0.6B 4-bit, int8 head 31.3 tok/s 30.7 tok/s 103.4 tok/s 13.9 GB/s (82%) Qwen3.5-0.8B int8 17.6 tok/s 16.3 tok/s 61.4 tok/s 14.5 GB/s (85%) Qwen3.5-0.8B 4-bit, int8 head 24.5 tok/s 23.3 tok/s 66.7 tok/s 14.1 GB/s (83%) Gemma 4 E2B 4-bit, int8 head 10.57 tok/s 10.53 tok/s 32.1 tok/s 15.6 GB/s (92%) Gemma 4 E2B 4-bit, 4-bit head 12.14 tok/s 12.09 tok/s 29.9 tok/s 15.5 GB/s (91%) LFM2-2.6B int8 6.05 tok/s 6.03 tok/s 21.4 tok/s 16.1 GB/s (94%) LFM2-2.6B 4-bit, int8 head 10.96 tok/s 10.93 tok/s 20.6 tok/s 15.8 GB/s (93%) SmolLM3-3B int8 5.00 tok/s 4.99 tok/s 21.1 tok/s 16.0 GB/s (94%) SmolLM3-3B 4-bit, int8 head 8.74 tok/s 8.72 tok/s 22.8 tok/s 15.7 GB/s (92%) Phi-4-mini (3.8B) int8 3.99 tok/s 3.98 tok/s 13.8 tok/s 16.0 GB/s (94%) Phi-4-mini (3.8B) 4-bit, int8 head 6.56 tok/s 6.55 tok/s 15.0 tok/s 15.8 GB/s (92%) Qwen3.5-2B int8 8.02 tok/s 8.00 tok/s 38.2 tok/s 16.0 GB/s (94%) Qwen3.5-2B 4-bit, int8 head 12.09 tok/s 12.03 tok/s 41.7 tok/s 15.8 GB/s (92%) Qwen3.5-4B 4-bit, int8 head 5.88 tok/s 5.87 tok/s 12.9 tok/s 15.7 GB/s (92%) Gemma 4 E4B int8, 4-bit head and down 0-23 3.78 tok/s 3.75 tok/s 14.8 tok/s 16.0 GB/s (94%) Measured on the card: the first three models on 2026-09-29 with the production image deploy_champ_e698dcd7. LFM2-2.6B, SmolLM3-3B and Phi-4-mini on 2026-09-30, and Qwen3.5-2B and 4B and Gemma 4 on 2026-10-01, with build B, deploy_fused133c_79c5707a, production since then.