AI Arithmetic on Chips
November 29, 2023

AI arithmetic on silicon becomes tractable when its core structures stay visible. Carry-save adders compress many operands, parallel-prefix adders resolve carries, and one-hot state machines control deeply pipelined dataflow.

Brian Greenforest joins those elements with power delivery, H-trees, FIFOs, clock-domain crossing, IEEE 754 fused multiply-add, bfloat16, and fully pipelined backpropagation.

Build the Arithmetic and the Movement Together

CSA networks keep partial sums in redundant form until the final stage. Parallel-prefix structures deliver fast addition, while FMA pipelines organize alignment, multiplication, accumulation, normalization, rounding, and exceptional values.

Power planes, decoupling, clock H-trees, FIFO depth, and CDC protocols determine whether the arithmetic can run reliably across a large chip.

Make Training a Physical Pipeline

Bfloat16 reduces storage and multiplier cost while retaining an eight-bit exponent range suited to neural training. Large-area systems such as Cerebras show how much parallel training structure a wafer-scale fabric can host.

Brian offers online teaching in fully pipelined backpropagation. Chip teams, students, and independent designers can learn the datapath and build the accelerator that their model needs.

Build the Arithmetic Beneath AI Arithmetic on Chips

The Bit-Serial Bubbles-Free Multiplier turns local switching, state, and scheduling into a complete arithmetic engine.

Why Open-Source ASIC IP Is Hard: FMA From RTL to GDSII · Bit-Serial Bubbles-Free Multiplier · Learn How Chips Multiply and Add

Originally posted on LinkedIn

Brian Greenforest · (2023-11-29 00:03:57 UTC)

Open the original LinkedIn post · LinkedIn activity 7135418069195116545

LinkedIn status when archived: Visible to anyone on or off LinkedIn.

To do AI arithmetic well on chips, all you need to master are CSA and PPA. Everything else is just Harel UML Statecharts, meaning you must already be an expert in one-hot FSM (if you're doing HW at all!) Don't forget power planes, ground-decoupled crossovers, H-trees, FIFO, and CDC, but those are trivial technicalities not much patent-protected and there was less intentional deception in education has being committed (thanks Intel, AMD, Nvidia, especially UofA 😅). Keep in mind IEEE 754 was considered the most complex datapath by its creators, don't let the fused multiply-add struggle to put your AI chip ambition down. Google Brain's BFLOAT16 format makes things easier to wrap your head around, and training in it was proven to converge for Transformers. PM if you want to learn more. I can teach you online fully pipelined backpropagation, if you got a huge chip area like Cerebras Systems #aichips #asic #dsp #bfloat16 #fma #mac #fpga #ai #silicon #chips

Original LinkedIn media

Media shown with Brian Greenforest's original LinkedIn post
Photo 1 of 2 attached to the original LinkedIn post. Open the saved full-resolution image.
Media shown with Brian Greenforest's original LinkedIn post
Photo 2 of 2 attached to the original LinkedIn post. Open the saved full-resolution image.