Learn How Chips Multiply and Add
November 9, 2023

Understanding chip arithmetic opens the door to designing the entire AI machine. Wallace trees compress partial products, Han-Carlson networks carry addition quickly, fused multiply-add preserves accuracy, and Newton-Raphson turns division and transcendental functions into convergent arithmetic pipelines.

Brian Greenforest calls engineers to learn those structures transistor by transistor and build their own accelerator.

Replace Schoolbook Steps With Parallel Structure

A Wallace tree reduces many partial-product rows through carry-save compression before one final adder resolves the result. A Han-Carlson parallel-prefix adder computes carries through a balanced network with a practical wiring tradeoff.

FMA joins multiplication and addition before rounding, improving both performance and numerical behavior. Newton-Raphson uses repeated multiply-add operations to refine reciprocals and other functions rapidly.

Put the Whole Neural Network on the Chip

These arithmetic blocks can form deeply pipelined tensor engines, activation units, normalization paths, and training hardware. Old patents and public literature expose a rich engineering lineage.

Study the structures, implement them in Verilog, verify their arithmetic, and carry the design through FPGA or ASIC. The path from one adder to a complete custom AI chip remains open.

Build the Arithmetic Beneath Learn How Chips Multiply and Add

The Bit-Serial Bubbles-Free Multiplier turns local switching, state, and scheduling into a complete arithmetic engine.

Why Open-Source ASIC IP Is Hard: FMA From RTL to GDSII · Bit-Serial Bubbles-Free Multiplier · Begin with Boolean logic, a visible circuit, and the FPGA toolchain

Originally posted on LinkedIn

Brian Greenforest · (2023-11-09 17:15:57 UTC)

Open the original LinkedIn post · LinkedIn activity 7128430023149027328

LinkedIn status when archived: Visible to anyone on or off LinkedIn.

If you don't know how to multiply and add numbers, you don't know anything. Learn, how IBM, Nvidia, Intel, AMD multiply numbers, transistor-by-transistor. All patents are open and old. Make your ASIC. Innovate. If you don't know what Wallace tree and Han-Carlson are, you don't know anything. Make your own FMA. The school arithmetic they taught you is too slow. Learn how Newton-Raphson can approximate any division and exponentiation to floating point precision fast without actual cycles and counters. Put your entire neural network on chip. Don't just train a LLaMA. Make your own. #chips #math #ai #innovation #gpu #gpgpucompute #fpga #asic #nvidia #tsmc #verilog #gdsii #diy #gpt

Comments added by Brian Greenforest on LinkedIn

This comment was also preserved verbatim from Brian Greenforest’s LinkedIn data export or the public post page.

And if any of this sounds amazing, I would love to help you with it, and I'm looking for a work ASAP.

View the LinkedIn post