The FPGA You Can Debug from a Mountain

Circuits, Cartilage, and reconfigurable hardware

A design for a portable and remote FPGA workbench that preserves exact sequential state, inspectable transitions, local safety, and reproducible snapshots.

I want to edit a finite-state machine while sitting on a mountain, run it in my own FPGA cluster, stop it on the wrong transition, and carry the complete experiment home.

The mountain is not the technical requirement. It is a test of ownership.

If the development system works only when I am physically attached to one workstation, logged into one vendor environment, and willing to treat the hardware as a black box between builds, then I do not quite own the instrument. A remote FPGA IDE should make the circuit’s state more tangible, not merely move the same opaque buttons into a browser.

I designed the workbench from the boundary inward: name the circuit, capture its running state, stop it locally, and bring home everything required to replay the exact transition.

The Circuit Has Three Bodies

The workbench needs three related but distinct bodies.

The design body

This is the circuit description: Boolean relations, state elements, clocks, ports, and named finite-state transitions. It must be inspectable without running.

The execution body

This is the configured hardware or exact emulator state at one moment. It includes register values, memories, pending inputs, and the clock or sampling phase.

The replay body

This is the material required to reproduce an observation: design version, configuration, initial state, input sequence, transition count, expected values, and captured result.

A conventional “program” file is not enough. If I cannot reconstruct the execution body, I cannot reproduce the bug.

The Browser Is a Workbench, Not the FPGA

The browser can provide an infinite canvas, waveforms, state diagrams, design editing, and a portable control surface. It can also run a bounded emulator.

GPU rendering suggested a particularly interesting emulator: encode many Boolean state elements in textures, evaluate local logic in parallel, and write the next state into a separate target. Desktop and mobile graphics devices expose different output widths and texture capabilities, so the mapping must be negotiated rather than assumed.

The browser gives me a deterministic laboratory for the semantics I choose to model. The GPU texture update and the FPGA are two execution targets with deliberately different timing models.

The physical FPGA remains a separate target. Its clock distribution, routing delays, I/O standards, metastability boundaries, and synthesis results are not produced by a texture update.

Keeping these bodies separate lets each do useful work without impersonating the other.

Sequential State Needs a Sampling Rule

Combinational logic continuously relates current inputs to outputs. A D flip-flop samples an input at a defined event and retains the result.

In a parallel emulator, I can preserve that distinction with double buffering:

  1. Read every value from the current-state surface.
  2. Calculate every next-state value from that same snapshot.
  3. Write results to a different surface.
  4. Swap the surfaces at the sampling boundary.

No element may observe a neighbor’s newly written value during the same modeled edge. Otherwise evaluation order leaks into the circuit.

A larger apparent flip-flop can also be represented as a cycle of smaller pipeline stages, but only if the sampling and latency semantics are stated exactly. The emulator must not call every pipeline delay a physical D flip-flop. It is a model of a chosen temporal relation.

Glitches require another decision. A sampled synchronous model may intentionally ignore combinational transitions between edges. A timing experiment must represent delay and repeated settling steps. The IDE should name which mode is active rather than presenting both as “the circuit.”

Stop Is Part of the Machine

Remote execution needs a local stop path.

The design may expose a halt bit, breakpoint condition, failed expectation, or completed transaction. The wrapper must be able to stop issuing modeled or physical steps without relying on a network round trip. If connectivity disappears, the hardware should enter a bounded, known behavior.

That creates two control planes:

The remote interface must not offer an unbounded “write anything forever” channel by accident. A run should have a configuration identity, resource boundary, allowed ports, maximum duration, and explicit owner.

The ability to debug from anywhere should not become the ability for anyone to reconfigure the device.

Append-Only Experiments

One early architecture used two storage streams. Read an experiment state from one device, process it in the FPGA, write the next state to another, swap roles, and repeat.

Writing complete successive states costs bandwidth, but it has a beautiful property: the experiment becomes append-only. I can inspect the state before and after a transition without asking the live machine to travel backward.

For large systems, full snapshots are too expensive. The same principle can use periodic snapshots plus an append-only event log:

snapshot N
+ input event
+ clock or sampling event
+ configuration change
+ observed assertion result
= reproducible state N+1

The IDE can move along that history, compare runs, and attach a failure to the exact transition that created it.

A Mountain-Sized Debugging Session

Imagine I carry a small device with switches, LEDs, and a local FPGA. The browser shows a state machine controlling a packet parser.

I load a bounded configuration. The local controller confirms its identity. I provide a recorded byte stream and an assertion: after the final byte, packet_valid must be high and error low.

The machine runs until the assertion fails. It stops locally. The browser retrieves:

I edit the transition rule, run the emulator against the same trace, then submit a new bounded hardware run. The mountain may have intermittent connectivity. The experiment remains coherent because execution and stopping do not depend on continuous remote control.

That is a portable instrument, not merely cloud compilation.

Show Me Causality

CPU, GPU, and FPGA performance comparisons often count unlike operations. A GPU texture update may process enormous width but require a dispatch and a global state swap. An FPGA counter may change locally on every clock. Neither number alone describes a useful workload.

The IDE should therefore expose causal units:

That view can make a small circuit more educational than a large benchmark.

Distance makes the design useful because a remote boundary forces the circuit, its state, its owner, and its run record to become explicit.

If I can debug the FPGA from a mountain, I can also explain exactly what I brought with me.