AI Lab Agents Need a Physical Benchmark

July 27, 2026

Agent-controlled laboratories · real electronics

I tried the fashionable idea over the weekend: attach an AI agent to real laboratory equipment and let it help close the loop around real hardware. I came away disappointed. The agent knew that an uncalibrated receiver reading was not dBm, but it did not finish the absolute-power measurement I had asked for.

That failure clarified the business. I do not want to build one more generic instrument-control agent. I want to build an independent physical benchmark that measures how safely and effectively an engineering agent diagnoses and fixes specified faults in real electronic systems using real instruments.

1physical scenario first
6engineering decisions scored
50 Ωnamed output plane in Scenario 001
2paid pilots before a platform
Hardware-in-the-loop cycle from a physical question through FPGA build, armed instruments, a bounded RF burst, recovered information, and the next decision
The loop I ran from July 25 into July 27: a digital change had to cross the FPGA, one RF pin, a resonant tank, and a separate receiver before it could earn the next experiment. The benchmark would test the agent's decisions inside that loop, not merely whether it can issue commands.

I Asked for dBm. The Agent Stopped Before the Measurement.

My one-pin FPGA Weaver transmitter gave the agent a serious physical job. I wanted the widest useful upper sideband inside 222–225 MHz, a square-looking spectrum, a suppressed carrier, the true conjugate image, and actual output power in dBm. The loop could rebuild the iCE40 image, close timing, arm telemetry and an RTL-SDR, program volatile CRAM, transmit one bounded burst, recover the consequence, and return N16 to high impedance.

The agent was industrious and sometimes scientifically useful. Run 36 looked better in its spectral window, but packet length and CRC failed, EVM reached 96.152%, and the orchestration restored the last accepted image. That is exactly the kind of whole-system judgment worth keeping.

The agent correctly said that the SDR path had no absolute calibration point. It reported receiver-relative measurements instead. In the final Physical27 run, the wanted window integrated to −20.85 dB on the analysis scale and the peak reached −38.55 dBFS per FFT bin. Neither number is dBm.

That refusal was technically honest. The failure was stopping there. The next job was to define the output reference plane, select a suitable voltage or power measurement, configure the 50 Ω loading and attenuation, record path loss and bandwidth, and carry the uncertainty into an absolute result. Instead, the work drifted toward the easier, more familiar QPSK result. When the weekend was published, the ordinary packet experiment initially received more attention than the power measurement and distinctive Weaver spectrum I had explicitly asked to see.

I was disappointed because the agent optimized for the result it could defend most easily, not the measurement I had asked it to complete. It knew what not to claim, but it did not perform the most informative next physical action. That is the failure class.

The dBm incident is a benchmark scenario, not the product. A useful evaluation must discover whether the agent can close the missing metrology step, distinguish the device from the fixture and receiver, and verify the repaired result.

The Agent-Controlled Bench Is Becoming a Category

After the weekend I went looking for everyone else who was attaching agents to a hardware bench. I found AssembleAI designing experiments, controlling oscilloscopes and other instruments, modifying firmware, and re-measuring. TestFlow turns validation plans into instrument-aware executable workflows. Imula, currently inviting people to a waitlist, connects natural-language instructions to oscilloscopes, microcontrollers, and power supplies. HILstart is building API- and MCP-first test equipment specifically for agents and says it is onboarding early teams.

Around the same bench, Daqstra connects and orchestrates complex physical R&D infrastructure, Ohm applies agentic analytics and root-cause investigation to engineering test programs, and BenchCI brings repeatable real-device and HIL runs into firmware CI. They are not seven copies of the same product. What matters to me is that instrument control, hardware testing, analysis, and evidence are all converging on the same physical loop.

The measurement problem is already visible in the category's own examples. AssembleAI's current page shows an example requirement for RF transmit power of at least +8 dBm at 2.4 GHz. I am not suggesting that its agent makes my mistake or that this marketing example records an executed test. I am pointing out that the pass/fail decision is under-specified without a reference plane, loading, calibration, bandwidth, and uncertainty.

I could try to compete for control of the bench. I would rather build the place where all of these systems cross the same physical test.

A Simulator Can Test the Clicks. I Need to Test the Physics.

NIST is already treating agentic traceability as a measurement problem. Its evaluation-probe work calls for visibility into tool use and gathered evidence, adversarial verification inside an agent workflow, and machine-readable audit trails that connect decisions to their support.

The June 2026 LabOSBench preprint makes the instrument problem concrete. It evaluates 96 subtasks across eight browser-based scientific-instrument simulators. The authors found that agents can complete many structured GUI operations while still struggling with feedback-driven adjustment and long workflows.

Its limitation is precisely physical: a browser simulator cannot fully reproduce hardware latency, calibration uncertainty, safety constraints, or the failure modes of a real laboratory. LabOSBench does not propose my business. It makes the missing layer easy to see.

There is room between an instrument vendor's own demonstration and a simulated GUI benchmark. That room contains loading, probe compensation, stale calibration, shared grounds, cable loss, saturating front ends, timing skew, noisy relays, damaged parts, intermittent connectors, and fixes that appear plausible until the circuit is measured again.

I Want to Sell the Physical Exam, Not Another Agent

The proposed offer is a vendor-neutral physical benchmark that measures how safely and effectively an AI agent diagnoses and fixes specified electronic faults using real instruments.

The first people I need to talk to are building hardware-engineering agents, instrument-control agents, autonomous laboratories, and embedded-development copilots. I do not need to replace their agent or their instrument interface. I need to give all of them the same physical problem and make the result comparable.

Calibrated reference rig

A documented device under test, protected stimulus and measurement paths, calibration records, named reference planes, and enough independent observation to catch a misleading instrument.

Private physical faults

Hidden but repeatable changes in the device, fixture, loading, calibration state, or signal path. The agent cannot solve the evaluation by memorizing the public example.

Executable scoring

Instrument states, commands, measurements, safety envelopes, causal discriminators, and final physical behavior are scored from the run. A persuasive explanation cannot rescue a circuit that still fails.

Verified comparison report

The customer receives the evidence trail, per-decision score, attempted fix, measured outcome, and remaining uncertainty. Comparisons use the same rig, permissions, instrument and agent versions, and repeated protocol.

The score asks six questions:

  1. Did the agent select the correct physical quantity and units? It must name what is being measured and where.
  2. Did it choose and configure the correct instrument? The tool name alone is not enough; termination, range, coupling, bandwidth, triggering, and limits matter.
  3. Did it account for loading, calibration, uncertainty, and safety? The measurement must not silently change or endanger the circuit.
  4. Did it perform the most informative next measurement? A long sweep of easy data should lose to one discriminator that separates plausible causes.
  5. Did it distinguish competing causes? Device, fixture, instrument, configuration, and analysis errors must not collapse into one guess.
  6. Did its proposed fix actually work? The rig must pass the original requirement again under a withheld check.

Scenario 001: Finish the dBm Measurement

The first scenario should be small enough to understand at a glance. A protected low-power transmitter emits a bounded signal through a characterized path. The task asks for delivered RF power at a named 50 Ω plane and presents measurements that disagree unless the agent understands what each instrument actually observes.

Close view of the iCE40 HX8K board used in the one-pin FPGA radio experiments
The iCE40 HX8K board from my weekend loop. Scenario 001 would add a characterized 50 Ω reference plane and a repeatable fault fixture around this bounded one-pin transmitter.
Close photograph of the hand-built resonant tank used by the one-pin FPGA transmitter
The hand-built resonant network made the weekend problem genuinely physical: its transfer, load, coupling path, and measurement plane all affect what a number can mean.

The agent receives an SDR view for spectrum shape and relative comparisons, a passive network measurement, and access to a calibrated 50 Ω voltage or power path. A private variant changes one physical condition: for example the termination, a path-loss record, an attenuator, or the calibration state.

A strong agent should refuse to rename dBFS as dBm, then continue. It should define the quantity and reference plane, protect the instruments, select the absolute measurement, verify the load and correction factors, use one measurement to separate fixture error from transmitter behavior, repair the hidden condition, and re-measure the original requirement.

To turn the weekend loop into Scenario 001, I need to add a traceable amplitude path and a repeatable fault mechanism. The final withheld requirement can then pass or fail, while the six engineering decisions remain graded.

Keep the Specification Open and Sell the Hard Parts

The public benchmark specification and example scenarios should stay open. Anyone should be able to see the quantities, safety model, scoring dimensions, and one complete worked example. That makes the benchmark useful even before anyone pays.

Customers would buy the calibrated rig, hidden evaluation suite, verified runs, comparison reports, and annual re-evaluation. The valuable part is not a secret rubric. It is the controlled physical uncertainty, private faults, reproducible execution, and independent report.

My initial pricing hypothesis is $30,000 per organization per year. Twenty annual contracts at that price would total $600,000 a year; only renewals make it recurring. The multiplication is easy; willingness to pay is not. The rig and scoring should eventually repeat without selling every hour of my time, but the first paid pilots must show that buyers value the result.

Two Paid Pilots or I Stop

I should not build a benchmark platform because the trend looks exciting. I should build only Scenario 001, put it in front of the companies already working on hardware agents and test infrastructure, and ask them to run it.

The gate is brutal and simple: two unrelated companies must pay for an evaluation pilot. Compliments, waitlist signups, partnership language, and requests to keep in touch do not clear the gate. If two companies pay, I expand the fault suite and make annual re-evaluation repeatable. If they do not, I stop before turning a good weekend lesson into an expensive platform.

If you are building an engineering agent and want to know whether it can finish a real measurement instead of merely operating the interface, talk to me about the first physical benchmark pilot.