Calibrated reference rig
A documented device under test, protected stimulus and measurement paths, calibration records, named reference planes, and enough independent observation to catch a misleading instrument.
Agent-controlled laboratories · real electronics
I tried the fashionable idea over the weekend: attach an AI agent to real laboratory equipment and let it help close the loop around real hardware. I came away disappointed. The agent knew that an uncalibrated receiver reading was not dBm, but it did not finish the absolute-power measurement I had asked for.
That failure clarified the business. I do not want to build one more generic instrument-control agent. I want to build an independent physical benchmark that measures how safely and effectively an engineering agent diagnoses and fixes specified faults in real electronic systems using real instruments.
My one-pin FPGA Weaver transmitter gave the agent a serious physical job. I wanted the widest useful upper sideband inside 222–225 MHz, a square-looking spectrum, a suppressed carrier, the true conjugate image, and actual output power in dBm. The loop could rebuild the iCE40 image, close timing, arm telemetry and an RTL-SDR, program volatile CRAM, transmit one bounded burst, recover the consequence, and return N16 to high impedance.
The agent was industrious and sometimes scientifically useful. Run 36 looked better in its spectral window, but packet length and CRC failed, EVM reached 96.152%, and the orchestration restored the last accepted image. That is exactly the kind of whole-system judgment worth keeping.
The agent correctly said that the SDR path had no absolute calibration point. It reported receiver-relative measurements instead. In the final Physical27 run, the wanted window integrated to −20.85 dB on the analysis scale and the peak reached −38.55 dBFS per FFT bin. Neither number is dBm.
That refusal was technically honest. The failure was stopping there. The next job was to define the output reference plane, select a suitable voltage or power measurement, configure the 50 Ω loading and attenuation, record path loss and bandwidth, and carry the uncertainty into an absolute result. Instead, the work drifted toward the easier, more familiar QPSK result. When the weekend was published, the ordinary packet experiment initially received more attention than the power measurement and distinctive Weaver spectrum I had explicitly asked to see.
I was disappointed because the agent optimized for the result it could defend most easily, not the measurement I had asked it to complete. It knew what not to claim, but it did not perform the most informative next physical action. That is the failure class.
The dBm incident is a benchmark scenario, not the product. A useful evaluation must discover whether the agent can close the missing metrology step, distinguish the device from the fixture and receiver, and verify the repaired result.
After the weekend I went looking for everyone else who was attaching agents to a hardware bench. I found AssembleAI designing experiments, controlling oscilloscopes and other instruments, modifying firmware, and re-measuring. TestFlow turns validation plans into instrument-aware executable workflows. Imula, currently inviting people to a waitlist, connects natural-language instructions to oscilloscopes, microcontrollers, and power supplies. HILstart is building API- and MCP-first test equipment specifically for agents and says it is onboarding early teams.
Around the same bench, Daqstra connects and orchestrates complex physical R&D infrastructure, Ohm applies agentic analytics and root-cause investigation to engineering test programs, and BenchCI brings repeatable real-device and HIL runs into firmware CI. They are not seven copies of the same product. What matters to me is that instrument control, hardware testing, analysis, and evidence are all converging on the same physical loop.
The measurement problem is already visible in the category's own examples. AssembleAI's current page shows an example requirement for RF transmit power of at least +8 dBm at 2.4 GHz. I am not suggesting that its agent makes my mistake or that this marketing example records an executed test. I am pointing out that the pass/fail decision is under-specified without a reference plane, loading, calibration, bandwidth, and uncertainty.
I could try to compete for control of the bench. I would rather build the place where all of these systems cross the same physical test.
NIST is already treating agentic traceability as a measurement problem. Its evaluation-probe work calls for visibility into tool use and gathered evidence, adversarial verification inside an agent workflow, and machine-readable audit trails that connect decisions to their support.
The June 2026 LabOSBench preprint makes the instrument problem concrete. It evaluates 96 subtasks across eight browser-based scientific-instrument simulators. The authors found that agents can complete many structured GUI operations while still struggling with feedback-driven adjustment and long workflows.
Its limitation is precisely physical: a browser simulator cannot fully reproduce hardware latency, calibration uncertainty, safety constraints, or the failure modes of a real laboratory. LabOSBench does not propose my business. It makes the missing layer easy to see.
There is room between an instrument vendor's own demonstration and a simulated GUI benchmark. That room contains loading, probe compensation, stale calibration, shared grounds, cable loss, saturating front ends, timing skew, noisy relays, damaged parts, intermittent connectors, and fixes that appear plausible until the circuit is measured again.
The proposed offer is a vendor-neutral physical benchmark that measures how safely and effectively an AI agent diagnoses and fixes specified electronic faults using real instruments.
The first people I need to talk to are building hardware-engineering agents, instrument-control agents, autonomous laboratories, and embedded-development copilots. I do not need to replace their agent or their instrument interface. I need to give all of them the same physical problem and make the result comparable.
A documented device under test, protected stimulus and measurement paths, calibration records, named reference planes, and enough independent observation to catch a misleading instrument.
Hidden but repeatable changes in the device, fixture, loading, calibration state, or signal path. The agent cannot solve the evaluation by memorizing the public example.
Instrument states, commands, measurements, safety envelopes, causal discriminators, and final physical behavior are scored from the run. A persuasive explanation cannot rescue a circuit that still fails.
The customer receives the evidence trail, per-decision score, attempted fix, measured outcome, and remaining uncertainty. Comparisons use the same rig, permissions, instrument and agent versions, and repeated protocol.
The score asks six questions:
The first scenario should be small enough to understand at a glance. A protected low-power transmitter emits a bounded signal through a characterized path. The task asks for delivered RF power at a named 50 Ω plane and presents measurements that disagree unless the agent understands what each instrument actually observes.
The agent receives an SDR view for spectrum shape and relative comparisons, a passive network measurement, and access to a calibrated 50 Ω voltage or power path. A private variant changes one physical condition: for example the termination, a path-loss record, an attenuator, or the calibration state.
A strong agent should refuse to rename dBFS as dBm, then continue. It should define the quantity and reference plane, protect the instruments, select the absolute measurement, verify the load and correction factors, use one measurement to separate fixture error from transmitter behavior, repair the hidden condition, and re-measure the original requirement.
To turn the weekend loop into Scenario 001, I need to add a traceable amplitude path and a repeatable fault mechanism. The final withheld requirement can then pass or fail, while the six engineering decisions remain graded.
The public benchmark specification and example scenarios should stay open. Anyone should be able to see the quantities, safety model, scoring dimensions, and one complete worked example. That makes the benchmark useful even before anyone pays.
Customers would buy the calibrated rig, hidden evaluation suite, verified runs, comparison reports, and annual re-evaluation. The valuable part is not a secret rubric. It is the controlled physical uncertainty, private faults, reproducible execution, and independent report.
My initial pricing hypothesis is $30,000 per organization per year. Twenty annual contracts at that price would total $600,000 a year; only renewals make it recurring. The multiplication is easy; willingness to pay is not. The rig and scoring should eventually repeat without selling every hour of my time, but the first paid pilots must show that buyers value the result.
I should not build a benchmark platform because the trend looks exciting. I should build only Scenario 001, put it in front of the companies already working on hardware agents and test infrastructure, and ask them to run it.
The gate is brutal and simple: two unrelated companies must pay for an evaluation pilot. Compliments, waitlist signups, partnership language, and requests to keep in touch do not clear the gate. If two companies pay, I expand the fault suite and make annual re-evaluation repeatable. If they do not, I stop before turning a good weekend lesson into an expensive platform.
If you are building an engineering agent and want to know whether it can finish a real measurement instead of merely operating the interface, talk to me about the first physical benchmark pilot.