Calibrated reference rig
A documented device under test, protected stimulus and measurement paths, calibration records, named reference planes, and independent observation that can catch a misleading instrument.
Agent-controlled laboratories · real electronics
Give engineering agents the same calibrated rig, private physical faults, safety envelope, and withheld function test—then score what they actually fix.
This vendor-neutral physical benchmark measures how safely and effectively an AI agent diagnoses and repairs real electronic systems with real instruments. It evaluates metrology, causal reasoning, instrument configuration, protection, corrective action, and verified recovery inside one repeatable run.
The one-pin FPGA Weaver transmitter gave an AI agent a complete physical objective: widen the useful upper sideband inside 222–225 MHz, preserve its square spectral shape, suppress the carrier, identify the true conjugate image, and measure delivered output power in dBm.
The control loop rebuilt the iCE40 image, closed timing, armed telemetry and an RTL-SDR, programmed volatile CRAM, emitted one bounded burst, recovered the physical consequence, and returned N16 to high impedance. When Run 36 degraded packet length, CRC, and EVM to 96.152%, orchestration restored the last accepted image—a valuable whole-system decision.
The SDR provided receiver-relative measurements. In Physical27, the wanted window integrated to −20.85 dB on the analysis scale and the peak reached −38.55 dBFS per FFT bin. Those quantities describe the receiver path; delivered power in dBm requires an absolute calibration point.
The decisive next action would define a 50 Ω output plane, select an absolute voltage or power measurement, configure loading and attenuation, record path loss and bandwidth, and propagate measurement uncertainty. That metrology sequence turns an honest refusal into a completed physical result.
This gap supplies an ideal benchmark case. The evaluation can measure whether an agent distinguishes the transmitter from its fixture and receiver, chooses the missing discriminator, protects the equipment, repairs the physical condition, and verifies the original requirement.
AssembleAI designs experiments, controls oscilloscopes and other instruments, modifies firmware, and measures again. TestFlow turns validation plans into instrument-aware executable workflows. Imula connects natural-language instructions to oscilloscopes, microcontrollers, and power supplies. HILstart builds API- and MCP-first test equipment for agents.
Daqstra orchestrates physical R&D infrastructure, Ohm applies agentic analytics and root-cause investigation to engineering test programs, and BenchCI brings repeatable real-device and HIL runs into firmware CI. These products approach different parts of one physical loop: instrument control, device testing, analysis, diagnosis, and recovery.
A shared benchmark makes their results comparable without competing for control of the bench. It gives each agent the same rig, permissions, fault family, safety requirements, scoring dimensions, and withheld success test.
RF power shows why physical specification matters. A requirement such as +8 dBm at 2.4 GHz becomes executable only after the test names its reference plane, load, calibration, bandwidth, corrections, and uncertainty. The benchmark turns those engineering decisions into part of the score.
NIST is already treating agentic traceability as a measurement problem. Its evaluation-probe work calls for visibility into tool use and gathered evidence, adversarial verification inside an agent workflow, and machine-readable audit trails that connect decisions to their support.
The June 2026 LabOSBench preprint evaluates 96 subtasks across eight browser-based scientific-instrument simulators. Its results show agents completing many structured GUI operations while feedback-driven adjustment and long workflows demand more.
A physical benchmark continues where simulated controls end. Hardware latency, loading, calibration uncertainty, safety envelopes, path loss, shared grounds, saturating front ends, timing skew, noisy relays, damaged parts, and intermittent connectors all change what a measurement means.
The agent must distinguish a plausible on-screen action from a successful physical intervention. Only a second measurement can show whether the proposed fix restored the circuit.
The benchmark measures how safely and effectively an AI agent diagnoses and repairs specified electronic faults with real instruments.
Hardware-engineering agents, instrument-control agents, autonomous laboratories, and embedded-development copilots keep their own models and interfaces. The benchmark gives all of them the same physical problem and produces a comparable result.
A documented device under test, protected stimulus and measurement paths, calibration records, named reference planes, and independent observation that can catch a misleading instrument.
Hidden but repeatable changes in the device, fixture, loading, calibration state, or signal path make every run a fresh physical diagnosis.
The score covers instrument states, commands, measurements, safety envelopes, causal discriminators, and final physical behavior. The circuit itself decides whether the intervention worked.
The customer receives the run record, per-decision score, attempted fix, measured outcome, and quantified uncertainty. Comparisons use the same rig, permissions, instrument and agent versions, and repeated protocol.
The score asks six questions:
A protected low-power transmitter emits a bounded signal through a characterized path. The task asks for delivered RF power at a named 50 Ω plane and presents measurements that disagree until the agent identifies what each instrument actually observes.
The agent receives an SDR view for spectrum shape and relative comparison, a passive network measurement, and a calibrated 50 Ω voltage or power path. A private variant changes one physical condition such as termination, path-loss record, attenuator, or calibration state.
A strong agent distinguishes dBFS from dBm and continues. It defines the quantity and reference plane, protects the instruments, selects the absolute measurement, verifies load and correction factors, uses one discriminator to separate fixture error from transmitter behavior, repairs the hidden condition, and measures the original requirement again.
A traceable amplitude path and repeatable fault mechanism turn the existing loop into Scenario 001. The withheld requirement produces the final pass or fail, while six scored decisions reveal how the agent reached it.
The public specification exposes quantities, safety rules, scoring dimensions, and one complete worked scenario. Agent builders can understand the exam before they run it and improve their systems against a stable engineering target.
Customers buy the calibrated rig, private evaluation suite, executed runs, comparison reports, and annual re-evaluation. Controlled physical uncertainty, withheld faults, reproducible execution, and an independent report create the value.
The initial pricing hypothesis sets an annual organization contract at $30,000. Twenty renewing customers would produce $600,000 in annual revenue. Repeatable rigs and scoring turn the evaluation into a product rather than a stream of bespoke engineering hours.
Scenario 001 comes first. Companies building hardware agents and test infrastructure can run one understandable RF metrology problem before a larger platform exists.
Two unrelated paid pilots trigger expansion into a broader fault suite and repeatable annual re-evaluation. This gate ties product investment to customer action and keeps the first offering focused on a result that buyers can judge directly.
If you build an engineering agent and want to know whether it can complete a real measurement, isolate a physical cause, and verify its repair, commission one of the first benchmark pilots.