Why your AI agent needs virtual, auditable hardware
Category: Whitepaper
Agents write firmware. Agents still cannot prove that the firmware runs. Agents need hardware that the agents can flash. Agents need hardware that the agents can watch. Agents need hardware that the agents can trust. The hardware is not a closed black box. LabWired is open-source (MIT). The virtual boards are under MCP. The same oracle is in the Playground. The same oracle is in CI.
Ask an agent for a UART driver. Or ask the agent for a clock tree. Or ask the agent for a FreeRTOS task. Or ask the agent for an I²C sensor bring-up. You usually get something that compiles. The next step is more difficult. The agent cannot flash your board. The agent cannot watch the pins. The agent cannot tell you if the OLED painted. The agent cannot tell you if the peripheral clock started.
On the web, agents run tests. On the web, agents read the failure. On embedded, agents stop at a clean build.
The feedback loop
Agents get better when the agents write code. The agents run the code. The agents read the output. The agents fix the failure. Without the run step, you have confident autocomplete.
Backend work closes that loop locally.
The agent runs npm test or pytest.
The agent reads the stack trace.
The agent patches the bug.
The agent runs the test again.
Embedded does not have an equivalent for “USART2 is at 115200 with this RCC path and pinmux.” Firmware runs on a chip. Until recently, that chip sat on a desk. A human held a probe. The agent has no hands.
Caption: The loop stops at the bench. A clean build still does not prove the firmware.
Most teams end the agent session after compile. Language models invent register names with the same confidence as real register names. Language models invent HAL calls with the same confidence as real HAL calls. The compiler does not catch the invented register names. The compiler does not catch the invented HAL calls. You see the failure only when the binary runs on a target. At that time, you watch the UART, the buses, the displays, or the faults. Agents have no access to that step.
What a virtual board must do
Teams had this problem with human engineers. This problem came long before agents. Boards are scarce. Boards are flaky. Boards are hard to put under CI. See why embedded tests flake. Simulation can take the place of a board. You run the same production code that you would flash. You get a pass or a fail that you can trust. You can do these runs in parallel.
For an agent, the requirement is higher.
The loop must look like pytest.
The loop loads the firmware.
The loop runs the firmware.
The loop reads structured results.
The loop iterates.
The loop does this every few seconds.
A human is not in the middle.
The contract is short.
The contract is machine-readable.
The contract includes MCP tools, result.json, UART logs, and traces.
The contract is not a Monitor shell.
The contract is not a pile of scripts.
LabWired is that substrate. LabWired has full boards. LabWired has multi-device systems. LabWired has deterministic runs. You can audit the MIT source. The Playground is for humans. MCP is for agents. CI is for the team. LabWired is one oracle.
Other methods
Hosted chip MCP services close hello-world. An example is Veecle Chiplab. These services build an ELF. These services run on a virtual chip. These services read UART. These services are useful. These services are often chip-scoped. These services are hosted-first. These services are UART-heavy. LabWired is open core. LabWired has full boards. The boards include displays, sensors, and buses. The same run has a Playground URL.
Some agents flash real silicon. Examples are BootLoop and Embedder. These agents keep hardware on the critical path. The path includes serial. The path includes flaky timing. A bad PWM can still fry a stage. LabWired keeps hardware off that path. Prove the binary on a deterministic virtual board first. Include this proof in CI without a HIL bench. Silicon can come later.
MCP connection
LabWired speaks MCP. Claude Code, Codex, Cursor, and other hosts can load boards. These hosts can run firmware. These hosts can read results. These hosts do not need a custom plugin. The install step is one line. Then the agent calls tools. The agent calls tools in the same way that the agent already reads files. The agent calls tools in the same way that the agent already runs shell commands.
Connecting MCP is the easy part.
The value is what answers.
What answers is a register-accurate virtual board.
The board runs the code.
The board returns UART, displays, traces, and result.json.
The result is a measured signal.
The result is not a prose guess.
LabWired rule
A claim that firmware works never rests on an LLM self-report.
The agent proposes.
The simulator decides.
See also the closed loop.
Board contents
Products are MCUs plus sensors, displays, power, and buses. Often several MCUs share CAN, SPI, or a fieldbus. The long bugs are in that wiring.
Caption: Each lap returns structured evidence in seconds.
LabWired models boards and multi-board systems.
- MCUs. ARM Cortex-M, RISC-V, and Xtensa (ESP) run the same code that you would flash to silicon.
- Parts. The parts are UART, SPI, I²C, CAN, GPIO, and timers. The displays are SSD1306, Nokia 5110, and e-paper. The sensors are BME280 and MPU6050. The industrial stacks are IO-Link and UDS.
- Determinism. The same binary and the same system description give the same result.
- Evidence. The evidence is
result.json, UART logs, VCD, PC history, and display frames. - Three surfaces. The three surfaces are the Playground, MCP, and CI.
Examples
1. ESP32 e-reader (Arduino + FreeRTOS + GxEPD2)
An ESP32-WROOM-32 drives a Waveshare 2.9″ e-paper display over SPI.
The agent builds with PlatformIO.
The agent sends firmware.elf to the simulator.
The agent checks the rendered frame.
See the Arduino e-paper tutorial.
2. Full UDS ECU diagnostic check (27 services)
Two STM32H5 nodes use a simulated FDCAN bus.
The tester sends ISO-14229 requests.
The ECU runs unmodified udslib firmware.
See 27-service ECU check.
3. IO-Link master station (Master + 2 Sensors over C/Q)
An STM32L476 IO-Link master drives two sensor nodes across modeled C/Q links. See full IO-Link master.
Build, run, inspect
The core contract has three operations.
- Build. Use PlatformIO,
make, orwest. The simulator accepts standard ELFs. - Run. Execute in the simulator through the CLI or MCP (
labwired test). - Inspect. Assert against UART logs,
g_service_resultsmemory addresses, or VCD trace files.
# Example: run a test script against an ELF
labwired test --script test.yaml --firmware build/firmware.elf
Practical changes
- No hardware bottlenecks. CI pipelines run 50 virtual boards at the same time. The pipelines do not wait for a USB bench slot.
- Zero HIL flakiness. Fixed lockstep instruction stepping removes clock drift. Fixed lockstep instruction stepping removes execution jitter.
- Auditable proof. Every test run outputs machine-readable evidence. The evidence is
result.jsonand VCD traces. The run pins the evidence to git commit SHAs.
Why open source
Hardware tools are traditionally proprietary. Hardware tools are traditionally expensive. Vendor dongles traditionally lock the tools. LabWired’s core engine is MIT licensed on GitHub.
Anyone can audit the CPU models. Anyone can extend peripheral models. Anyone can run the engine locally. The local run does not need cloud dependencies. The local run does not need license servers.
Why LabWired
- Complete systems: One modeled canvas holds MCUs, buses, displays, and sensors.
- Zero-setup CI: The run on free GitHub Actions runners takes less than 2 minutes.
- Agent native: LabWired includes native Model Context Protocol (MCP) integration out of the box.