Playing the simulated core from a browser, over the bridge.
Twelve minutes of wall time per half minute of gameplay is workable for captures and miserable for iteration. This log is about the harness work that followed, and it changed nothing in the RTL and nothing in the guest binary. Only the shell changed.
The legacy flow ran Verilator in timing mode, wrote each frame pixel by pixel through Verilog file IO, and rebooted 120 million cycles on every run. The new flow keeps the same RTL and the same guest ELF and replaces the shell: a synchronous testbench with no timing constructs (tb_doom_live.v), a hand-written C++ loop (sim_main.cpp), and an SDL player on top for live keyboard and video.
| Run | Legacy | New | Factor |
|---|---|---|---|
| 3 M-cycle probe | 1.62 MHz | 6.24 MHz | 3.9x |
| Boot to first frame, 120.3 M cycles | 50.9 s | 23.8 s | 2.1x |
| 8-frame scripted play, same keyfile | 57.7 s | 25.6 s | 2.25x |
| Replay 8.6 M gameplay cycles from snapshot | about 24 s boot | about 2 s | boot skipped |
The gain came from dropping timing mode, staying single threaded, and the x-initial 0 and x-assign fast flags. Threading was measured and lost: 175 kHz on two threads against 6.2 MHz on one, because synchronization cost more than the evaluation it parallelized on a 17-module design. Snapshots leave throughput alone and remove the 24-second boot from every run after the first, which is what makes a one-line change testable in seconds instead of a coffee break.
This is a correctness tool as much as a speed tool, and the checks are the reason the hardware bring-up was debuggable. Both flows ran the same 20-event keyfile and all 8 captured frames are byte identical. Totals match to the cycle, with the old testbench stopping 2 cycles late through its delays. A snapshot taken at frame 1 replays 6 frames with identical bytes and identical commit cycles, frame 1 at cycle 122,127,592 and instret 144,551,704 on both runs. Frame 0 stays pixel identical to the reference, the instruction suite matches 46 of 46, and the M1 to M18 micros pass.
Three failures along the way, all in the shell, with the core untouched:
- A frame watcher that sampled one posedge late reported one frame dumped while writing no file. It showed up at all only because MM_DUMP and MM_EXIT dual-issued in the same cycle, which is exactly the class of bug the harness exists to catch.
- A throughput metric that included the snapshotted prefix in its wall time read 76 MHz. Counting only simulated cycles gives 5.1 MHz.
- Two DPI lessons: calling a DPI export from C++ aborts unless svSetScope points at the exporting scope first, and Verilator rejects 4-state [31:0] on exports, so every DPI signature uses plain int.
Live play came next. Keys arrive through a 16-deep FIFO posted from C++, which is what makes live input possible at all, since production and consumption run at different rates. The guest transcript drains live instead of at exit, and the results file records the guest exit code and any dropped keys.
The real-time budget, measured rather than assumed: 35 tics per second at about 1.61 million cycles per level tic needs about 56.7 million cycles per second of simulation, and the harness delivers about 5.1. Scripted runs therefore play at about 10 percent of real time, against about 4 percent before. A level tic costs about 316 ms of wall time, a cheap menu tic about 76 ms. Input latency is 1 to 2 tics: 320 to 630 ms of wall time in a level, 80 to 150 ms in menus. In game time that is 1 to 2 tics, 28.6 to 57 ms, exactly as on hardware, so it feels like DOOM on a slow 386 rather than lag on a fast machine. Cycle-exact real time needs the FPGA, which this project gets in the next log.
The last piece is the browser bridge, and it exists because the SDL player needs a display on the machine running the simulation, which is not always where you are. A Python script (bridge.py, standard library only) holds the TCP connection to the sim, polls for the latest frame, and serves the page, a JSON state object, and one POST endpoint for keys. The page (ui.html, no dependencies) paints the 64,000 palette indices onto a 320x200 canvas and turns keydown and keyup into key bytes. It ignores key auto repeat and releases every held key if the tab loses focus, so a lost keyup cannot wedge a key down.
The protocol is four message types: G asks for the latest frame, K carries one key byte plus pressed or released, F carries a frame number with cycle and instret stamps plus the 64,000 index bytes, and C carries a console chunk. Sends from the sim are best effort and non-blocking, so a slow browser sees fewer frames while the simulation never waits. The sim takes a single TCP client; any number of browsers may poll the bridge.
One behaviour to know: the bridge holds one key until the guest acknowledges it, so a human tap shorter than one wall tick can vanish if no tick boundary falls between press and release. That is a property of playing at one tenth speed, not a bug in the input path.
At the end of this log the system runs DOOM correctly in simulation, plays it live at a tenth of real time, and can be driven from a browser. The board work starts in the next log.
M_Jagadeesh97
Discussions
Become a Hackaday.io Member
Create an account to leave a comment. Already have an account? Log In.