Close

Making simulation a bit faster?

A project log for Retro Gaming Console on RV32IM CPU (DE0 Nano FPGA)

A 32-game console on a custom dual-issue RV32IM CPU and a DE0-Nano FPGA, running DOOM, CP/M 2.2 and Chip-8 bare-metal

mjagadeesh97M_Jagadeesh97 • 2 hours ago•0 Comments

Playing the simulated core from a browser, over the bridge.

Twelve minutes of wall time per half minute of gameplay is workable for captures and miserable for iteration. This log is about the harness work that followed, and it changed nothing in the RTL and nothing in the guest binary. Only the shell changed.

The legacy flow ran Verilator in timing mode, wrote each frame pixel by pixel through Verilog file IO, and rebooted 120 million cycles on every run. The new flow keeps the same RTL and the same guest ELF and replaces the shell: a synchronous testbench with no timing constructs (tb_doom_live.v), a hand-written C++ loop (sim_main.cpp), and an SDL player on top for live keyboard and video.

RunLegacyNewFactor
3 M-cycle probe1.62 MHz6.24 MHz3.9x
Boot to first frame, 120.3 M cycles50.9 s23.8 s2.1x
8-frame scripted play, same keyfile57.7 s25.6 s2.25x
Replay 8.6 M gameplay cycles from snapshotabout 24 s bootabout 2 sboot skipped

The gain came from dropping timing mode, staying single threaded, and the x-initial 0 and x-assign fast flags. Threading was measured and lost: 175 kHz on two threads against 6.2 MHz on one, because synchronization cost more than the evaluation it parallelized on a 17-module design. Snapshots leave throughput alone and remove the 24-second boot from every run after the first, which is what makes a one-line change testable in seconds instead of a coffee break.

This is a correctness tool as much as a speed tool, and the checks are the reason the hardware bring-up was debuggable. Both flows ran the same 20-event keyfile and all 8 captured frames are byte identical. Totals match to the cycle, with the old testbench stopping 2 cycles late through its delays. A snapshot taken at frame 1 replays 6 frames with identical bytes and identical commit cycles, frame 1 at cycle 122,127,592 and instret 144,551,704 on both runs. Frame 0 stays pixel identical to the reference, the instruction suite matches 46 of 46, and the M1 to M18 micros pass.

Three failures along the way, all in the shell, with the core untouched:

Live play came next. Keys arrive through a 16-deep FIFO posted from C++, which is what makes live input possible at all, since production and consumption run at different rates. The guest transcript drains live instead of at exit, and the results file records the guest exit code and any dropped keys.

The real-time budget, measured rather than assumed: 35 tics per second at about 1.61 million cycles per level tic needs about 56.7 million cycles per second of simulation, and the harness delivers about 5.1. Scripted runs therefore play at about 10 percent of real time, against about 4 percent before. A level tic costs about 316 ms of wall time, a cheap menu tic about 76 ms. Input latency is 1 to 2 tics: 320 to 630 ms of wall time in a level, 80 to 150 ms in menus. In game time that is 1 to 2 tics, 28.6 to 57 ms, exactly as on hardware, so it feels like DOOM on a slow 386 rather than lag on a fast machine. Cycle-exact real time needs the FPGA, which this project gets in the next log.

The last piece is the browser bridge, and it exists because the SDL player needs a display on the machine running the simulation, which is not always where you are. A Python script (bridge.py, standard library only) holds the TCP connection to the sim, polls for the latest frame, and serves the page, a JSON state object, and one POST endpoint for keys. The page (ui.html, no dependencies) paints the 64,000 palette indices onto a 320x200 canvas and turns keydown and keyup into key bytes. It ignores key auto repeat and releases every held key if the tab loses focus, so a lost keyup cannot wedge a key down.

The protocol is four message types: G asks for the latest frame, K carries one key byte plus pressed or released, F carries a frame number with cycle and instret stamps plus the 64,000 index bytes, and C carries a console chunk. Sends from the sim are best effort and non-blocking, so a slow browser sees fewer frames while the simulation never waits. The sim takes a single TCP client; any number of browsers may poll the bridge.

One behaviour to know: the bridge holds one key until the guest acknowledges it, so a human tap shorter than one wall tick can vanish if no tick boundary falls between press and release. That is a property of playing at one tenth speed, not a bug in the input path.

At the end of this log the system runs DOOM correctly in simulation, plays it live at a tenth of real time, and can be driven from a browser. The board work starts in the next log.

Discussions