Close
0%
0%

Retro Gaming Console on RV32IM CPU (DE0 Nano FPGA)

A 32-game console on a custom dual-issue RV32IM CPU and a DE0-Nano FPGA, running DOOM, CP/M 2.2 and Chip-8 bare-metal

Similar projects worth following
55 views
0 followers
This is a retro console built on a RISC-V CPU I designed and a Terasic DE0-Nano (Altera Cyclone IV E). It boots into a launcher with 32 games across three runtimes: DOOM Episode 1 and a Wolfenstein-style raycaster written in C against the console framebuffer, CP/M 2.2 software including Microsoft BASIC-80 under a bare-metal RunCPM port, and 19 Chip-8 titles from the 1977 COSMAC VIP era.

The whole stack is custom. The CPU is a dual-issue in-order RV32IM core with a hybrid branch predictor, a 256-entry branch target buffer and a return address stack, measured at CPI 0.727 on a 498-million-instruction DOOM run. The FPGA side adds an SDRAM arbiter with a 1K-word I-cache, a 4 KB UART TX FIFO with a hardware frame commit path, an input decoder, and an autonomous I2C master that drives a physical 128x64 OLED. The test harness runs the same RTL in Verilator about 4x faster than the original flow, adds snapshot/restore so runs skip the 24-second boot, and exposes the framebuffer and keyboard

1. Overview

The 32-game launcher on the host viewer.

The 32-game launcher on the host viewer.

This project is a multi-system retro console on a Terasic DE0-Nano. It runs software from three periods on one board: DOOM (1993), CP/M 2.2 and Microsoft BASIC-80 (late 1970s and 1980s), and Chip-8 (1977). All three runtimes sit behind one bare-metal launcher, on a RISC-V core written for this project. Nothing runs on top of an operating system.

The work divides into four parts: the CPU, the FPGA machine built around it, the test harness used to develop both, and the console software. The sections below follow that order.

Terasic DE0-Nano, front. Cyclone IV E EP4CE22F17C6N, 32 MB SDRAM, 50 MHz oscillator, USB-Blaster for JTAG.

Terasic DE0-Nano, front. Cyclone IV E EP4CE22F17C6N, 32 MB SDRAM, 50 MHz oscillator, USB-Blaster for JTAG.

2. At a glance

ItemValue
BoardTerasic DE0-Nano, Altera Cyclone IV E EP4CE22F17C6N
CPUCustom dual-issue in-order RV32IM at 50 MHz, CPI 0.727 measured on DOOM
System memory32 MB SDRAM, 16-bit bus
Memory layout158.6 KB firmware, 12 MB heap, 512 KB stack, 5.11 MB asset blob
Device utilization17,743 logic elements (79 percent), 6,984 flip-flops, 54 M9K, 12 DSP
Game catalog32 titles: 7 native and 3D, 6 CP/M 2.2, 19 Chip-8
Primary display64x32 RGB332 framebuffer at 45 FPS over 921,600 baud UART
Hardware display0.96 inch I2C OLED, 128x64, 32.6 FPS, hardware 2x2 scaler
InputHost keyboard over UART with press and release events, decoded in RTL
Load5.27 MB image over serial in about 57 seconds, CRC32 verified
TimingSlow corner worst slack -3.998 ns; nominal silicon runs on the bench

3. The bandwidth problem and the 64x32 decision

DOOM renders to a native Mode 13h buffer of 320x200, which is 64,000 bytes. Frame logic completes in about 28 ms (roughly 35.7 FPS internally), but the serial link is the ceiling. At 921,600 baud the theoretical maximum is 92.16 KB/s, and a 64,000-byte frame takes 694 ms on the wire. The game logic ran at full speed while the display ran at 1.44 FPS.

The intermediate step was 160x100: 16,000 bytes per frame, 174 ms, about 5.75 FPS. Playable, but still behind human reflex speed.

The console settled on a uniform 64x32 framebuffer, 2,048 bytes per frame, which crosses the wire in 22.2 ms and gives a steady 45 FPS. The choice also holds up on other axes: 64x32 is the native resolution of the 1977 Chip-8 virtual machine, it maps cleanly to fixed-point raycasting columns, and it scales 2x2 to a 128x64 monochrome OLED with no geometric distortion.

Resolution vs wire time at 921,600 baud (92.16 KB/s):

320x200 (64,000 bytes):  694 ms per frame  ->  1.44 FPS
160x100 (16,000 bytes):  174 ms per frame  ->  5.75 FPS
 64x32   (2,048 bytes):   22 ms per frame  -> 45.00 FPS

In the memory map the framebuffer occupies a fixed 2 KB window at 0xFFF00000 to 0xFFF007FF. Each byte is an 8-bit RGB332 pixel for serial streaming. The OLED controller interprets any non-zero byte as a lit pixel, so the same buffer drives both displays without conversion.

Code, from games/common/console.h:

#define FB_W        64
#define FB_H        32
#define FB          ((volatile uint8_t *)0xFFF00000u)

#define MMIO_BASE   0xFFFFF000u
#define MM_KEY      (*(volatile uint32_t *)(MMIO_BASE + 0x014u))
#define MM_DUMP     (*(volatile uint32_t *)(MMIO_BASE + 0x028u))

void dump_tiny(void) {
    MM_DUMP = 2u; /* Trigger 64x32 tiny mode dump in mem_top_fpga */
}

4. Frame commit path: 154 microseconds of CPU time per frame

When a frame is ready, the game writes 2 to MM_DUMP (0xFFFFF028). This replaces the earlier multi-byte host command protocol (0xFC, 0xFD, 0xFE) with a single MMIO write.

In mem_top_fpga.v, that write makes the SDRAM arbiter read 512 words (2,048 bytes) from the framebuffer backing store at FB_SDR_WORD (0x7FC000). Reading 512 words takes about 7,680 system clock cycles, about 153.6 microseconds at 50 MHz. The FPGA has a 4,096-byte TX FIFO, so the whole frame fits in one burst. As soon as the last word lands in the FIFO, the arbiter releases the gated core clock and the CPU continues.

The CPU is stalled for about 154 microseconds per frame. The autonomous UART feeder drains...

Read more »

  • 1 × DE0-Nano FPGA board Terasic DE0-Nano, Altera Cyclone IV E EP4CE22F17C6N, 22,320 LEs, 32 MB SDRAM, 50 MHz oscillator, USB-Blaster
  • 1 × CP2102 USB-to-TTL serial module Silicon Labs CP2102, 3.3 V logic, 921,600 baud, carries upload, framebuffer, console and keyboard
  • 1 × .96 inch I2C OLED module 0128x64 monochrome, SSD1306 or SH1106 or SSD1315 class controller, address 0x3C
  • 1 × Mini-USB cable Board power and USB-Blaster JTAG programming
  • 1 × Jumper wires OLED: VCC to JP1 pin 29, GND to pin 12, SCL to pin 6, SDA to pin 8 | Serial: module TXD to JP1 pin 2, module RXD to JP1 pin 4, GND to pin 12

View all 7 components

  • Fixed Point Ray Casting and Making the Final Version

    M_Jagadeesh97 • an hour ago • 0 comments

    Sixty-four column rays per frame, 22.10 fixed point, no floating point unit.

    The last set of games to add were the ones that test the machine hardest: real-time 3D in software, and a virtual machine from 1977.

    The raycaster. Written in C against the 64x32 framebuffer, the first prototype ran at about 3 FPS for two separate reasons.

    Floating point was the first. Sine and cosine through soft float choked the dual-issue core, so all math moved to 22.10 fixed point with FP_SHIFT at 10 and FP_ONE at 1024, and the trigonometry became a 256-entry lookup table generated at startup from a Bhaskara rational approximation.

    Fisheye distortion was the second. Measuring Euclidean distance straight from the player to the wall intersection bends the image along wall edges. The fix is to use the perpendicular distance to the camera plane:

    int32_t perp_wall_dist = (side == 0) ?
        (side_dist_x - delta_dist_x) : (side_dist_y - delta_dist_y);
    int line_height = (int)(((SCREEN_H) << FP_SHIFT) / perp_wall_dist);
    

    The column step is a DDA over the map grid, accumulating side distance per axis and stepping the smaller one each iteration:

    while (!hit) {
        if (side_dist_x < side_dist_y) {
            side_dist_x += delta_dist_x; map_x += step_x; side = 0;
        } else {
            side_dist_y += delta_dist_y; map_y += step_y; side = 1;
        }
        if (world_map[map_y][map_x] > 0) hit = 1;
    }
    

    The division budget is where this game meets the hardware. Each frame casts 64 column rays and wall heights need one division per column. The FPGA's iterative divider takes 33 cycles, so the column pass costs under 4,000 divider cycles in total. The whole pass, raycasting plus texture mapping plus sprite sorting, runs in under 400,000 cycles per frame, which fits the 45 FPS budget at 50 MHz with headroom to spare.

    Chip-8. The interpreter implements all 35 opcodes: 16 8-bit registers, a 16-bit index register, a 16-level call stack, 60 Hz delay and sound timers tied to the millisecond tick, and XOR sprite drawing with collision detection.

    Two Chip-8 ROMs from the active list were moved to the reserve vault for a reason worth recording. Guess the Number and Kaleidoscope were written for the 1977 COSMAC VIP hex keypad, a 4x4 matrix of keys 0 through F. With a modern keyboard and no on-screen instructions, pressing keys produced zero reaction, which reads as a broken emulator rather than a period-correct input scheme. Duplicate variants such as Pong 2 and vertical Brix, and system test ROMs, went the same way, making room for native titles with clear keyboard layouts.

    The vault (games/vault/) holds those ROMs along with extra CP/M adventures such as ELIZA and Oregon Trail, plus a document describing the swap. Swapping a title in is one table entry and one build. A slightly shorter catalog in which every game responds to the keyboard is a better console than a longer one where some titles look broken.

    Native games. Snake (circular-buffer tail management with an attract-mode AI), 2048 (matrix shifting, merging, pseudo-random spawn), Tetris (4x4 tetromino rotation, line clear, wall kicks), Minesweeper and Flappy Bird, all drawing directly into the framebuffer and polling the hardware key register.

    That completes the catalog: 7 native and 3D titles, 6 CP/M 2.2 titles and 19 Chip-8 games, 32 in total, all behind one launcher on a board that fits in a hand, running on a CPU designed for this project. The next work is about cutting the last cord to the laptop: booting the bitstream and image from the on-board configuration flash, a VGA output on the GPIO header, and a local controller.

  • RunCPM Prompt Support

    M_Jagadeesh97 • an hour ago • 0 comments

    The RunCPM prompt with DIR, ASM, DDT and MBASIC.

    The console runs original CP/M 2.2 software, including Microsoft BASIC-80, through a port of RunCPM to bare metal. The Z80 decoder supplies the register file, a 64 KB virtual address space, and the zero page vectors at 0x0000 for warm boot and 0x0005 for the BDOS entry.

    Standard RunCPM expects a hosted POSIX filesystem to expose its virtual drives. There is none here, so the disk is packed ahead of time: a script scans a directory, converts filenames into 11-byte CP/M form with 8.3 space padding, and writes the disk blob that goes into the asset container. The BDOS translation layer traps the calls and routes them:

    • Console output, functions 1, 2, 6 and 9, goes to the hardware UART and to a four-line by sixteen-character virtual terminal rendered into the 64x32 framebuffer with a 4x6 bitmap font, which is what makes the phosphor-green screen in the capture above possible on a monochrome display.
    • File access, functions 14, 15, 16, 20 and 21, reads the in-memory RAM disk arena. Writes update it in place, so file changes persist for the session.
    • AUTOEXEC.TXT is synthesized into the RAM disk at boot so the internal CCP can auto-launch a title such as MBASIC STARTRK without a human typing at the prompt.

    One bug is worth recording in full, because its symptom was a hang rather than an error.

    Files larger than 16 KB are split across CP/M extents. The record offset was being computed from the record field alone, so at record 128 the address wrapped back to the start of the file and the read looped forever. The correct linear offset uses both fields:

    offset = ((fcb->ex * 128) + fcb->cr) * 128
    

    MBASIC.COM is 24 KB, so this hit on the first program that mattered. A naive implementation does not crash; it spins, which costs an afternoon before the penny drops.

    ESC during blocking console input needed its own path for the same reason the launcher did. The character input loop polls the hardware key register every 22 ms and unwinds to the launcher on keycode 27, so a CP/M program that blocks on input can still be exited with one key, and the console does not have to be rebooted to escape a text adventure.

    Speed: at 50 MHz with SDRAM access latencies, the Z80 interpreter runs at roughly a 2 to 4 MHz equivalent, which is the range of the machines that ran this software originally. Star Trek, Hamurabi and Lunar Lander are period-appropriate at that speed rather than merely tolerable.


    Star Trek (1971)

    Hamurabi (1968)

    Hunt the Wumpus

    Porting this was the most retrocomputing part of the project so far: original 8080-era binaries, unmodified, running on a RISC-V core I designed, with the disk image packed at build time and the console rendered into a 64x32 framebuffer.

  • Re-running games to catch a few bugs

    M_Jagadeesh97 • an hour ago • 0 comments

    Both displays live: Snake on the host viewer and on the OLED.

    Viewing games on a laptop screen proved the console worked, but a console needs its own display. The DE0-Nano was expanded with a 0.96-inch 128x64 monochrome I2C OLED on the GPIO header, driven by an RTL master rather than by the CPU, and two bugs followed that are worth writing down in full because each looked like a different problem than it was.

    The wiring is four pins. VCC to JP1 pin 29, ground to pin 12, SCL to GPIO_0[2] which is FPGA pin A3, and SDA to GPIO_0[4] which is FPGA pin B4. Weak pull-ups were enabled on both I2C pins in the Quartus settings file to complement the 4.7 kOhm resistors on the breakout, and the controller drives both lines open drain: it never drives the bus high, only low or high-Z, and lets the pull-ups bring the line up.

    Bug 1: random static across 7 of 8 screen rows.

    The display powered up with turquoise noise and black dots, and once a game started only a narrow 8-pixel band at the top updated. The driver had been written assuming the SSD1306 horizontal addressing mode, where streaming 1,024 continuous data bytes wraps from column 127 of page 0 to column 0 of page 1 and advances through all eight pages.

    The module on the bench was an SH1106 class controller. That part does not implement horizontal addressing or the column and page range commands; it is permanently in page addressing mode. A 1,024-byte burst therefore wrote 128 bytes to page 0, wrapped to the start of page 0 and repeated eight times. Pages 1 through 7 never received anything and stayed in their power-on state, which is why the failure looked like a dead region rather than a protocol mismatch.

    The fix streams page by page. For every frame the state machine iterates all eight pages: set the page address with 0xB0 | page, reset the column with 0x00 and 0x10, then stream 128 data bytes.

    MODE_PAGE_CMD: begin
        case (cmd_idx)
            5'd0: begin tx_byte <= 8'h00; cmd_idx <= 5'd1; state <= S_SEND_BYTE; end
            5'd1: begin tx_byte <= {4'hB, 1'b0, page_idx}; cmd_idx <= 5'd2; state <= S_SEND_BYTE; end
            5'd2: begin tx_byte <= 8'h00; cmd_idx <= 5'd3; state <= S_SEND_BYTE; end
            5'd3: begin tx_byte <= 8'h10; cmd_idx <= 5'd4; state <= S_SEND_BYTE; end
            default: state <= S_STOP;
        endcase
    end
    

    Bug 2: the display froze on page 0.

    With page streaming in place, the panel updated one page and stopped. The column counter was declared as seven bits. After column 127 it incremented to 128, wrapped silently back to 0, and the terminal check against 128 never fired. The state machine looped inside page 0 forever, and because there is no timeout in the FSM, it looked like the display had died mid-frame.

    Widening the counter to eight bits let 128 be represented and the transition to STOP fired. Two lines of Verilog, roughly two days of bench time, and a lesson about counters that must be able to hold the value one past their last valid index.

    The scaler is a 2x2 integer scale in hardware. The panel is 128x64 organized as eight pages of eight vertical rows, and the console framebuffer is 64x32, so each console row bit is duplicated twice vertically and each column byte is sent twice horizontally:

    wire [5:0]  col_x        = col_idx[6:1];
    wire [7:0]  oled_scaled_byte = {px3_r, px3_r, px2_r, px2_r, px1_r, px1_r, px0_r, px0_r};
    

    Result: eight pages of 133 bytes per frame at a 312.5 kHz Fast-Mode clock, 30.6 ms per frame, 32.6 FPS, for 320 logic elements and no additional M9K blocks. The host viewer at 45 FPS and the OLED at 32.6 FPS run from the same framebuffer at the same time, which is the point: the panel is the console's own screen, not a debug output.

    One design decision worth stating: acknowledgement sampling is unconditional. If the display is unplugged mid-frame, the master does not wait for an ACK that will never arrive, so a missing display can never lock the memory arbiter and stall the CPU.

  • Making a Retro Gaming Console

    M_Jagadeesh97 • an hour ago • 0 comments

    The launcher: 32 titles, categories, selection cursor.

    The board ran DOOM, and only DOOM. Switching games meant restarting the upload script, transferring the image over serial and waiting, so the machine was a demo rather than a console. This log is what changed, and it is the point where the project becomes a console.

    The catalog is 32 titles in three runtime categories:

    • Native and 3D, 7 titles: DOOM Episode 1, a Wolfenstein-style raycaster, Minesweeper, Flappy Bird, Snake, 2048, Tetris.
    • CP/M 2.2, 6 titles: Star Trek, Ladder, Hamurabi, Lunar Lander, Hunt the Wumpus, and a RunCPM prompt with DIR, ASM, DDT and MBASIC.
    • Chip-8, 19 titles: Space Invaders, Pong, Brix, Tank, Blitz, Missile Defense, UFO, Cave, Lunar Descent, Airplane, Connect 4, Tic-Tac-Toe, 15 Puzzle, Logic Puzzle, Merlin, Hidden Pairs, Blinky, Maze, Wipeoff.

    All of it is one firmware image: 158.6 KB of code plus a 5.11 MB asset blob, uploaded in about 57 seconds. The blob is a single aligned container with a header, a ROM table, the CP/M virtual disk, the DOOM IWAD and the Chip-8 ROMs, built by one script so that adding a game never means hand-editing offsets.

    The first problem was memory. Consolidating several runtimes into one binary created layout conflicts immediately: DOOM assumed the heap could grow above .bss, which overwrote the regions where ROMs and CP/M images lived. The linker script now binds every section to a fixed boundary on the 32 MB SDRAM:

    0x00000000 - 0x00027FFF   158.6 KB  firmware (.text, .rodata, .data, .bss)
    0x00028000 - 0x00C27FFF     12 MB   dynamic heap
    0x00C28000 - 0x00CA7FFF    512 KB   execution stack
    0x00CA8000 - 0x014A7FFF      8 MB   asset window (menu.blob, 5.11 MB used)
    0x014A8000 - 0x01FEFFFF   11.3 MB   unallocated
    0x01FF0000 - 0x01FFFFFF     64 KB   framebuffer SDRAM backing
    

    and ends with a build-time guard that replaces the 64 MB simulation boundary:

    ASSERT(_wad_end < 0x02000000, "Firmware and assets exceed 32 MB SDRAM window")
    

    That assert is the direct descendant of the paste-over bug from log 3. A layout mistake now fails the build instead of surfacing as missing assets halfway through an upload.

    The second problem was getting back. Once a game started, the only way to the menu was pressing the board's reset button, which cleared the SDRAM controller and forced a 57-second re-upload before a different game could be played. Bare-metal RISC-V has no process model, so a game loop deep inside its own call stack cannot be unwound with ordinary returns.

    The solution is a freestanding setjmp and longjmp in games/common/crt0.S. The jump buffer is 14 words, 56 bytes: the return address, the stack pointer and the twelve saved registers.

    setjmp:
        sw ra,  0(a0)
        sw sp,  4(a0)
        sw s0,  8(a0)
        ...
        sw s11, 52(a0)
        li a0, 0
        ret
    

    The launcher establishes the jump target before dispatching, so any game can return by calling longjmp with the same buffer:

    g_in_game = 1;
    if (setjmp(menu_jmp_buf) == 0) {
        launch_game(sel);
    }
    g_in_game = 0;
    key_flush();
    draw_menu(sel, top);
    

    Re-entrancy needed its own pass. Entering DOOM, leaving with ESC and selecting DOOM again crashed with heap corruption, because the zone allocator retained stale allocation pointers from the previous run. Now doom_main resets its zone on every entry, resets static player state, and flushes the key register before drawing the first frame. ESC is disabled inside DOOM's own menu so the key means the same thing across all 32 titles.

    The result is that ESC returns to the launcher from any game in well under a second, with no reset and no re-upload. Combined with the launcher, this is what makes the machine a console rather than a set of demos: switch from DOOM to a CP/M adventure to a Chip-8 game and back, on a device that fits in one hand.

  • A Decent Working Prototype on FPGA

    M_Jagadeesh97 • an hour ago • 0 comments

    DOOM on the board, frames arriving over the serial link.

    The game was alive and the only output was tic lines scrolling past in a terminal. A VGA DAC board would have solved the display in an afternoon, but none was on the desk, and the serial wire was already there. DOOM renders into a contiguous 64 KB buffer of palette indices. The question was whether the frames could simply come down the wire.

    The first calculation was what 64,000 bytes at 92,160 bytes per second does to a game loop. The answer is 0.7 seconds per frame, and the question underneath is what the game sees during those 0.7 seconds. DOOM reads time from its tick counter, and the tick counter is the cycle counter divided by a thousand. Letting the core run while the UART drains would inject two dozen phantom tics between renders: physics advancing in large steps, controls sampled with stale timing, monster thinkers skipping animation frames. The frames would arrive, and the game logic would be wrong.

    So the frame commit freezes the game clock for the duration of the transfer. When the engine writes MM_DUMP, the memory system drops the core clock enable, and the game clock, which divides the gated cycle counter, reads zero elapsed milliseconds:

    uint32_t DG_GetTicksMs(void) { return MM_CYC_LO / 1000u; }
    

    Zero cycles elapse and zero milliseconds pass, and the engine resumes with its state exactly as it left it. Momentum, monster thinkers and sound timers all freeze for the transfer and continue where they stopped. That is the same mechanism the console later replaced with the FIFO burst in the final build, and this log is where it was invented.

    While frozen, a readout engine takes over the SDRAM controller and burst reads the framebuffer into the 4 KB synchronous UART FIFO the console already uses. Burst reads outrun the wire by orders of magnitude, so a watermark paces them, pausing before any overflow while the UART drains the FIFO at line rate:

    wire fifo_has_space = (fifo_count < 13'd4090);
    

    Next is framing. Console bytes and pixels share one stream with no side channel, so the host has to find frame starts inside what looks like a text log. Every frame opens with a 10-byte header: four magic bytes (0x55, 0xAA, 0x5A, 0xA5), a mode byte, a 16-bit little-endian frame number, a 16-bit width and an 8-bit height. The viewer scans for the magic, prints everything ahead of it as console text, and only treats a header as real after the width and height sanity check passes, so a stray 0x55 in a log line costs one skipped byte instead of a torn frame. It keeps the last three bytes of each read unflushed for the same reason: a magic split across two USB packets must still be found.

    while len(self.rx_buf) >= 10:
        idx = self.rx_buf.find(self.MAGIC)
        if idx < 0:
            # No magic in buffer: all is console/terminal text
            # Keep last 3 bytes in case magic is split across reads
            text_bytes = bytes(self.rx_buf[:-3])
            self.rx_buf = self.rx_buf[-3:]
    

    Two modes trade clarity for rate. Full 320x200 streams all four bytes of every word on all 200 rows, 64,000 bytes, which is 1.44 FPS. Fast 160x100 keeps bytes 0 and 2 of each word on even rows only, 16,000 bytes, a quarter of the pixels for four times the rate, about 5.75 FPS. Both modes are live switchable from the viewer with single command bytes, along with a pause command.

    Input travels the same wire back, and the first movement test failed in a familiar-feeling way: the player ran into the nearest wall and stayed there, stride animation looping, ignoring every further key. The cause is that terminals send an event on press and nothing on release, so the key register held each code until something replaced it and the engine treated every key as held forever. The fix is a break protocol: the viewer watches its own key table and, on release, sends 0xF0 ahead of the code.

    def on_key_release(self, event: tk.Event):
        key = event.keysym
        keycode = DOOM_KEYS.get(key) or DOOM_KEYS.get(event.char)
     if keycode is not None and keycode...
    Read more »

  • Fixing a few bugs with FPGA Bring-up

    M_Jagadeesh97 • an hour ago • 0 comments

    GPIO-0 and GPIO-1 headers. The serial link uses pin 2, pin 4 and pin 12.

    Programming succeeded, but the uploader got no reply. The PC side looked fine: pyserial installed, COM7 present, both files on disk, and the info command reporting a clean layout. But the handshake timed out: no bootloader answer (is the bitstream running?).

    Two symptoms narrowed it down. LEDs 0 and 1 were blinking in turn, heartbeat plus transmit activity, which on this top means the bootloader is running and sending its ready beacon every second. And a raw terminal at 2 Mbaud showed this, repeating:

    P¶�Z�P¶�Z�P¶�Z�P¶�Z�
    

    A repeating roughly six-byte pattern where BLRDY1, six bytes, belongs. The FPGA was running and transmitting; the bytes were being corrupted between the pin and the PC. That split the problem in half: not the bitstream, not the wiring continuity, but the sampling itself, either baud accuracy or signal integrity at 2 Mbaud.

    The adapter answered it. Device Manager named it a classic CP2102, VID 10C4 and PID EA60, and the datasheet gave the ceiling: the part divides an internal 48 MHz clock, and the maximum supported rate is 921,600 baud. Asking Windows for 2 Mbaud aliased to an unsupported divider, so the two ends never agreed on a bit period and every byte framed wrong. The original 2 Mbaud plan had been written against an assumption about the module that was never checked.

    Both ends retuned to 921,600. The receiver uses parametrized oversampling, each bit sampled at its middle after a two-flop synchronizer:

    module uart_rx #(
        parameter CLK_FREQ = 50_000_000,
        parameter BAUD     = 115_200,
        parameter OVR      = 16,
        parameter DIVT     = CLK_FREQ / BAUD / OVR
    ) (
    

    With oversampling 6, the ideal divider is 50,000,000 divided by 921,600 times 6, about 9.042, so integer 9 gives exactly 54 system clocks per bit and a synthesized 925,926 baud, 0.47 percent above nominal. The CP2102 synthesizes 923,077 baud from its own 48 MHz clock. The two ends are within a fraction of a percent of each other, which is what a UART needs.

    Handshake green, the next upload printed its banner and froze. The progress bar never drew a single block, and within two seconds the script gave up on the first chunk: lost ack at offset 0.

    Failing at offset zero means the transfer never started, and the natural suspect is a crashed bootloader. But the bootloader had just answered the handshake, which exercises the same receive path, the same FIFO and the same transmitter the transfer would use. The receive path worked. The question was what the FPGA was waiting for.

    Reading the two sides against each other found it. The bootloader emits one dot per 4,096 payload bytes retired into SDRAM, counted on write acknowledges as words land in memory. The host counted wire bytes: it packed headers and payload into one blob, cut it into 4,096-byte slices, and demanded a dot per slice. The blob opens with a 10-byte segment header, so the first slice carried 10 bytes of framing and only 4,086 bytes of payload. The counter sat at 4,086 against a 4,096 threshold, the dot never fired, and the host waited for an acknowledgement the hardware was not scheduled to send, holding back the very next slice that contained the missing 10 bytes. A deadlock on chunk zero, and every boundary after it would have skewed the same way.

    The fix keeps framing and payload on separate tracks. Headers go out as their own writes, and the loop measures only payload against the dot interval, asking for the dot exactly where the hardware sends it:

    data_off = 0
    while data_off < len(data):
        rem_to_dot = ACK_EVERY - (sent_payload % ACK_EVERY)
        chunk_size = min(len(data) - data_off, rem_to_dot)
        ser.write(data[data_off:data_off + chunk_size])
        data_off += chunk_size
        sent_payload += chunk_size
        if sent_payload % ACK_EVERY == 0:
            ack = ser.read(1)
            if ack != b'.':
                raise SystemExit(f'lost ack at payload offset {sent_payload} (got {ack!r})')
    

    The...

    Read more »

  • FPGA Bring UP

    M_Jagadeesh97 • an hour ago • 0 comments

    DE0-Nano, back. The SDRAM and the 50 MHz oscillator sit on this side of the board.

    The bitstream programmed successfully and the game booted on hardware. It printed its headers, rendered frame 0, and hung.

    This is the part of the project where simulation stops helping. The same RTL that ran 500 million instructions in Verilator now stopped in the first second of gameplay, and the first question was whether the port or the core was at fault.

    A Python cycle tracer sampling the pipeline every edge, plus a full SDRAM dump taken at the hang, pointed at the cause. The game timestamp lasttime read -4,203 ms where the reference read 134. The tic computation produced nonsense, and the game never got past the frame 0 wipe.

    Tracing back from the corrupted timestamp led to DG_GetTicksMs, which is a load into a multiply-high into a shift, the exact sequence the game runs to turn cycle counts into milliseconds. The tracer showed the mechanism clearly enough to name it:

    • The load and the MULHU issued as a pair.
    • The dependency squashed slot 1 for replay.
    • On the same edge, an older M operation released.

    The old flush term cleared ID/EX and discarded slot 0, while the replay PC only re-fetched slot 1. The consumer therefore read a stale register and the timestamp came out corrupted.

    // before: release flushes even across a squash, dropping slot 0
    assign flush_ex = hz_flush_ex || md_release;
    

    The replacement qualifies the release with the squash, so ID/EX advances slot 0 in exactly that case, while redirect flushes still win:

    // NOTE (squash/release fix): ID/EX must still ADVANCE slot 0 when a
    // squash coincides with md_release. Flushing here drops the slot-0
    // instr (it never executes and the replay PC only re-fetches slot 1),
    // so its rd goes stale for the consumer. Redirect flushes still win.
    wire flush_ex_idex = hz_flush_ex || exr_valid || (md_release && !squash_s1);
    

    The bug predated the timing rewrite and went unnoticed in simulation because it needed a specific coincidence: an M operation releasing on the same edge that a dependent pair squashed. DOOM is simply the program that eventually lined up that coincidence. This is the case for running real software on the real board rather than only benchmarks in a simulator.

    The regression test that captures the coincidence is nine instructions. A divide stalls the pipe, a load plus add pair arrives as the next pair and squashes at the release edge, and the test demands the load survive with exact values:

    elif which=='M15':  # squash@release REGRESSION: divu stalls, Q=(lw,add) squashes at release; lw must survive
        dmem=[0]*2048; dmem[0x100]=1000
        prog=li(5,100)+li(6,7)+li(27,0x400)+[nop()]       # idx0-3 (pad: divu must be even)
        prog+=[divu(7,5,6),nop()]                          # idx4-5: M1 pair (divu slot0)
        prog+=[lw(8,27,0),add(9,8,6)]                      # idx7-8: Q pair, squashes at release
        prog+=[ecall()]                                    # idx9
        emit(prog,[(7,14),(8,1000),(9,1007)])
    

    The pre-fix core fails it with the exact drop signature: x8 reads 0 instead of 1000, and x9 reads 7 instead of 1007. With the fix, the game runs ten frames end to end at a steady two million cycles per frame. Frame 0 is pixel identical to the reference still, the full instruction suite matches 46 of 46 with identical halts, and all six C test programs halt identically.

    Heap-level memory comparison answered the last question the long runs raised: 156 words differed between the hardware run and the reference, and every one traced to thinker state pointers and a small time skew between the runs. The state records themselves are byte identical, and the board run is simply further ahead in tick AI, which means the game is advancing normally rather than the core misbehaving.

    The second half of this log is a duller failure with a useful lesson. Moving the retimed core from the verification tree to the Windows Quartus tree, one file at a time, produced six Error 12002s:

    Error (12002): Port "csr_stall" does not exist in macrofunction "hz"
    Error (12002): Port "rdata_q" does not...
    Read more »

  • Making simulation a bit faster?

    M_Jagadeesh97 • an hour ago • 0 comments

    Playing the simulated core from a browser, over the bridge.

    Twelve minutes of wall time per half minute of gameplay is workable for captures and miserable for iteration. This log is about the harness work that followed, and it changed nothing in the RTL and nothing in the guest binary. Only the shell changed.

    The legacy flow ran Verilator in timing mode, wrote each frame pixel by pixel through Verilog file IO, and rebooted 120 million cycles on every run. The new flow keeps the same RTL and the same guest ELF and replaces the shell: a synchronous testbench with no timing constructs (tb_doom_live.v), a hand-written C++ loop (sim_main.cpp), and an SDL player on top for live keyboard and video.

    RunLegacyNewFactor
    3 M-cycle probe1.62 MHz6.24 MHz3.9x
    Boot to first frame, 120.3 M cycles50.9 s23.8 s2.1x
    8-frame scripted play, same keyfile57.7 s25.6 s2.25x
    Replay 8.6 M gameplay cycles from snapshotabout 24 s bootabout 2 sboot skipped

    The gain came from dropping timing mode, staying single threaded, and the x-initial 0 and x-assign fast flags. Threading was measured and lost: 175 kHz on two threads against 6.2 MHz on one, because synchronization cost more than the evaluation it parallelized on a 17-module design. Snapshots leave throughput alone and remove the 24-second boot from every run after the first, which is what makes a one-line change testable in seconds instead of a coffee break.

    This is a correctness tool as much as a speed tool, and the checks are the reason the hardware bring-up was debuggable. Both flows ran the same 20-event keyfile and all 8 captured frames are byte identical. Totals match to the cycle, with the old testbench stopping 2 cycles late through its delays. A snapshot taken at frame 1 replays 6 frames with identical bytes and identical commit cycles, frame 1 at cycle 122,127,592 and instret 144,551,704 on both runs. Frame 0 stays pixel identical to the reference, the instruction suite matches 46 of 46, and the M1 to M18 micros pass.

    Three failures along the way, all in the shell, with the core untouched:

    • A frame watcher that sampled one posedge late reported one frame dumped while writing no file. It showed up at all only because MM_DUMP and MM_EXIT dual-issued in the same cycle, which is exactly the class of bug the harness exists to catch.
    • A throughput metric that included the snapshotted prefix in its wall time read 76 MHz. Counting only simulated cycles gives 5.1 MHz.
    • Two DPI lessons: calling a DPI export from C++ aborts unless svSetScope points at the exporting scope first, and Verilator rejects 4-state [31:0] on exports, so every DPI signature uses plain int.

    Live play came next. Keys arrive through a 16-deep FIFO posted from C++, which is what makes live input possible at all, since production and consumption run at different rates. The guest transcript drains live instead of at exit, and the results file records the guest exit code and any dropped keys.

    The real-time budget, measured rather than assumed: 35 tics per second at about 1.61 million cycles per level tic needs about 56.7 million cycles per second of simulation, and the harness delivers about 5.1. Scripted runs therefore play at about 10 percent of real time, against about 4 percent before. A level tic costs about 316 ms of wall time, a cheap menu tic about 76 ms. Input latency is 1 to 2 tics: 320 to 630 ms of wall time in a level, 80 to 150 ms in menus. In game time that is 1 to 2 tics, 28.6 to 57 ms, exactly as on hardware, so it feels like DOOM on a slow 386 rather than lag on a fast machine. Cycle-exact real time needs the FPGA, which this project gets in the next log.

    The last piece is the browser bridge, and it exists because the SDL player needs a display on the machine running the simulation, which is not always where you are. A Python script (bridge.py, standard library only) holds the TCP connection to the sim, polls for the latest frame, and serves the page, a JSON state object, and one POST endpoint...

    Read more »

  • Getting thing work resolving bugs

    M_Jagadeesh97 • an hour ago • 0 comments

    934 rendered frames from one uninterrupted run: walk, run, turn, fire, open a door, walk through.

    With pixels and time working, the next problem was input. There is no keyboard in simulation, so the testbench injects keys from a schedule file of cycle and keycode pairs, pressing and releasing through a 9-bit path where bit 8 marks a release. The guest acknowledges each key by writing to the key register at 0x014.

    The first version of the bridge decoded every MMIO write, except that it only looked at port 0. The core dual-issues stores, so an acknowledge landing in slot 1 was silently dropped. The key_valid bit never cleared, and the guest re-read the same key on every poll.

    The symptom was strange enough to be worth describing. Movement half worked, by luck of pairing, and firing never registered at all. A DOOM player who can walk but not shoot reads like a game bug, and it took a while to look one layer down.

    The fix is to decode both ports in the MMIO task, servicing the older instruction first so paired UART stores also keep their order:

    // mem_top_dg.v, after the fix: an MMIO store may arrive on either
    // issue port, because the core dual-issues stores.
    always @(posedge clk) begin
        if (dg_enable && mem_we0 && mmio0)
            mmio_write(mem_addr0[11:0], mem_wdata0);
        if (dg_enable && mem_we1 && mmio1)
            mmio_write(mem_addr1[11:0], mem_wdata1);
    end
    

    On a board the same rule applies, and it is the kind of thing that has to be right before anything else can be tested: a peripheral bridge must accept stores from both issue ports or the second one vanishes. Every action after this point depends on that fix.

    Debugging routes came next. The player's own state is the thing that explains why a route fails, so the port prints a per-tic trace over the UART: leveltime, ready weapon, ammo, x, y, angle and the use button.

    printf("[tk%05d] rw=%d ammo=%d x=%d y=%d ang=%d u=%d\n",
           leveltime, (int)pl->readyweapon, (int)pl->ammo[am_clip],
           pl->mo ? (int)(pl->mo->x >> FRACBITS) : 0,
           pl->mo ? (int)(pl->mo->y >> FRACBITS) : 0,
           pl->mo ? (int)(pl->mo->angle >> 24) : 0,
           pl->usedown);
    

    In simulation the UART is a 1 MiB buffer in the bridge flushed at exit; on a board it is ordinary printf to a serial console. This trace is what turned the door work from guesswork into measurement.

    Input arrives as one action at a time to keep the captures readable: raise the pistol, walk forward, run forward, walk backward, turn left, turn right, strafe both ways, fire, switch weapons, punch, open the door. Thirteen scripted actions, each its own short run and its own clip.

    The longer take is a single uninterrupted run of 934 rendered frames, about 26 seconds of game time at 35 tics per second. It walks out of the spawn room, runs the eastern corridor, turns into the dirt area, drops the trooper blocking the lane, pushes into the door recess, opens the door and walks through into the imp on the other side. Weapon raise, running, gunfire and the door sequence all flow out of one boot and one input file.

    The schedule is a plain text file you can read and diff, and the simulation is deterministic given the ELF, the WAD and the schedule, so re-running the capture reproduces the clip. That determinism is what made the hardware bring-up comparable later: a frame produced on the board can be diffed against the frame the prototype produced from the same input.

    Cost, stated honestly: one tic is about 1.62 million cycles, so one second of game time is about 57 million cycles, and the Verilator build sustained roughly 1.7 million cycles per second of wall time during these runs. One second of gameplay therefore costs about thirty seconds of wall clock, and the 26-second take is about twelve minutes of simulation. Live play against the prototype is slow motion, so the practical deliverable at this stage was offline recordings driven by precomputed schedules. The gap is a property of the verifier, not the design: a prototype is slower than the thing it prototypes.

  • The Verilator Prototype in Simulation

    M_Jagadeesh97 • an hour ago • 0 comments

    DOOM running against the simulated board, before any hardware existed.

    The core boots an ELF from a hex file and prints cycle counts. It has no display, no input and no clock. DOOM needs all three. There was no FPGA on the desk at this point, so the decision was to prototype the board in Verilator first, and that decision shaped everything that came after.

    The method: every component the board would carry gets a testbench stand-in that presents the same interface to the core, so the guest cannot tell which world it is in.

    • Crystal oscillator: a clock generator in the testbench, half period 5 time units.
    • Reset button: an initial block holding reset for 20 time units.
    • Program flash or BRAM: the instruction memory array, loaded with $readmemh from hex.
    • Data DDR or BRAM: the data memory array, a 64 MiB window in the DOOM build.
    • Keyboard controller: a schedule-driven injector latching presses and releases into the KEY register.
    • Display controller: the framebuffer array plus a commit handler writing one PGM per frame.
    • UART to a console: a 1 MiB accumulation buffer flushed at exit.
    • Microsecond timer: the cycle counter registers.
    // tb_doom.v: the oscillator and reset the board would provide
    clk = 0;
    forever #5 clk = ~clk;
    
    initial begin
        rst = 1;
        #20;
        rst = 0;
    end
    

    Step 1 was the display, because without it there was no way to see anything. The port defines CMAP256 so the game renders one byte per pixel, and the resolution is left at the native 320x200. DG_DrawFrame copies the 64,000-byte screen into the framebuffer window and stores once to the commit register at 0x028. The bridge watches that store on either issue port and dumps the framebuffer to a P5 PGM file stamped with the exact cycle and retired instruction counts. The guest side of the entire display path is one line:

    void DG_DrawFrame(void)
    {
        MM_DUMP = 1u;   /* one MMIO write, one frame_<n>.pgm on the host */
    }
    

    An offline tool pairs the PGM bytes with PLAYPAL from the WAD and produces the PNG stills and GIFs in these logs. The first still, the E1M1 view with the status bar, was the proof that the render path worked end to end.

    Step 2 was the WAD, and it produced the most expensive bug of the entire project. The IWAD is linked into data memory as a blob with the heap below it. The first boot died with W_GetNumForName: PNAMES not found even though the bytes were demonstrably present.

    The cause was the data memory index slice:

    // data_mem.v: the index slice is clog2(W) bits wide, so the
    // reachable window is exactly 2*W bytes and nothing more.
    wire [$clog2(`DATA_MEM_WORDS)-1:0] idx0 =
            offset0[$clog2(`DATA_MEM_WORDS)+1:2];
    

    With a 40 MB heap, the WAD's lump directory landed at 0x04490B34, past the 64 MiB window, and silently wrapped to an alias that read zeroes. Nothing faulted; the bytes were simply not there. The fix is a heap size in the linker script that keeps the WAD end inside the window, plus a build-time check in doom/build.sh that fails the build if _wad_end ever crosses 0x04000000 again. That check later saved the same mistake in a different form on the board.

    Step 3 was time. DOOM asks the platform for milliseconds, and the answer comes from the cycle counter, divided by the measured cycles per millisecond. The port runs in singletics mode so exactly one tic elapses per rendered frame, which keeps the frame stream, the tic stream and the key schedule in a fixed relationship: one captured frame is one tic, whether the counter is fed by a board oscillator or by the simulator.

    Three inputs is all the guest actually relies on: the key register, the cycle counter and reset. Interrupts are implemented in the CSR file but never enabled, because nothing here needs them.

    The milestone for this log is a rendering DOOM booting in simulation with the display path, the WAD in memory and a clock, all of it before any board work.

View all 12 project logs

  • 1
    BLOG

    BLOG INCLUDES IMPLEMENTATION TIMELINE IN A DETAILED WAY

View all instructions

Enjoy this project?

Share

Discussions

Similar Projects

Does this project spark your interest?

Become a member to follow this project and never miss any updates