- 32-bit RISC-V ISA
- pipelined with hazard handling and branch prediction
- cached data memory
Program counter, control unit, register file, ALU, hazard handling unit, instruction memory ROM, data memory RAM, data cache, and the pipeline registers between stages.
The single-cycle version is the straightforward part. Pipelining is where it gets interesting: forwarding paths for data hazards, stall logic for the loads forwarding cannot rescue, static branch prediction with the flush machinery to undo speculatively fetched instructions on a miss, and a data cache that turns the memory hierarchy from an idea into a set of hit and miss latencies you can measure.
Verified end to end against assembly test programs rather than per-module, which is the only way to catch hazards that emerge from interactions between individually correct stages.
Building the machine that runs the code is what gave me a concrete model of why code is fast or slow: branch misprediction cost, cache-miss latency, and the memory-hierarchy trade-offs that govern everything from a tight C++ loop to a GPU kernel.