Go back

Castlevania: Symphony of the Night running natively on the ESP32-S3

The XIAO ESP32S3 handheld running Castlevania: Symphony of the Night in the Alchemy Laboratory

In the Metal Gear Solid article I showed the handheld console I built around a XIAO ESP32S3 Sense. Castlevania: Symphony of the Night (SOTN) is the other PlayStation game it runs, and the one that came first: it was running at 60 fps on a Waveshare ESP32-S3 development board before the handheld existed, and later moved onto it.

It was possible for the same reason as Metal Gear Solid. SOTN’s community decompilation builds as portable C, so the game can be compiled for the ESP32-S3 instead of emulated.

Gameplay on the handheld: watch it on YouTube.

This article covers how the port is structured, how the game fits into the SoC’s memory, how the software rasterizer reached 60 fps, the bugs the original hardware had been hiding, and a handheld console I built around a Seeed XIAO ESP32S3 Sense.

Why a native port is possible

Emulating a PlayStation on an ESP32-S3 is out of reach. An emulator has to interpret a MIPS CPU and a GPU instruction by instruction, which needs several times the performance the SoC has.

Porting is different. The SOTN decompilation project has reconstructed the game as C source that compiles to a PC executable. The game logic is ordinary C, and the only thing tying it to the console is Sony’s SDK: the library calls that draw polygons, upload textures, read the controller and play sounds. The PC build replaces that SDK with psyz, a reimplementation on top of SDL.

The job, then, has three parts:

  1. Compile the game for Xtensa.
  2. Replace psyz’s PC backend with one for the SoC.
  3. Make everything fit.

The target SoC is the ESP32-S3:

Architecture

The layers, from the game down to the hardware:

  1. Game: decompiled engine, player and weapon code, stages
  2. psyz: SDK reimplementation with PSYZ_RENDERER=soft
  3. Platform layer (esp32/main): VSync, root counters, input, FAT files, audio, LCD scanout
  4. ESP32-S3 hardware: LX7 cores, SRAM, PSRAM, flash, SPI LCD

The layers of the port and the two cores: the game and the rasterizer on core 0, LCD scanout and audio on core 1, and the panel and microSD sharing SPI2 on the handheld

Two design decisions shaped everything else.

VRAM stays the game’s own buffer. The PlayStation has 1 MB of video memory, 1024×512 pixels of 16-bit color. The decomp already models it as a plain array. The rasterizer draws straight into it (in PSRAM), and the LCD is fed from it. There are no shadow copies and no format conversions except the final one to the panel’s RGB565.

Stages are linked statically. On a PC, each castle area is a DLL loaded on demand. The ESP32-S3 has no dynamic linker, so every area is compiled into the firmware, and a small table maps a stage name to its init function. That decision caused problems twice, as described in “Bugs the PlayStation forgives” and “Adding a stage” below.

Fitting a PlayStation game into 512 KB

Alucard next to one of the laboratory's statues

The first link failed by 3.34 MB of internal RAM. Most of that was not the game’s fault:

That last one has a trap: some “read-only” data is written. The SDK updates the sound bank headers in place when a bank is opened. Declared const, they live in flash, and writing to flash through the cache is a hard fault (“Dbus write to cache”). The pattern that works is to keep a const master in flash, copy it to a PSRAM buffer at startup, and let the game write the copy.

After 28 rounds of linking, the numbers were:

The software rasterizer, and getting to 60 fps

The laboratory's flames: semi-transparent effects drawn by the software rasterizer

SOTN is a 2D game, and its primitive mix is small: textured and Gouraud quads, 16×16 sprites for the tile layers, rectangles, and lines. There is no 3D and no geometry transform engine, which made a software rasterizer realistic.

The rasterizer is bit-exact against the GPU reference build. I compared frame checksums between the PC build and the board after every optimization.

Most of the frame time went to memory access. Every textured pixel reads a texel and a palette entry from PSRAM and writes a pixel to PSRAM. The optimizations that mattered all reduce that traffic:

StepResult
Triangle edge stepping: divides, then 64-bit accumulators, then 32-bit 16.1664-bit was slower on a 32-bit core; 16.16 shipped
Scanout moved to core 1, fed by a notifyThe game never waits for SPI
Palette cache in internal RAM, raster loops in IRAMWarp Room at 31 fps
Tile layer fast path: 4 texels per read, pixel pairs as 32-bit stores10.4 ms to 6.1 ms per frame
Scan out from the buffer the game just finished drawingNo copy, one frame less latency
PSRAM and flash from 80 to 120 MHz19.4 ms to 16.4 ms per frame: 60 fps

I also tried a texture-row cache in the triangle rasterizer and reverted it. It gained nothing, because GCC had already hoisted the loads.

The change to scan out from the finished buffer needs some explanation. The obvious source for the display is the SDK’s current display area, but it lags one frame behind: at VSync it names the buffer the game is about to draw into. Scanning from it meant the DMA and the rasterizer were fighting over the same half of VRAM for the whole frame. Reading the display origin from the game’s own buffer struct makes the two work on opposite halves by construction.

Bugs the PlayStation forgives

The PlayStation has no memory protection. A stray read returns whatever is there, and a stray write lands in memory that often nobody checks. The PC build mostly hides these bugs too, because its statics are large and forgiving. The ESP32-S3 has an MMU, tight memory, and neighbors that matter, so porting to it turned out to be a very good way to find latent bugs.

Fighting in the Alchemy Laboratory on the handheld

How I found them

A panic backtrace names the victim, not the culprit. The tool that worked was OpenOCD over the SoC’s built-in USB-JTAG, together with GDB:

  1. Reset and halt the SoC through OpenOCD.
  2. Set a breakpoint on the panic handler.
  3. When it hits, decode the real PC from the panic frame.
  4. Always decode against the exact ELF that was flashed, because addresses shift on every rebuild.

Palette animation out of bounds

The Alchemy Lab requests palette animation (tileset & 0xFF) + 0x7FFF | 0x4000, which indexes entry 2 of the stage’s palette table. The table has one entry. The out-of-bounds read landed on the adjacent sprite-bank table, which was then registered as a palette-animation descriptor and written every frame.

The fix rejects descriptors whose extent exceeds the palette buffer. The same undefined behavior exists upstream; the PC absorbs it in a large BSS.

Two stages, one HitDetection

With stages linked statically, 122 global symbols were defined in more than one stage. Shared headers implement things like collision and entity updates, and on the PlayStation only one stage was ever loaded. Static archives let the linker silently resolve each name to one definition. The Alchemy Lab ran the Warp Room’s HitDetection against its own entity tables, and corruption appeared a layer later.

The fix is a generated list of every symbol two stages share, renamed per stage with -D defines.

Rooms that mutate their own map

Doors and breakable walls write to the room’s tile map. With the generated maps in flash, the first door crashed the game. The fix is a single point where the active room’s foreground layer is copied to a writable PSRAM buffer.

The enemy name

The enemy-name box at the bottom of the screen, drawn by the same BottomCornerText code that used to overrun the stack

With the Faerie Scroll relic, the game prints the name of the enemy you hit. BottomCornerText scans the string for the PlayStation’s FF 00 terminator. PC strings are plain C strings, so the scan ran off the end and overwrote a 64-byte stack buffer. The symptom was “it freezes as soon as I touch an enemy.” The fix bounds the scan.

A race that didn’t exist on the PlayStation

Audio runs on the second core, and the audio pull was advancing the root counters. Those counters fire the game’s VSync handler, texture uploads and GPU queue callbacks. On the PlayStation those are interrupts on the same CPU. Here they ran concurrently with the game on the other core.

The fix accumulates counter ticks on core 1 and fires the handlers on the game thread.

Input that never arrived

The report was “I can’t jump.” Three bugs were stacked:

The handheld: hardware bring-up on a XIAO ESP32S3 Sense

The finished handheld from above: ILI9341 panel, drone-gimbal stick, buttons and the XIAO on perfboard

The back of the handheld: point-to-point wiring on the perfboard and the panel's own SD slot (unused). The XIAO's camera module, removed, sits beside it

The front of the handheld next to the XIAO's camera module, which the port doesn't use and was taken off

The goal was a handheld console. It is the same hardware that later ran Metal Gear Solid: SOTN was the first game on it. The parts:

The XIAO has eleven usable pins. The panel needs five (SCK, MOSI, MISO, CS, DC) plus a reset line, and the stick needs two ADC pins. That leaves three for eight buttons.

Six of the buttons share one ADC pin through a resistor ladder: a 10k pull-up and, per button, a resistor to ground (0 Ω, 2k2, 4k7, 10k, 22k, 47k). Each button produces a different voltage. The limitation is that two buttons held together read as the lower one. The pair each game holds together (attack and jump in SOTN) gets the two remaining dedicated pins.

The full wiring diagram, parts list and button map are in the hardware guide in the repository.

Wiring of the handheld, drawn with Velxio's part art: XIAO ESP32S3 Sense, ILI9341 panel, analog stick, six-button resistor ladder and two direct buttons

Bring-up was its own list of lessons. Before touching the game, I wrote a separate shakedown firmware that draws color bars and shows every button and the stick on screen:

The handheld: display, input and a crash caught by a watchpoint

Everything on the SD card

8 MB of flash can’t hold a 3 MB app next to 3.7 MB of game data, so on the handheld all the data comes from the microSD card. That card shares the SPI bus with the panel, which later caused its own crash (see “The shared SPI bus” below).

Two traps specific to this board

GPIO43 is both the panel’s data/command pin and UART0 TX. The Waveshare code installed the UART driver for its serial console. Doing that silently hands the pin back to the UART, and the panel stops distinguishing commands from pixels. The handheld uses native USB for its console only.

printf over native USB blocks when nobody is reading. With a PC attached you never notice. Unplugged, the transmit buffer fills in a few seconds and the game stops inside a printf, on a device that is meant to run unplugged. All output, including ESP-IDF’s own logging, now goes through a wrapper that writes with a zero timeout and drops what nobody reads.

Flicker, then sliding rectangles

The first build flickered. I assumed the SPI clock was too slow and raised it from 40 to 80 MHz. That helped a little, but instrumenting the scanout showed what was really happening:

Timeline: the game draws a frame every 16.7 ms into buffers A and B while the panel takes about 24 ms to receive one, so the game redraws a buffer the DMA is still sending

This was a buffer coherence problem. The fix was the copy path the Waveshare port had kept as a safety net: copy the finished frame, then send the copy. My first version of it had two bugs:

Even after both fixes the rectangles stayed. Dropping the SPI clock back to 40 MHz made them vanish: the jumper wires couldn’t hold 80 MHz, and a byte slip shifts a whole 8-line band. The S3’s SPI clock divides an 80 MHz source, so the only options are 80 and 40 MHz, with nothing in between.

With the wires as the bottleneck, the remaining option was to send less data. Each 8-line band is fingerprinted as it is converted to RGB565, and bands identical to what the panel already shows are skipped. Dropped frames went from about 40% to 6%. That gives roughly 55 fps on screen when the camera is still, while the game logic stays at full speed.

The crash a hardware watchpoint caught

The next report was “I pressed START and it rebooted.” Earlier there had been another: “text appeared, then the screen went white, then it restarted.”

The panic was in ClearOTag, called from the main loop with an ordering-table pointer of 0x4b8. The main loop advances with g_CurrentBuffer = g_CurrentBuffer->next, so working backwards, something had written 0x44 into the next link of one of the two GPU buffers.

The ESP32-S3 has two hardware watchpoints. I armed both, one on each buffer’s next field, two seconds after boot, well after the only legitimate write. Pressing START stopped the CPU on the culprit’s store:

Debug exception reason: Watchpoint 1 triggered
AddPrim ← MenuDrawImg ← MenuDrawChar ← MenuDrawStr ← MenuDrawStats ← MenuDraw

The pause menu takes the next free sprite with &g_CurrentBuffer->sprite[g_GpuUsage.sp] and never checks the bound. sprite[] is the last field of a GPU buffer, and the second buffer starts right after it, beginning with its next link. Drawing the stats text pushed past 512 sprites, and AddPrim wrote a primitive header over the link. The same game checks g_GpuUsage.sp < MAX_SPRT_COUNT elsewhere, but the menu didn’t. Two lines fixed it.

The two GPU buffers sit back to back in memory; the pause menu's 513th sprite lands on the next buffer's next link, which a hardware watchpoint caught

The shared SPI bus

Right after that, a different crash appeared on stage loads: an assertion in the SPI driver, running_cmd == 0. The card started a command while a panel DMA transfer was still on the wire. On the Waveshare board the card only streamed optional music; on the handheld it carries everything.

The fix is a mutex:

Instrumenting without misleading yourself

Two lessons from this stretch:

Running each build without flashing the board

After several hours of trial and error, flashing the board after every change was the slowest part of the loop: write the image over USB, wait for the reboot, reconnect the serial port without resetting the board again, read the log, repeat. So to run and check each build I used velxio-cli, the command-line client of Velxio. It takes the same .bin my toolchain produced, runs it on Velxio’s simulated ESP32-S3 for a fixed amount of simulated time and hands back the serial output, so I could read counters and timings without touching the hardware. The board only got a new image when a build was worth it.

# velxio.toml, next to the build: board = "xiao-esp32-s3", firmware = the merged .bin
velxio-cli run --timeout 5000 --timeout-exit-code 0 --serial-log-file serial.log .

Installing it and doing a first run takes a few minutes: see the velxio-cli quickstart.

Adding a stage: the memory budget, again

Walking out of the Alchemy Lab leads to the Castle Entrance (NP3). Adding it overflowed internal RAM by 160 KB. The stage’s compressed graphics, palettes, tile maps and sprite tables were declared writable, so they were copied into RAM at boot. Marking them const moved them to flash. The upstream PC port had already added NP3, and that commit came across cleanly.

Two more discoveries:

The generated files are regenerated from each user’s disc, so the const changes live in a script (esp32/const_stage_tables.py) rather than in the generated sources.

Advice for a similar port

Status and next steps

Another fight in the Alchemy Laboratory, on the handheld

The code is on GitHub at davidmonterocrespo24/sotn-decomp (branch esp32-port). The wiring diagram and parts list are in the hardware guide. You need your own copy of the game; the repository contains no game data.

Thanks to the SOTN decompilation contributors, and to Xeeynamo for psyz. This project is built on their work.


Share this post on:
David Montero

Written by

David Montero

Creator of Velxio, the open-source circuit and Arduino simulator.

GitHub velxio.dev

Related posts


Next Post
Metal Gear Solid running natively on the ESP32-S3