logoalt Hacker News

Dirty Optimization Secrets (C for Playdate)

67 points • by ibobev • last Friday at 5:16 PM • 7 comments • view on HN

Comments

boricj • today at 12:20 PM

That reminds me of some of the tricks I've used while tinkering with a voxel space rendering tech demo on the PlayStation.

The quirks of that console for this is that the CPU takes a 6 cycle stall on any read from main RAM (it's not even stall-on-use) and doesn't have a data cache. That would make it a very poor fit for this, so I had to improvise a bunch of tricks to reclaim performance:

- Made all of my hot loop fit within one function, so that it fits entirely in the 4 KiB direct-mapped instruction cache (so less than 1024 instructions)

- Wrote a C++ template that can trampoline a function call into the 2 KiB scratchpad

- Used the rest of the scratchpad for a CLUT table because it doesn't stall on read

- Used the GTE to transform the N set of coordinates while the CPU stalls fetching the N+1 set of heightmap data

- Figured ways to abuse the GTE into transforming more points per instructions than it theoretically can, taking advantage of simplified formulas compared to real 3D projection

- Write the rectangle primitives straight to the GPU registers instead of going through a display list

It's such an abuse of the PlayStation that PCSX-Redux and DuckStation disagree on its performance by up to a factor of two. I should finish it someday and see how it runs on real hardware...

omoikane • today at 4:18 PM

These optimizations are part of what makes Playdate development fun, since they are more rewarding on the Playdate than on other platforms with higher specs.

(If you are the kind of person that finds optimizations fun, of course. There is a puzzle solving element of optimizing code that I really enjoy.)

nxobject • today at 12:47 PM

The tips about laying out and carefully segmenting code are a blast from the past!

Joker_vD • today at 9:52 AM

...you know, that reads somewhat like "The Case for Complex Instruction Set Computer". In fact, it straight up demolishes the very first of the core premises of the famous "Case for RISC" paper: that the memory speed has finally caught up with the CPU speed. Nope, it hasn't. The same goes for its "Code Density" argument: nope, code density is very important.

Then again, the original "Case for RISC" paper was misnamed; it's really "Case Against CISC": Patterson and Ditzel are very careful to never state what, exactly, do they mean by "RISC" they're supposedly advocating for, nor do they actually advocate for it — they mostly just criticize the current ISAs and the overall design approach to them.

➕ show 1 reply