brainfuck is unpleasant to write directly - e.g. the language doesn't have variables, so you need to manually do the bookkeeping of which memory offset is storing what 'variable'. & if you need to refactor your program slightly, in a way that changes the memory layout, maybe you need to manually rework the absolute & relative offsets. So I can appreciate why the author didn't roll up their sleeves to directly write BF - that's neither a productive nor interesting exercise.
Interesting to see how the author decomposed the problem:
- C raytracer https://github.com/mTvare6/rayfuck/blob/master/ray.c
~~ LLM refactor of the C code ~~>
- SSA-style C raytracer code https://github.com/mTvare6/rayfuck/blob/master/ray_ssa.c
~~ c2dsl.py helper script (compiler) ~~>
- DSL raytracer https://github.com/mTvare6/rayfuck/blob/master/ray.dsl
~~ dsl2bf.py helper script (another compiler) ~~>
BF raytracer https://github.com/mTvare6/rayfuck/blob/master/ray.bf (~22 mb of unreadable nonsense)
The dsl2bf compiler has a bunch of examples of implementing slightly higher level abstractions atop BF primitives. E.g. "go" to move the pointer to a different offset, destructive & non-destructive copies, all the way up to things like division -- BF only natively offers unary addition/subtraction.
If we have a read of the code of the final compiler, dsl2bf.py, the abstractions used in that code are relatively simple: global variables, local variables, lists, dicts, for loops, function definitions & function calls. It is feasible to implement a simple compiler like dsl2bf in BF itself, with sufficient head scratching. Again, quite unpleasant to try it directly in BF, but a next step could be to implement the dsl2bf compiler in the DSL itself - extending it if necessary, then compiling it with itself to produce a dsl2bf compiler implemented in BF.
That does sound like a fun step, I'd already begun experimenting with some optimizations after having received suggestions in reddit to add fork/join primitives. Adding a compiler with these added performance gains sounds reasonable and something which will run quickly. dicts certainly involve some thought there.
I hadn't considered self-hosting the compiler, but having put it into works, I probably will.
This was the render the speed up version gave: https://paste.c-net.org/SpikingCarbs