Nice idea actually, I always wondered why there was no actual programming in ML!
Homie just had some tokens to burn at the end of the month and “in Rust” is pure HN clickbait.
You can just tell when the idea itself was generated
How would it compare to current state of the art tested state space layers like Kimi Delta Attention?
Isn't this just RNN and nearly nothing about this is actually new?
“In Rust” man of all the cliche title baits, I hate this one the most.
Why did you choose Rust? Why does that matter?
Nothing wrong with Rust. Lots wrong with the bandwagon that “in Rust” somehow adds value.
vibe coded vibe coder!
OP, you should not have written this in Rust. It should be in PyTorch, which is by far the most popular. We can’t tell if this architecture is good or whether there is a problem in your implementation.
You can test the whole thing for free on a GPU with Google Colab. Test both the transformer and your new architecture on a larger dataset. Something that maxes out the GPU for an hour each run.
Also, the readme mentions keeping the same optimizer schedule which sounds nice at first but they are completely different architectures. The loss is high on the transformer, did you try raising the learning rate on it?
In general I’m interested in parameter efficient architectures. I don’t think transformers are optimal, and indeed many improvements have been made to vanilla transformers. But if you have an idea for something better you need to show it.