It’s interesting and worthy of genuine applaud for being a good starting point for further work.
That said, I am more interested in what size model this could manage while producing “just enough” tokens per second to work at a conversational rate. What are models in that class capable of doing for me?