logoalt Hacker News

roadside_picniclast Monday at 5:35 PM1 replyview on HN

In addition to all this, I also feel we have been getting so much progress so fast down the NN path that we haven't really had time to take a breath and understand what's going on.

When you work closely with transformers for while you do start to see things reminiscent of old school NLP pop up: decoder only LLMs are really just fancy Markov Chains with a very powerful/sophisticated state representation, "Attention" looks a lot like learning kernels for various tweaks on kernel smoothing etc.

Oddly, I almost think another AI winter (or hopefully just an AI cool down) would give researchers and practitioners alike a chance to start exploring these models more closely. I'm a bit surprised how few people really spend their time messing with the internals of these things, and every time they do something interesting seems to come out of it. But currently nobody I know in this space, from researchers to product folks, seems to have time to catch their breath, let along really reflect on the state of the field.


Replies

bawislast Monday at 7:48 PM

> we haven't really had time to take a breath and understand what's going on.

The field of Explainable AI (or other equivalent names, interpretable AI, transparent AI etc) is looking for talent, both in academia and industry.