I was digging into this recently with ChatGPT. I’ve loosely followed the progression of ML and NN over the past 20 years, but struggled to put it into context of where an LLM lives. The big inflection point was the 2017 Attention Is All You Need paper [1].
Artificial Intelligence
|
+-- Symbolic / rule-based AI
| +-- expert systems
| +-- search / planning
| +-- logic / knowledge representation
|
+-- Machine Learning
|
+-- classical statistical ML
| +-- regression
| +-- decision trees
| +-- SVMs
| +-- Bayesian methods
|
+-- Neural Networks / Deep Learning
|
+-- computer vision
+-- speech
+-- Natural Language Processing
|
+-- Transformers
|
+-- Large Language Models
|
+-- chat systems
+-- multimodal models
+-- tool-using systems
+-- agents
[1] https://en.wikipedia.org/wiki/Attention_Is_All_You_NeedAIAYN was very influential but as its title implies its contribution was mostly about simplifying the architecture, and "attention" blocks were already known before, but their message was that you can build a model pretty much by just stacking those (and MLPs). The parallel trainability vs the rollout needed with recurrent nets (like LSTMs) made this much more scalable. But besides the architecture, what was equally important is the increase in available data, and compute. The other inflection point before that was around 2008-2012 when GPGPU (general purpose GPU programming) took off through CUDA (before that, GPGPU was much more tedious as you had to formulate your task as a graphics task about 3d meshes and pixel shaders, but people did that anyway, I had a college class on that in the 2000s).
Also a lot of the vision and speech ideas cross pollinated with the NLP field. One big trend that enabled faster progress is bringing all this onto a common platform. First via Deep Learning and backprop, formulating everything as some vector input, some model architecture, some vector output, some loss, and then gradient descent optimization. This replaced the specialized optimization tricks people used to develop for their own little niche tasks. Before DL, papers usually derived their own math for how to solve their own specific formulation of a task, so it was hard to reuse ideas.
(Reuse was also hard because platforms like GitHub didn't exist, the Python ecosystem wasn't nearly close to what we have, code sharing wasn't as common, and anyway the code was some mess in MATLAB, not in a sane language.)
The second thing that allowed converging these fields was the transformer architecture that allowed turning everything into tokens and throwing it all into the same transformer architecture, making multimodal models that can learn from everything and do everything, instead of having to make specialized models for each little task.
Essentially because attention introduced a way to scale un/self-supervised learning to the level of data out there, and learnable inference time 0-shot feature selection. Impressively in a autoregressive, unidirectional manner.
My interpretation is that before the transformer, most everything under the domain of 'AI' was either an academic curiosity or only applicable in very narrow fields. GPT-3 was when the 'magic' that people had always dreamed of with AI began to emerge, and it's only really this year that we are starting to be seriously confronted with the possibility of a general intelligence emerging from LLMs (albeit, not quite the same thing as 'true' AI which would necessarily be more of a biological exercise).