logoalt Hacker News

dparkyesterday at 5:48 PM1 replyview on HN

These models are trained on way more than just books. GPT-3 was trained on about half a terabyte of filtered plaintext and the training corpuses have grown significantly by then by all accounts.


Replies

pfdietzyesterday at 5:52 PM

I imagine that compresses by ~90%, and current top commercial models have a couple of trillion parameters, don't they?

show 1 reply