Yeah, we’re in a funny position. By all accounts it is fair use (at least in the US) to train models (and build search indexes, e.g. Google Books), but sharing the books dataset itself is absolutely forbidden (clear non-transformative copying).
Anyone that wants to train a model needs to procure and destroy their own physical copy of each book!
> Anyone that wants to train a model needs to procure and destroy their own physical copy of each book!
Why? Couldn't they resell or give away the books after scanning them?
How is this in any way, shape, or form a promotion of the arts and sciences anymore?
This could very easily be turned into a preservation and archiving operation with just a tweak of the laws, or a carve-out.
And make the bank once, and make it legal to train on? How people in the bank get compensated is a different question, but -while almost impossible to settle on an individual basis- could be settled reasonably in bulk by some form of mandatory implied statutory contract?