Yeah, we’re in a funny position. By all accounts it is fair use (at least in the US) to train models (and build search indexes, e.g. Google Books), but sharing the books dataset itself is absolutely forbidden (clear non-transformative copying).
Anyone that wants to train a model needs to procure and destroy their own physical copy of each book!
Yeah, we’re in a funny position. By all accounts it is fair use (at least in the US) to train models (and build search indexes, e.g. Google Books), but sharing the books dataset itself is absolutely forbidden (clear non-transformative copying).
Anyone that wants to train a model needs to procure and destroy their own physical copy of each book!