Do you expect any of the labs to have an accompanying data dump with: here’s every book ever written, newspaper article, song lyric, Disney movie, GitHub repo, etc. Oh, and we obviously never paid for any of this.
Even if you did, I doubt training is bit-for-bit reproducible, so you will always have to take someone’s word for the final artifact.
Sure. The models are still not open source and should not be called open source.