Another thought: openAI literally has an (almost) entire copy of the public internet they use for their training dataset! Why cannot they create an internal version of it that doesn’t require accessing public servers? They have the data already