> Data moats are gone if you can simulate the data with AI.
This is a hilarious premise if you work in a domain where it matters even a little bit whether the data is correct or not.
There’s 2 arguments.
1. He’s talking about training new models and at one point, having data was valuable. Now synthetic data is being used to train models
2. Companies like SalesForce who’s moat is having all your customer data so you’ll be locked in. You could extract it but you’d have to clean it and then change it to your new schema. With LLMs, you can do that in minutes and even use SalesForces MCP or API to get all your data and leave.
It’s exactly why companies like Figma are gate keeping their MCP. They know that swapping their MCP with Paper’s or any new one is easy.
The moats are evaporating as we speak. Distribution is one of the smaller ones left, but the personal software trend might eat that too.
I work in science, the data is becoming far bigger of a moat than it ever was before because of this.
Every company can now apply the latest and greatest analysis. Data generation is where the cost is. It's where the time was spent, time that can never ever be retrieved at any cost.
AI won't solve biology, make a pathogenic virus, etc, without tons and tons of data, of both types we know and types we have not yet figured out how to generate.
Perhaps the area where AI has the most to help bio is in figuring out novel measurement technology. But it's not going to be able to reason or deep-net its way to figuring out systems for which we can't even measure the parts.
You are assuming that getting the AI to generate correct (or accurate, representative) data will be very difficult. I would agree, but I think it will become possible in the long run.
(Edit: alternatively you just use AI to get rid of the need for data to solve a problem, like Jev did for traditional classification models)
I think current incentives definitely go against any efforts to build this. It's very hard to build this and be rewarded for it by, say, investors or your boss, because you can't really prove that your system is non-sloppy while your competitor's is (even if being non-sloppy is all that matters), because by definition your novel results are not verifiable or else the model labs will have already trained it into their model.
But the same is true for high-quality AI systems in general. In general, I think AI model advancements will make the systems easier and easier to build until some small guy accountable to no one but themselvs can build it, and then it will actually be built.