> If I were OpenAI or Anthropic I’d be racing to make their products as sticky as possible bc ppl will flock to what’s cheapest otherwise.
My approach would be per-user (or per-project) "memory".
I don't mean a tack-on like a RAG / document store with "notes" added to it in the background, but an actual medium- and long-term memory. Something like an extension of the KV cache stored in High Bandwidth Flash (HBF), or a subset of the weights trained "online", similar to LoRA.
This would not be transferable to any other base model, so would be excellent "lock in".
The downside is that the memories likely wouldn't be transferable to new models either, but I can imagine solutions to that too. I.e.: Train an MLP to "translate" from the old memory weight space to the new one.