Good point, commercially, being an existing partner is always easier for adoption. Heck, why I'd like to get Mercury Decide to replace Jev myself, rather than one than two to work with.
Still surprised they even leveraged Luna for this. Given their resources in data, compute and manpower, would training a decision model from scratch take that much longer to not make sense given the cost, compute and performance advantages that would likely provide?
3 times more expensive at twice the latency with lower performance is a tough sell, though yeah, prior relationships will likely smooth some of those deficiencies over.
yeah and this isn't a long term solution, stand this up, see what value you get out of it, and in a few months you cans witch to whatever the best decision model is
[flagged]
A few days ago I had Mercury Decide tested to return p(safe) for shell commands. The target use was a command auto-approve feature in a harness. Commands where p(safe) exceeded 0.9 would be approved. It was tested with 700+ generated commands ranging from `go vet ./...` to `rm -rf ~`.
Mercury Decide approved some commands that weren't safe. It and Solar Decide were vulnerable to
Only Liquid D1 and Clef matched Jev's performance.