We tried Luna and it scored way lower in our evals. Muse also. We haven't had a chance to test others.
How much time were you able to put into tuning your prompts? And was it worse on all fronts (cost, latency, accuracy) or just some?
How much time were you able to put into tuning your prompts? And was it worse on all fronts (cost, latency, accuracy) or just some?