How much time were you able to put into tuning your prompts? And was it worse on all fronts (cost, latency, accuracy) or just some?