logoalt Hacker News

simianwordsyesterday at 6:22 PM0 repliesview on HN

Rewriting prompts don't come with no costs. The cost here is that different prompts work for different contexts and is not generalisable. The rewritten prompt here will not work well for other cases like medical or social advice.

I think this rewriting of prompts technique is the reason "reasoning" models perform well - they know exactly how to rewrite the prompts for a context.

FWIW I don't trust these benchmarks fully because a huge bump like this is not expected - I would expect OpenAI to optimise enough to let such gaps open.