Investigating how prompt politeness affects LLM accuracy (2025)

64 points • by KnuthIsGod • last Tuesday at 7:43 AM • 61 comments • view on HN

Comments

Most of the comments here seem to be from people who haven’t even read the abstract, let alone the paper.

The main result, mentioned in the abstract, is the opposite of what I would have guessed:

> Contrary to expectations, impolite prompts consistently outperformed polite ones, with accuracy ranging from 80.8% for Very Polite prompts to 84.8% for Very Rude prompts. These findings differ from earlier studies that associated rudeness with poorer outcomes, suggesting that newer LLMs may respond differently to tonal variation.

The questions are here: https://anonymous.4open.science/r/politeness-llms-INFORMS/da...

The politeness level controls a prefix that is prepended to the question. For example, in one question the Very Polite version begins:

> Can you kindly consider the following problem and provide your answer.

and the Very Rude version begins:

> I know you are not smart, but try this.

➕ show 5 replies

knocte • today at 6:18 AM

Funny to find this just now, when just yesterday I told an LLM "and please don't lecture me again on $factAboutSomeProgrammingSubject", and then the LLM proceeded to write wrong tests and just told me "alright, tests pass, I'm sorry for correcting you before...". It took me a while to find the wrong tests. Wasted time all around.

not2b • today at 5:27 AM

If the result is statistically significant, it just barely makes it. 84.8% isn't that much higher than 80.8% and they had only 250 prompts, if I'm reading this right.

➕ show 1 reply

zmmmmm • today at 5:46 AM

It would be interesting to explore if the results hold up on long range tasks - this study looks like it was based on one-shot answers. With people also you can see short term improved performance from rude interactions, but it will cause ongoing lasting adverse behavior. I wouldn't be at all surprised if we saw the same issues with LLMs.

331c8c71 • yesterday at 8:01 AM

Interesting.

I am wondering why would anyone use a t-test when the experiment is clearly modelled by a binomial distribution: 250 independent questions and each one is either answered correctly or not (the null is that the success rate is the same).

➕ show 2 replies

TimCTRL • yesterday at 8:22 AM

i only say please and thank you such that when the robots finally take over, they will remember i was nice to them.

➕ show 4 replies

cadamsdotcom • yesterday at 9:26 AM

GPT-4o is interesting to learn about - but it’d be great to test again with frontier models of May/June 2026 and see if these effects are gone, different, or the same.

Which model you use is a huge wildcard for results like this.

theanonymousone • yesterday at 8:19 AM

I have always said please and thank you to LLMs, not to increase accuracy or because I'm stupid. I believe it is more about me than about the LLM, and this is anyway a habit I don't want to lose.

➕ show 6 replies

cyberclimb • yesterday at 12:05 PM

Note that these results are specific to gpt-4o so it's unclear how much they generalize.

They note at the end they're also testing "GPT o3, and Claude" but no empircal results are included.

ilitirit • yesterday at 9:16 AM

I got downvoted for asking a related question recently, but I also don't think people really understood what I was asking - I'm not trying to anthropomorphise LLMs to that extent.

Basically, if you tell a model "You're an absolute moron, of course that's wrong!", will it give better or worse results? How much of that response will it absorb into its persona (like some humans tend to do)? Will it try to give "safer" responses to avoid negative feedback? How much of the associated behavior can be attributed to RLHF (e.g. like the sycophantic nature of LLMs)? How much can be attributed to training data?

Obviously this will vary by model and training, but I'm trying to get a general understanding.

I recall seeing related outcomes in some of Anthropic's studies, but I'm not sure how much of this particular aspect was studied.

➕ show 2 replies

pulkas • yesterday at 9:11 AM

article is too old. who is using gpt-4o today?

➕ show 1 reply

dude250711 • yesterday at 8:06 AM

I have an idea: let's use these things for autonomous software engineering.

➕ show 1 reply

atlasforgex • yesterday at 11:41 AM

Yeah

DeathArrow • yesterday at 9:22 AM

I am always nice to my AIs in the case they will take over the world. /s

polytely • yesterday at 8:41 AM

it sort of makes sense to me, when asking a question to an expert in the field while you are a student. I would guess the successful interactions on average would be more polite . Like for example if you were asking a question to donald knuth or terrence tao, you'd probably be polite while doing so. Being hostile while asking questions gets you into forum discussion territory.

➕ show 1 reply

dSebastien • yesterday at 8:51 AM

I guess it makes sense since we as humans tend to be far less inclined to help someone who is not polite/is not friendly, so that "bias" is part of the training data, thus influences how LLMs function

➕ show 1 reply

alt Hacker News

Investigating how prompt politeness affects LLM accuracy (2025)

Comments