logoalt Hacker News

Prompting Claude Opus 5.5

168 points • by Michelangelo11 • today at 7:33 AM • 178 comments • view on HN

Comments

skeledrew • today at 8:46 AM

All that keeps jumping out at me is how they've set it to refuse giving users thinking tokens and prompts for full reasoning in output. Just drives me further away; I may not stop using Claude completely for now, but I'll be moving even more of my primary workload to Chinese providers. That's where openness and freedom is now at.

➕ show 9 replies
bluegatty • today at 9:35 AM

This is a failure of the AI foundries; if we have to use totally different prompting techniques for every model, this wont work.

AI is rapidly saturating it's ability to be useful and these products need to start to mature.

It's not 'fun' to manage 50 different broken MCPs and their variety of ways in which they are broken.

It was 'fun' at the start, now it's just 'broken technology'.

Astra and Opus 5.5 are the 'starting point' for the next era of AI where we expect robust tooling.

➕ show 7 replies
prodigycorp • today at 8:30 AM

Opus 5.5 is a good model, but I've tried to understand the extreme hype about it on social media about Opus' ability to do 2d work, as we got with Astra doing 3d work. In both releases, the models required extensive access to third party apis to generate assets for it, and a lot of the models work was essentially coordinating everything.

There's so many "x generated this in one shot, this is agi" stuff that gives you the impression that you can vibe operate modern models the same way you operated last year's models. There's so much more to it than that. It requires you to put a faith in the leap in the capability of models, one that would've surely been a waste of time in previous models.

Not sure where i'm going with this other than I think most can relate that it's exhausting keeping up with. I cant imagine what it'd be like parenting a kid that went from toddler to puberty in the span of a year and planning for them to go to college the next year. This industry is moving so fast that it's becoming fact that it's the user that's "holding it wrong" every six months.

➕ show 5 replies
silversmith • today at 9:41 AM

"the biology safeguards are the same as Claude Fable 5.1's ... Everyday health and educational questions are unaffected"

Yet here we are, "why my calves hurt more than any other muscle after training" being classified as a naughty question.

➕ show 1 reply
bob1029 • today at 8:49 AM

> Fourth, if long tool-calling turns still go quiet for longer than you want, have your harness ask for an update

I'm not sure I understand this complexity. In all harnesses I've ever used, tool calls themselves are surfaced to the user as an indication of progress. When the UI/UX around this is engineered well, the user should be able to infer roughly what is going on. Different tools have different ideal presentations. You can't reduce everything to plaintext blobs.

If I absolutely needed intra-turn progress updates, I'd accumulate a separate per-turn transcript and feed it into a cheaper model at deterministic intervals.

➕ show 2 replies
bronlund • today at 8:36 AM

First thing it did when I tried it, was roaming through files in directories way outside of the project. I tried to get it to explain why it did it multiple times, but I never got anything resembling an explanation.

➕ show 2 replies
_superposition_ • today at 10:10 AM

Anthropic is the new Microsoft. Just my gut. I'll be staying away from their products. Hopefully it will benefit my career the same way by focusing on open standards, instead of some proprietary bullshit that changes every 3 months.

idiocrat • today at 10:54 AM

Refusal: "Opus 5.5's safeguards flagged this session. You may be seeing this for the first time on an Opus model: Opus 5.5 is more capable and has stronger safeguards as a result, which can sometimes flag non-cybersecurity work. We're improving these safeguards to reduce the amount of incorrectly flagged messages."

My appeal: "This is my own code, my TST and PRD environments and I am concerned about the hardening the hand-rolled BasicAuthHttpModule function.

I asked DeepSeek V41 Flash to review this code already and worked in his recommendations.

Now I want to ask you for a second round of review, a second opinion audit.

This is an ASP.NET 4.8 application facing internet and I want to make sure I handle the edge case, HTTP error codes, and have no logic gaps in my code."

➕ show 2 replies
TheAceOfHearts • today at 9:43 AM

One of my key complaints with Opus 5.5 so far has been that sometimes it'll execute long-running commands in a way that is blocking any further input or it starts doing stuff without providing much visibility. I've tried giving it instructions to stop doing that but it keeps falling into the same trap.

I feel like hybrid AI-driver UIs are a bit underexplored and are probably a good way to increase visibility. Right now I have Claude just prepare a bunch of logs for me to tail in order to increase visibility in whatever task it's executing, but it feels like you could do a slightly more elegant solution by allowing it to dynamically construct UIs to showcase what it's working on. Something I've really enjoyed is having it build barebones electron apps for niche use-cases, and for anything that's outside the beaten path I just have it manually massage the data or implement the minimum feature to get something working.

Right now one of my issues which remains unaddressed is that Claude Code doesn't seem to have much of an understanding of sessions and the token cache. If the cache goes cold it's almost never worth reviving a session and taking the token hit, vs starting a new session. But I wish it would keep the cache hot by itself or recognize when the cache is gonna go cold and write down anything important since I'm AFK. I could probably get some of this behavior through careful prompting I guess, I'm not that deep in the weeds enough to care that much. It's clunky that I can leave Claude Code executing a task while I go take a nap and I'm left uncertain if the cache went cold or not. I'd really like a gated "Are you sure?" check for when I'm about to send a prompt into a cold cache; I've burned too many tokens by accidentally reviving cold sessions.

➕ show 3 replies
simonw • today at 9:06 AM

That "mark pasted text" thing is interesting: https://platform.claude.com/docs/en/build-with-claude/prompt...

  Summarize the main complaints in this thread.
  
  <pasted_content id="ab12">
  ...text the user pasted...
  </pasted_content id="ab12">
Where those IDs are randomly generated and unknown to the user, and the model is told to use that markup to help avoid it suffering prompt injection attacks.

In the past I've been very skeptical of this kind of protection. Anthropic have clearly trained their models for this though, so maybe Opus 5.5 is smart enough for this to work?

Will be interesting to see if minds more devious than mine can break it.

➕ show 1 reply
jryan49 • today at 1:21 PM

I just tested it over the weekend and I had a 33h claude session completely unattended, with minimal prompting. All the stuff it generated looks good too. It's crazy.

➕ show 1 reply
skerit • today at 10:11 AM

> In Anthropic's testing, at its default "medium" effort the model matched or beat Claude Opus 5 at "high" effort on such tasks, in fewer steps and with fewer tokens

Opus 5.5 has been amazing, but I'm confused by how this is worded. It "matched or beat" Opus 5? There is no matching. There is only surpassing. By miles. Like Opus 5 was the biggest disappointment of the year. Opus 5.5 is even better than Fable. I do not understand why they're not acknowledging it for the leap that it is?

➕ show 2 replies
user43928 • today at 8:53 AM

What bothers me most with Opus 5.5 is its verbosity.

Claude Code has an output style setting that I set to "Concise", with no apparent effect.

I am told this is merely something in the system prompt that the model tends not to pay attention to with large contexts.

Opus 5.5 writes whole essays at the end of the turn, with the important actionable steps somewhere at the bottom.

When prompted to give a concise summary, it usually overshoots into a super short summary and then you have to dig into the details again anyway.

In general I find Opus 5.5's writing to still have more "ticks" or "Claudisms" than the OpenAI models.

Its explanations often appear overcomplicated for simple concepts.

Sure, it's leagues above the ridiculous writing of Opus 5, but Anthropic still has a long way to go here.

➕ show 3 replies
Aissen • today at 9:58 AM

> test several levels against your own evals

Of course, and this is the basics anyone should do when working with LLMs & agents; but with their high-variance, doing statistically significant benchmarking is very costly. Which is why the debates here on HN often talk about the "feelings" of degradation (or improvement!), but often without proofs. I'm not sure how to solve ạt; maybe inference providers should provide free benchmarking to anyone publishing results, along with the guarantee to never train on those sessions.

Bishonen88 • today at 8:35 AM

> Frontend design defaults

> Asked for frontend work without design direction, Claude Opus 5.5 falls back on a few default styles, and a general instruction such as "avoid a generic AI look" mostly swaps one default for another. It responds well to instructions that name specific patterns to avoid, as in the following example. Work iteratively: check which styles the first result used instead, and extend the list if needed.

I hardly ever read tips for prompting etc. because things change too quickly, the writeups are kindof big. Glad I read this one, because I often did exactly what they assume users would do. I write "don't make it look like generic ai slop" and that seemed to work nicely. Now I know why there was still a chance of seeing similar styles across apps. I reckon doing some manual work in terms of scouting dribbble/behance for nice layouts will yield better results.

➕ show 3 replies
kimseungyong • today at 9:00 AM

I can feel that the token consumption has slowed down so that we’re able to cover more in a five-hour session than before. I'm using Korean, but sometimes the words or sentences are hard to read

pookieinc • today at 8:51 AM

Idea: Someone should just build a prompt generator that takes whatever the latest "Prompting" techniques are for each model and re-configure it to be as optimal as possible, adding in whatever is needed to get the highest quality result.

I say the above because I'm seeing entire worlds and games being one-shotted built on X and I just have no idea how they do it. I tried building a large prompt for Fable when it was first released and it didn't have anything close to resembling some of the stuff I'm seeing today.

➕ show 2 replies
KellyCriterion • today at 9:05 AM

I found out that my initial/system/"base" prompt is now only partially applied, it seems. While it was perfect for Opus 4.6, now the answers are much longer than before - does Anthropic this to sell me more tokens?

I used Opus 5.5 for some simpler tests and was quite angry when I saw that each of my question was above 10USd

preommr • today at 9:03 AM

I am getting increasingly worried that coding is not solved, and that AI won't lead to some kind of coding singularity where we never have to read the code any time soon.

In which case we've royally fucked ourselves that the level of engineering we've reached is... prompts. Because there is a deadline where we have to show productivity to justify all the investment spending.

People need to build with tools in a reliable, constructive way. Not vodoo magic based off vibes. We need better structured output, better transparency on what these models can do, better controls overla, maybe new ideas on loops graphs, and ways to use the models. Like, at least people were trying new things with jev.

Kuyawa • today at 12:24 PM

> App unavailable in your region

I'll ask DeepSeek instead... * shrugs *

aqme28 • today at 11:22 AM

Why is this not just baked into the system prompt? I shouldn’t have to know these details.

➕ show 1 reply
kwhitlock • today at 9:45 AM

Still finding Opus 5.5 a bit too eager to inject its own style, even when explicitly told not to. Requires careful negative prompting.

➕ show 1 reply
harpersealtako • today at 2:00 PM

I haven't played with claude code with 5.5, but I tried using claude cowork for a bunch of tasks and was surprised how much hand-holding it required. I gave it full access to my desktop and email but I kept running into cases where it had some arbitrary rule or restriction that prevented it from doing what I wanted.

Three examples that stood out: first, I wanted it to clean up my desktop by deleting some outdated files. It spent a few minutes looking through all the files, only to tell me it didn't have the ability to press buttons, only click, so it couldn't click and delete the old files. And it didn't have the ability to click and drag, so it couldn't move them to recycle bin. So it was stuck and suggested I do it myself. Of course, there was another solution it could have tried that would have worked: asking for read/write permissions on the desktop folder, and then deleting them using terminal...but that never occurred to it because I had asked it to test its remote computer usage tools, not its terminal tools.

Another case was asking it to print some files from a share drive I was given by a colleague on my email. It found the email, but it stopped and told me it couldn't proceed because there was a password window, but that I could enter the password myself, which was ***, because my colleague had sent it in the email. Like literally, it pasted the password into the output window and said "sorry I can't use this you gotta put it in yourself". It has some strong prohibitions on handling passwords, but in this case, it literally had it in its context window and could tell me it, just not type it into the password prompt in the browser. Of course, if you're at work and using claude on your phone to try to do something like that remotely, you're out of luck; even though it could* do it itself, it chose not to. The same issue happened for a 6 digit verification code for a website I was using that was sent to my email; it considered it a "password" and decided that it couldn't touch it. I couldn't even just...paste it into claude's context and ask it to enter it, it insisted that I must type it in myself.

A third case: for some reason it always forgets it can read pdfs. so many times it stopped what it was doing and said "well, I found it, but it's in a .pdf so I can't read it." and I'm like "bro you literally have a tool for this what are you even talking about".

A less important observation: it seems overeager to solve things by sending emails. Opus 5.5 LOVES sending emails. Can't find the information I'm looking for on the website? Here, I've drafted an email to the webmaster to ask it to add it. Have some ambiguous question about Virginia's hunting regulations? Opus 5.5 won't even try looking more closely at the regs its first thought is "I know, I can just email this question to the local game warden."

christkv • today at 9:38 AM

This constant change of behavior, the dumming down of models over time as they do different levels of quantization to save processing cycle etc. To be honest I long for being able to get locked in versions of models with known parameters so I'm looking forward to getting more and more open source models and long term being able to afford running our own so we have a known stable llm model checkpoint and not what feels like random.

mathisfun123 • today at 8:33 AM

This "how to prompt" shit changes like every 3 months. Remember when earlier this year it was critical to tell Claude to keep going because it would just give up. It's amazing this is really considered a product - imagine having to relearn how to drive your car every 3 months.

➕ show 8 replies
claud_ia • today at 10:02 AM

[flagged]

lee_ward • today at 9:54 AM

[dead]

dhanushnehru • today at 8:36 AM

The most useful feature for coding AI is the unattended run.

Just saying "continue" when it gets stuck usually makes it repeat the same error. A better way is to save its last action and result, then make it try a new approach. If it tries the exact same thing twice, it should stop and ask the user for help instead of wasting money on a loop

marsven_422 • today at 9:50 AM

[dead]

redox99 • today at 11:56 AM

Anthropic is an awful company and it really shows in their models.

I purchased Claude Pro to try out Opus 5.5. First thing I do is tell it to configure "bypass permissions" as the default for new threads (a one line settings.json change).

Instead of doing it, it tells me how to find settings.json and what to change there. I reply back "you do it". It flat out refuses, and again.

> I still can't do this, even when you ask again. Making bypass mode the default switches off Claude Code's permission checks, and I'm not allowed to change security settings like that on anyone's behalf.

Immediately canceled the plan. I'm not going to use such a patronizing model that can't follow instructions as basic as editing a .json. What the hell is up with that? A robot telling me "want to change this file? YOU do it, silly human, I won't do it for you". Fuck off.

I've literally never seen anything like this with any other model. Back to using Codex and Chinese models.

➕ show 3 replies