logoalt Hacker News

thornewolftoday at 4:18 PM3 repliesview on HN

Write a prompt, evaluate the prompt, understand that is succeeds 95% of the time.

Write a new prompt, evaluate, it now succeeds 99% of the time. Measure what changes between prompt #1 and prompt #2, understand what contributed to the performance jump.

Write a third prompt, this one succeeds 100% of the time. Increase the size of your evaluation set, find a 1/5000 error-class and a 1/10000 error-class, add some explicit code to correct for this cases.

Roll out to production, collecting usage metrics. You make some tweaks to your harness, your prompts. Eventually you have confidence that your system has fewer mistakes than 1 in 100k.

Now, multiply this iteration across all your different prompts and different ways that they might interact with one another.


Replies

kfsonetoday at 6:14 PM

There's a reason engineers are prissy about people coming along and saying "I write code, I'm an engineer" that people periodically try to sand-paper away.

Engineers don't just tie a sheet to a rock and throw it off a cliff and call themselves aerospace engineers.

They do full diligence on the theory, math, physics, material science, fluid dynamics, etc, and plan a controlled series of tests specifically designed to verify/challenge/disprove their concept and the theories behind it.

Sure, there's a team member ultimately responsible throwing half a dozen rocks off a cliff in the first test.

A technician.

The guy who throws the rock off the cliff is a technician.

show 1 reply
badnewtoday at 4:44 PM

It's not engineering if you're just guessing as to what is degrading the performance and what might improve it.

show 4 replies
kevin_thibedeautoday at 4:24 PM

95% is shit tier engineering. Would you be satisfied if your keyboard randomly failed 5% of the time.

show 2 replies