logoalt Hacker News

arstructincyesterday at 12:50 AM0 repliesview on HN

I think reproducing existing software and building a new product are quite different tasks.

In a benchmark, there are usually clear tests and a correct reference. In actual product development, requirements are often unclear, and we do not always know what the correct result is.

Passing tests is also not enough to confirm security, maintainability, or operability. I would like to see a benchmark where an AI continues changing the same product for several months. It would be interesting to see whether the architecture remains understandable and safe after many changes.