logoalt Hacker News

hedgehogyesterday at 3:37 PM1 replyview on HN

Have you quantified the performance on any particular benchmark?


Replies

ac-cianoyesterday at 3:44 PM

I thought about it but where the tool shines are large undefined tasks, which are complex to quantify and test (not as easy as implementing a simple bug fix that you can test directly). Even if I were to find such test dataset it would probably require a lot of money to reach statistically significant results. Either way, the framework is inspired a lot from Matt Pocock, and is pretty much established.

show 1 reply