logoalt Hacker News

viccistoday at 6:31 PM1 replyview on HN

Tbh the extreme discrepancies between very confidently made statements about "X is better than Y" a sign a to me this is something where peoples' biases are out of control. Not saying you are wrong. Just that these things are so hard to evaluated apples to apples and that things like under-the-hood routing changes muddy the waters so much that it's really hard to get reasonable assessments.


Replies

numandinatoday at 7:48 PM

My tests are very visual based and easy to notice when the seed/temp or the routing or the A/B test comes into play. And it's been months of working with these models so comparison is straightforward and robust. Fable is the first time I can use AI to code with good results while everyone has been doing it for ages before me.