logoalt Hacker News

TN1ck • today at 1:39 PM • 1 reply • view on HN

I just did a run with a benchmark I just used to test other models against. (It's about detecting irony in german soccer tweets). On my M5 Pro with 48GB it took over 30min to decide on just 100 tweets, the thinking definitely takes long.

It performed quite below Jev, but above other open decision models I tested (68 correct vs 79 correct for Jev - see [1]). I'm running it for the moderation benchmark as well, but that will probably take a few hours on my machine.

[1] https://tn1ck.com/blog/jevdit


Replies

TN1ck • today at 5:57 PM

Update: Jeeves took about 2 hours to moderate 394 data points and performed really well. It’s not as good as Jev, but it’s super close! In general, it’s super cool that you can tune how strict you want content moderation to be with these models.

➕ show 1 reply