logoalt Hacker News

brianyu8 • yesterday at 5:00 AM • 1 reply • view on HN

I tried reproducing your examples as predicate questions on the Decisions API (https://gist.github.com/by-openai/7b376d7866e36a4f38581f73ce...) and got:

  1. Title: "Mortgage calculator: estimate your monthly payment (Bankrate)"
     Tag:   "house hunting"
     Decisions API (2 runs):  0.99, 0.99

  2. Title: "S&P 500 index: live chart and news (Bloomberg)"
     Tag:   "investing"
     Decisions API (2 runs):  1.0, 1.0

If you have other examples of requests with unexpected outputs, feel free to email me at [email protected] and we can try to get to the bottom of it. Thanks for trying out the API!

Replies

Topfi • yesterday at 11:06 AM

Happy to follow up with you, will write a mail with some of my tasks. For the record, just ran again on OpenRouter and got 0.66 [0]. Screenshots for transparency.

Also reran Jev [1] and Mercuy Decide [2] (which got 0.83) with same input, for reference.

Also, also, used the example via OpenRouter exactly as you did with the only change being my original input (which has my slightly odd tagging system and multiple tags in input though requests a noul as output) and got a 0.47 [3] on investing both times.

If Jev did well but both Mercury Decide and Luna failed, I'd chuck that up to my use/prompting, but Mercury Decide does well here so it seems Luna specific. Will add that I did also try Clef, happy to share that data if anyone wants, but quickly dropped Clef due to pricing vs Jev.

[0] https://imgur.com/a/nAkCZiG

[1] https://imgur.com/a/scIa9nF

[2] https://imgur.com/a/M7Mm1l7

[3] https://gist.github.com/Topfi/d77e503c7d1f6d11fc32d1b2174ec0...