logoalt Hacker News

mritchie712today at 10:50 AM1 replyview on HN

in short: it's faster, cheaper, smart structured output.

each "question" is answered in parallel instead of a sequential (like an LLM). so if you have an input like:

    {"is_it_hotdog": noul, "is_it_apple", noul}

it answers is_it_hotdog and is_it_apple in parallel and gives a probability.

Replies

satvikpendemtoday at 11:16 AM

Can't I just parallelize my LLM calls myself for each question?

show 2 replies