The paper explains absolutely everything as if it was a tutorial "how to made your own modern agentic LLM". They even tell how they made their dataset. https://aleph-alpha.com/downloads/tech-report.pdf ; It's the first time I see this level of openness.
In my opinion not being open about which data is ingested and trained on, and trying to make that a repeatable thing for a third party, is not worth being called "open". Glad they did that.
The paper mentioned is here: https://tej.as/blog/aleph-alpha-kolibri.
(This comment was originally posted to https://news.ycombinator.com/item?id=49943034, but we're merging the threads.)
Such a crazy change from the times of Luminous, when they published a three pager with a claim that the model is similar good as „gpt 3“ (which??) with some graphs without y axis.
Bravo team!
Yeah the pdf alone is awesome as a learning tool.
This alone makes it much more valuable than many high-profile releases despite not quite performing at the same level.
Hopefully this becomes the new standard.
It’s seemed crazy to me that anyone thought these could stay closed or even SHOULD be closed source.
His point about regulation and innovation is great and I wish more people thought like that.
One of humanity’s biggest problems here is we don’t know how to do moderation.
We have two modes. One is a brick taped to the accelerator and damn all consequences, driven by national pride or corporate greed or egos. The other is a brick taped to the brake driven by histrionic doomers and anti-everything pessimists.
The extremes are loud and fit in a tweet. Nuance is quiet and contemplative and usually requires an essay or a book. It’s also dynamic. Nuanced positions evolve over time as new things are learned. Extremes tend to be fixed and rigid. All this, I think, gives them higher memetic fitness in the discourse.
I don’t think this is new. Look at nuclear power, a largely pre-Internet example. You had pro nukes who minimized and hand waved away any risk and anti nukes that wanted it utterly outlawed. Nobody said “hey this is a great zero carbon source of energy but we really need to think it through carefully and manage it well.” Or if they did they were drowned out by the loud screaming extremes.
I worked on Kolibri, in particular pre-training data and mid-training. We strive to be as open as possible. Glad you like it.