I worked on Kolibri, in particular pre-training data and mid-training. We strive to be as open as possible. Glad you like it.
How do you cleanse the data at this scale?
How do you cleanse the data at this scale?