There was some paper about routing at training bio-knowledge into a particular region of the model, which you then can cutoff when serving. But you probably lose some efficiency since maybe you sized that region too small/too big.
That's a very clever approach. Any idea about the papers title or authors? I'd love to look it up.
That's a very clever approach. Any idea about the papers title or authors? I'd love to look it up.