interesting bench list, what about benchmark against smaller or bigger models? 9B looks too huge for small like laya, and too small for llm-level decisions.