Appreciate the detail in this and the previous post on creating internal benchmarks!
Have you all attempted finetuning smaller OSS models on your repos for coding?