The problem with this idea is that knowing Python makes the model a better Swift programmer, as does a higher-number of parameters during training. So you'd be so much better off with a 90B general purpose model trained on everything anyway.
Except assuming a fixed budget of parameters, there is clearly stuff that's better for programming than others. E.g. Qwen 27b is a better coding model than Gemma 4 31B. More params doesn't automatically win. Perhaps what makes a good swift programming model is a ton of python training, so those two things can't be separated - but that doesn't negate the idea of loading a model that's good at the specific task you want - or has specific knowledge of the libraries and tools for the language you're using at the expense of the ones you're not using.
Except assuming a fixed budget of parameters, there is clearly stuff that's better for programming than others. E.g. Qwen 27b is a better coding model than Gemma 4 31B. More params doesn't automatically win. Perhaps what makes a good swift programming model is a ton of python training, so those two things can't be separated - but that doesn't negate the idea of loading a model that's good at the specific task you want - or has specific knowledge of the libraries and tools for the language you're using at the expense of the ones you're not using.