I'm getting about 3k tok/s on prefill so it heavily depends on the input size. I hooked it up into a coding agent for shell permission checks, estimated shell execution time (buckets), goal pre-screening, subagent routing etc. and I have RTTs between 100-200ms. It can take noticeably longer for huge contexts because of the prefill speed but it can be cached so subsequent calls will be much faster.
I tried smaller decision models like laya (also with custom finetunes) but the accuracy was not really good (for the things I tested). Also I don't have any VRAM left on this GPU so i had to decide whether to host laya or qwen-3.8-27b but not both at the same time. Running decision models on a CPU will also be noticeably slower so I went down this route to have both combined with shared base weights.
I'm getting about 3k tok/s on prefill so it heavily depends on the input size. I hooked it up into a coding agent for shell permission checks, estimated shell execution time (buckets), goal pre-screening, subagent routing etc. and I have RTTs between 100-200ms. It can take noticeably longer for huge contexts because of the prefill speed but it can be cached so subsequent calls will be much faster.
I tried smaller decision models like laya (also with custom finetunes) but the accuracy was not really good (for the things I tested). Also I don't have any VRAM left on this GPU so i had to decide whether to host laya or qwen-3.8-27b but not both at the same time. Running decision models on a CPU will also be noticeably slower so I went down this route to have both combined with shared base weights.