what is the easyest way to use the chinese models and which harness does work with them well?
opencode with model inference on cheaperinference.com has been working well for me - glm-5.3-flash is shockingly cheap (i've spent a total of a few dollars over several weeks of heavy usage), fast and capable for cyber tasks
OpenCode harness with their subscription would be my recommendation.