I think the race is actually between locally optimized rigs with extremely strict context management, and the rest.
Those rigs are running one of the open Chinese models though, right?
Those rigs are running one of the open Chinese models though, right?