I see that Muse Spark 1.3 (max) beats GPT-6 Astra on some benchmarks:
Test: Muse Spark 1.3 / GPT 6 Astra
DeepSWE v1.1: 75.4% / 74.1%
AutomationBench: 49.4% / 41.4%
Is that enough to bring this discussion down to earth again?
The DeepSWE one is interesting. Even Gemini 3.8 Flash is only 0.5% behind Astra. Maybe DeepSWE is saturated at 75%.
The DeepSWE one is interesting. Even Gemini 3.8 Flash is only 0.5% behind Astra. Maybe DeepSWE is saturated at 75%.