The math and cybersecurity improvements this year don’t look like plateauing to me. They’re clearly improving?
But I think they’re becoming more specialized. Luna is working fine for me for ordinary web development, but I was impressed by Sol tracking down an OS-level bug that was causing my Playwright tests to be flaky. Previous models couldn’t figure it out.
A lot of the improvements might be on questions that ordinary users aren’t normally asking.