logoalt Hacker News

moritz • today at 3:33 PM • 0 replies • view on HN

You‘ll notice a lot of "I found"s and "my hunch"s and similar in the comments.

There aren’t many comparative benchmarks, and obviously the design of such a benchmark is difficult, but, e.g. https://arxiv.org/abs/2508.09101

In this benchmark, models can correctly solve Rust problems 61% on first pass — A far cry from other languages such as C# (88%) or Elixir (97%, no static typing whatsoever).