logoalt Hacker News

Jackson__ • today at 6:32 PM • 3 replies • view on HN

I've just tested Strata on a simple 50 image vision benchmark. The task is to output the exact coordinates of a requested object. The result via Strata had a median error distance of 154.8 pixels, avg of 168.8. Running the exact same GGUF and vision adapter weights on llama.cpp gives me a median error of 46.5, avg 81.4.

To put that into perspective, here are some more numbers from other models via llama.cpp:

Median/Average

Qwen 3.5 9B BF16: 46.5 / 193.3

Qwen 3.6 35B Q4 K XL: 38.4 / 76.4

Qwen 3.5 122B Q3 K M: 32.9 / 68.6

The difference in vision performance is as large as the jump from a 9B model to a 35B model. All tests were performed at temp=0.

I have done no further testing, as these results line up perfectly with my expectations.


Replies

Xenograph • today at 9:59 PM

Is this test available somewhere? Would like to test it out on my models.

throwaway219450 • today at 9:46 PM

What does SAM3 get on the same test set?

NamlchakKhandro • today at 8:44 PM

Tldr, strata is a waste of time.