logoalt Hacker News

NalNezumi • today at 1:39 PM • 1 reply • view on HN

Two aspects: industry & research (worked in both).

Research is using a lot of VLM (or whatever Jim Fan decided to call them this year) when it comes to manipulation (robot arm). For navigation it is still a little mixed bag, best to see self driving advancement for that. On the control side (cool jumping humanoid) it's a mixture of RL/IL and traditional control(MPC) and running a slow ass VLM/LLM is not an option.

SOTA is hard to say in the sense that there's just so much variation to the task you want to measure. This can be considered SOTA in one aspect https://youtu.be/3aQWvdCac9o As it's an robot that can lift a heavy unknown object without information about it, RL trained in simulation, but at the same time trained eye will see that the task is made simple by fridge being put on a pallet making the bottom of it exposed for grabbing. On the other hand you have PI (Chelsea Finn / Sergey Levine) https://youtu.be/cRZNwgvcWUg that can do nothing of the thing above but can cook food and fold shirts. Deformable objects and tools interaction is hard so their demo is more focused on understanding and interacting with the world compared to the fine motion like the Boston dynamics one. The latter is also teleoperated, and teleoperation tend to look good but have pretty sad deployment success %

Industry is... Depends on startup vs established. VLA/LLM is flashy but really even 95% success rate is not enough for factory/logistics automation if the failure is unrecoverable, and VLM/LLM is far, faaaaaaaaaaaar from that success rate. In controlled settings traditional pipeline of detection + motion planning often outperform VLM/LLM hybrid models by a mile. Established ones have run humanleess logistic centers for many years now and they're seldom VLA stuff, although there is an interest in using it for the 5% failure cases. But they're not interested in SOTA.


Replies

mohamedkoubaa • today at 1:47 PM

Even if VLA reaches 100% success rate it still isn't viable. No factory operator wants a cloud hosted model in the critical path. And nobody is (yet) building a server rack that can be air gapped in the factory to run VLA, and these models don't run on a Mac Mini. Probably an opportunity for something like Oxide to pick up, but they have more profitable segments to go after.