Pretty cool!
Only 10B parameters if I correctly understand some of those histograms, which is tiny in LLM world! It bodes well for the future.
The training videos seem to show human hands using the robot grippers to execute the tasks, which must have been a fun but tedious process.
Yes, it's truly remarkable how small modern VLAs are.
Even Tesla's FSD AI stack is small by modern LLM standards. Once the hardware is there to enable real time 100B inference on endpoint SBCs? We'd see gains in AI robotics from that alone.