Explicit Language Memory for Long-Horizon Planning in Vision-Language-Action Models
arXiv:2608.04765v3 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models provide a unified paradigm for connecting visual perception, language understanding, and robotic control. However, existing VLA models still face major challenges in
arXiv cs.AI··Updated just now·38 sightings