Skip to content
THE AI WIREINTELLIGENCE THAT MATTERS
Research

M2Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models

arXiv:2609.18259v2 Announce Type: cross Abstract: Recent advancements have successfully adapted autoregressive language models to process multimodal signals, such as images and actions. Since raw action signals are continuous, effective tokenization is essenti

arXiv cs.AI··Updated just now·38 sightings