Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data
arXiv:2609.18842v1 Announce Type: new Abstract: The scaling laws hold that a language model grows more capable with more parameters and more training data, and Mixture-of-Experts (MoE) architectures have ridden these laws to remarkable results, activating only
arXiv cs.AI··Updated just now·34 sightings