GrainSpeech: Less Context, More Detail for Compact Speech Synthesis
arXiv:2609.18856v1 Announce Type: cross Abstract: Compact acoustic models face a challenging quality-capacity trade-off. We investigate two factors in this regime: encoder context and Mel-spectrogram supervision. A receptive-field-scaling study shows that expa
arXiv cs.AI··Updated just now·34 sightings