Here is training code for the MIMO foundation model pre-trained on 100 million hours audio and corresponding amounts of text. would-be-great-to-fine-tune-for-music-generation, voice-acting or sound-effect-generation. @JulienBlanchon@realmrfakenamegithub.com/XiaomiMiMo/MiM…
15% of the Emilia datasets are annotated with emotion scores + captions. This is more than 30,000 hours. But we need more GPUs to finish it completely. We also have 700,000 hours of permissively licensed speech snippets to be transcribed and annotated. We need more GPUs! :)