Even though they're technically trained for music, it might be worth testing Suno/Lyria to see if they're able to do this. You might have to isolate it afterwards, but seems like it could be viable
Another HN user suggested embedding sound effects. The lead author of MOSS-TTS emailed me that the next version will have a SOTA voice designer. Poke around the voice design space, you'll find text-to-effects here and there. I certainly have my eyes on this space.
Maybe a dumb idea, but how good are image models with manipulating spectrogram images? Then a workaround could be: convert input audio to spectrogram -> pass spectrogram + prompt to an image model -> convert modified spectrogram back to audio waveform.
Could be interesting, though spectrogram voice gen is a step back in terms of advancement. Vocoder models that take in Mel Specs suffer from a lot of issues like hissing due to Griffin Lim issues. Would be curious to see if diffusion models can work on neural codecs directly.
I think Google had one called riffusion (the first version was designed for specs)
Even though they're technically trained for music, it might be worth testing Suno/Lyria to see if they're able to do this. You might have to isolate it afterwards, but seems like it could be viable
https://github.com/OpenMOSS/MOSS-TTS
KeenLore is my locally hosted full-cast emotive audiobook generator I'm developing for my novel:
https://www.youtube.com/watch?v=WAeHgE94rVo
Another HN user suggested embedding sound effects. The lead author of MOSS-TTS emailed me that the next version will have a SOTA voice designer. Poke around the voice design space, you'll find text-to-effects here and there. I certainly have my eyes on this space.
KeenLore is really cool.
Daydream Music - https://daydream.live
The DEMON realtime engine is open source. The hosted service has a number of integrations with popular audio tools. (I'm on the team)
Maybe a dumb idea, but how good are image models with manipulating spectrogram images? Then a workaround could be: convert input audio to spectrogram -> pass spectrogram + prompt to an image model -> convert modified spectrogram back to audio waveform.
https://huggingface.co/docs/transformers/model_doc/audio-spe...
I wouldn't call it a dumb idea, but there's soooo much subtlety to sound that wont be visible in any reasonably sized image.
Could be interesting, though spectrogram voice gen is a step back in terms of advancement. Vocoder models that take in Mel Specs suffer from a lot of issues like hissing due to Griffin Lim issues. Would be curious to see if diffusion models can work on neural codecs directly.
I think Google had one called riffusion (the first version was designed for specs)
Apparently the AudioX and AudioLDM(2) models do this but I think you've found a genuine gap.
Plenty of models take in text + audio and spits out audio. It's the format of most newer generation accent conversion/voice cloning models.
What's your exact use case?
Sound effects are completely different to voice, which those models are trained to output.