11 comments

  • moonu an hour ago

    Even though they're technically trained for music, it might be worth testing Suno/Lyria to see if they're able to do this. You might have to isolate it afterwards, but seems like it could be viable

  • thangalin an hour ago

    https://github.com/OpenMOSS/MOSS-TTS

    KeenLore is my locally hosted full-cast emotive audiobook generator I'm developing for my novel:

    https://www.youtube.com/watch?v=WAeHgE94rVo

    Another HN user suggested embedding sound effects. The lead author of MOSS-TTS emailed me that the next version will have a SOTA voice designer. Poke around the voice design space, you'll find text-to-effects here and there. I certainly have my eyes on this space.

  • jallmann 37 minutes ago

    Daydream Music - https://daydream.live

    The DEMON realtime engine is open source. The hosted service has a number of integrations with popular audio tools. (I'm on the team)

  • xg15 2 hours ago

    Maybe a dumb idea, but how good are image models with manipulating spectrogram images? Then a workaround could be: convert input audio to spectrogram -> pass spectrogram + prompt to an image model -> convert modified spectrogram back to audio waveform.

    • 0x20cowboy an hour ago
    • Buttons840 2 hours ago

      I wouldn't call it a dumb idea, but there's soooo much subtlety to sound that wont be visible in any reasonably sized image.

    • narrationbox 37 minutes ago

      Could be interesting, though spectrogram voice gen is a step back in terms of advancement. Vocoder models that take in Mel Specs suffer from a lot of issues like hissing due to Griffin Lim issues. Would be curious to see if diffusion models can work on neural codecs directly.

      I think Google had one called riffusion (the first version was designed for specs)

  • chr15m an hour ago

    Apparently the AudioX and AudioLDM(2) models do this but I think you've found a genuine gap.

  • narrationbox 2 hours ago

    Plenty of models take in text + audio and spits out audio. It's the format of most newer generation accent conversion/voice cloning models.

    What's your exact use case?

    • chr15m an hour ago

      Sound effects are completely different to voice, which those models are trained to output.