6 comments

  • yoloakki 2 minutes ago

    You definitely need independent evals by Datapoint AI or someone who can verify your claims about TTS quality

  • mowmiatlas 19 minutes ago

    Cool, I’ve released something to the same beat of the dr this weekend as well

    https://github.com/loudreader/loudkit

    I think real time natural tts should be possible everywhere soon

  • rahimnathwani 41 minutes ago

    For some reason it switched voices half way through a 33 second clip.

    For OP the clip name is nari-nina-01a0a12f-980a-765e-8029-fa56bd23210d.wav

  • asaiacai an hour ago

    This is really cool work! I'm curious like what do you see as the biggest lever for speeding up TTS models or from a technical perspective that this was a promising direction in the first place to push on. If I were to guess, some distillation but I'm certain there are probably TTS model aware architectural changes that just make inference wayyyy faster?

  • ipsum2 an hour ago

    If you're going to announce a TTS model, service, or whatever, you really need demos.

  • meatmanek an hour ago

    > and Qwen3-ASR

    Is the ASR inference engine open source as well?