Ask HN: Are there AI models for generating sounds based on a text and reference?
33 points by onemiketwelve 2 days ago
33 points by onemiketwelve 2 days ago
I've been having a hard time finding a solution. Is there really no commercialized model that I can feed in a reference sound and text instruction and get another sound out?
Right now having a multimodal inputs to image or text output is a commodotized, solved problem. IE you can put a prompt for some image and use a reference image to guide the model on what you want. After all, a picture is worth a thousand words right?
Ive also used a text and audio input in order to get a text description or classification out.
I cannot for the life of me find a solution for Audio + text -> Audio
My usecase is that I'm trying to generate new sound effects based on a source that doesn't have that many clean examples and thought this would be just another solved workflow but I'm not finding much. Tried elevenlabs SFX but it's only text->audio and without the reference, it's hard to guide with just text and recreate a sound with any accuracy. I tried the Stable Audio 3 model which is supposed to do what I want but the results were terrible. This modality seems to be stuck where image generation was in 2016. Is there something else or is this just not a common need?
You should look into using AI to generate code that synthesizes sounds. Not exactly what you asked for, but I do think it is an approach worth considering: https://m.youtube.com/watch?v=1-i45X5aj94 I've had some luck with that approach - giving a frontier LLM a reference sound, asking it to replicate the sound via code (basically asking it to create a physical model, I guess) and being able to steer the output by talking with the LLM. I was working with very simple elements, but I was surprised by some of the outputs. I would have also suggested Waves Illugen, but it turns out that is text-only, you can't give it an audio reference. https://github.com/OpenMOSS/MOSS-TTS KeenLore is my locally hosted full-cast emotive audiobook generator I'm developing for my novel: https://www.youtube.com/watch?v=WAeHgE94rVo Another HN user suggested embedding sound effects. The lead author of MOSS-TTS emailed me that the next version will have a SOTA voice designer. Poke around the voice design space, you'll find text-to-effects here and there. I certainly have my eyes on this space. Even though they're technically trained for music, it might be worth testing Suno/Lyria to see if they're able to do this. You might have to isolate it afterwards, but seems like it could be viable Not exactly what you asked for, but are you aware of CLAP?
https://github.com/LAION-AI/CLAP This generates audio embeddings - much like CLIP does for visual inputs. Maybe a dumb idea, but how good are image models with manipulating spectrogram images? Then a workaround could be: convert input audio to spectrogram -> pass spectrogram + prompt to an image model -> convert modified spectrogram back to audio waveform. I wouldn't call it a dumb idea, but there's soooo much subtlety to sound that wont be visible in any reasonably sized image. Could be interesting, though spectrogram voice gen is a step back in terms of advancement. Vocoder models that take in Mel Specs suffer from a lot of issues like hissing due to Griffin Lim issues. Would be curious to see if diffusion models can work on neural codecs directly. I think Google had one called riffusion (the first version was designed for specs) Hi! I’m unsure if this is what you were looking for, QWEN3 TTS has a “clone” feature. You can use the local version, give it a voice reference and text and it will read the text with your instructions with the sample voice you selected. I used this combo to make a customised “audiobook” for my mother. It worked quite well. Daydream Music - https://daydream.live The DEMON realtime engine is open source. The hosted service has a number of integrations with popular audio tools. (I'm on the team) Apparently the AudioX and AudioLDM(2) models do this but I think you've found a genuine gap. Plenty of models take in text + audio and spits out audio. It's the format of most newer generation accent conversion/voice cloning models. What's your exact use case? Sound effects are completely different to voice, which those models are trained to output.
soundworlds - 18 hours ago
SyneRyder - 12 hours ago
thangalin - 19 hours ago
moonu - 20 hours ago
bobosha - 8 hours ago
xg15 - 20 hours ago
Buttons840 - 20 hours ago
narrationbox - 19 hours ago
jajazheng - 8 hours ago
jallmann - 19 hours ago
chr15m - 20 hours ago
narrationbox - 21 hours ago
chr15m - 20 hours ago