Ask HN: Are there AI models for generating sounds based on a text and reference?
I've been having a hard time finding a solution. Is there really no commercialized model that I can feed in a reference sound and text instruction and get another sound out? Right now having a multimodal inputs to image or text output is a commodotized, solved problem. IE you can put a prompt for some image and use a reference image to guide the model on what you want. After all, a picture is worth a thousand words right? Ive also used a text and audio input in order to get a text description or classification out. I cannot for the life of me find a solution for Audio + text -> Audio My usecase is that I'm trying to generate new sound effects based on a source that doesn't have that many clean examples and thought this would be just another solved workflow but I'm not finding much. Tried elevenlabs SFX but it's only text->audio and without the reference, it's hard to guide with just text and recreate a sound with any accuracy. I tried the Stable Audio 3 model which is supposed to do what I want but the results were terrible. This modality seems to be stuck where image generation was in 2016. Is there something else or is this just not a common need? 0 comments on Hacker News.
I've been having a hard time finding a solution. Is there really no commercialized model that I can feed in a reference sound and text instruction and get another sound out? Right now having a multimodal inputs to image or text output is a commodotized, solved problem. IE you can put a prompt for some image and use a reference image to guide the model on what you want. After all, a picture is worth a thousand words right? Ive also used a text and audio input in order to get a text description or classification out. I cannot for the life of me find a solution for Audio + text -> Audio My usecase is that I'm trying to generate new sound effects based on a source that doesn't have that many clean examples and thought this would be just another solved workflow but I'm not finding much. Tried elevenlabs SFX but it's only text->audio and without the reference, it's hard to guide with just text and recreate a sound with any accuracy. I tried the Stable Audio 3 model which is supposed to do what I want but the results were terrible. This modality seems to be stuck where image generation was in 2016. Is there something else or is this just not a common need?
I've been having a hard time finding a solution. Is there really no commercialized model that I can feed in a reference sound and text instruction and get another sound out? Right now having a multimodal inputs to image or text output is a commodotized, solved problem. IE you can put a prompt for some image and use a reference image to guide the model on what you want. After all, a picture is worth a thousand words right? Ive also used a text and audio input in order to get a text description or classification out. I cannot for the life of me find a solution for Audio + text -> Audio My usecase is that I'm trying to generate new sound effects based on a source that doesn't have that many clean examples and thought this would be just another solved workflow but I'm not finding much. Tried elevenlabs SFX but it's only text->audio and without the reference, it's hard to guide with just text and recreate a sound with any accuracy. I tried the Stable Audio 3 model which is supposed to do what I want but the results were terrible. This modality seems to be stuck where image generation was in 2016. Is there something else or is this just not a common need? 0 comments on Hacker News.
I've been having a hard time finding a solution. Is there really no commercialized model that I can feed in a reference sound and text instruction and get another sound out? Right now having a multimodal inputs to image or text output is a commodotized, solved problem. IE you can put a prompt for some image and use a reference image to guide the model on what you want. After all, a picture is worth a thousand words right? Ive also used a text and audio input in order to get a text description or classification out. I cannot for the life of me find a solution for Audio + text -> Audio My usecase is that I'm trying to generate new sound effects based on a source that doesn't have that many clean examples and thought this would be just another solved workflow but I'm not finding much. Tried elevenlabs SFX but it's only text->audio and without the reference, it's hard to guide with just text and recreate a sound with any accuracy. I tried the Stable Audio 3 model which is supposed to do what I want but the results were terrible. This modality seems to be stuck where image generation was in 2016. Is there something else or is this just not a common need?
Hacker News story: Ask HN: Are there AI models for generating sounds based on a text and reference?
Reviewed by Tha Kur
on
October 06, 2026
Rating:
No comments: