Question Local LLM For Voice Over With Exertion Modulation

kim05

Active member
Joined
Aug 12, 2023
Messages
118
Reactions
31
Have anyone experimented with any LLMs with emotion control that can be run locally with no censorship? I really want to make a scene with a female trying to talk normally whist getting banged, so that would entail a relatively normal script for the machine to read, but relatively extensive emotion control like gasps and other inflections of exertion. So far Qwen3 TTS (under 8Gb VRAM) and IndexTTS2 had made me feel like I've wasted my time installing them. I'd be delighted if someone could prove me wrong and explain how it's done with those models or advise something else.

Please, leave comments like "AI = bad" on reddit, I'm aware.
 
Yeah, man. You have Voxta and some plugins out there, but running locally also means you have the HW power to do VAM and LLMs.
 
Upvote 0
Yeah, man. You have Voxta and some plugins out there, but running locally also means you have the HW power to do VAM and LLMs.
No, I mean an LLM with a audio file output. I don't need chatbot with VaM integration, those still sound very much like it wants to sell me some sort of corpoproduct or a subscription, so Voxta is not really for me. What I thought of is more like a triggered voice line type use. Line 1 on Pose 1, Line 2 and 3 on Pose 2 etc. I was surprised by how Qwen3TTS can clone voices so they don't sound robotic, but it cannot do emotion modulation and breath/exertion sounds to that or at least I haven't succeeded at getting such a result. I was curious if anyone tried dealing with this kind of generation for the purposes of AI simply voicing the script with emotion and heavy breathing, not interacting with me directly.
 
Upvote 0
There's some plugins that do X or Y things, and vamx also has some sort of AI features. I don't know what they do or not do.

For my use I do AI cloned voice audio sentences files and play them on my scenes during a story intro or more commonly during the saucy sex animations. They have some emotion depending on the source material and the model used, currently Index-TTS and Higgs, but I create them all beforehand, not in real time with "script interpretation".
 
Upvote 0
I had some limited success with coqui xtts2 cloning locally and using samples of audio from solo porn vids, asmr audio pron, etc.

That's probably a bit dated these days though and something better might now be out but I've not really bothered looking in to it again. it was mostly out of curiosity.

Impressive results to say the least but In the end it was too much work involved creating a good enough dialogue
 
Upvote 0
If you're looking for better and up to date tools to create AI voices or AI cloned voices audio files, look no further:
11labs can suck it.
 
Upvote 0
Even the most advanced cloning services can't pull that off. Highly unlikely you'd find a free open source one able to do that properly.
 
Upvote 0
Rather than going the TTS route, why not put your script through video+audio gen (with nsfw LoRA) like LTX or MiniMax H3? The speech might be a little more natural when driven by a video.
 
Upvote 0
Even the most advanced cloning services can't pull that off. Highly unlikely you'd find a free open source one able to do that properly.
Gemini (aistudio.googledotcom) did all I asked to specs and left me impressed to be honest, when it comes to emotion and inflection. It only has presets for voice though. Qwen3TTS cloning was extremely accurate, but no emotion control for cloned voice. IndexTTS2 has a lot of promise, but I haven't had any good yield from trying since it has higher recommended VRAM requirement than I can afford. I am trying to figure out if there is some formula, a chain of model out there that would permit to finally output something exactly to my specified instruction. I do not hope to find the philosopher's stone, you understand, I'm just looking at the current open-source possibities and pondering the future if anything. And if I will find a sufficiently effective chain or a wonder model than that's even better. That's the point of asking to see the things I haven't yet.
 
Upvote 0
If you're looking for better and up to date tools to create AI voices or AI cloned voices audio files, look no further:
11labs can suck it.
This is exactly why I've asked. Wonderful suggestion, I'll investigate with pleasure. Thank you!
 
Upvote 0
Rather than going the TTS route, why not put your script through video+audio gen (with nsfw LoRA) like LTX or MiniMax H3? The speech might be a little more natural when driven by a video.
Haven't thought of that, honestly. Perhaps it is worth looking into. My logic was, initially, that I don't need video, so that idea just went over my head completely. Thanks!
 
Upvote 0
I had some limited success with coqui xtts2 cloning locally and using samples of audio from solo porn vids, asmr audio pron, etc.

That's probably a bit dated these days though and something better might now be out but I've not really bothered looking in to it again. it was mostly out of curiosity.

Impressive results to say the least but In the end it was too much work involved creating a good enough dialogue
This one I've seen and it does seem like it is outperformed by Qwen3 and IndexTTS2, so I did not try. Maybe worth trying if I'm at my wit's end in the future.
 
Upvote 0
Back
Top Bottom