Here is my plan:
1. Make a script contextually tied to the scene.
2. Generate AI text-to-speech audio file using that script.
3. Play that audiofile through atom head audio.
I have successfully created a scene like that already. It is up to my specifications, it is near perfect with emotion, delivery, inflection, exertion etc. The issue is I have made it using Gemini that only has a limited number of voices that can be used. It is also limited in amount of tokens and sends everything for analysis to Google. I am not too worried about the contents of my scripts or breach of privacy, but willingly sending shit like "fuck my this and lick my that, mmm so hot bla bla bla" feels a bit like standing naked in Google HQ and reading fanfiction for employees. SO, what I need is an offline model that can:
1. Read a script, preferably with a voice it cloned from reference. Especially useful if I will have an idea for a scene with look-a-like.
2. Modulate it to sound like sexual exertion. Gasps, inhale, exhales, the works.
3. Output it in a singular audio file or a series of audio files that can be later imported into VaM and used with head audio on an atom.
At no point I wish to have an LLM active whilst VaM is booted up and playing a scene. Voxta is something I'm willing to look at later, if I put my mind to different kind of experience in VaM, but from what I've seen it solves a different problem with a different approach, aside from mainly providing services I don't need.
Something like that, yes.