About us
Visionary Minds builds conversational voice assistants that carry a specific person’s knowledge, judgment, and way of talking. Instead of a generic chatbot, each assistant is trained to sound and reason like a particular expert — so someone can talk through a hard decision, get grounded advice, or learn from a domain leader as if that person were on the other end of the line. We’ve built assistants across a few different areas — personal wellbeing and recovery, organizational leadership, and executive decision-making among them — and we’re onboarding new clients who each want their own.
Every new client means standing up a current assistant from their material, end to end: the language model and the voice. That’s the work.
What you’d be doing
You’d own the model side of getting a new client’s assistant from “we have their training data” to “it’s live, it sounds like them, and you can talk to it.” In practice:
- Building the fine-tuning pipeline for each persona — turning transcripts, writing, talks, and interviews into a model that answers the way that person would. We use a mix (Llama, Mistral, and API models depending on the client), LoRA/QLoRA where it fits and full fine-tunes where it earns it, with RAG when the knowledge moves faster than we want to retrain.
- Training and adapting the voice — the TTS that makes each assistant sound like the actual person, and the ASR that lets it understand whoever’s talking to it. This is real speech work: voice adaptation from limited audio, streaming synthesis,
and keeping it intelligible across accents and phone-quality audio.
- Owning the real-time loop. These get spoken to, so you’ll be living in time-to-first-byte, streaming ASR, and turn-taking — endpointing and barge-in that don’t cut people off — against a sub-second budget end to end.
- Writing the eval that says an assistant is ready — not just word-error-rate and benchmark numbers, but “does this actually sound like the person, does it stay in its lane, does it handle a hard conversation well.” You’ll help us build the tooling around that judgment.
- Running the rollout. You’ll set and hold the timeline for each new client — what happens, in what order, when it goes live — and flag early when a date’s going to slip.
- Improving the ones already in production, off the back of real conversations: what to retrain, re-prompt, re-record, or rethink.
What we’re looking for
- 2–3+ years building ML systems that shipped and stayed shipped. You’ve owned a model in production and dealt with what happens after.
- Real LLM fine-tuning depth — you can talk through LoRA vs. a full fine-tune vs. better retrieval, and why.
- Genuine speech experience:
you’ve trained or fine-tuned TTS or ASR models and you’re comfortable with the vocabulary of the field — neural codecs, endpointing, streaming inference, the usual latency traps.
- Strong Python and the training stack (PyTorch, Hugging Face). Comfort near the metal for real-time serving — CUDA, streaming, quantization — is a real plus.
- You can run your own timeline. This is a contract role on a small team; nobody’s building your Gantt chart for you.
- You write and explain clearly, including to the clients themselves, who aren’t engineers.
Nice to have, not required experience
- Voice cloning or speaker adaptation from small amounts of audio.
- Making a model sound like a specific person — persona conditioning, character consistency across long conversations.
- Real-time serving optimization: KV-cache management, speculative decoding, streaming codec encode/decode.
- RLHF or preference tuning.
- You’ve delivered for clients directly, not just handed work over the wall.
The logistics This is an independent contract engagement, fully remote and based in India, working with a US team.
Pay is USD$3,500/month.
Because the team runs on US hours, we need a few hours of daily overlap — realistically your evening, to catch US mornings.
To apply
Send a note to
[email protected] with something you’ve actually built — a repo, a model, a writeup of a fine-tune or a voice you’re proud of. Skip the cover letter. Tell us the hardest thing you’ve gotten a model to do and how you knew it worked.
📌 Machine Learning Engineer (Visionary Minds Company) (India)
🏢 Titan Holdings
📍 India