indigo.ai logo

Software Engineer, Voice - Milan

Hiring from
Italy
Work type
Remote
Posted
Oct 4, 2026
Is this job info correct?

If you are here, it is because you know that we are looking for a Software Engineer, Voice for our Product team.

indigo.ai is the leading platform in Italy for building next-generation AI Agents that transform the way companies communicate with their customers. Since 2016, we’ve been helping enterprises in industries such as finance, insurance, utilities, retail, and e-commerce to evolve their Customer Experience through conversational AI. We don’t just “sell software”: we enable a shift in how organizations interact with people, automating millions of conversations every year, reducing operational costs, improving conversion rates, and creating more personalized, scalable, and compliant customer journeys.

Backed by a recent €10 million investment from Azimut, we are on a mission to take this technology global. This is a unique opportunity to join a well-funded, highly ambitious team and play a direct role in shaping the future of enterprise AI.

To make this happen, we are looking for a Software Engineer, Voice to support our Chief Product Development Officer in making talking to AI on the phone feel human.

What are we looking for?

We are looking for a Software Engineer, Voice to own the real-time voice layer of our AI Agents. Voice is where conversational AI is being decided right now, and our voice agents already handle production phone traffic for large companies. The bar is moving fast: we want to make the leap from "works reliably" to "feels human on a real phone line", and we want one person to own that leap. This is a specialist role with end-to-end ownership: the architecture, the model and provider choices, the latency budget, the way a conversation feels. You'll join our Product Engineering team, reporting to our Chief Product Development Officer, as our first full-time engineer dedicated to voice. And if voice grows the way we believe it will, you'll shape the team that grows around it.

Key Responsibilities:

  • Own the real-time voice pipeline end-to-end. From audio ingress on the telephony edge, through streaming STT and turn-taking, to the agent brain and back out through streaming TTS. Every millisecond in between is yours.

  • Engineer how fast the agent feels. Semantic end-of-turn detection, preemptive generation on partial transcripts, eager TTS, filler and backchannel strategies that mask tool calls. All measured on real 8kHz phone audio, not in a browser demo.

  • Make turn-taking human. Barge-in that survives noisy lines. Endpointing policies that know the dialog state, so a caller never gets cut off mid-IBAN. The difference between an IVR and a conversation lives here.

  • Raise voice quality on the channel that actually ships: the phone. Benchmark and A/B STT and TTS providers on real G.711 calls (Italian first: WER, naturalness, numbers and codes read right), exploit wideband/HD voice where the carrier allows it, and experiment with context-aware TTS and conversational speech models as they mature.

  • Build the evaluation harness. Turn "this voice sounds better" into numbers we trust: per-stage latency budgets, turn-taking metrics, regression suites on recorded calls, quality gates before anything reaches a client.

  • Keep production boringly reliable. Per-stage observability, live-call incident debugging (dead air, stuck turns, provider hiccups), graceful degradation when a vendor blinks.

  • Track a weekly-moving ecosystem and turn it into strategy. New STT/TTS/speech-to-speech releases land every month. You decide what we integrate, what we self-host for EU compliance and data residency, and what we skip. And you make provider swaps cheap.

You will need:

The filter is not your degree, and it's not years-of-experience arithmetic. It's having built it. Tell us about a real-time voice or audio system you designed and shipped: the latency budget, where it broke, and what you changed to make it feel right. That tells us more than any title.

  • Real-time audio systems, shipped. You've built voice agents, telephony systems, conferencing or live-streaming products that ran in production. You know what it means to move audio over WebSockets/WebRTC/SIP, through codecs (G.711/μ-law, Opus), against a latency budget.

  • The modern voice AI stack, hands-on. Streaming STT and TTS, VAD and turn detection, voice orchestration frameworks (Pipecat, LiveKit Agents or equivalent), speech-to-speech models. You have opinions on the trade-offs, grounded in things you've actually built, not blog posts.

  • Strong software engineering. TypeScript/Node.js and/or Python, and the maturity to own a production service end-to-end: containers, cloud infrastructure, CI/CD, observability.

  • A latency obsession. You think in milliseconds per stage, you instrument before you optimize, and you know the difference between measured and perceived latency, and how to exploit it.

  • A product ear. You can hear the difference between a demo and a conversation, and you can translate what you hear into engineering priorities and measurable evals.

  • An AI-native way of working. You use agentic coding tools (e.g. Claude Code) daily and you're good at directing them: setting up the problem, judging the output.

  • Language Skills: Fluent English.

We will really like (but they are not required):

  • Italian: our voice market is Italian-first, and you'll be tuning pronunciation, prosody and evals for it every week.

  • Contact-center / CCaaS ecosystem experience: SIP trunking, SBCs, enterprise telephony platforms.

  • ML audio experience: evaluating or fine-tuning ASR/TTS models, working with speech datasets.

  • Elixir: our agent platform is built on it.

  • Open source: contributions to open-source voice/audio projects.

Our values:

At indigo.ai we prize Vision (curiosity and courage to challenge the status quo), Connection (empathy, candor, and trust), Responsibility (ownership, reliability, and follow-through), and Excellence (the habit of raising the bar and refining until it’s right). If you naturally think ahead, build strong relationships, take accountability, and obsess over the quality of what you deliver, we’re probably a great match.

What do we offer?

  • A key role in one of Europe’s fastest-growing AI scale-ups, backed by a €10M investment from Azimut.

  • A competitive salary in the range of 40-70k RAL, commensurate with experience + a performance-based Bonus.

  • A flexible, remote-friendly work environment.

  • Meal vouchers and Welfare programs to support your everyday life.

  • Access to a dedicated education budget for continued learning and growth.

  • Career Development Plan, ensuring a clear path for both personal and professional growth.

  • Top-grade equipment, which may include a MacBook Air, iPhone, and other top-tier devices.

  • Unlimited coffee.

  • Company retreats in stunning locations throughout the year.

Where is the job?

This position is fully remote, so there's no need to be in a specific location to do your work. That said, we have an amazing office at SPACES, Piazza Gae Aulenti 1/Torre B in Milan, available to anyone who wants to use it. We also love getting together and organize various retreats and meetups throughout the year to stay connected.

Why join us?

At indigo.ai you’ll be part of a fast-growing SaaS company where your ideas and work can turn into products used by millions. Join us and you’ll:

  • Work with the most advanced AI technologies, shaping how enterprises across industries engage with their customers.

  • Be part of a dynamic and passionate team, where everyone has a direct impact on company growth.

  • Grow in a culture that values transparency, continuous learning, and career development, with clear paths for personal and professional progression.

  • Enjoy a flexible, people-first environment, with remote-friendly policies, stunning retreats, and a strong focus on well-being.

  • Contribute to our mission of reshaping how companies and people communicate worldwide.

Similar jobs

Apply for this job