Clera logo

Senior Voice AI Engineer

Salary
$96K
USD per year
Hiring from
Worldwide
Work type
Remote
Posted
Oct 2, 2026
Is this job info correct?

About the Role

As a founding engineer on a small conversational AI team, you will own the real-time voice layer, from incoming speech through AI reasoning to spoken responses. You will help make natural, responsive voice interactions work reliably in production, with a focus on end-to-end latency.

What You'll Do

  • Build and own streaming speech-to-text, LLM turn-taking, text-to-speech, and telephony or WebRTC transport.

  • Measure and reduce latency, targeting first audio under 800 milliseconds on real calls.

  • Address interruptions, barge-in, silence detection, overlapping speech, poor audio, accents, and mid-sentence changes.

  • Build an evaluation harness from recorded calls, transcripts, and scored turns to detect regressions and guide product decisions.

  • Compare voice providers and models through evidence-based testing, and make changes based on results.

  • Instrument production systems for turn latency, transcription confidence, drop-offs, and cost per minute.

  • Work directly with founders and make technical decisions in a fast-moving team.

What We're Looking For

  • At least 5 years building production software, including 2 or more years shipping voice, speech, or real-time audio systems.

  • Experience building and shipping end-to-end real-time voice pipelines, including streaming speech recognition, LLM turn-taking, speech synthesis, and telephony or WebRTC.

  • Strong Python or TypeScript skills and comfort working in both.

  • Hands-on experience with an audio stack such as LiveKit, Pipecat, Vapi, Twilio Media Streams, Daily, or a custom WebSocket implementation.

  • Experience debugging audio at the frame level, including sample rates, codecs, jitter, and voice activity detection thresholds.

  • Experience building LLM evaluation harnesses, optimizing latency against real-world targets, and using evaluation results to make product decisions.

  • Clear written English for asynchronous communication. Experience with speech model serving or fine-tuning, SIP, telephony, or LLM orchestration frameworks is a plus.

Compensation & Benefits

Compensation is $96,000 USD annually, regardless of location. Visa sponsorship is not available.

Location

Fully remote, anywhere in the world. Core team overlap is 13:00 to 17:00 UTC.

Similar jobs

Apply for this job