Senior Applied AI Engineer, Speech and Voice
- Salary
- $65K–$90KUSD
- Hiring from
- Singapore
- Work type
- Remote
- Posted
- Sep 29, 2026
Location: Remote. Compensation: USD 65,000 to 90,000 base, plus equity.
About Talkpush
Talkpush is the AI-powered applicant tracking system and recruitment CRM built for high-volume hiring. Founded in Hong Kong about ten years ago, we help employers in retail, customer experience, healthcare, transportation and finance attract, screen, interview and hire thousands of candidates at once, over WhatsApp, SMS, social channels, chat and voice. Millions of people have been hired through the platform, mostly across Asia-Pacific, Latin America and other emerging markets.
Mission
Give every candidate a fair hearing by voice, in their own accent and language, at the scale of high-volume hiring.
You will build the speech and voice AI that runs across Talkpush: our own AI interviewer, real-time conversation analysis, candidate scoring, and support for new languages. You will be a senior speech specialist on the team, deciding with product what to build next and shipping it. You will get context, data and support, and little technical hand-holding, by design.
The problem and the data
Our candidates speak on real phone lines and in browsers, often on 8 kHz audio, in noisy rooms, answering unscripted questions. Many are Filipino, Indian, Latin American and South African English speakers, and many switch between languages mid-sentence (Tagalog and English, for example). Off-the-shelf speech models were mostly trained on clean, scripted audio and a narrow set of reference accents, and they struggle here.
We run our own voice stack and keep our own recordings. That gives us a large, rare dataset: tens of thousands of labelled audio clips from real interviews, many linked to recruiter ratings and hiring outcomes.
First 90 days
- Listen to a lot of real calls. Map the recordings, ratings and hiring outcomes, and check how reliable the human labels are.
- Build a gold evaluation set, held out and stratified by market, accent and score band, joined to hiring outcomes. This becomes the bar every model must clear.
- Fine-tune an open speech encoder (Whisper, WavLM or wav2vec2) into a first speech scorer for fluency, pronunciation and intelligibility. Run it silently alongside live traffic and publish honest results, including where it fails.
- Write a dated plan for the year and share it.
First year
- Add a transcript and content branch and fuse it with the acoustic model.
- Calibrate scores to real hiring outcomes, with light adapters for different clients and roles.
- Distil the model into a fast version for real-time analysis during and right after each call, so the AI interviewer can adapt mid-conversation.
- Extend to new languages (candidates include Spanish, Brazilian Portuguese, French, Turkish and Urdu) and to code-switched speech.
- Publish an accent-robustness benchmark showing that matched-proficiency speakers get matched scores across accent groups.
Responsibilities
- Own speech and voice models end to end: data, training, evaluation, deployment and monitoring.
- Shape the voice AI roadmap with product, engineering and client-facing colleagues. Write the specs, then ship them.
- Keep measurement honest: held-out test sets, per-accent error analysis, test-retest checks, no vanity metrics.
- Make fairness across accents a design requirement from day one.
- Work in production with our engineers on latency, cost and reliability.
- Own outcomes end to end, commit to dates, and write short, clear updates.
Must-haves
- You have shipped speech ML models to production that real users depend on.
- You have fine-tuned Whisper, WavLM, wav2vec2 or similar encoders on your own data and can explain what worked and what did not.
- You have worked with messy real-world audio: telephony, 8 kHz, noise, spontaneous speech.
- You have built pronunciation, fluency, speaker or paralinguistic models, or ASR for accented or low-resource speech.
- You have designed evaluation sets and reported results you did not like, and changed course because of them.
- You have owned a product area or a technical roadmap, including telling people what you would not build.
Nice-to-haves
- Experience with code-switched or multilingual speech.
- Work on fairness, bias audits or calibration of scores against human raters or downstream outcomes.
- Streaming or real-time inference, and model distillation.
- Forced alignment tools (Montreal Forced Aligner, WhisperX) and public L2 corpora (speechocean762, L2-ARCTIC).
- Using audio-LLMs as weak labellers, or TTS to create synthetic training data.
- Papers at Interspeech or ICASSP, or well-used open-source speech work.
- Experience in HR tech, contact centres or language testing.
Interview process
Our process: a 10-minute voice pre-screen with our AI interviewer, a take-home on real audio (about a day of work across a week), a case study with our CEO, a technical deep dive with a peer engineer, then references. Every hire needs a unanimous yes.
Compensation and location
- Location: Remote.
- Compensation: USD 65,000 to 90,000 base, plus equity.