Senior Voice AI Engineer
Clera · remote
- Location
- remote
- Funding
- $3M
- Posted
- Oct 2, 2026
Clera is hiring a Senior Voice AI Engineer based in remote. Every apply link on Engg.space goes straight to the company's own careers page - no recruiter middleman, no generic job-board form.
Apply directly at CleraRole details
ABOUT THE ROLE As a founding engineer on a small conversational AI team, you will own the real-time voice layer, from incoming speech through AI reasoning to spoken responses. You will help make natural, responsive voice interactions work reliably in production, with a focus on end-to-end latency. WHAT YOU'LL DO - Build and own streaming speech-to-text, LLM turn-taking, text-to-speech, and telephony or WebRTC transport. - Measure and reduce latency, targeting first audio under 800 milliseconds on real calls. - Address interruptions, barge-in, silence detection, overlapping speech, poor audio, accents, and mid-sentence changes. - Build an evaluation harness from recorded calls, transcripts, and scored turns to detect regressions and guide product decisions. - Compare voice providers and models through evidence-based testing, and make changes based on results. - Instrument production systems for turn latency, transcription confidence, drop-offs, and cost per minute. - Work directly with founders and make technical decisions in a fast-moving team. WHAT WE'RE LOOKING FOR - At least 5 years building production software, including 2 or more years shipping voice, speech, or real-time audio systems. - Experience building and shipping end-to-end real-time voice pipelines, including streaming speech recognition, LLM turn-taking, speech synthesis, and telephony or WebRTC. - Strong Python or TypeScript skills and comfort working in both. - Hands-on experience with an audio stack such as LiveKit, Pipecat, Vapi, Twilio Media Streams, Daily, or a custom WebSocket implementation. - Experience debugging audio at the frame level, including sample rates, codecs, jitter, and voice activity detection thresholds. - Experience building LLM evaluation harnesses, optimizing latency against real-world targets, and using evaluation results to make product decisions. - Clear written English for asynchronous communication. Experience with speech model serving or fine-tuning, SIP, telephony, or LLM orchestration frameworks is a plus. COMPENSATION & BENEFITS Compensation is $96,000 USD annually, regardless of location. Visa sponsorship is not available. LOCATION Fully remote, anywhere in the world. Core team overlap is 13:00 to 17:00 UTC.
More roles at Clera
- Software Engineer (Voice AI / Backend)6 days agoBerlin$3M
- Member of Technical Staff6 days agoUnited States$180,000 to $250,000 USD$3M
- Member of Technical Staff25-09-2026San Francisco$150,000 to $250,000 USD$3M
- Founding Engineer25-09-2026Berlin$3M
- Senior/Staff Software Engineer (Platform)25-09-2026San Francisco$160,000 to $250,000$3M
- Founding Engineer25-09-2026remote$3M
- Full Stack Engineer25-09-2026San Mateo$3M
- Founding Engineer25-09-2026San Francisco$150,000 to $250,000 USD$3M
- Senior Software Engineer, Infrastructure25-09-2026Singapore$150,000 to $250,000 USD$3M
- Software Engineer25-09-2026San Francisco$150,000 to $300,000 USD$3M
- Backend Engineer25-09-2026Berlin$3M
- Full Stack Engineer25-09-2026Munich$3M