Expertshub.ai · Remote
Role Overview Avashya is an AI-first, platform-led services company founded by engineers and specialists with experience at AWS and Microsoft. This Voice AI Engineer role owns the real-time voice platform and every client implementation built on it—closing the gap between a demo and a production-grade voice AI product. Responsibilities • Extend the real-time voice framework; when configuration reaches its limits, work directly with the runtime. • Turn lessons from each client build into reusable platform capabilities for the next. • Take client use cases—such as outbound negotiation, multilingual support, and appointment booking—from scoping through go-live. • Tune ASR, LLM, TTS, RAG, tool-calling, and orchestration to the client, languages, and call volume rather than relying on a default stack. • Remain the technical owner after launch and diagnose production behaviour quickly. • Set latency budgets for every stage of a conversation and maintain them as load and providers change. • Gate releases on measured evaluations; investigate degraded call quality before clients need to report it. • Make client-specific tradeoffs between speed and polish; review code and mentor as the team grows. Required Depth Architecture and real-time voice systems: • Deep, hands-on knowledge of Pipecat, LiveKit Agents, or comparable orchestration frameworks, including their internals and limits. • Ability to select cascaded ASR/LLM/TTS pipelines or speech-to-speech approaches based on control, swapability, latency, and prosody needs. • Strong turn-taking expertise: VAD tuning, endpointing, barge-in, and natural handling of interruptions. • WebRTC and SIP telephony experience, including jitter buffers, PSTN versus browser tradeoffs, and Opus, PCM, and mu-law codecs. Latency, retrieval, and tool use: • Ability to decompose end-to-end latency across ASR partials, LLM time-to-first-token, TTS time-to-first-audio-byte, network hops, and tool-calling round trips. • Experience streaming ASR, LLM, and TTS together, using approaches such as fillers, early commits, or speculative synthesis when tool calls are slow. • Ability to use RAG and live tool calls while preserving a natural, live conversation. Evaluation and guardrails: • Voice-specific evaluation of ASR accuracy (WER), TTS naturalness, and end-to-end conversation quality. • Track P50, P95, and P99 latency by stage and provider over time. • Implement content safety, hallucination control, PII handling for transcripts and recordings, consent, and data-residency practices for recorded calls. Experience Depth matters more than years of experience. This role calls for senior-level ownership of production voice AI systems. Compensation Top-of-market; pay reflects depth, not tenure.