Senior Software Engineer (Voice Backend)
Workato · State of Telangāna, India
Workato · State of Telangāna, India
# **About Workato** Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise MCP and trusted by 50% of the Fortune 500, Workato's cloud-native architecture connects every application, data source, and process to power real-time orchestration at scale. With enterprise-grade security and continuous innovation at its core, Workato provides the trusted foundation for organizations to automate with confidence and operationalize AI across the business. To learn more, visit www.workato.com # **Why join us?** Ultimately, Workato believes in fostering a **flexible, trust-oriented culture that empowers everyone to take full ownership of their roles**. We are driven by **innovation** and looking for **team players** who want to actively build our company. But, we also believe in **balancing productivity with self-care**. That's why we offer all of our employees a vibrant and dynamic work environment along with a multitude of benefits they can enjoy inside and outside of their work lives. If this sounds right up your alley, please submit an application. We look forward to getting to know you! Also, feel free to check out why: - Business Insider named us an "enterprise startup to bet your career on" - Forbes' Cloud 100 recognized us as one of the top 100 private cloud companies in the world - Deloitte Tech Fast 500 ranked us as the 17th fastest growing tech company in the Bay Area, and 96th in North America - Quartz ranked us the #1 best company for remote workers # **Responsibilities** We are looking for an experienced and exceptional **Senior Software Engineer (Voice Backend)** to join our growing team. In this role, you will design, build, and scale the backend systems and real-time infrastructure that power our voice AI agents — the low-latency media pipelines, streaming orchestration, and high-concurrency services behind every live conversation. The ideal candidate pairs strong distributed-systems and backend engineering skills with hands-on experience running real-time audio and ML inference in production. If building voice infrastructure from zero to one that scales to serve millions of conversations excites you, you are in the right place. We are building organization-aligned voice agents that are consistent and reliable at scale, and we are looking for passionate builders who care deeply about latency, reliability, and clean systems design to reimagine the agent control and data plane. In this role, you will also be responsible to: - Design and build the real-time voice backend: low-latency audio streaming pipelines (WebRTC, SIP/telephony, RTP) and the orchestration layer that wires streaming ASR LLM TTS with turn-taking, endpointing, and barge-in under strict end-to-end latency budgets. - Own session and conversation-state management for thousands of concurrent live calls, including context propagation across turns and seamless, reliable handoff and escalation to human agents. - Build high-throughput, horizontally scalable services for serving and routing speech and language models in real time, optimizing for tail latency, concurrency, and cost. - Integrate and operate streaming ASR and TTS engines (and speech-to-speech models) as backend services — handling streaming protocols, audio codecs, voice activity detection, and graceful degradation. - Engineer reliability into the voice stack: failover, retries, backpressure, load shedding, and observability and distributed tracing focused on per-turn and end-to-end latency. - Design the APIs, event-driven / Pub-Sub interfaces, and data plane that connect the voice runtime to the broader Workato platform and downstream systems. - Lead technical initiatives to integrate the voice backend into existing products and services, and partner with ML engineers to productionize and serve new models. # **Requirements** ### **Qualifications / Experience / Technical Skills** - Bachelor's degree in Computer Science, or a related field, or equivalent practical experience. - 5+ years building backend services and distributed systems using modern programming languages (e.g., Python (strongly preferred!), Golang or Java). - Hands-on experience building real-time, low-latency systems — streaming media or audio (WebRTC, SIP, RTP), WebSocket/gRPC streaming, or comparable high-throughput event-driven services. - Experience integrating and operating streaming ASR and TTS (or speech-to-speech) engines as production services, including streaming protocols, audio codecs, voice activity detection / endpointing, and barge-in handling. - Experience serving ML/LLM models in production with a focus on tail latency, concurrency, and throughput (inference serving, batching, GPU utilization). - Working understanding of how LLMs, ASR, and TTS behave in production — their latency characteristics, streaming behavior, and failure modes — without needing to train models from scratch. - Experience with dialogue state tracking and context management for reliable multi-turn voice and chat interactions. - Deep knowledge of REST API design, Pub-Sub and event-driven architectures, and integration patterns. - Strong understanding of software architecture, scalability, security, and system design for high-availability services. - Experience with Docker, Kubernetes, and deploying to cloud environments (AWS, GCP, or Azure). - Experience working with PostgreSQL and ClickHouse, or similar relational and analytical databases (added advantage but not a necessary skill). - Experience with A/B testing and experimentation frameworks. - Familiarity with telephony/voice platforms (e.g., Twilio, LiveKit, Pipecat, Asterisk/FreeSWITCH) is a strong plus. ### **Soft Skills / Personal Characteristics** - Strong communication abilities to explain technical concepts - Collaborative mindset for cross-functional team work - Detail-oriented with a st