Senior Software Engineer, SRE
Roku · Bengaluru, India
Free to search · AI fit score against your CV · tailor your résumé in one click
Roku · Bengaluru, India
Teamwork makes the stream work. Roku is changing how the world watches TV Roku is the #1 TV streaming platform in the U.S., Canada, and Mexico, and we've set our sights on powering every television in the world. Roku pioneered streaming to the TV. Our mission is to be the TV streaming platform that connects the entire TV ecosystem. We connect consumers to the content they love, enable content publishers to build and monetize large audiences, and provide advertisers unique capabilities to engage consumers. From your first day at Roku, you'll make a valuable - and valued - contribution. We're a fast-growing public company where no one is a bystander. We offer you the opportunity to delight millions of TV streamers around the world while gaining meaningful experience across a variety of disciplines. What does the team work on? The Platform Infrastructure team ensures that all Roku systems run smoothly. These systems support over 100M+ users and billions in transaction revenue per year. We are a group of highly skilled infrastructure and software engineers who help build and operate systems at internet scale, including Platform (Kubernetes, Istio, Envoy, operators, and more) and Observability (OSS/CNCF-supported observability projects). We engage with multiple teams to achieve company-impacting results. What is the role? We are seeking a talented and experienced SRE (Site Reliability Engineering) Senior Software Engineer to help architect, build, and operate large-scale systems that stay reliable, secure, and cost-effective at internet scale. The ideal candidate takes end-to-end ownership of outcomes—treating reliability, security, cost, operability, and supportability as part of the job, not just delivering code. They bring calm, decisive incident leadership, separating mitigation from root-cause investigation and running blameless reviews that produce lasting improvements. Strong judgment and prioritization are essential, balancing roadmap delivery against operational debt, security, and compliance while clearly explaining trade-offs. This engineer pairs deep technical depth in distributed systems with broad systems thinking, and turns ambiguous objectives into executable roadmaps, epics, and backlogs. Just as important is the ability to build influence through credibility and sound reasoning, coach other engineers, and raise the operational capability of the whole team. If you enjoy solving intriguing system challenges, are innovative at heart, and thrive on making a measurable impact across teams, this role might be a great fit for you. How will I use AI at Roku? At Roku, we don't just use AI, we work with it. AI agents and smart tools help power drafts, analysis, and repetitive workflows, while our people bring direction, judgment, and accountability. We value curious, adaptable builders, who can show how they've used AI, agents, or automation to move faster, improve quality, and scale their impact. Strong candidates know how to frame problems, guide AI-assisted work, check the output, and learn quickly. Above all, they bring curiosity, adaptability, and sound judgment. What are the responsibilities of the role? Ownership & Incident Leadership • p]:inline" data-streamdown="list-item">Take responsibility for service outcomes end to end, including reliability, security, cost, operability, and supportability • p]:inline" data-streamdown="list-item">Lead major incidents with composure when information is incomplete, separating mitigation from root-cause investigation and communicating impact, status, risks, and next steps without speculation • p]:inline" data-streamdown="list-item">Facilitate comprehensive, blameless post-incident reviews that identify root causes and contributing factors, and follow corrective actions through to completion • p]:inline" data-streamdown="list-item">Track incident trends to surface systemic issues and prioritize reliability improvements • p]:inline" data-streamdown="list-item">Implement chaos engineering, game days, and disaster recovery exerc