C

Sr.Specialist, Software Development & Engineering - Core Enterprise Services

Charles Schwab · State of Telangāna, India

8–16 yrs experiencePosted 3w ago
Apply now →

Job description

Hyderabad, Telangana **Requisition ID** 2026-123607 **Category** Engineering & Software Development **Position type** Regular ## **Your opportunity** At Charles Schwab, our purpose is simple: we championclients’goals with passion and integrity. Guided by honesty, mutualrespectand a commitment to doingwhat’sright, we bring innovation, education, and service together to help shape financial futures. Our people are the foundation of our success – they approach their work with curiosity and collaboration, coming together to create solutions that make a meaningful impact for clients and communities. As we expand into India, we are bringing this same culture of inclusion, learning, and opportunity to new talent. Joining us means becoming part of a global team where your work matters and your future can take shape. Our Hyderabad location is central to Schwab’s growth, bringing together talented people and technology to drive innovation,scale,and efficiency. Here, you will work alongside teams who create solutions that support millions of clients every day. The work you do is more than daily operations –it’sa chance to experiment, learn, and build within avalue–driven, supportive environment. This is a unique opportunity to be part of our early growth phase and shape something new, backed by the stability and strength of a Fortune 500 company. Your impact begins on day one, and your contributions will help define our future in the region Weareseeking an AI Site Reliability Engineer to join a forward-thinking engineering team that builds intelligent observability, monitoring, and deployment automation solutions using AI-augmented development practices. This is not traditional production support. You will build software solutions for operational challenges,leveragingGen AI to accelerate workflows, automate incident response, and drive systems toward five 9s (99.999%) availability while ensuring capacity planning, redundancy, and the highest standards of reliability, security, and scalability. This is a hands-on engineering role where you will actively contribute to architecture, code, AI-driven tooling, and modern DevOps practices. Ideal for engineers who are curious, adaptable, and excited about working at the intersection of software engineering, AI, and cloud-native DevOps. KeyResponsibilities: High Availability & Resilience - Design and implement architectures that achieve and sustain 99.999% uptime across critical systems - Define, measure, and track SLOs, SLIs, and error budgets - Build AI-powered self-healing systems with automated failover, redundancy, and graceful degradation - Perform AI-assisted capacity planning, demand forecasting, and load testing - Conduct chaos engineering practices enhanced with AI-driven failure prediction Observability & Monitoring - Design, build, andmaintainAI-enhanced observability platforms covering metrics, logs, traces, and intelligent alerting - Implement AI-powered anomaly detection, predictive alerting, and proactive system health management - Leverage Gen AI to auto-generate and refine dashboards, alert rules, and runbooks - Build real-time availability dashboards with AI-driven trend analysis tracking 99.999% targets Root Cause Analysis - Build AI-accelerated root cause analysis processes with thorough postmortems and actionable remediation - Develop AI-powered diagnostic tools that automatically correlate logs, metrics, and traces - Use Gen AI to analyze incident patterns, predict recurring failures, and recommend preventive actions - Build AI agents that automate initial triage and preliminary RCA for common incident types - Continuously reduce MTTD and MTTR through AI-assisted workflows Deployment Automation & DevOps (GitHub-Centric) - Design andmaintainAI-enhanced CI/CD pipelines using GitHub and GitHub Actions as primary DevOps platforms - Implement workflows for build, test, deploy, and release automation using GitHub Actions - Enable progressive delivery with canary releases, blue-green deployments, and automated rollbacks triggered by AI-driven anomaly detection - Leverage Gen AI to generate andoptimize: GitHub Actions workflows oDeployment scripts oInfrastructure-as-Code (IaC) templates - Establish branching strategies, PR governance, and code review standards to ensure quality and reliability - Integrate GitHub ecosystem capabilitiesincluding: oGitHub Advanced Security (code scanning, secret scanning) oGitHub Packages/Artifacts for build distribution - Build AI agents to automate repetitive DevOps and release management tasks - Drive end-to-end DevOps observability across pipelines, deployments, and production systems AI-Driven Development - Use GenAI tools (GitHub Copilot) for coding, debugging, documentation, and operations - Developer ownership is non-negotiable all code, whether human or AI-generated, must be reviewed, tested, and understood before merging - Design and develop AI agents for incident response, log analysis, capacity management, and operational automation Define agent goals, tool use, memory, and orchestration logic for multi-step SRE workflows - spec-driven development and continuously evaluate emerging Gen AI tooling Testing & Quality - Leverage Gen AI to generate tests,identifycoverage gaps, and create edge-case scenarios - Drive AI-powered security scanning, performance testing, and reliability validation early in development - Embed quality gates within GitHub Actions pipelines to enforce standards Modernization & Collaboration - Modernize existing monitoring, alerting, and deployment systems using AI-assisted workflows - Identifytechnical debt and propose AI-accelerated remediation strategies with measurable outcomes - Participate in architecture discussions, design reviews, and code reviews - Develop andmaintainprompt engineering guidelines and AI usage standards - Share learnings, patterns, and tooling insights across the organization ## **What you have** **Required Qua