Z

Cloud DevOps -Operations Support-Puppet

Zensar Technologies · State of Karnataka, India

3–9 yrs experiencePosted Yesterday
Apply now →

Job description

**Job Description** **Site Reliability Engineer (SRE – SaaS Platform Operations)** **Position Summary:** Experienced Azure-based SRE required to support a large-scale enterprise SaaS platform, with strong Microsoft SQL Server expertise, future PostgreSQL readiness, production operations, patching, automation, troubleshooting, and incident response capability. **Role Type** Site Reliability Engineering / SaaS Platform Operations **Primary Technology Focus** Microsoft Azure, Microsoft SQL Server, PostgreSQL, Windows/Linux, Monitoring, Automation **Future Scope** Upcoming releases include database movement from MSSQL to PostgreSQL; PostgreSQL operational support is expected to become increasingly important. **Experience Level** 4+ years relevant experience; strong MSSQL production support background; PostgreSQL exposure preferred **Working Model** 24x7 production support environment; US time zone support/night shifts as required **Audience** Client, delivery leadership, hiring panel, and internal stakeholders Position Overview We are seeking an experienced Site Reliability Engineer (SRE) to support , a large-scale enterprise SaaS platform operating in cloud and high-availability environments. This role is responsible for maintaining infrastructure reliability, availability, performance, and operational excellence across Microsoft Azure, Microsoft SQL Server, PostgreSQL readiness, Windows, Linux, monitoring, automation, and incident response functions. This is not a first-line helpdesk role. The role requires a hands-on engineer who can independently investigate complex technical issues, collaborate with engineering and platform teams, and provide clear technical communication to internal and client stakeholders. **Future Scope:** In upcoming product releases, selected databases are expected to move from Microsoft SQL Server to PostgreSQL. The role should therefore include PostgreSQL awareness and operational readiness in addition to current MSSQL responsibilities. Role Focus Areas **Focus Area** **Expected Capability** **Business Outcome** **Cloud Operations** Manage and support Azure infrastructure, monitoring, storage, compute, and patching activities. Stable, secure, and scalable platform operations. **Database Reliability** Administer MSSQL workloads today and support PostgreSQL readiness for future releases, including tuning, maintenance, backups, and HA/DR. Improved database performance, resiliency, and recoverability. **Patching & Maintenance** Plan and support SQL/database patching and operating system patching across Windows and Linux environments. Improved security compliance and reduced operational risk. **Incident Response** Perform initial analysis, support P1/P2 triage, contribute to RCA, and improve runbooks. Reduced MTTR and stronger production readiness. **Automation & Observability** Use PowerShell/Python, alert tuning, logging, dashboards, and Infrastructure-as-Code practices. Better operational efficiency and proactive issue detection. **Engineering Collaboration** Reproduce issues, validate defects, and escalate with evidence to product engineering. Faster defect resolution and improved customer experience. Key Responsibilities Infrastructure & Platform Reliability - Maintain highly available, reliable, and scalable cloud infrastructure in Microsoft Azure. - Monitor platform health, review technical logs, and proactively address performance and availability issues. - Improve infrastructure monitoring, alerting, and logging to support proactive reliability management. - Plan, coordinate, and support operating system patching and maintenance activities across Windows and Linux servers. - Ensure security updates, compliance patches, and platform upgrades are executed in line with change management processes while minimizing service disruption. - Drive operational excellence through standardization, automation, and continuous improvement. Database Administration & Performance Management - Manage and support Microsoft SQL Server databases across production and non-production environments on Azure. - Support PostgreSQL operational readiness and future PostgreSQL database support as selected databases move from MSSQL to PostgreSQL in upcoming releases. - Troubleshoot and tune queries, SQL jobs, indexing, CPU, memory, I/O, and storage utilization. - Implement and maintain backup, restore, high availability, disaster recovery, and routine database maintenance processes. - Perform SQL Server and PostgreSQL database patching, upgrades, maintenance, and version lifecycle activities under enterprise change management processes. - Use T-SQL and scripts for investigations, reporting, and controlled production data fixes under change control. Incident Management & Technical Troubleshooting - Investigate complex infrastructure, application, database, Windows, Linux, and Azure-related issues. - Participate in incident response, perform initial technical analysis, and contribute to root cause analysis documentation. - Reproduce customer-reported issues in staging or lab environments where required. - Escalate verified product defects to engineering teams with clear technical evidence, logs, and impact analysis. Automation, DevOps & Observability - Create and maintain operational automation using PowerShell, Python, or equivalent scripting languages. - Support Infrastructure-as-Code and configuration management practices using tools such as Terraform and Puppet. - Collaborate on CI/CD and operational tooling improvements using platforms such as GitHub and Jenkins. - Enhance monitoring and alert management practices for SaaS product operations. Stakeholder & Customer Communication - Collaborate closely with onsite DB SREs, application teams, platform teams, product support, and engineering. - Provide clear and timely communication on issue status, technical findings, risks, and next steps.