Site Reliability Engineer
Autodesk · APAC - India - Pune - ABIL Boulevard
Free to search · AI fit score against your CV · tailor your résumé in one click
Autodesk · APAC - India - Pune - ABIL Boulevard
Job Requisition ID # 26WD98218 Position Overview An exciting opportunity is available for a Site Reliability Engineer to join Autodesk's Product Design and Manufacturing Solutions (PDMS) Platform Site Reliability Engineering (SRE) team. In this role, you will wear multiple hats, including first responder, performance analyst, system architect, capacity planner, and monitoring expert. You will bring strong technical and communication skills, a passion for learning new technologies, and a problem-solving mindset. You will help build and operate reliable, scalable, secure, and high-performing cloud infrastructure that supports Autodesk products and customers. Responsibilities • Architect and implement hosting solutions for highly dynamic Software as a Service (SaaS) web applications, ensuring reliability, scalability, and performance • Design, implement, and maintain Infrastructure as Code (IaC) solutions to support scalable, reliable, and secure global environments • Develop and maintain well-documented engineering standards, processes, and best practices • Implement infrastructure and application security best practices, including system hardening and the principle of least privilege • Use modern infrastructure management tools such as Docker, Terraform, Amazon Web Services (AWS) CloudFormation, and AWS Cloud Development Kit (CDK) to manage and deploy containers and virtual machines • Collaborate with Development, Quality Assurance, and Documentation teams throughout the product development lifecycle to ensure quality and reliability • Automate operational processes and integrate new technologies to improve efficiency, reliability, and scalability • Define and monitor Service Level Objectives (SLOs) and Service Level Indicators (SLIs) and manage error budgets to ensure reliability goals are achieved • Partner with stakeholders to align technical strategies with business requirements • Participate in on-call support and incident management, ensuring timely resolution and clear stakeholder communication • Conduct blameless post-incident reviews to identify root causes, document learnings, and drive continuous improvement • Take ownership of initiatives and contribute to a culture of continuous learning, operational excellence, and continuous improvement Minimum Qualifications • 5+ years of experience in DevOps, Site Reliability Engineering (SRE), or a related role supporting cloud-based applications • Bachelor's degree in Computer Science or a related technical field • Advanced hands-on experience with Linux administration, including monitoring, troubleshooting, reliability, performance, and security • Experience managing large-scale cloud infrastructure, preferably on Amazon Web Services (AWS) • Strong scripting skills using languages such as Bash, Python, or Perl • Expert-level knowledge of AWS services, including Amazon Elastic Compute Cloud (EC2), Elastic Container Service (ECS), Elastic Kubernetes Service (EKS), AWS Lambda, Elastic Load Balancing (ELB), Amazon Simple Storage Service (S3), Identity and Access Management (IAM), Virtual Private Cloud (VPC), Amazon DynamoDB, and Amazon Relational Database Service (RDS) • Hands-on experience with Docker, Kubernetes, and container technologies • Proficiency with Infrastructure as Code (IaC) tools such as Terraform and AWS CloudFormation • Experience with Continuous Integration and Continuous Deployment (CI/CD) tools and technologies such as Jenkins, JFrog Artifactory, and Git • Experience with logging, monitoring, and observability tools such as Amazon CloudWatch, Splunk, Dynatrace, New Relic, and Grafana • Experience with relational database technologies such as MySQL, PostgreSQL, and Microsoft SQL Server, along with Structured Query Language (SQL) • Excellent analytical and problem-solving skills with the ability to work independently • Excellent written and verbal communication skills Preferred Qualifications • Experience using Artificial Intelligence (AI)-assisted engineering tools and development practices • Experience designing and operating highly available, distributed cloud-native systems • Experience with Site Reliability Engineering practices, including observability, capacity planning, incident management, and error budget management • Experience automating infrastructure and operational processes at scale • Knowledge of cloud security, infrastructure hardening, and compliance best practices • Experience working in globally distributed engineering teams The Ideal Candidate The ideal candidate is a technically strong and highly motivated Site Reliability Engineer who combines cloud infrastructure expertise, automation skills, and a reliability-first mindset to build and operate secure, scalable, and resilient systems • Demonstrates strong technical expertise in cloud infrastructure, Linux admi