Job description

We're Nagarro, we are a Digital Product Engineering company that is scaling in a big way! We build products, services, and experiences that inspire, excite, and delight. We work at scale across all devices and digital mediums, and our people exist everywhere in the world (17500+ experts across 39 countries, to be exact). Our work culture is dynamic and non-hierarchical. We are looking for great new colleagues. That is where you come in! **REQUIREMENT:** - We are hiring for Devops Architected and managed scalable, highly available, secure, and cost-optimized cloud environments on Google Cloud Platform (GCP) while supporting multi-cloud deployments across AWS and Azure. - Designed and implemented cloud-native infrastructure solutions to support enterprise applications and digital transformation initiatives. - Built, automated, and optimized CI/CD pipelines, reducing deployment time and improving software delivery efficiency. - Developed Infrastructure as Code (IaC) frameworks using automation tools to enable repeatable and consistent cloud deployments. - Managed enterprise Kubernetes platforms including GKE, with operational exposure to AKS and EKS, ensuring governance, security, and performance optimization. - Configured and maintained Kubernetes networking, RBAC policies, cluster security controls, autoscaling, and workload management. - Led complex infrastructure troubleshooting efforts involving compute, networking, storage, platform services, and security layers. - Implemented DevOps best practices across development and operations teams to improve release quality and operational stability. - Designed and deployed observability solutions using Prometheus, Grafana, and ELK Stack to enhance monitoring, alerting, and incident response capabilities. - Established centralized logging, monitoring, and operational dashboards to improve platform visibility and reliability. - Designed and managed GPU-enabled infrastructure supporting AI/ML workloads, model training, and inference platforms. - Built and maintained centralized Model Control Plane (MCP) and model-serving environments for scalable AI platform operations. - Collaborated closely with engineering, product, and security teams to ensure reliable, compliant, and secure cloud service delivery. - Authored architecture designs, technical documentation, SOPs, operational runbooks, and implementation guidelines. - Mentored DevOps and cloud engineering teams on modern infrastructure, cloud architecture, automation, and reliability engineering best practices. **RESPONSIBILITIES:** - Writing and reviewing great quality code - Understanding functional requirements thoroughly and analysing the clients needs in the context of the project - Envisioning the overall solution for defined functional and non-functional requirements, and being able to define technologies, patterns and frameworks to realize it - Determining and implementing design methodologies and tool sets - Enabling application development by coordinating requirements, schedules, and activities. - Being able to lead/support UAT and production roll outs - Creating, understanding and validating WBS and estimated effort for given module/task, and being able to justify it - Addressing issues promptly, responding positively to setbacks and challenges with a mindset of continuous improvement - Giving constructive feedback to the team members and setting clear expectations. - Helping the team in troubleshooting and resolving of complex bugs - Coming up with solutions to any issue that is raised during code/design review and being able to justify the decision taken - Carrying out POCs to make sure that suggested design/technologies meet the requirements