N

Staff Engineer (Datadog Engineer)

Nagarro · Hyderabad, Telangana, India

8–15 yrs experiencefull_timePosted 2 days ago
Apply now →

Job description

**👋🏼We're Nagarro.** We are a Digital Product Engineering company that is scaling in a big way! We build products, services, and experiences that inspire, excite, and delight. We work at a scale — across all devices and digital mediums, and our people exist everywhere in the world (18500+ experts across 40 countries, to be exact). Our work culture is dynamic and non-hierarchical. We are looking for great new colleagues. That is where you come in! **Requirements** - Experience : 5.5+ years - Strong experience in enterprise monitoring, observability, or cloud infrastructure engineering. - Strong hands-on experience with **Datadog** across Infrastructure Monitoring, APM, Synthetic Monitoring, Database Monitoring, Real User Monitoring (RUM), and Log Management. - Expertise in **Microsoft Azure DevOps Pipeline Strategy**, including designing and implementing CI/CD pipelines for monitoring and observability solutions. - Strong experience with **Terraform** for Infrastructure as Code (IaC), automating Datadog configuration and deployment. - Strong knowledge of **ITIL** processes, including Incident, Problem, Change, and Service Management. - Hands-on experience in incident management, root cause analysis, troubleshooting, and production support. - Experience creating and maintaining Datadog dashboards, monitors, alerts, log pipelines, and service-level monitoring. - Strong understanding of cloud platforms such as AWS, Azure, or Google Cloud Platform. - Experience with automation tools and scripting to streamline monitoring deployment and operational processes. - Knowledge of application performance monitoring, distributed tracing, infrastructure monitoring, and observability best practices. - Familiarity with ServiceNow and ITSM workflows is preferred. - Exposure to other enterprise monitoring and observability tools is an advantage. - Strong analytical, troubleshooting, and problem-solving skills with the ability to resolve complex production issues. - Excellent verbal and written communication skills with the ability to collaborate across cross-functional teams. - Ability to manage multiple priorities in a fast-paced, enterprise environment. **Responsibilities** - Design, implement, and manage Datadog monitoring solutions across infrastructure, applications, databases, synthetic monitoring, and Real User Monitoring (RUM). - Configure and maintain Datadog dashboards, monitors, alerts, log pipelines, and observability frameworks to provide comprehensive operational visibility. - Develop and maintain Terraform modules to automate Datadog configuration, deployment, and infrastructure provisioning. - Design and implement Azure DevOps CI/CD pipelines for monitoring configuration, automation, and continuous delivery. - Collaborate with application, infrastructure, cloud, and DevOps teams to ensure end-to-end monitoring coverage across enterprise platforms. - Implement monitoring standards, instrumentation, and best practices for cloud-native and enterprise applications. - Support incident, problem, and change management processes while ensuring adherence to ITIL best practices. - Perform root cause analysis, troubleshoot monitoring issues, and optimize platform performance to improve service reliability. - Develop monitoring strategies for cloud environments across AWS, Azure, and hybrid infrastructure. - Configure log collection, parsing, enrichment, and routing to support operational monitoring and analytics. - Build automated monitoring and alerting solutions to proactively identify performance, availability, and infrastructure issues. - Integrate Datadog with enterprise tools, ITSM platforms, and automation frameworks to improve operational efficiency. - Maintain technical documentation, monitoring standards, operational procedures, and deployment guides. - Participate in production support, release activities, platform upgrades, and continuous improvement initiatives. - Work closely with stakeholders to enhance observability capabilities, optimize monitoring coverage, and improve overall platform reliability and operational excellence. Bachelor’s or master’s degree in computer science, Information Technology, or a related field.