Search 100,000+ live jobs across India

Free to search · AI fit score against your CV · tailor your résumé in one click

Job description

Description Key Responsibilities Core DevOps Operations – Manage and upgrade EKS clusters — version upgrades, node group rotations, add-on compatibility, and zero-downtime rollouts – Troubleshoot and resolve Kubernetes issues: pod failures, OOMKills, scheduling problems, networking, ingress/ALB misconfigurations – Handle Jenkins administration — pipeline debugging, plugin upgrades, controller/agent connectivity, build failure triage – Manage AWS infrastructure: EC2 instance rehydration, ASG self-healing, ALB/ELBv2 target health, Route53 DNS, CloudTrail/CloudWatch investigation – Perform instance rehydration and capacity management across auto-scaling groups – Respond to and investigate infrastructure incidents; perform root cause analysis and post-incident documentation – Maintain and improve Helm chart deployments; manage helm upgrade releases across environments – Manage IAM roles, cross-account access (STS AssumeRole), and security group configurations – Maintain SSL/TLS certificate lifecycle across internal and external endpoints Platform & Tooling Development – Write and maintain Python tooling that integrates with internal systems (Kubernetes API, AWS, Jira, Confluence, Jenkins, LDAP/AD) – Contribute to an internal AI-powered operations platform — extending tool integrations, debugging issues, and adding capabilities as needed – Maintain REST APIs built on FastAPI; write clean, typed Python (3.10+) following team standards – Write and maintain tests using pytest; follow code quality standards (black, isort, mypy) Collaboration & Process – Document runbooks, operational procedures, and architecture decisions – Review infrastructure and code changes; enforce operational best practices — Required Skills DevOps & Infrastructure (Mandatory) – 5+ years of hands-on Kubernetes experience — EKS preferred; cluster upgrades, node management, networking, RBAC – Strong AWS skills — EC2, ASG, ALB/ELBv2, Route53, IAM, CloudWatch, CloudTrail, STS – Helm — chart authoring, release management, values file strategy – Jenkins — administration, pipeline authoring (Declarative/Scripted), build failure investigation – Docker — image builds, multi-platform (linux/amd64), container registry (ECR) – Infrastructure-as-code mindset; comfortable reading and modifying YAML-heavy configs Python (Mandatory) – Strong Python 3.10+ proficiency — type hints, pydantic v2, async patterns, httpx – Experience building or maintaining REST APIs (FastAPI or similar) – Comfortable writing scripts and tools that integrate with AWS (boto3), Kubernetes (kubernetes client), and internal REST APIs – Able to read, extend, and debug an existing Python codebase without full prior context Systems Integration – REST API integration with enterprise tooling (Jira, Confluence, Jenkins, or similar) – Authentication patterns: Bearer tokens, Basic Auth, AWS IRSA, cross-account IAM General – Strong incident investigation and debugging skills across distributed systems – Able to work independently, manage priorities, and communicate blockers early – PostgreSQL — basic query and operational familiarity — Nice to Have – LangChain / LangGraph — multi-agent graphs, ReAct pattern, stateful graph checkpointing – Any LLM API experience (Anthropic, OpenAI, or similar) — tool use and prompt construction – Experience contributing to AI-assisted operations or internal developer tooling — What You Don’t Need – ML or data science background – Prior experience with LangGraph or AI agent frameworks — willingness to learn from existing code is sufficient – Full-stack frontend experience Requirements About the Role We are looking for a Senior DevOps Engineer to join our operations team. You will own day-to-day operations of our cloud infrastructure — Kubernetes cluster management, CI/CD pipelines, AWS infrastructure, and incident response — while also contributing to an internal AI-powered operations platform built in Python. The AI platform work is additive; your primary responsibility is keeping production infrastructure healthy and reliable Job responsibilities Key Responsibilities Core DevOps Operations – Manage and upgrade EKS clusters — version upgrades, node group rotations, add-on compatibility, and zero-downtime rollouts – Troubleshoot and resolve Kubernetes issues: pod failures, OOMKills, scheduling problems, networking, ingress/ALB misconfigurations – Handle Jenkins administration — pipeline debugging, plugin upgrades, controller/agent connectivity, build failure triage – Manage AWS infrastructure: EC2 instance rehydration, ASG self-healing, ALB/ELBv2 target health, Route53 DNS, CloudTrail/CloudWatch investigation – Perform instance rehydration and capacity management across auto-scaling groups – Respond to and investigate infrastructure incidents; perform root cause analysis and post-incident documentation – Maintain and improve Helm chart deployments; manage helm upgrade releases across environments – Manage IAM roles, cross-account access (STS AssumeRole), and security group configurations – Maintain SSL/TLS certificate lifecycle across internal and external endpoints Platform & Tooling Development – Write and maintain Python tooling that integrates with internal systems (Kubernetes API, AWS, Jira, Confluence, Jenkins, LDAP/AD) – Contribute to an internal AI-powered operations platform — extending tool integrations, debugging issues, and adding capabilities as needed – Maintain REST APIs built on FastAPI; write clean, typed Python (3.10+) following team standards – Write and maintain tests using pytest; follow code quality stan

More jobs at GlobalLogic

All GlobalLogic jobs (410)

Engineering jobs in Hyderabad

Engineering jobs in Hyderabad (7,575)

Other Engineering jobs in India

All Engineering jobs in India (44,397)