DevOps IRC302483
GlobalLogic · Hyderabad, Telangana, India
Free to search · AI fit score against your CV · tailor your résumé in one click
GlobalLogic · Hyderabad, Telangana, India
Description Key Responsibilities Core DevOps Operations – Manage and upgrade EKS clusters — version upgrades, node group rotations, add-on compatibility, and zero-downtime rollouts – Troubleshoot and resolve Kubernetes issues: pod failures, OOMKills, scheduling problems, networking, ingress/ALB misconfigurations – Handle Jenkins administration — pipeline debugging, plugin upgrades, controller/agent connectivity, build failure triage – Manage AWS infrastructure: EC2 instance rehydration, ASG self-healing, ALB/ELBv2 target health, Route53 DNS, CloudTrail/CloudWatch investigation – Perform instance rehydration and capacity management across auto-scaling groups – Respond to and investigate infrastructure incidents; perform root cause analysis and post-incident documentation – Maintain and improve Helm chart deployments; manage helm upgrade releases across environments – Manage IAM roles, cross-account access (STS AssumeRole), and security group configurations – Maintain SSL/TLS certificate lifecycle across internal and external endpoints Platform & Tooling Development – Write and maintain Python tooling that integrates with internal systems (Kubernetes API, AWS, Jira, Confluence, Jenkins, LDAP/AD) – Contribute to an internal AI-powered operations platform — extending tool integrations, debugging issues, and adding capabilities as needed – Maintain REST APIs built on FastAPI; write clean, typed Python (3.10+) following team standards – Write and maintain tests using pytest; follow code quality standards (black, isort, mypy) Collaboration & Process – Document runbooks, operational procedures, and architecture decisions – Review infrastructure and code changes; enforce operational best practices — Required Skills DevOps & Infrastructure (Mandatory) – 5+ years of hands-on Kubernetes experience — EKS preferred; cluster upgrades, node management, networking, RBAC – Strong AWS skills — EC2, ASG, ALB/ELBv2, Route53, IAM, CloudWatch, CloudTrail, STS – Helm — chart authoring, release management, values file strategy – Jenkins — administration, pipeline authoring (Declarative/Scripted), build failure investigation – Docker — image builds, multi-platform (linux/amd64), container registry (ECR) – Infrastructure-as-code mindset; comfortable reading and modifying YAML-heavy configs Python (Mandatory) – Strong Python 3.10+ proficiency — type hints, pydantic v2, async patterns, httpx – Experience building or maintaining REST APIs (FastAPI or similar) – Comfortable writing scripts and tools that integrate with AWS (boto3), Kubernetes (kubernetes client), and internal REST APIs – Able to read, extend, and debug an existing Python codebase without full prior context Systems Integration – REST API integration with enterprise tooling (Jira, Confluence, Jenkins, or similar) – Authentication patterns: Bearer tokens, Basic Auth, AWS IRSA, cross-account IAM General – Strong incident investigation and debugging skills across distributed systems – Able to work independently, manage priorities, and communicate blockers early – PostgreSQL — basic query and operational familiarity — Nice to Have – LangChain / LangGraph — multi-agent graphs, ReAct pattern, stateful graph checkpointing – Any LLM API experience (Anthropic, OpenAI, or similar) — tool use and prompt construction – Experience contributing to AI-assisted operations or internal developer tooling — What You Don’t Need – ML or data science background – Prior experience with LangGraph or AI agent frameworks — willingness to learn from existing code is sufficient – Full-stack frontend experience Requirements About the Role We are looking for a Senior DevOps Engineer to join our operations team. You will own day-to-day operations of our cloud infrastructure — Kubernetes cluster management, CI/CD pipelines, AWS infrastructure, and incident response — while also contributing to an internal AI-powered operations platform built in Python. The AI platform work is additive; your primary responsibility is keeping production infrastructure healthy and reliable Job responsibilities Key Responsibilities Core DevOps Operations – Manage and upgrade EKS clusters — version upgrades, node group rotations, add-on compatibility, and zero-downtime rollouts – Troubleshoot and resolve Kubernetes issues: pod failures, OOMKills, scheduling problems, networking, ingress/ALB misconfigurations – Handle Jenkins administration — pipeline debugging, plugin upgrades, controller/agent connectivity, build failure triage – Manage AWS infrastructure: EC2 instance rehydration, ASG self-healing, ALB/ELBv2 target health, Route53 DNS, CloudTrail/CloudWatch investigation – Perform instance rehydration and capacity management across auto-scaling groups – Respond to and investigate infrastructure incidents; perform root cause analysis and post-incident documentation – Maintain and improve Helm chart deployments; manage helm upgrade releases across environments – Manage IAM roles, cross-account access (STS AssumeRole), and security group configurations – Maintain SSL/TLS certificate lifecycle across internal and external endpoints Platform & Tooling Development – Write and maintain Python tooling that integrates with internal systems (Kubernetes API, AWS, Jira, Confluence, Jenkins, LDAP/AD) – Contribute to an internal AI-powered operations platform — extending tool integrations, debugging issues, and adding capabilities as needed – Maintain REST APIs built on FastAPI; write clean, typed Python (3.10+) following team standards – Write and maintain tests using pytest; follow code quality stan