Senior Cloud Platform Engineer - CL
Endava · Bengaluru, Karnataka, in
Endava · Bengaluru, Karnataka, in
Technology is our how. And people are our why. For over two decades, we have been harnessing technology to drive meaningful change.   By combining world-class engineering, industry expertise and a people-centric mindset, we consult and partner with leading brands from various industries to create dynamic platforms and intelligent digital experiences that drive innovation and transform businesses.   From prototype to real-world impact - be part of a global shift by doing work that matters. About the Role You will be the primary architect, operator, and developer-enablement partner for the client’s AWS engineering platform. This is a hands-on role responsible for designing, securing, scaling, and supporting the infrastructure, CI/CD services, developer tools, and observability capabilities that allow engineering teams to deliver software reliably. You will work across AWS, EKS/Kubernetes, Linux, networking, identity, automation, and shared engineering tools while mentoring teams on cloud-native practices.   Responsibilities ● Design and implement secure, multi-account structures and landing zones across AWS. ● Manage and evolve multi-cloud IAM, SSO, and Role-Based Access Controls (RBAC) to ensure least-privilege access. ● Enforce tagging standards, resource hierarchies, and cost-optimization strategies (rightsizing, idle resource elimination) to maintain fiscal accountability. ● Lead the deployment, scaling, and management of Kubernetes clusters (EKS, GKE, or self-managed). Manage CNI plugins, ingress controllers, and service meshes (Istio/Linkerd). ● Administer and optimize Linux (Ubuntu, Amazon Linux, RHEL) and Windows Server environments, ensuring hardened configurations and automated patching. ● Manage the intersection of cloud services and traditional OS-level dependencies, including Active Directory integration and file system performance tuning. ● Develop and maintain modular templates using Terraform, CloudFormation, or Pulumi. ● Build, orchestrate, and support CI/CD pipelines using Jenkins and Bitbucket, including webhooks, Jenkins agents, build/test stages, artifact publishing, approvals, deployments, rollback, notifications, and troubleshooting. Maintain GitOps workflows using GitHub Actions, GitLab CI, Flux, or ArgoCD where applicable. ● Design security controls including encryption at rest/transit (KMS), VPC Service Controls, and audit logging to meet SOC2, HIPAA, or FedRAMP standards. ● Leverage AI-native development tools (e.g., Cursor, GitHub Copilot) and LLM-powered agents to accelerate Infrastructure-as-Code (IaC) authoring, automate complex root-cause analysis, and proactively optimize cloud utilization through predictive anomaly detection. ● Install, configure, administer, upgrade, and integrate shared engineering tools such as Jenkins, SonarQube, Polaris, and Trino. Configure authentication, permissions, databases or supporting services, backups, health checks, and CI/CD integrations; manage secure, role-based, auditable infrastructure access using StrongDM. ● Enable developers to onboard applications, create and manage jobs, consume standard pipeline templates, and troubleshoot build, test, scan, deployment, and environment issues. ● Provision and operate AWS infrastructure, EC2 instances, EKS clusters, worker nodes, networking, load balancers, storage, databases, and supporting services using Infrastructure as Code. ● Design and maintain SSO, IAM, Kubernetes RBAC, service accounts, secrets, and read/write access models with least-privilege controls and auditable access reviews. ● Define operational ownership for the engineering platform, including tool availability, capacity, patching, upgrades, backups, disaster recovery, certificates, plugin lifecycle, and documented runbooks. ● Establish monitoring and observability for infrastructure, EKS, CI/CD, shared tools, databases, and applications. Progress from availability monitoring and metrics to centralized logs, dashboards, tracing, SLI/SLO-based alerting, incident response, and root-cause analysis. ● Monitor service health indicators such as uptime, CPU/memory/disk, network, latency, errors, throughput, pod and node health, Jenkins queue and job failure rates, scan failures, Trino coordinator/worker health, database connectivity, storage capacity, certificates, security events, and access failures. ● Partner with developers and security teams to define onboarding standards, quality gates, vulnerability policies, deployment controls, alert thresholds, and platform documentation. Qualifications AWS specific (Must have) Client-specific must-have capabilities Jenkins administration and Jenkins file/Groovy pipeline development Bitbucket webhooks, pull-request integration, and source-control workflows SonarQube administration and CI/CD integration Polaris or equivalent application-security scanning platform administration Trino deployment, configuration, access control, and operational support Observability using CloudWatch, Prometheus, Grafana, ELK/OpenSearch, Loki, OpenTelemetry, or equivalent tools Developer enablement, platform onboarding, runbooks, and production troubleshooting SSO, IAM, Kubernetes RBAC, read/write permissions, secrets, and auditability • AWS Organizations & Landing Zones • IAM, SSO, RBAC • VPCs, Transit Gateway, networking • EKS administration • Terraform • CloudFormation • GitOps (ArgoCD/Flux) • Linux administration • Kubernetes security • Cost optimization • Multi-account governance • SonarQube • StrongDM GCP specific (nice to have) • GKE • VPC Service Controls • IAM • Or