Walk-in || Linux Engineer
Tata Consultancy Services · Gurugram, Haryana, India
Tata Consultancy Services · Gurugram, Haryana, India
**Role & responsibilities** Provide L2/L3 operations for customers UNIX/Linux estate across onprem + AWS + Azure + VMware, with strong focus on OS currency (EOL remediation), patch compliance, monitoring/event management integration, and operational stability Ownership of end-to-end Linux tower delivery service and transformation — driving Linux standardization, EOL burn-down strategy, patch automation adoption, DR readiness governance alignment, observability maturity, and SLA-led operations, emphasizing event management, patching automation, risk registry, and operating model maturity. Tools / Platforms ITSM / Event Mgmt / AIOps: ServiceNow, ServiceNow ITOM Infra Monitoring: SolarWinds Observability / Logs: Grafana (Prometheus stack), Loki/Fluentd Patching: AWS SSM direction for patching onprem + cloud; SCCM/Intune Automation/Orchestration: ServiceNow + AWX/Ansible + Terraform Batch: ControlM / Tidal (where Linux workloads rely on scheduling dependencies) Provide advanced troubleshooting and resolution for Linux incidents (availability, performance, filesystem, user/access, services, kernel-related issues). Support operational activities across onprem, AWS, and Azure Linux estates, aligning to the broader customer operating model for infrastructure operations and administration. Execute OS version upgrades and remediation for EOL/EOS Linux servers, prioritizing critical systems and high-risk populations (e.g., EOL VMs) Maintain compliance with “supported version levels” expectations by ensuring planned upgrades and reducing vendor support risk exposure. Perform patch planning/execution and support the customer’s direction to extend AWS SSM for patching across onprem and cloud, including pre-check/post-check automation and maintenance coordination. Operate within the approved tool landscape for patch/image management (SCCM/Intune) where relevant to server operations and standardization. Integrate Linux monitoring into Customer’s ServiceNow ITOM event management approach and downstream monitoring tools like SolarWinds, plus observability via Grafana/Loki where used. Improve alert quality by participating in tuning thresholds and reduction of noisy events, consistent with the customer event management direction. Support DR readiness activities for Linux workloads, aligning to the TDAF callout that periodic DR testing/governance must be established (execution support, runbook validation, recovery assistance). Provide operational inputs into capacity planning and proactive risk identification (CPU/memory/storage growth trends). Maintain Linux SOPs/runbooks and update CMDB/service records as required through ServiceNow, aligned to the tool strategy emphasizing ITSM/CMDB integration. Define the Linux OS standard/version baselines and a sequenced plan to reduce EOL exposure (including EOL VM populations highlighted in inventory) Drive estate rationalization across platforms (VMware/AWS/Azure) and ensure “supported release” posture to reduce operational and security risk. Establish patch policy, cadence, and compliance reporting; operationalize customer’s direction to extend AWS SSM patching for onprem and cloud, including automation of checks and standardized runbooks. Align patch/image practices to Customer’s retained tool ecosystem (SCCM/Intune). Ensure Linux signals (metrics/logs/events) are integrated into ServiceNow ITOM event management and correlated with downstream tooling (SolarWinds) and observability dashboards (Grafana/Loki). Sponsor alert-quality improvements (threshold strategy, noise reduction) Implement Linux tower contribution to DR readiness, aligned to the TDAF callout that DR drills/process governance must be mandated and executed, ensuring Linux workloads meet agreed RTO/RPO expectations by criticality. Define Linux tower SLAs/OLAs, measurement methods, and reporting cadence. Ensure Linux tower maintains and governs a risk register and operational documentation discipline via ServiceNow/CMDB integration. Serve as the Linux tower escalation owner for P1/P2 issues, stakeholder communications, and cross-tower coordination (Compute/Network/DB/Storage).