L2 Support System Administrator
HCLTech Β· Pune, Maharashtra, India
HCLTech Β· Pune, Maharashtra, India
Key Responsibilities - Handle L1/L2 support tickets triage, classify, and resolve customer-reported issues on product. - Diagnose problems across product planes: agent runtime failures, LLM/MCP Gateway errors, authentication issues, and policy violations. - Analyse platform logs, telemetry, and usage data to identify root causes. - Guide customers on product onboarding, configuration, and best practices. - Document troubleshooting steps, known issues, and workarounds in the knowledge base. - Escalate complex issues to L3/Engineering with clear reproduction steps and evidence. - Participate in on-call rotations for critical incident response. Required Skills & Qualifications Education B.E. / B.Tech in Computer Science, IT, or related field Experience 1β2 years in software product support or related technical role Cloud / Infra Basic understanding of containerisation (Docker/Kubernetes) and cloud deployments APIs & Auth Familiarity with REST APIs, API gateways, OAuth / token-based authentication AI / LLMs Awareness of LLM concepts; experience with AI tools is a plus Scripting Python scripting for log parsing and automation Communication Strong written and verbal English; ability to communicate technical issues clearly to customers Problem-Solving Methodical approach to debugging; comfortable working with logs and error traces Good to Have - Exposure to agentic frameworks (LangChain, CrewAI, AutoGen, etc.) - Experience with observability tools (Grafana, Prometheus, ELK stack) - Familiarity with OpenShift or enterprise Kubernetes distributions - Understanding of vector databases or graph databases Senior Software Engineer Band: 3β4 Years Experience | Sub-Band : 2.2 | Joining Location : Pune Key Responsibilities - Own end-to-end resolution of complex, multi-component issues spanning product Runtime, Connectivity, Governance, and Model planes. - Conduct deep-dive root cause analysis on agent failures, LLM routing anomalies, policy engine misconfigurations, and memory/context service issues. - Act as the escalation bridge between customers and the engineering team β filing detailed bug reports with reproduction environments. - Advise enterprise customers on product architecture, agent design patterns, security hardening, and cost optimisation using Billing plane controls. - Lead post-incident reviews and contribute to platform runbooks and support playbooks. - Mentor junior support engineers on troubleshooting methodologies and product internals. - Identify patterns in support tickets to proactively surface product gaps and influence roadmap priorities. Required Skills & Qualifications Education B.E. / B.Tech in Computer Science, IT, or related field Experience 3β4 years in product support, platform engineering, or SRE for enterprise software Cloud / Infra Hands-on with Kubernetes/OpenShift, containerised microservices, and enterprise cloud environments APIs & Security Deep understanding of API gateway architecture, OAuth, JWT, mTLS, and RBAC/ABAC policy models AI / LLM Ops Working knowledge of LLM APIs (OpenAI, Azure OpenAI, Watsonx, etc.), prompt engineering, and agent frameworks Scripting / Dev Proficient in Python; comfortable reading application code to identify defects Communication Excellent written and verbal English; ability to run executive-level incident bridges and write clear RCA documents Good to Have - Experience with agentic frameworks: LangChain, AutoGen, CrewAI, or similar - Knowledge of MCP (Model Context Protocol) architecture - Familiarity with vector databases (Pinecone, Weaviate, pgvector) or graph databases (Neo4j) - Exposure to enterprise governance frameworks for AI (responsible AI, regulatory compliance)