Job description

Role: AI/Gen AI Engineer Level: Senior Associate Experience: 5-10 years Location: Bangalore and Hyderabad **The role** You own the reliability of the deployed fleet of agents and automations. You design the evaluation and guardrail strategy, lead agent-behavior incident response, and are the technical authority on whether an agent is behaving in production --- and on when it is safe to give an agent more autonomy. AgentOps is what keeps non-deterministic systems trustworthy at scale. It is distinct from Observability \& SRE: AgentOps owns the behavior and safety of our agents --- evals, guardrails, drift, and cost --- while Observability \& SRE owns the health of the client's estate. As the fleet grows, this discipline --- not headcount --- is what lets a lean crew operate an ever-larger automated footprint. **About PwC AIOps** PwC's AIOps runs the client's operations with intelligence built in: we deploy PwC's agentic platforms and automation accelerators into the client's environment and operate them in production --- taking cost out of the run while improving service quality, measured and sustained year over year. This is broader than the classic "AI for IT operations"; it is operating a workforce of agents and automations that do the work, with the engineering discipline to keep them reliable and to prove the value. Forward Deployed Engineers (FDEs) are embedded with the client --- high-agency, ship-to-production engineers who own outcomes rather than tickets, sitting close to the work and building what is needed. AIOps is delivered across three disciplines that work as one loop. Build \& Deploy ships the work-doers --- AI agents for judgment- and language-heavy tasks, and deterministic automation for high-volume, rules-based work; the two go hand in hand, chosen by whichever is the faster, safer path to takeout. Operations (AgentOps) keeps the deployed fleet trustworthy through evaluation, guardrails, drift detection, and incident response --- what lets a lean core safely operate an ever-growing footprint. Observability \& SRE instruments the estate to find the toil, cut the noise, and quantify the efficiency delivered. Build creates capability, AgentOps keeps it reliable, and Observability finds the next opportunity and proves the last. Efficiency compounds because we run a flywheel, not a project: observe to locate toil and cost, build or automate it away, operate it reliably so it scales without adding heads, prove the takeout, and reinvest the freed capacity into the next opportunity --- so value scales with the fleet, not headcount. Operate is not "keep the lights on": a disciplined innovation cadence runs inside delivery --- daily intake, value-based prioritization, rapid build, and measured results, governed jointly with the client --- making the operate layer the client's continuous-improvement and step-change engine. Every role contributes to, and is measured against, that year-over-year takeout. **What you'll do** **Evaluation \& reliability** * Own the eval-in-prod strategy: offline suites, online evals, and regression testing of agent behavior. * Design drift detection, guardrails, and blast-radius containment for the deployed fleet. * Own the autonomy gates --- define and enforce the criteria that move an agent from shadow β†’ assist β†’ supervised β†’ autonomous. * Own cost and token monitoring and optimization. **Operate** * Lead incident response for agent-behavior issues; run blameless postmortems and drive systemic fixes. * Define telemetry and tracing standards across deployed agents. * Build the human-in-the-loop and escalation patterns the run organization operates against. **Close the loop** * Partner with Build \& Deploy to fix reliability gaps at the source, not just in production. * Feed failure patterns into better build standards and into the team's innovation cadence. **People** * Mentor Associate FDEs on operating and evaluating agentic systems. **Your toolkit** * Engineering: strong Python; comfort across logs, traces, and metrics. * Tracing \& telemetry: observability for LLM systems --- Langfuse, LangSmith, Datadog, OpenTelemetry. * Evaluation: frameworks and methodology for non-deterministic systems. * Reliability: guardrail/safety tooling, incident management, on-call, and SLO practice. **What you'll bring** * 6 years in SRE, ops, or ML engineering, including 2 operating production ML or LLM systems. * A strong grasp of evaluation, monitoring, and reliability for non-deterministic systems. * Incident-management and on-call leadership experience. * An instinct for why agentic systems fail differently --- and how to contain that. **Nice to have** * AIOps platform experience or ML monitoring at scale. * Exposure to safety, eval, or guardrail tooling and research. **What success looks like (first 6--12 months)** * Incidents on the deployed fleet trend down even as the fleet grows. * There's a living eval and guardrail framework the whole team trusts. * The run organization can operate agents safely because of the patterns you built. *You'll thrive here if making AI reliable in production excites you more than building the next shiny thing --- and you want to define a discipline that's still being written.* **Working at PwC** You will work within PwC's quality, risk, and independence standards, alongside a team that invests in your growth and a practice that is building the future of intelligent operations. PwC is an equal opportunity employer.