P

Member of Technical Staff, Production Engineering Lead

Pure Storage · Bangalore, India

~₹65L (est.)8–15 yrs experiencePosted Today
Apply now →

Job description

We’re in an unbelievably exciting area of tech and are fundamentally reshaping the data storage industry. Here, you lead with innovative thinking, grow along with us, and join the smartest team in the industry. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE We are seeking an experienced Software Engineer with a strong background in Platform Engineering to join our AI Applications team. You will be responsible for the "Inner Loop" of AI development—ensuring that our LLM-based applications, RAG pipelines, and agentic workflows are built, tested, and deployed with the same rigor as traditional software. You will bridge the gap between high-velocity AI experimentation and stable, enterprise-grade production infrastructure WHAT YOU'LL DO • Platform Self-Service: Create internal APIs and abstractions that allow Software Engineers to provision AI-ready environments (complete with model weights, vector DBs, and event streams) with a single command. • Agentic Automation: Use Golang and Python to build internal tools and AI Agents that automate root-cause analysis of infrastructure failures and proactively optimize infrastructure.. Event-Driven AI Integration: Leverage Kafka or RabbitMQ to build asynchronous AI processing pipelines (e.g., long-running document ingestion for RAG systems). • Reproducibility: Ensure "Golden Path" deployments for AI models using Docker and Kubernetes, ensuring that the model, the prompt, and the code are all perfectly synced across environments. WHAT YOU BRING Required Technical Skills • Experience: 8+ years in Platform Infrastructure Engineering with a shift toward AI Application development. • Languages: Strong proficiency in Python (for AI logic) and Go (for platform tools). • AI Literacy: Practical experience with RAG (Retrieval-Augmented Generation), Prompt Engineering, and integrating LLM APIs (OpenAI, Anthropic, or local models via Ollama). • Event-Driven Scaling: Design and implement high-performance Go services that listen to Kafka/RabbitMQ streams to trigger dynamic infrastructure scaling based on real-time AI model demand. • Distributed State & Storage: Manage and optimize the infrastructure for Vector Databases and distributed caches (Redis), ensuring high availability for RAG data. • Go-Based Tooling: Replace brittle shell scripts with robust, type-safe Internal Tooling in Go for automated environment provisioning and disaster recovery. • Infrastructure: Deep experience with Kubernetes, specifically managing GPU workloads and specialized storage for vector databases. • Data/Events: Proven experience with Kafka or RabbitMQ for managing high-volume data streams. • Observability: Using Prometheus and Grafana to monitor not just system health, but AI-specific metrics like token latency and model "drift." Technical Problem Solving • A "Code-First" Infrastructure Mindset: You don't just click in consoles; you view every infrastructure problem as a software challenge. You bring the ability to write clean, maintainable, and testable Go code to manage complex cloud environments. • Expertise in High-Concurrency Systems: You understand how to use Go’s goroutines and channels to handle thousands of concurrent events, making you an expert at managing the massive data throughput required by Kafka-driven AI pipelines. • Deep Orchestration Knowledge: You bring a builder’s perspective to Kubernetes. You don't just deploy apps; you understand how to extend the K8s API with custom controllers to make the cluster "AI-aware." • The Bridge Between Data & Ops: You understand the unique infrastructure needs of AI (GPU memory management, high-speed NVMe storage, and vector retrieval) and can translate those requirements into stable, scalable systems. • Operational Resilience & Reliability: With your background in Monitoring (Grafana/Prometheus) and Log Management (ELK), you bring a "Zero-Downtime" philosophy, ensuring that even under heavy AI inference loads, the system remains performant and observable. • Event-Driven Strategy: You bring a sophisticated understanding of asynchronous architecture, knowing exactly when to use Kafka for high-volume streaming versus RabbitMQ for complex task routing. • Security & Compliance Guardrails: You bring a "Security-as-Code" approach, ensuring that data privacy (vital for AI) is baked into the infrastructure layer rather than bolted on at the end. A mindset of collaboration, reliability, and continuous improvement to strengthen team productivity and delivery speed. #LI-ONSITE <div class="setting-item-core__grid-container" data-test-setting-item="form-closed" data-test-promoti