Artificial Intelligence Performance Analyst
Ibm · Bengaluru/Bangalore, Karnataka
Ibm · Bengaluru/Bangalore, Karnataka
AI Performance Analyst Introduction At IBM Infrastructure \& Technology, we design and operate the systems that keep the world running. From high-resiliency mainframes and hybrid cloud platforms to networking, automation, and site reliability. Our teams ensure the performance, security, and scalability that clients and industries depend on every day. Working in Infrastructure \& Technology means tackling complex challenges with curiosity and collaboration. You'll work with diverse technologies and colleagues worldwide to deliver resilient, future-ready solutions that power innovation. With continuous learning, career growth, and a supportive culture, IBM provides the opportunities to build expertise and shape the infrastructure that drives progress. The IBM Z Software Performance team is responsible for designing, executing, and analyzing stress workloads and benchmarks on IBM Z and LinuxONE systems to ensure these platforms meet stringent customer expectations for reliability, scalability, and performance. The team develops and maintains a suite of tools and automation scripts to support: Performance environment setup and configuration Performance measurements and data capture Data storage in centralized repositories Presentation and visualization of performance metrics Analysis and comparison of captured data across multiple scenarios Your role and responsibilities Your Role and Responsibilities\* As a AI Performance Analyst you will be responsible for: Design and implement benchmarks and stress workloads for AIU IO Cards, ensuring they remain current and relevant. Set up benchmarks and stress workloads, including configuring the underlying AIU IO Card for various performance scenarios. Automate performance measurements and streamline data collection processes for benchmarks and stress workloads. Develop and enhance data collection and analysis tools to improve efficiency and accuracy. Execute performance benchmarks and stress workloads to validate system performance. Analyze performance measurements and collected data to identify bottlenecks and resolve performance issues. Collaborate with development teams across the stack (IBM Z Hardware, IBM Research, IBM AIU application stack, Middleware/Applications) to guide and support performance optimization efforts related to configurations. Required education Bachelor's Degree Preferred education Master's Degree Required technical and professional expertise Overall Experience: 2 to 5 years in performance measurement, analysis, and system testing. Education: Bachelor's degree in Computer Science or Information Science. Technical Skills AI/ML Knowledge: Basic understanding of ML/AI model architecture, training, and inferencing. 2 years of experience with PyTorch, Tensorflow, vLLM Development \& Automation: Proficiency in source code repository systems (e.g., Git). Strong scripting and test automation skills. System \& Containerization: Basic Linux administration skills. Hands-on experience with Docker and Podman containers. Solid understanding of Operating System fundamentals and Computer Architecture concepts. Basic experience in performance analysis of applications/systems and familiarity with performance tools. Hands-on experience in functional and performance testing of multi-tiered applications. Additional Skills: Exposure to Agile methodologies and ability to apply agile concepts effectively. Strong presentation skills and ability to communicate technical concepts clearly. Collaborative team player with excellent interpersonal and communication skills. Programming Languages and scripts: Python, C/C , Bash, Ansible, Java Preferred technical and professional experience Master's degree in Information Technology, Computer Science, or Computer Engineering. AI/ML \& Model Serving Know‑how in Transformer model design and modification (architecture tuning, fine‑tuning, optimization). Hands-on experience with TensorFlow and model inference serving using TensorFlow Serving, NVIDIA Triton Inference Server, and vLLM. Systems Performance \& Observability Performance profiling and tracing (e.g., Linux perf, flame graphs, instrumentation). Advanced Linux administration skills (networking, storage, process management, kernel parameters, automation). Hardware \& Accelerators Experience in hardware design and debugging (board bring-up, driver interactions, performance counters). Working knowledge of AI accelerator architectures---GPU, TPU, AMX (capabilities, memory hierarchies, scheduling/tiling considerations). Programming Languages CUDA (kernel development, memory optimization, streams/concurrency) and Java (services, tooling, SDK integration). Years of Experience: 2 - 5 ABOUT BUSINESS UNIT IBM Systems helps IT leaders think differently about their infrastructure. IBM servers and storage are no longer inanimate - they can understand, reason, and learn so our clients can innovate while avoiding IT issues. Our systems power the world's most important industries and our clients are the architects of the future. Join us to help build our leading-edge technology portfolio designed for cognitive business and optimized for cloud computing. YOUR LIFE @ IBM In a world where technology never stands still, we understand that, dedication to our clients success, innovation that matters, and trust and personal responsibility in all our relationships, lives in what we do as IBMers as we strive to be the catalyst that makes the world work better. Being an IBMer means you'll be able to learn and develop yourself and your career, you'll be encouraged to be courageous and experiment everyday, all whilst having continuous trust and support in an environment where everyone can thrive whatever their personal or professional background. Our IBMers are growth minded, always staying curious, open to feedback and learning new information and skills to constantly transform themselves and our company. They are trusted to provide on-going feedback to help other IBMers grow, as wel