Architect - Data Engineer
TransUnion · Hyderabad
TransUnion · Hyderabad
TransUnion's Job Applicant Privacy Notice Team Overview The team is focused on building and evolving job manager and distributed processing solutions using Apache Spark and cloud technologies.

This is a hybrid position and involves regular performance of job responsibilities virtually as well as in-person at an assigned TU office location for a minimum of two days a week.Role Overview And Core Responsibilities Architecture & Solution Design Responsibilities Design scalable and resilient data processing solutions using Apache Spark and modern data platform technologies. Create architecture designs, reference patterns, and technical specifications for distributed data platforms. Guide engineering teams on architecture, design, and implementation decisions. Ensure solutions align with enterprise architecture standards, security requirements, and engineering best practices. Evaluate technology choices, frameworks, and platform capabilities to address evolving business needs. Identify architectural risks and recommend mitigation strategies. Contribute to platform modernization and cloud transformation initiatives. Apache Spark & Data Platform Leadership Responsibilities Provide technical leadership for Apache Spark-based batch and streaming data processing solutions. Guide development teams on Spark architecture, performance optimization, and operational best practices. Define design patterns for scalable, fault-tolerant, and maintainable Spark applications. Assist teams in resolving complex technical challenges involving distributed processing and large-scale data workloads. Drive performance improvements through effective partitioning, query optimization, resource management, caching, and tuning strategies. Establish standards for observability, monitoring, reliability, and operational excellence. Solution Architecture & Engineering Guidance Responsibilities Collaborate with engineering teams throughout the software development lifecycle. Conduct architecture and design reviews for new features, data pipelines, and platform enhancements. Support engineering teams in implementing scalable and maintainable solutions. Promote adoption of CI/CD, automated testing, infrastructure-as-code, and DevOps best practices. Contribute to technical decision-making for platform enhancements and modernization initiatives. Stakeholder Collaboration Responsibilities Work closely with product managers, engineering leads, data scientists, and platform teams. Translate business requirements into scalable technical solutions and architecture designs. Communicate architecture decisions, technical trade-offs, and implementation approaches to stakeholders. Participate in roadmap discussions to ensure alignment between business goals and technology strategy. Collaborate across teams to drive successful implementation of data platform capabilities. Technical Mentoring Responsibilities Mentor engineers on distributed systems design, Spark development, and data engineering best practices. Provide guidance through architecture reviews, design discussions, and technical workshops. Share knowledge on emerging data platform technologies and architectural patterns. Required Knowledge And Experiences Process & Quality Responsibilities Ensure architecture and design activities follow established SDLC and governance standards. Participate in code reviews, design reviews, and architecture assessments. Promote reliability, maintainability, scalability, and performance considerations throughout solution development. Drive continuous improvement in architecture practices and engineering standards. Required Knowledge & Experience Technical Expertise 10+ years of experience in software engineering, data engineering, or distributed systems development. Strong hands-on expertise with Apache Spark (Spark SQL, Structured Streaming, DataFrames, Dataset APIs). Experience designing and building large-scale distributed data processing systems. Strong knowledge of Spark optimization techniques including: • Partitioning Strategies • Shuffle Optimization • Join Optimization • Memory Management • Resource Utilization • Performance Tuning Proficiency in Scala, Java, or Python. Experience with technologies such as: • Hadoop • Hive • Iceberg • AWS EMR • AWS Glue • GCP Dataproc • BigQuery Experience designing and operating batch and streaming data pipelines. Understanding of cloud-native architecture and distributed systems principles. Experience deploying Spark solutions on AWS and/or GCP platforms. Scope & Positioning Individual contributor architecture role focused on distributed data platforms and Spark-based solutions. Provides architecture leadership and technical guidance across one or more engineering teams. Responsible for solution architecture quality, technology selection, design governance, and technical direction within the domain. Partners with Engineering Managers and Technical Leads to deliver scalable and reliable platform capabilities. Serves as a key technical advisor for Spark architecture, distributed systems design, and cloud-native data platforms. Preferred Qualifications • Experience with AWS services such as EMR, Glue, S3, Lambda, EKS,