AWS Data Engineer
EXL Service · State of Mahārāshtra, India
Free to search · AI fit score against your CV · tailor your résumé in one click
EXL Service · State of Mahārāshtra, India
Job Description: We are seeking a highly skilled AWS Data Engineer with deep expertise in AWS cloud architecture, big data processing, real-time streaming, and modern data lake technologies. The ideal candidate will have strong hands-on experience in Spark (PySpark), Iceberg, EMR, Starburst/Trino, and event-driven architectures, along with experience building real-time and API-driven data applications who can design and build generic solutions for one of our Fortune 500 Client programs in the realm of Financial Master & Reference Data Management . This is high visibility, fast-paced key initiative will integrate data across internal and external sources, provide analytical insights, and integrate with the customer’s critical systems. Responsibilities: Key Responsibilities • Design and implement scalable, secure, and cost-optimized AWS data architectures . • Develop and maintain ETL pipelines using AWS Lambda and AWS Glue ETL . • Configure and manage AWS Glue Crawlers, Glue Data Catalog, and schema evolution. • Build, optimize, and unit test applications on the Apache Spark framework using PySpark. • Design and optimize data lakes using Apache Iceberg on AWS, including table compaction and Iceberg performance tuning. • Work extensively with data formats such as Avro, Parquet, JSON, XML , and CSV . • Orchestrate event-driven workflows using AWS Step Functions and Amazon EventBridge. • Connect and integrate Starburst from Lambda and Glue ETL jobs for federated querying. • Implement CI/CD pipelines for automated testing and deployment. • Perform unit testing using PyTest , and performance tuning of Spark and Python applications • Qualifications: Strong understanding of AWS architecture best practices , scalability, security, and cost optimization strategies. • Strong hands-on experience with AWS services including Lambda, Glue ETL, Athena, S3, DynamoDB, Step Functions, EventBridge, SNS, and SQS. • Deep experience in Apache Spark (PySpark/Scala) development, unit testing, and performance optimization. • Strong Python programming skills using libraries such as pandas, requests, json, and awswrangler. • Experience on Apache Kafka and Confluent Kafka . • Experience designing and optimizing data lakes using Apache Iceberg , including compaction and Iceberg optimization techniques.