AWS Big Data Engineer (PySpark & EMR)
Tata Consultancy Services · Hyderabad, Telangana, India - Indore, Madhya Pradesh, India - Pune, Maharashtra, India
Tata Consultancy Services · Hyderabad, Telangana, India - Indore, Madhya Pradesh, India - Pune, Maharashtra, India
**Key Responsibilities** - Design, develop, and maintain large-scale data processing pipelines using PySpark and AWS services. - Build and optimize Spark-based ETL solutions for enterprise data platforms. - Work with AWS services including EMR, S3, IAM, Lambda, SNS, and SQS. - Develop efficient data transformation and processing logic using Python and PySpark. - Analyze and optimize Spark jobs using Spark UI, debugging techniques, and performance tuning methodologies. - Write and optimize SQL queries involving joins, subqueries, CTEs, and complex transformations. - Support data integration, ingestion, and reporting requirements across multiple systems. - Participate in solution design using Data Vault, Data Mesh, and Data Fabric architectural concepts. - Collaborate with business users, architects, and technical teams to understand requirements and deliver scalable solutions. - Participate in Agile ceremonies and contribute to project planning, estimation, and execution. - Troubleshoot production issues and implement performance improvements across data platforms. **Required Skills** - 4-7 years of experience in Data Engineering or Big Data development. - Strong hands-on experience with PySpark and Apache Spark. - Experience working with AWS EMR and AWS cloud services. - Strong Python programming and scripting skills. - Experience with AWS S3, IAM, Lambda, SNS, and SQS. - Good understanding of Big Data architecture and distributed computing. - Strong SQL skills including joins, subqueries, and CTEs. - Experience with database technologies and data processing frameworks. - Knowledge of Spark optimization, debugging, and performance tuning. **Preferred Skills** - Understanding of Data Vault architecture. - Knowledge of Data Mesh and Data Fabric concepts. - Experience with large-scale cloud data migration projects. - Exposure to Hadoop and related Big Data ecosystems. - Agile/Scrum project experience. **Desired Candidate Profile** - Strong analytical and problem-solving skills. - Excellent communication and stakeholder management abilities. - Experience working directly with business and IT teams. - Ability to estimate effort, plan work, and deliver projects effectively. - Self-motivated, proactive, and adaptable to fast-paced environments.