Big Data Engineer- SPARK, SCALA, AWS
Coforge · Noida, Uttar Pradesh, India - Hyderabad, Telangana, India - Pune, Maharashtra, India
Free to search · AI fit score against your CV · tailor your résumé in one click
Coforge · Noida, Uttar Pradesh, India - Hyderabad, Telangana, India - Pune, Maharashtra, India
Role- Big data engineer (Spark, Scala, AWS) We are looking for a Data Engineer with strong experience in designing, developing, and maintaining scalable, secure, and high-performance ETL/data processing pipelines. The ideal candidate should have hands-on expertise in Apache Spark, Scala, AWS EMR, and Amazon S3, along with a strong understanding of big data technologies and cloud-based data platforms. Key Responsibilities • Design, develop, and maintain scalable batch and big data processing pipelines using Apache Spark and Scala. • Build, deploy, and manage large-scale data processing solutions on AWS EMR. • Utilize Amazon S3 as the primary data lake and storage layer for ingesting, storing, and processing structured and unstructured data. • Optimize Spark applications for performance, scalability, reliability, and cost efficiency. • Develop reusable frameworks and components for data ingestion, transformation, validation, and processing. • Collaborate with data architects, business analysts, application teams, and stakeholders to understand data requirements and deliver robust data solutions. • Ensure data quality, integrity, security, and governance across all data pipelines. • Monitor, troubleshoot, and resolve production issues related to data workflows and processing jobs. • Support deployment, release management, and ongoing maintenance of data platforms. • Follow best practices related to coding standards, testing, documentation, version control, and CI/CD processes. Required Skills & Qualifications • Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline. • Strong hands-on experience in Scala programming. • Extensive experience with Apache Spark for large-scale distributed data processing. • Proven experience working with AWS EMR for big data workloads. • Strong knowledge of Amazon S3 and modern data lake architectures. • Solid understanding of ETL/ELT design principles and data pipeline development. • Strong SQL skills and experience working with large and complex datasets. • Good understanding of cloud-based architecture, performance tuning, job monitoring, and operational support. • Experience with version control systems such as Git/GitHub. • Excellent analytical, problem-solving, communication, and stakeholder management skills.