PySpark+ Databricks(7-14yrs)
Infosys · Bengaluru, Karnataka, India - Chennai, Tamil Nadu, India - Hyderabad, Telangana, India
Free to search · AI fit score against your CV · tailor your résumé in one click
Infosys · Bengaluru, Karnataka, India - Chennai, Tamil Nadu, India - Hyderabad, Telangana, India
Roles and responsibility • Design, develop, and maintain scalable data pipelines using PySpark and Databricks. • Build and optimize ETL/ELT processes for large-scale data processing. • Develop data transformation and data integration workflows in Databricks. • Work with structured and unstructured data from various sources. • Optimize Spark jobs for performance, scalability, and cost efficiency. • Implement Delta Lake architecture and manage data quality. • Collaborate with data analysts, data scientists, and business stakeholders. • Troubleshoot and resolve data pipeline and performance issues. • Ensure adherence to data governance and security standards. • Participate in code reviews, testing, and deployment activities. Required Skills • Strong experience in PySpark and Apache Spark. • Hands-on experience with Databricks workspace and notebooks. • Proficiency in Python and SQL. • Experience with Delta Lake, Spark SQL, and Spark Optimization techniques. • Knowledge of data warehousing concepts and ETL frameworks. • Experience with cloud platforms such as Azure, AWS, or GCP. • Understanding of data modeling and big data technologies. • Experience with version control tools like Git.