PySpark+ Databricks(7-14yrs)
Infosys · Chennai, Tamil Nadu, India - Hyderabad, Telangana, India - Pune, Maharashtra, India
Infosys · Chennai, Tamil Nadu, India - Hyderabad, Telangana, India - Pune, Maharashtra, India
**Roles and responsibility** - Design, develop, and maintain scalable data pipelines using PySpark and Databricks. - Build and optimize ETL/ELT processes for large-scale data processing. - Develop data transformation and data integration workflows in Databricks. - Work with structured and unstructured data from various sources. - Optimize Spark jobs for performance, scalability, and cost efficiency. - Implement Delta Lake architecture and manage data quality. - Collaborate with data analysts, data scientists, and business stakeholders. - Troubleshoot and resolve data pipeline and performance issues. - Ensure adherence to data governance and security standards. - Participate in code reviews, testing, and deployment activities. Required Skills - Strong experience in **PySpark** and **Apache Spark**. - Hands-on experience with **Databricks** workspace and notebooks. - Proficiency in **Python** and **SQL**. - Experience with **Delta Lake**, Spark SQL, and Spark Optimization techniques. - Knowledge of data warehousing concepts and ETL frameworks. - Experience with cloud platforms such as **Azure, AWS, or GCP**. - Understanding of data modeling and big data technologies. - Experience with version control tools like Git.