16 Aug
|
Infosys
|
Bengaluru
Roles and responsibility
- Design, develop, and maintain scalable data pipelines using PySpark and Databricks.
- Build and optimize ETL/ELT processes for large-scale data processing.
- Develop data transformation and data integration workflows in Databricks.
- Work with structured and unstructured data from various sources.
- Optimize Spark jobs for performance, scalability, and cost efficiency.
- Implement Delta Lake architecture and manage data quality.
- Collaborate with data analysts, data scientists, and business stakeholders.
- Troubleshoot and resolve data pipeline and performance issues.
- Ensure adherence to data governance and security standards.
- Participate in code reviews, testing, and deployment activities.
Required Skills
- Solid experience in PySpark and Apache Spark.
- Hands-on experience with Databricks workspace and notebooks.
- Proficiency in Python and SQL.
- Experience with Delta Lake, Spark SQL, and Spark Optimization techniques.
- Knowledge of data warehousing concepts and ETL frameworks.
- Experience with cloud platforms such as Azure, AWS, or GCP.
- Understanding of data modeling and big data technologies.
- Experience with version control tools like Git.
📌 PySpark+ Databricks (Bengaluru)
🏢 Infosys
📍 Bengaluru