28 Sep
|
Infosys
|
Chennai
Technology->Analytics - Solutions->SQL Server - Analytics Technology->Big Data - Data Processing->PySpark Technology->Data Engineering->Databricks
Key Responsibilities
- Lead the development of end-to-end data pipelines on Databricks using PySpark for batch and incremental processing.
- Design scalable data models and curated datasets to support analytics and downstream consumption.
- Write and optimize advanced SQL for transformations, validations, and performance-critical queries.
- Implement robust data quality checks, reconciliation logic, and monitoring to ensure trusted datasets.
- Tune Spark jobs for performance and cost efficiency (partitioning, caching, file formats, cluster sizing).
- Establish coding standards, reusable frameworks, and review practices to improve maintainability.
- Collaborate with stakeholders to translate requirements into technical designs and delivery plans.
- Troubleshoot production issues, perform root-cause analysis, and drive preventive improvements.
- Mentor team members and provide technical guidance across design, implementation, and optimization.
Minimum
Qualifications:
- BTECH, MTECH, MCA, or MSC in Computer Science, Information Technology, or a related field.
- 6–8 years of overall experience in data engineering or large-scale data processing roles.
- Strong hands-on experience with PySpark for distributed data processing and transformation logic.
- Strong hands-on experience with Databricks for building, running, and managing data workloads.
- Proficiency in Advanced SQL including complex joins, window functions, and query optimization.
- Experience building reliable pipelines with robust focus on data quality, performance, and stability Preferred Qualifications:
- Experience designing lakehouse-style architectures and organizing curated layers for analytics readiness.
- Strong experience with Spark optimization techniques and handling large-scale datasets efficiently.
- Ability to build reusable PySpark utilities/frameworks for ingestion, transformation, and validation patterns.
- Experience with orchestration and scheduling approaches for dependable pipeline execution and recovery.
- Proven track record of technical leadership: mentoring, conducting reviews, and driving engineering best practices.
📌 Databricks, Pyspark (Chennai)
🏢 Infosys
📍 Chennai