26 Aug
|
Tekskills
|
Chennai
Ro
- We are seeking a highly experienced Data Engineer with over 10+ years of expertise in Data Engineering, Databricks, Azure, SQL, PySpark, and Python.
- The ideal candidate will collaborate with Business and functional teams and create the data pipelines, apply CICD and DevOps practices, and optimize query performance and data transformations.
- This role requires a deep understanding of data governance, quality, and access control policies, as well as the ability to support production data pipelines and contribute to architectural decisions.
Key Responsibilities:
- Apply CICD and DevOps practices to automate data workflows and deployments (e.g., with GitHub Actions, Jenkins, Terraform).
- Optimize query performance and data transformations using advanced SQL.
- Implement and uphold data governance, quality, and access control policies.
- Support production data pipelines and respond to issues and performance bottlenecks.
- Contribute to architectural decisions around data strategy and platform scalability.
- Develop and maintain Databricks notebooks and jobs for large-scale batch and streaming data processing.
- Write modular, production-grade PySpark and Python code, including reusable functions and libraries for data transformation.
- Implement streaming data ingestion and Structured Streaming in Databricks for near real-time data solutions.
- Apply performance tuning techniques in Spark, including job optimization, caching, and partitioning strategies.
- Utilize data quality frameworks and testing practices.
- Manage data governance, access controls, and lineage tracking using Unity Catalog.
- Requirements: Proven expertise in Databricks, Delta Lake, and Apache Spark (PySpark preferred).
- Deep understanding of Unity Catalog for fine-grained data governance and lineage tracking.
- Proficiency in SQL for large-scale data manipulation and analysis.
- Solid understanding of CICD, infrastructure automation, and DevOps principles.
- Familiarity with cloud platforms (Azure) and data services.
- Strong problem-solving, debugging, and system design skills.
- Excellent communication and collaboration abilities in cross-functional teams.
- Solid hands-on experience with Delta Lake, including table management, schema evolution, and implementing ACID-compliant pipelines.
- Knowledge of performance tuning techniques in Spark, including job optimization, caching, and partitioning strategies.
Basic understanding of Unity Catalog for managing data governance, access controls, and lineage tracking from a developers perspective. le & responsibilities
Preferred candidate profile
📌 Data Engineer (Chennai)
🏢 Tekskills
📍 Chennai