03 Sep
|
Tata Consultancy Services
|
Gandhinagar
03 Sep
Tata Consultancy Services
Gandhinagar
Job Title: Data Engineer - Databricks (AWS) | PySpark | Delta Lake
Location : Bengaluru, Noida, Gandhi Nagar, Pune, Chennai
Experience : 4+ Years
Job Summary
We are looking for a talented Data Engineer with robust expertise in Databricks on AWS, PySpark, Delta Lake, and modern data engineering practices . The ideal candidate will be responsible for building scalable data pipelines, implementing data governance frameworks, optimizing platform performance, and supporting enterprise analytics initiatives within the Healthcare/Life Sciences domain.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines using Databricks on AWS.
- Build and support batch and streaming data ingestion frameworks.
- Implement data quality, governance, lineage, and security controls using Unity Catalog and Databricks best practices.
- Develop and optimize ETL/ELT processes using PySpark and SQL.
- Design and implement Medallion Architecture (Bronze, Silver, Gold) data solutions.
- Build and maintain CDC-based ingestion pipelines using AWS and Databricks services.
- Optimize Spark performance, cluster utilization, and platform costs.
- Implement and support Delta Lake features to ensure data reliability and governance.
- Manage Structured Streaming, Auto Loader, and real-time data processing requirements.
- Develop CI/CD processes using Git, Databricks Asset Bundles, and deployment automation tools.
- Monitor production environments, troubleshoot issues, and ensure SLA compliance.
- Collaborate with business stakeholders, analysts,
and engineering teams to deliver high-quality data solutions.
- Mentor junior engineers and contribute to data engineering best practices.
Required Skills & Qualifications
Technical Skills
- 4+ years of experience in Data Engineering.
- 4+ years of hands-on experience with Databricks on AWS.
- Strong expertise in PySpark and Spark optimization techniques.
- Experience with Delta Lake and Lakehouse architecture.
- Strong SQL programming skills.
- Experience with Unity Catalog and Databricks governance frameworks.
- Knowledge of dimensional data modeling and Medallion Architecture.
- Hands-on experience with:
- Structured Streaming
- Auto Loader
- Change Data Capture (CDC)
- Experience with AWS services:
- Amazon S3
- IAM
- AWS DMS
- Kinesis and/or MSK
- AWS Glue
- Experience with Git, CI/CD pipelines, and Databricks Asset Bundles.
- Experience supporting production environments, performance tuning, and cost optimization.
Domain Experience
- Healthcare or Life Sciences domain experience preferred.
- Understanding of data governance, analytics, privacy, and regulatory compliance requirements.
Education
- Bachelor's or Master's Degree in Computer Science, Information Technology, Engineering, or related field.
- Flexible for candidates with strong relevant experience.
Certifications (Preferred)
- Databricks Certified Data Engineer Associate/Professional
- AWS Certified Data Engineer
- AWS Certified Solutions Architect
- Microsoft Azure Certified Developer
- Google Professional Cloud Developer
📌 AWS Data Engineer (Gandhinagar)
🏢 Tata Consultancy Services
📍 Gandhinagar