02 Sep
|
Birlasoft
|
Pune
Area(s) of responsibility
Key Responsibilities
•
Design and develop scalable ETL/ELT pipelines using PySpark, Spark SQL, Python, and Databricks.
Develop complex transformation logic across Bronze, Silver, and Gold layers.
Implement reusable transformation frameworks and common data processing utilities.
Optimize transformation workloads for performance, scalability, and cost.
Apply data quality, validation, and business-rule checks throughout the pipeline.
Enterprise Data Modelling
Design scalable data models using:
Dimensional modelling
Star/Snowflake schemas
Data Vault 2.0
Canonical data models
Develop curated data products and analytical datasets using Databricks SQL and Delta Lake.
Support enterprise-wide data harmonization across clinical, regulatory, safety, and commercial domains.
Experience supporting data platforms across:
Clinical trials – EDC, CDMS, CTMS
Regulatory and submission systems
Pharmacovigilance and safety
Commercial analytics
Real-World Evidence (RWE)
Ensure adherence to relevant regulatory and data governance requirements including GxP, 21 CFR Part 11, HIPAA/GDPR, and ALCOA+ principles.
Required Qualifications
10–15 years of experience in Data Engineering and enterprise data platforms.
Solid hands-on experience with Databricks, Apache Spark, PySpark, Python, and SQL.
4+ years of experience with Databricks/Spark-based data engineering preferred.
Solid experience in:
Data migration and modernization
ETL/ELT development
Batch and streaming data pipelines
Delta Lake
Databricks Workflows
Unity Catalog
Data modelling
Data quality and reconciliation
Experience with cloud platforms – AWS/Azure/GCP.
Robust understanding of CI/CD, Git, Jenkins, and DevOps practices.
Experience working in Life Sciences/Pharma environments preferred.
📌 Technical Lead Data Engg Pune
🏢 Birlasoft
📍 Pune