01 Sep
|
Birlasoft
|
Pune
Area(s) of responsibility
Key Responsibilities
•
- Design and develop scalable ETL/ELT pipelines using PySpark, Spark SQL, Python, and Databricks.
- Develop complex transformation logic across Bronze, Silver, and Gold layers.
- Implement reusable transformation frameworks and common data processing utilities.
- Optimize transformation workloads for performance, scalability, and cost.
- Apply data quality, validation, and business-rule checks throughout the pipeline.
- Enterprise Data Modelling
- Design scalable data models using:
- Dimensional modelling
- Star/Snowflake schemas
- Data Vault 2.0
- Canonical data models
- Develop curated data products and analytical datasets using Databricks SQL and Delta Lake.
- Support enterprise-wide data harmonization across clinical, regulatory, safety, and commercial domains.
- Experience supporting data platforms across:
- Clinical trials – EDC, CDMS, CTMS
- Regulatory and submission systems
- Pharmacovigilance and safety
- Commercial analytics
- Real-World Evidence (RWE)
- Ensure adherence to relevant regulatory and data governance requirements including GxP, 21 CFR Part 11, HIPAA/GDPR, and ALCOA+ principles.
Required Qualifications
- 10–15 years of experience in Data Engineering and enterprise data platforms.
- Solid hands-on experience with Databricks, Apache Spark, PySpark, Python, and SQL.
- 4+ years of experience with Databricks/Spark-based data engineering preferred.
- Strong experience in:
- Data migration and modernization
- ETL/ELT development
- Batch and streaming data pipelines
- Delta Lake
- Databricks Workflows
- Unity Catalog
- Data modelling
- Data quality and reconciliation
- Experience with cloud platforms – AWS/Azure/GCP.
- Strong understanding of CI/CD, Git, Jenkins, and DevOps practices.
- Experience working in Life Sciences/Pharma environments preferred.
📌 Technical Lead-Data Engg (Pune)
🏢 Birlasoft
📍 Pune