Background: Bioinformatics / Life Sciences preferred
Role Summary
We are looking for a Data Engineer with strong Databricks experience to build and integrate clinical, biomarker and multi-omics datasets into a governed, analysis-ready data platform. The candidate should have exposure to Bioinformatics/Life Sciences data and experience developing scalable data pipelines on Databricks.
Key Responsibilities
- - Build Bronze → Silver → Gold data pipelines on Databricks.
- Develop ingestion and transformation pipelines using PySpark, Python and SQL.
- Build and maintain Delta Lake tables and analytical datasets.
- Integrate and harmonize structured and semi-structured scientific datasets.
- Implement data-quality, validation and reconciliation rules.
- Develop incremental loading, exception handling and reprocessing mechanisms.
- Support Unity Catalog, RBAC, metadata and data lineage.
- Work with the Data Architect and Bioinformatics SME on domain-specific transformations and data models.
- Develop Gold-layer curated datasets for analytics and downstream AI use cases.
- Support testing, UAT, documentation and production readiness.
Must Have
- - 4 years of Data Engineering experience.
- Robust hands-on Databricks, PySpark, Python and SQL.
- Spark, Delta Lake and Medallion Architecture.
- ETL/ELT and data-pipeline development.
- Data modelling and data-quality implementation.
- Cloud object storage, preferably AWS S3.
- Git and CI/CD experience.
- Bioinformatics/Life Sciences data exposure.
Preferred
- - Experience with genomics, biomarker or multi-omics datasets.
- Exposure to RNA-seq, WES/WGS, VCF, CNV or similar scientific datasets.
- Unity Catalog and Databricks governance.
- Databricks certification.
📌 Data Engineer (Hyderabad)
🏢 Clovertex
📍 Hyderabad
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.