Data Engineer with 5+ years of experience:
We are seeking a Data Engineer with strong dual competency in data and life science to build
and scale data pipelines and curated datasets that power scientific analytics, AI/ML, and data
products across R&D.; You will work in a multidisciplinary environment, partnering closely with
scientists, data scientists, data architects, and product owners to translate scientific workflows
into reliable, reusable, and governed data assets.
This role is hands-on and delivery-focused: you will design, develop, and operate data ingestion
and transformation pipelines, optimize performance and reliability, and ensure data is
discoverable and trusted through strong metadata, lineage, and access controlswithin the
scientific data and AI ecosystem for development and RWE assets.
Key Responsibilities:
Build & Operate Data Pipelines:
Design and implement scalable ingestion and transformation pipelines for diverse
scientific datasets (structured and unstructured), using Databricks as a core execution
workplace.
Develop robust ETL/ELT workflows, including batch and (where relevant) incremental
processing, with production-grade practices (testing, monitoring, reliability).
Collaborate with platform and security teams to ensure pipelines follow access,
governance, and operational standards.
Build and maintain data structures and curated layers to support analytics, reporting,
and downstream data products.
Optimize performance and cost through good engineering practices (e.g., efficient
transformations, workload tuning, data lifecycle considerations).
Scientific Context & Business:
Work effectively with researchers and scientists to understand workflows and ensure
pipelines preserve scientific meaning, context, and traceability.
Support analytical environments (e.g., Power BI, Tableau) by delivering data in formats and
structures fit for scientific exploration and enterprise connectivity.
Data Quality, Metadata, Lineage
📌 Exciting Opportunity (Delhi)
🏢 Excelra
📍 Delhi