08 Aug
|
Excelra
|
Hyderabad
Data Engineer with 5+ years of experience:
We are seeking a Data Engineer with strong dual competency in data and life science to build
and scale data pipelines and curated datasets that power scientific analytics, AI/ML, and data
products across R&D.; You will work in a multidisciplinary environment, partnering closely with
scientists, data scientists, data architects, and product owners to translate scientific workflows
into reliable, reusable, and governed data assets.
This role is hands-on and delivery-focused: you will design, develop, and operate data ingestion
and transformation pipelines, optimize performance and reliability, and ensure data is
discoverable and trusted through strong metadata, lineage, and access controlswithin the
scientific data and AI ecosystem for development and RWE assets.
Key Responsibilities:
Build & Operate Data Pipelines:
Design and implement scalable ingestion and transformation pipelines for diverse
scientific datasets (structured and unstructured), using Databricks as a core execution
environment.
Develop robust ETL/ELT workflows, including batch and (where relevant) incremental
processing, with production-grade practices (testing, monitoring, reliability).
Collaborate with platform and security teams to ensure pipelines follow access,
governance, and operational standards.
Build and maintain data structures and curated layers to support analytics, reporting,
and downstream data products.
Optimize performance and cost through good engineering practices (e.g., efficient
transformations, workload tuning, data lifecycle considerations).
Scientific Context & Business:
Work effectively with researchers and scientists to understand workflows and ensure
pipelines preserve scientific meaning, context, and traceability.
Support analytical environments (e.g., Power BI, Tableau) by delivering data in formats and
structures fit for scientific exploration and enterprise connectivity.
Data Quality, Metadata, Lineage & Governance:
Embed data quality checks and validation steps into pipelines to ensure data is reliable
and fit for scientific decision-making, and follow data strategy.
Implement and maintain metadata and lineage practices so that datasets are
discoverable and governed (e.g., cataloging, traceability, access controls).
Contribute to governance-by-design by working with architects and governance leads on
standards and reusable patterns.
Collaboration & Delivery:
Partner with Data Architects and product teams to implement agreed data models and
patterns in production pipelines.
Support sprint delivery by estimating, implementing, and continuously improving
pipelines and datasets based on feedback and evolving use cases.
Governance, Quality and Alignment:
Ensure products comply with data governance, quality, privacy, and compliance
requirements in collaboration with Data Governance and Architecture teams.
Align product decisions with architectural standards and platform capabilities, balancing
short-term needs with long-term sustainability.
Promote reuse and consistency across R&D; data products.
Qualifications & Experience:
Bachelors or Master’s degree (or equivalent experience) in Computer Science, Data
Engineering, Bioinformatics, Biomedical Sciences, Life Sciences, or a related field.
5+ years in data engineering / data platform engineering (or equivalent).
Strong hands-on experience with operating data pipelines in a production workplace.
Proven ability to work across scientific and technical domains collaborating credibly with
scientists and technical teams.
Experience working with clinical and RWE datasets and workflows.
Experience with metadata/catalog and lineage practices in enterprise environments.
Required Skills & Competencies:
Data pipeline engineering, transformation design, performance optimization.
Expertise with ETL/ELT, collaborative engineering workflows like Snowflake or
Databricks.
Experience with data modelling, data curation.
Experience with data visualization and analytics.
Strong SQL + Python (and comfort supporting scientific tooling needs such as R where
relevant).
Expertise in FAIR data principles, meta data, data catalogues...
Previous experience with semantic, ontology and knowledge graph is a plus.
Experience working in regulated environments (e.g. GxP, data privacy) is a plus.
📌 Exciting Opportunity - Data Engineer (Life Sciences) (Hyderabad)
🏢 Excelra
📍 Hyderabad