Responsibilities
- Build reliable batch, micro-batch, and event-driven pipelines on AWS and Databricks
- Develop reusable ingestion frameworks for REST APIs, FHIR Bulk Export, HL7 interfaces, databases, SFTP/file exchange, JSON/NDJSON, CSV, XML, PDFs, and clinical text
- Implement scalable Spark/PySpark and SQL transformations, including schema inference/evolution, checkpointing, idempotency, retries, backfills, and replay
- Design and maintain Delta Lake and Apache Iceberg tables, including physical design, partitioning/clustering, compaction, file sizing, incremental reads/writes, performance tuning, and cost optimization
- Build curated healthcare data models and transformations using FHIR, HL7, OMOP CDM, and clinical terminology mappings
- Implement data-quality frameworks: schema validation, referential-integrity checks, business-rule testing, anomaly detection, reconciliation, completeness checks, and data-quality observability
- Build metadata and lineage capture across source systems, pipeline runs, code versions, transformation rules, mappings, and published data products
- Implement privacy-aware data processing for PHI, including access controls, masking, tokenization/pseudonymization, de-identification, and auditable handling patterns
- Deliver CI/CD pipelines, automated unit/integration/data tests, Terraform or CloudFormation, containerized services, monitoring, alerting, runbooks, and incident-recovery procedures
- Integrate and optimize governed data consumption through Databricks SQL, Athena, Snowflake, Trino, or equivalent engines
- Comply with all applicable Company policies, procedures, and business directives,
changes including those relating to work location, team assignments, work schedules, and flexible work arrangements
Eligibility
- The candidate should have completed 12 months in the current role
- The candidate should not be on any active CAP/ PIP
- The performance review of the candidate must be ME Above in the last common review
Required Qualifications
- Graduate degree or equivalent experience
- 5+ years of production data-engineering experience
- Hands-on AWS experience with S3, IAM, VPC, ECS/EKS or Lambda, Step Functions, EventBridge, CloudWatch, Secrets Manager, and KMS
- Experience with Git, pull requests, CI/CD, Docker, Terraform/CloudFormation, automated testing, observability, and incident response
- Experience processing semi-structured data and building resilient ingestion pipelines with quality controls, error handling, and replay capability
- Production Databricks experience with Auto Loader, Delta Lake, Workflows, Unity Catalog, notebooks/jobs, SQL Warehouses, and cluster/job optimization
- Advanced Python, SQL, and Apache Spark/PySpark; robust understanding of distributed processing and performance tuning
- Practical Apache Iceberg knowledge, including tables, catalogs, snapshots, schema/partition evolution, compaction, and interoperability with query engines
- Healthcare data knowledge: FHIR R4 and/or HL7 v2, OMOP CDM, clinical terminologies, PHI, HIPAA-aligned engineering controls, and de-identification concepts
- Demonstrated responsible use of AI coding tools and the ability to critically review, test, and productionize generated code
- Healthcare or regulated-data experience
📌 Senior Data Engineering Lead (Pune)
🏢 Optum
📍 Pune