Overall 5+ YRS of experience in Data Engineering
Advanced Python, SQL, and Apache Spark/PySpark; solid understanding of distributed processing and performance tuning.
Production Databricks experience with Auto Loader, Delta Lake, Workflows, Unity Catalog, notebooks/jobs, SQL Warehouses, and cluster/job optimization.
Hands-on AWS experience with S3, IAM, VPC, ECS/EKS or Lambda, Step Functions, EventBridge, CloudWatch, Secrets Manager, and KMS.
Practical Apache Iceberg knowledge, including tables, catalogs, snapshots, schema/partition evolution, compaction, and interoperability with query engines.
Experience with Git, pull requests, CI/CD, Docker, Terraform/CloudFormation, automated testing, observability, and incident response.
Experience processing semi-structured data and building resilient ingestion pipelines with quality controls, error handling, and replay capability.
Healthcare data knowledge: FHIR R4 and/or HL7 v2, OMOP CDM, clinical terminologies, PHI, HIPAA-aligned engineering controls, and de-identification concepts.
Demonstrated responsible use of AI coding tools and the ability to critically review, test, and productionize generated code.
📌 Senior Data Engineering Lead Chennai (India)
🏢 Optum
📍 India