- Overall 5+ YRS of experience in Data Engineering
- Advanced Python, SQL, and Apache Spark/PySpark; solid understanding of distributed processing and performance tuning.
- Production Databricks experience with Auto Loader, Delta Lake, Workflows, Unity Catalog, notebooks/jobs, SQL Warehouses, and cluster/job optimization.
- Hands-on AWS experience with S3, IAM, VPC, ECS/EKS or Lambda, Step Functions, EventBridge, CloudWatch, Secrets Manager, and KMS.
- Practical Apache Iceberg knowledge, including tables, catalogs, snapshots, schema/partition evolution, compaction, and interoperability with query engines.
- Experience with Git, pull requests, CI/CD, Docker, Terraform/CloudFormation, automated testing, observability, and incident response.
- Experience processing semi-structured data and building resilient ingestion pipelines with quality controls, error handling, and replay capability.
- Healthcare data knowledge: FHIR R4 and/or HL7 v2, OMOP CDM, clinical terminologies, PHI, HIPAA-aligned engineering controls, and de-identification concepts.
- Demonstrated responsible use of AI coding tools and the ability to critically review, test, and productionize generated code.
📌 Senior Data Engineering Lead (Chennai)
🏢 Optum
📍 Chennai