22 Aug
|
Innovaccer
|
Noida
Job Summary
Engineering at Innovaccer
With every line of code, we accelerate our customers success, turning complex challenges into innovative solutions. Collaboratively, we transform each data point we gather into valuable insights for our customers. Join us and be part of a team thats turning dreams of better healthcare into reality, one line of code at a time. Together, were shaping the future and making a meaningful impact on the world.
Role Overview
As a Senior Software Engineer on the Lakehouse team, you will build and operate the data pipelines at the heart of Innovaccers on-premise platform: Spark ingestion jobs landing raw healthcare data into Apache Iceberg, and Trino SQL transforms building the layered tables thatpower analytics and applications. You will work hands-on across the full pipeline surface, from file validation and quarantine at ingestion to query performance and table health in serving.
Responsibilities
- Build Spark ingestion jobs that land high-volume raw files into Iceberg tables with schema handling, bad-record quarantine, and idempotent batch replay.
- Develop and operate Trino SQL transform pipelines across data layers: validation and typing, business-rule transforms, MERGE-based deduplication, and aggregate builds.
- Port existing warehouse SQL workloads to Trino and Spark SQL dialects, and validate results against source outputs.
- Automate Iceberg table maintenance: compaction, snapshot expiry,
and orphan-file cleanup as scheduled workflows.
- Tune query and pipeline performance: partitioning strategy, file sizing, statistics, and resource-group behavior.
- Instrument pipelines with data-quality checks, reconciliation reports, and alerting, and participate in per-dataset validation during rollout phases.
Qualifications
- B.E., B.Tech., M.Sc. degree in Computer Science or a related technical field.
- 5+ years of data engineering experience building production pipelines at scale.
- Strong SQL skills and hands-on experience with Apache Spark for batch processing.
- Experience with Trino/Presto (or a comparable distributed SQL engine) and open table formats: Iceberg preferred, Delta Lake or Hudi acceptable.
- Working knowledge of S3-compatible object storage and columnar file formats (Parquet).
- Experience with workflow orchestration tools (Airflow or equivalent) and CI/CD for data pipelines.
- Career development experience with Python and/or Java.
- Healthcare data formats (HL7, CCDA, claims and regulated-environment experience are pluses).
Skills
- Apache Spark
- SQL
- Trino/Presto
- Iceberg
- Delta Lake
- Parquet
- Airflow
- Python
- Java
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Software Development Engineer-III (Data Engineer) (Noida)
🏢 Innovaccer
📍 Noida