HiLabs is an AI-driven healthcare technology company solving one of the biggest challenges in the US healthcare industry, dirty data. Our AI-powered platform processes billions of healthcare records and helps payers, providers, and life sciences organizations unlock accurate insights at scale.
At HiLabs, we combine Artificial Intelligence, Data Engineering, and Healthcare expertise to build innovative products that transform healthcare data quality and enable better business and patient outcomes.
Role Overview
We are looking for a passionate and experienced Senior/Lead Data Engineer to build and scale modern data platforms that power AI and analytics solutions. You will work on designing high-performance data pipelines, processing large-scale datasets, and building reliable cloud-native data solutions.
You will collaborate closely with Data Scientists, Product Managers, and Software Engineers to deliver scalable and production-ready data engineering solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL and ELT pipelines for large-scale data processing.
- Build reliable and efficient data ingestion, transformation, and orchestration workflows.
- Develop high-performance data pipelines using Python and Apache Spark.
- Optimize SQL and NoSQL databases for performance, scalability, and reliability.
- Work with cloud platforms such as AWS, GCP, or Azure to build cloud-native data solutions.
- Build and optimize data warehouse solutions using Snowflake, Redshift, BigQuery, or similar technologies.
- Implement data quality checks, validation frameworks, monitoring, and alerting.
- Collaborate with cross-functional teams to support AI, analytics, and business reporting requirements.
- Improve pipeline performance, scalability, and operational efficiency.
- Follow best practices for data governance, security, version control, and CI/CD.
Required Skills
- 5 to 10 years of hands-on experience in Data Engineering.
- Strong experience building production-grade ETL and ELT pipelines.
- Excellent programming skills in Python. Experience with Scala or Java is a plus.
- Strong SQL skills and experience with relational and NoSQL databases.
- Hands-on experience with Apache Spark (PySpark or Scala Spark).
- Experience with workflow orchestration tools such as Apache Airflow.
- Experience with cloud platforms such as AWS, GCP, or Azure.
- Good understanding of cloud data services such as S3, Redshift, BigQuery, EMR, Glue, or equivalent.
- Experience with data warehousing solutions like Snowflake, Redshift, or BigQuery.
- Familiarity with Git, CI/CD pipelines, and Agile development practices.
- Strong analytical, debugging, and problem-solving skills.
Preferred Skills
- Experience with Hadoop, Kafka, or other distributed data technologies.
- Exposure to MLOps or AI data pipelines.
- Experience working in a product-based or SaaS company.
- Healthcare or healthcare analytics experience is an added advantage.
Why Join HiLabs?
- Build AI-powered healthcare products that process billions of healthcare records.
- Work on large-scale data engineering challenges using modern cloud technologies.
- Collaborate with highly talented engineers, data scientists, and AI researchers.
- Fast-paced product engineering workplace with significant ownership and learning opportunities.
- Competitive compensation, career growth, transport facility, and complimentary meals.
📌 Lead Data Engineer (Pune)
🏢 HiLabs
📍 Pune