HiLabs builds AI-driven data solutions for US healthcare payers, solving large-scale data quality problems across provider, claims and clinical datasets. Our engineering teams in Pune build the platform that powers this at production scale.
About the Role
We are hiring a Software Engineer II to build scalable, reliable, production-grade backend services for our core data platform. The work spans Python backend engineering, distributed data processing with PySpark, event-driven systems on Kafka and Redis, and ML inference integration — all within a secure AWS environment. You will work closely with Data Science, Data Engineering, DevOps/SecOps and Product teams.
What You Will Do
- Design, build and maintain Python backend services and microservices using FastAPI/Flask.
- Build scalable PySpark pipelines for ingestion, classification, linkage, feature engineering and schema generation.
- Develop event-driven and asynchronous services using Kafka and Redis Streams — producers, consumers, consumer groups, topics, partitioning, offset management and retry handling.
- Use Redis for caching, rapid lookups, TTL-based state and distributed coordination.
- Build and deploy containerized services on AWS EKS / Kubernetes.
- Integrate with AWS S3, EMR, EventBridge, PostgreSQL, Snowflake, Kafka and Redis.
- Deploy batch scoring and ML inference services using versioned model artifacts.
- Implement reliability patterns: idempotency, retries, timeouts, dead-letter handling and failure recovery.
- Design efficient database schemas, indexes, queries and data-access layers.
- Implement logging, metrics,
monitoring and alerting; write unit and integration tests; participate in code reviews.
- Troubleshoot across application, data, messaging and infrastructure layers.
Must Have
- 3–5 years of hands-on backend software engineering experience.
- B.E./B.Tech/M.Tech/MCA in Computer Science or a related field from a Tier-1 institute (IIT, NIT, BITS, IIIT or equivalent). This is a mandatory requirement for this position.
- Strong Python — OOP, data structures, design principles.
- PySpark / Apache Spark and distributed data processing at scale.
- Apache Kafka — producers, consumers, consumer groups, partitions and offset management.
- Redis — caching, TTL, Pub/Sub or Streams.
- FastAPI or Flask, REST APIs, microservices.
- AWS — S3, IAM, CloudWatch, and EKS or EMR.
- Docker, with working knowledge of Kubernetes.
- PostgreSQL / SQL — schema design, indexing, query optimization.
- Git, CI/CD, unit and integration testing.
Good to Have
- Amazon MSK or Kafka on Kubernetes; Redis Cluster or ElastiCache.
- AWS EMR and Spark at scale; Snowflake.
- Parquet and schema management.
- ML inference, model serving, embeddings, risk-scoring or MLflow.
- GPU-based workloads.
- Healthcare, claims or clinical data; PHI and HIPAA awareness.
Ideal Candidate A strong backend and distributed-systems engineer comfortable across application, data and infrastructure layers, with real ownership of production services — debugging depth, system-design fundamentals, and the instinct to build fault-tolerant, observable systems.
Location: Pune, Kharadi (on-site). Candidates open to relocating to Pune may apply.
📌 Software Engineer II - Backend & Data Platform (Python, Kafka, PySpark) (Pune)
🏢 HiLabs
📍 Pune