Role & responsibilities
8+ years experiences on relevant field below (Internship, prototype, and personal projects won't be counted)
Coding is required . (Ideally Python or Java)
Own end to end lifecycle (From development to deployment to production setting)
- Experience in building or deploying solution in the cloud.
- Either Cloud Native (Serverless) : S3, Lambda, AWS Batch, ECS
- Or Cloud Agnostic: Kubernetes, Helm Chart, ArgoCD, Prometeus, Grafana.
- CICD experience: Github action or Jenkin.
- Infrastructure as code : e.g., Terraform
And experience in at least one of this focus area:
- Big Data: Building Big data pipeline or Platform to process petabytes of data: (PySpark, Hudi, Data Lineage, AWS Glue, AWS EMR, Kafka, Schema Registry)
- Or GraphDB : Ingesting and consuming data in Graph Database such as Neo4J, AWS Neptune, JanusGraph or DGraph
Preferred candidate profile
- Specifically highlight Kafka expertise - include details like:
- Experience with Kafka cluster management and configuration
- Stream processing with Kafka Streams or KSQL
- Schema Registry implementation and management
- Kafka Connect for data integration
- Put significant focus on PySpark skills:
- Experience building and optimizing PySpark jobs for batch processing
- Stream processing with Spark Structured Streaming
- Familiarity with Delta Lake, Hudi, or Iceberg for lakehouse implementation
- Highlight data engineering skills that complement these technologies:
- Data pipeline design and implementation
- Experience with data quality, validation, and lineage tracking
- Performance optimization for large-scale data processing
📌 Data Engineer (Pune)
🏢 Calsoft
📍 Pune