- Design, build and maintain real-time and batch data pipelines for market data ingestion
- Build streaming pipelines using Apache Kafka — topics, partitioning, consumer groups, schema management
- Write distributed processing jobs in PySpark for large-scale transformation and aggregation
- Ensure pipelines are fault-tolerant, idempotent and recover cleanly from failures
Data lake and storage
- Design and maintain the data lake, including partitioning strategy, file formats and retention
- Model data for both analytical querying and low-latency platform reads
- Manage schema evolution without breaking downstream consumers
Data quality and reliability
- Build validation and reconciliation checks so bad or missing market data is caught before it reaches subscribers
- Monitor pipeline health, set up alerting, and own incident resolution and root-cause analysis
- Maintain lineage and documentation for critical datasets
DevOps and deployment
- Own deployment of data services using Docker and Kubernetes
- Build and maintain CI/CD pipelines for data jobs
- Manage cloud infrastructure (AWS) for data workloads, with an eye on cost
- Work with the DevOps team on infrastructure-as-code and observability
Requirements
Requirements
- 3+ years in data engineering or a closely related role
- Apache Kafka — production experience with streaming pipelines
- PySpark — batch and streaming, with an understanding of how to tune a job
- Data lake architecture — partitioning, formats such as Parquet, and query performance
- DevOps practices — Docker, Kubernetes, CI/CD
- Solid Python and SQL
- Cloud platform experience, preferably AWS (S3, EMR, Glue or equivalent)
- Bachelor's degree in Computer Science, Information Technology or related, or equivalent practical experience
📌 Data Enginner (Delhi)
🏢 STOCK BAZAAR
📍 Delhi
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.