03 Oct
|
STOCK BAZAAR
|
India
03 Oct
STOCK BAZAAR
India
Data pipelines and streaming
• Design, build and maintain real-time and batch data pipelines for market data ingestion
• Build streaming pipelines using Apache Kafka — topics, partitioning, consumer groups, schema management
• Write distributed processing jobs in PySpark for large-scale transformation and aggregation
• Ensure pipelines are fault-tolerant, idempotent and recover cleanly from failures
Data lake and storage
• Design and maintain the data lake, including partitioning strategy, file formats and retention
• Model data for both analytical querying and low-latency platform reads
• Manage schema evolution without breaking downstream consumers
Data quality and reliability
• Build validation and reconciliation checks so bad or missing market data is caught before it reaches subscribers
• Monitor pipeline health, set up alerting, and own incident resolution and root-cause analysis
• Maintain lineage and documentation for critical datasets
DevOps and deployment
• Own deployment of data services using Docker and Kubernetes
• Build and maintain CI/CD pipelines for data jobs
• Manage cloud infrastructure (AWS) for data workloads, with an eye on cost
• Work with the DevOps team on infrastructure-as-code and observability
Requirements
Requirements
• 3+ years in data engineering or a closely related role
• Apache Kafka — production experience with streaming pipelines
• PySpark — batch and streaming, with an understanding of how to tune a job
• Data lake architecture — partitioning, formats such as Parquet, and query performance
• DevOps practices — Docker, Kubernetes, CI/CD
• Solid Python and SQL
• Cloud platform experience, preferably AWS (S3, EMR, Glue or equivalent)
• Bachelor's degree in Computer Science, Information Technology or related, or equivalent practical experience
📌 Data Enginner (India)
🏢 STOCK BAZAAR
📍 India