09 Aug
|
NPCI
|
Hyderabad
The Opportunity
To design and implement real-time data streaming pipelines ensuring 24×7 availability of high-volume transactional data.
Your role includes building fault-tolerant, scalable systems using modern big data technologies, integrating data from multiple sources into data lakes/lakehouses, and enabling downstream analytics. You will collaborate with internal teams (Data Science, Product, and Technology) to deliver reliable data solutions, optimize performance, and maintain compliance with NPCI standards.
Job Details
- Job Title: Data Engineer (Streaming)
- Division: Data Analytics
- Years of Experience: 3–8 years
- Education: Graduation in Computer Science/IT (preferably BE/B.Tech) or equivalent; advanced degrees are a plus
- Employment Type: Full-time, Permanent
- Location: Mumbai & Hyderabad
Key Responsibilities
- Design and develop real-time data pipelines using Apache Kafka and stream processing frameworks (Spark Structured Streaming / Apache Flink).
- Ensure 24×7 data availability with fault-tolerant, highly reliable systems.
- Implement ingestion, transformation, and load (ITL) patterns for data lakes/lakehouses (S3/MinIO, HDFS).
- Work with table formats like Iceberg, Hudi, or Paimon for ACID transactions and schema evolution.
- Optimize SQL queries on Trino/Hive for large-scale analytics.
- Develop orchestration workflows using DBT, Dagster, or Airflow for data transformations.
- Write productive code in Python, Scala, and Java for data processing and automation.
- Collaborate with cross-functional teams to understand requirements and deliver high-quality data solutions.
- Monitor pipeline health, manage checkpoints, and implement observability for streaming jobs.
- Ensure compliance with security, governance, and audit standards.
Requirements
Key Skills and Experience Required
- Mandatory Technical Skills:
- Apache Kafka (topics, partitions, offsets, reliability)
- Stream processing (Spark Structured Streaming or Apache Flink)
- Preferred Technical Skills:
- Data Lake / Lakehouse (S3, MinIO, HDFS)
- Table formats: Iceberg, Hudi, Paimon
- SQL engines: Trino, Hive
- Orchestration tools: DBT, Dagster, Airflow
- Programming: Python, Scala, Java
- NoSQL databases (MongoDB, Cassandra, Redis)
- CI/CD (Jenkins, GitHub Actions), Linux basics
- BI tools: Superset, Tableau
- Data Quality frameworks and observability practices
- Other Requirements:
- Strong analytical and problem-solving skills
- Ability to work in 24×7 environments
- Quick adaptability to new technologies
📌 Senior Associate Data Engineering (Hyderabad)
🏢 NPCI
📍 Hyderabad