20 Aug
|
Passageway Tech
|
Jaipur
20 Aug
Passageway Tech
Jaipur
Senior Data Engineer – CDC, Schema Independence & GenAI Experience: 5–8 Years
Project: Enterprise Data Platform (EDP) – Smart Metering & Energy
Role Overview We are looking for a strong Senior Data Engineer to work on a large-scale Enterprise Data Platform for the smart metering/energy domain.
The role will primarily own Change Data Capture (CDC), schema-independent ingestion and metadata-driven data pipelines, while also contributing to GenAI and AI-ready data platform capabilities.
Key Responsibilities
- Design and maintain scalable CDC and incremental ingestion pipelines.
- Work with PostgreSQL WAL, logical replication, LSN and replication slots.
- Build schema-independent, metadata-driven ingestion frameworks.
- Automatically discover tables, columns, keys, data types and schema changes.
- Handle schema evolution with minimal manual intervention.
- Ensure reliable CDC recovery with no data loss or duplicate processing.
- Build reusable frameworks for onboarding new databases and schemas through configuration rather than custom code.
- Develop high-volume data pipelines using Python, SQL and Kafka.
- Implement data validation, reconciliation, monitoring and error handling.
- Contribute to AI-ready metadata and semantic capabilities for the EDP.
GenAI Responsibilities
- Use LLMs/GenAI for schema understanding and metadata enrichment.
- Generate business-friendly descriptions of technical data structures.
- Support AI-assisted source-to-canonical schema mapping.
- Help identify relationships between business entities and datasets.
- Contribute to semantic layer / business ontology development.
- Support RAG, Text-to-SQL, copilots and AI-agent use cases.
- Implement validation and guardrails for LLM-generated outputs.
Must-Have Skills
- 5–8 years of Data Engineering / Database Engineering experience
- Solid PostgreSQL, Python and SQL
- Hands-on experience with CDC / Logical Replication
- PostgreSQL WAL, LSN and replication slots
- Kafka or equivalent streaming technology
- Schema evolution and metadata-driven pipelines
- Large-scale data ingestion and production troubleshooting
- REST APIs, Linux and Git
GenAI Skills
- LLM fundamentals and API integration
- Prompt engineering
- RAG and embeddings
- Structured LLM outputs / tool calling
- Vector search / Vector DB fundamentals
- Exposure to LangChain, LangGraph or LlamaIndex
- Understanding of Text-to-SQL, semantic search and AI agents
Good to Have
- Debezium / Kafka Connect
- Airflow
- HDFS / Hive / MinIO
- Spark / ClickHouse
- Docker / Kubernetes
- Knowledge Graphs / Ontologies
- Smart Metering, Energy or IoT domain experience
Preferred Candidate
- We are looking for someone who can take ownership of an existing CDC and schema-independent ingestion framework, troubleshoot production issues, improve its scalability and reliability, and gradually help evolve the platform toward an AI-ready semantic data architecture.
📌 Data Engineer (Jaipur)
🏢 Passageway Tech
📍 Jaipur