Job Description:
Senior Data Engineer (GCP)
Onsite: Navi Mumbai
Share your CV at
[email protected]
Role Overview
The Data Engineer is a hands-on implementation specialist responsible for building the actual "plumbing" of the database-to-database pipelines. This role focuses on the technical nuances of connectivity, memory-efficient transformations, and robust error handling to ensure zero data loss.
Key Responsibilities
- Pipeline Development:
Build secure JDBC connections and develop the worker memory logic required for complex schema mappings (e.g., transforming Oracle DATE types to BigQuery DATETIME).
- DLQ & Fault Tolerance:
Implement Dead Letter Queues (DLQ) within the Dataflow architecture to capture and preserve data that fails validation, ensuring a "Fail" path exists for future reprocessing without stopping the pipeline.
- Runtime State Management:
Develop and maintain the BigQuery-based runtime state store for watermark persistence, ensuring that extraction points are accurately tracked and advanced only upon successful load.
- Audit & Metadata Enrichment:
Programmatically inject audit metadata including extraction_timestamp and source_system_id to maintain full traceability.
- Reliability Engineering:
Configure retry and restart behaviors that are secure for Raw layer appends, preventing data duplication during mid-stream failures.
Must-Have Technical Skills
- Data Processing:
Strong Java or Python skills for Cloud Dataflow development.
- Storage & SQL:
Proficient in BigQuery SQL and managing state in BigQuery metadata tables.
- Python :
Strong proficiency in Python (3.x), with experience writing clean, modular, and testable code.
- Security:
Experience applying "Least Privilege" principles using GCP IAM service accounts.
- Connectivity:
Practical experience configuring JDBC drivers for Oracle and SAP DB.
- ETL Framework Design
- Data Migration Architect
Minimum Experience
- 6 years in data engineering.
- Proven experience implementing Oracl
📌 Senior Data Engineer (Mumbai)
🏢 EliteSquad.AI
📍 Mumbai