17 Sep
|
Entiovi Technologies
|
India
17 Sep
Entiovi Technologies
India
Senior Data Engineer
About the Role The Senior Data Engineer in our AI & Data team will be responsible for designing and building scalable data platforms, enterprise-grade data architectures, and high-performance data ingestion frameworks across Azure, Snowflake, Databricks, and Lakebase.
This role requires a highly technical engineer capable of solving complex data platform challenges involving large-scale API integrations, distributed data processing, cloud-native architectures, and AI-enabled data platforms.
The ideal candidate is not only a Data Engineer but also a strong Python engineer with expertise in data architecture, system design, API engineering, concurrent processing, and enterprise scale platform development.
Main Responsibilities
- Design and implement scalable enterprise data architectures supporting AI, analytics,
reporting, and operational workloads.
- Build and optimize large-scale ELT/ETL pipelines using Databricks, Snowflake, Azure Data
Factory, and Azure services.
- Design and implement Medallion Architectures (Bronze/Silver/Gold), CDC frameworks,
lineage models, and master data management workflows.
- Develop high-performance Python-based ingestion frameworks supporting large-scale extraction from internal and external data sources.
- Build and maintain REST API and GraphQL integrations including OAuth/OAuth2
authentication, token lifecycle management, pagination, retry mechanisms, rate limiting,
and automated error recovery.
- Design parallelized ingestion solutions using multi-threading, asynchronous processing, and distributed execution patterns.
- Develop and maintain Databricks pipelines, Delta Lake architectures, and Snowflake analytical data platforms.
- Build and support real-time and near-real-time data processing solutions using event-driven architectures.
- Implement data quality, reconciliation, lineage, observability, and governance frameworks across the platform.
- Design scalable data models supporting analytics, machine learning, feature engineering,
and AI workloads.
- Monitor production environments, troubleshoot pipeline failures,
optimize platform performance, and drive cost optimization initiatives.
- Collaborate closely with AI Engineers, Architects, Product Teams, and Business Stakeholders.
Required Skills & Experience
Python (Mandatory)
- Strong production-grade Python development experience.
- Object-Oriented Programming (OOP).
- Modular framework development.
- Logging, exception handling, testing, and debugging.
- Performance optimization and profiling.
- Experience building reusable ingestion and transformation frameworks.
Advanced API Engineering (Mandatory)
- REST APIs.
- GraphQL APIs.
- OAuth/OAuth2.
- JWT Authentication.
- API Pagination.
- Rate Limiting.
- Retry Logic.
- Token Refresh Handling.
- Error Handling Frameworks.
- High-volume API ingestion architecture.
Data Architecture & System Design (Mandatory)
- Medallion Architecture.
- Data Warehouse Architecture.
- Data Lakehouse Architecture.
- Master Data Management (MDM).
- Golden Record Design.
- Data Lineage.
- Change Data Capture (CDC).
- Historical Data Management.
- Enterprise Data Modeling.
Data Pipeline Design (Mandatory)
- Design and build scalable, fault-tolerant enterprise data pipelines.
- Solid experience with batch, near real-time, and event-driven processing architectures.
- Expertise in designing ingestion, transformation, validation, reconciliation, and serving layers across modern data platforms.
- Experience implementing Medallion Architecture (Bronze, Silver, Gold) and data lakehouse patterns.
- Strong understanding of Change Data Capture (CDC), incremental processing, watermarking strategies, and SCD Type 1/Type 2 implementations.
- Design audit frameworks, lineage tracking, reconciliation controls, monitoring,
alerting, and observability solutions.
- Ability to build high-volume ingestion pipelines from APIs, databases, files, data streams,
and external systems.
- Experience designing resilient pipelines with retry mechanisms, checkpointing, idempotent processing, restartability, and failure recovery.
- Strong understanding of throughput optimization, parallel processing, concurrency, and workload orchestration.
- Experience designing data movement patterns across Snowflake, Databricks, Azure services,
and downstream analytics platforms.
Databricks (Mandatory)
- Databricks Workflows.
- Delta Lake.
- Unity Catalog.
- PySpark.
- Spark SQL.
- Databricks Performance Tuning.
- Distributed Data Processing.
Snowflake (Mandatory)
- Snowpipe.
- Streams.
- Tasks.
- Dynamic Tables.
- Time Travel.
- Zero-Copy Cloning.
- RBAC.
- Query Optimization.
- Warehouse Optimization.
- Cost Management.
SQL & Data Modeling (Mandatory)
- Advanced SQL.
- CTEs.
- Window Functions.
- Stored Procedures.
- MERGE.
- Star Schema.
- Snowflake Schema.
- SCD Type 1 & Type 2.
- Dimensional Modeling.
Azure (Mandatory)
- Azure Data Factory.
- ADLS Gen2.
- Azure Synapse.
- Event Hubs.
- Azure Integration Services.
Strongly Preferred
- Apache Airflow.
- dbt.
- Kafka.
- Event Driven Architecture.
- Terraform.
- Infrastructure as Code.
- Azure OpenAI.
- AI/ML Data Pipelines.
- Feature Stores.
Git & DevOps (Mandatory)
- Git Branching Strategies.
- Pull Requests.
- Git Rebase.
- Cherry-picking.
- Merge Management.
- Repository Governance.
- CI/CD Pipelines.
- GitHub Actions / Azure DevOps.
- Secret Management.
- Git History Cleanup and Recovery.
Experience
- 5+ years of hands-on Data Engineering experience.
- Proven experience designing production-grade enterprise data platforms.
- Strong experience working directly with business stakeholders and translating requirements into scalable technical solutions.
- Experience leading architecture discussions and solving complex technical problems independently.
📌 Senior Data Engineer (India)
🏢 Entiovi Technologies
📍 India