Principal Data Engineer – Real-time Data Platform (India)

Principal Data Engineer – Real-time Data Platform (India)

29 Aug
|
Algoworks
|
India

29 Aug

Algoworks

India

Principal Data Engineer – Real-time Data Platform

Location: India, Remote

Experience: 10+ Years

Algoworks www.algoworks.com

About the company

Algoworks is an award-winning artificial intelligence, engineering services and experience transformation firm with offices across the United States, Europe, South America and India. We bring together a global team of engineers, architects, designers, researchers and operators united by rigor, accountability and a commitment to delivering measurable results.

For over 20 years, Algoworks has partnered with Fortune 500 organizations across the Americas, Europe and Asia to define, build and run technology that drives meaningful business outcomes. Our work combines human-centered design, engineering excellence and AI-powered capabilities to solve complex challenges with clarity and precision. Innovation, particularly in the responsible application of AI, is embedded in how teams approach problem-solving and continuous improvement.

At Algoworks, growth is continuous and closely tied to impact. Teams collaborate across geographies and disciplines, strengthening outcomes through shared insight and collective expertise. The culture values transparency, open dialogue and an environment where every voice is heard and contribution is recognized.

Through collaboration, accountability and a focus on results, Algoworks operates at the intersection of technology and people, building not only advanced systems but strong global teams that elevate performance and create lasting impact.

Follow the video below to know about us! Clipchamp

Role overview

We are looking for a hands-on Lead Data Engineer to provide technical leadership for high-volume data ingestion and processing, with a strong focus on real-time CDC, Databricks, SQL Server, Debezium and Azure Event Hubs.

This role will own the technical architecture and engineering direction across Full Load / Batch and Real-Time CDC pipelines, with real-time streaming expected to become the primary long-term ingestion pattern.

The ideal candidate is a strong technical leader and architect who remains hands-on, can guide and mentor engineers and can design production-grade ingestion solutions for scalability, resiliency, performance and data integrity.

Key responsibilities:

- Lead the architecture and technical implementation of batch, full-load,



incremental and real-time CDC pipelines.
- Design high-volume ingestion from SQL Server using CDC and Debezium.
- Build scalable event-driven pipelines using Azure Event Hubs and Databricks.
- Design and optimize Databricks pipelines for large-scale data ingestion and transformation.
- Implement robust error handling, retry, replay, checkpointing, recovery and idempotency.
- Design solutions for schema drift and schema evolution without disrupting downstream processing.
- Design and optimize Delta Lake / Delta Tables, including partitioning, compaction, data layout and performance optimization.
- Optimize pipeline throughput, latency, parallelism, resource utilization and processing windows.
- Establish monitoring and observability for CDC lag, connector health, consumer lag, pipeline failures, throughput and processing latency.
- Implement reconciliation and data-quality controls to ensure source-to-target completeness and accuracy.
- Provide technical direction, perform design/code reviews, mentor engineers and establish engineering best practices.
- Drive technical readiness for scaling ingestion across significantly more clients, databases, tables and data volumes.

Required skills and qualifications:

- Bachelor’s or master's degree in computer science, Information Technology, Business, or related field (or equivalent practical experience).
- 10+ years of Data Engineering / Software Engineering experience.
- Robust hands-on experience with Databricks and Delta Lake.
- Strong experience designing and operating Databricks data pipelines at scale.
- Deep understanding of:
- Pipeline design and orchestration
- Error handling and recovery
- Schema drift
- Schema evolution
- Idempotent data processing
- Delta Tables
- Data partitioning and optimization
- Performance tuning
- Strong hands-on experience with SQL Server CDC, transaction logs, LSNs and high-volume transactional databases.




- Experience with Debezium SQL Server Connector, including configuration, offsets, snapshots, recovery and schema changes.
- Strong experience with Azure Event Hubs, including partitioning, consumer groups, scaling, throughput and checkpointing.
- Deep understanding of batch, micro-batch, streaming and event-driven data architectures.
- Strong experience with Python/PySpark, SQL, Azure Data Lake and distributed data processing.
- Experience designing production-grade solutions for retry, replay, fault tolerance, duplicate handling, reconciliation and observability.
- Strong performance engineering and troubleshooting skills across large-scale data pipelines.
- Ability to provide technical leadership, architecture guidance, mentoring and hands-on engineering support.

Must have skills:

- 10+ years of Data Engineering / Software Engineering experience.
- 3+ years working with production-scale CDC or real-time streaming architectures.
- Strong production experience with Databricks and Delta Lake.
- Experience processing millions to billions of records.
- Experience with multi-client or multi-tenant ingestion architectures.
- Experience implementing Medallion / Bronze-Silver-Gold architectures.

Good to have skills:

- Experience with Apache Kafka / Kafka Connect and streaming ecosystems.
- Knowledge of Azure Data Factory, Azure Functions and Azure Monitor.
- Experience with Infrastructure as Code (Terraform/ARM/Bicep) and CI/CD for data platforms.
- Familiarity with Unity Catalog, Databricks Workflows and advanced Spark optimization.

Key success criteria: The person in this role should be able to:

- Establish a scalable architecture for batch and real-time ingestion.
- Scale pipelines across substantially more databases, clients, tables and data volumes.
- Improve Databricks pipeline performance and processing windows.
- Deliver reliable high-volume CDC without sustained lag, duplication, or data loss.
- Handle schema changes and schema drift without destabilizing ingestion.
- Ensure pipelines are idempotent and safely recoverable/replayable following failures.
- Optimize Delta Tables and downstream processing for performance and scalability.
- Provide clear technical leadership and mentoring for the ingestion engineering team.

Interview process 2 rounds of discussion

📌 Principal Data Engineer – Real-time Data Platform (India)
🏢 Algoworks
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: principal data engineer – real-time data platform (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: principal data engineer – real-time data platform (india) / india