Distributed MPP Database Engine (HPC/C++) (Bengaluru)

Distributed MPP Database Engine (HPC/C++) (Bengaluru)

22 Aug
|
Circana
|
Bengaluru

22 Aug

Circana

Bengaluru

Job Title Principal engineer, Senior Staff Engineer, Staff Engineer – Distributed MPP Database Engine (HPC/C++)

Location: Hybrid – Bengaluru

Department: Liquid Platform Architecture & Modeling

Position Overview

We are seeking Principal/Senior Staff Engineer/Staff Engineers to drive design, evolution, and execution of our next-generation, massively parallel processing (MPP) clustered database engine (Whitebox) that currently powers large-scale retail and consumer analytics solutions across Circana clients in 22+ countries and serving millions of queries every day in real-time against massive data sets . This is a highly technical role for an expert who bridges the gap between High-Performance Computing (HPC) and Big Data systems engineering and low-latency machine learning infrastructure and involves database internals, distributed systems, performance engineering, data processing at scale.

In this role, you will be one of the primary designers and developer responsible for building low-latency, high-throughput distributed database kernels. You will design bare-metal optimized software layers, ensuring our database product fully exploits modern multi-core, vectorized hardware and distributed network fabrics to process terabyte/petabyte-scale datasets and where appropriate, seamlessly executes highly performant analytics and in-database ML inference directly on relational and the occasional vector datasets.

Core Responsibilities

- Architecture & Design: Play a key role in core architectural choices and roadmap for our distributed database kernel, including query execution engines, custom/hybrid columnar storage layers, distributed transaction managers, and cluster coordination protocols.
- Low-Level Development: Write, optimize, and maintain highly performant, production-ready backend code utilizing modern C++ (C++20/23).
- Parallel Execution Planning: Architect MPP execution engines that distribute, schedule, and execute queries across several dozens to hundreds of nodes using MPI and hybrid OpenMP models.
- Hardware Optimization: Profile and eliminate execution bottlenecks by implementing explicit SIMD vectorization (AVX-512, ARM Neon, SVE) and ensuring absolute cache locality (L1/L2/L3).
- Distributed Memory & Networks: Design ultra-low latency cluster communication subsystems utilizing RDMA, RoCE,



or kernel-bypass tech (DPDK).
- In-Database ML Framework Design: Architect and implement a native, highly performant, zero-copy ML execution framework directly inside the MPP engine to eliminate data serialization and transport bottlenecks between the database and external AI pipelines.
- Functional & High-Performance Algorithms: Develop mathematically rigorous, cache-oblivious, and highly vectorized algorithms for real-time statistical computations, matrix operations/manipulations on clustered data
- Vector & Embedding Infrastructure: Design and optimize or embed/reference distributed high-dimensional vector storage subsystems, custom SIMD-accelerated indexing structures (e.g., HNSW, IVF-PQ), and parallelized approximate nearest neighbor (ANN) search algorithms.
- Be comfortable working in a cross country/multi time zone team of elite systems engineers, implementing as well as conducting rigorous peer design and development reviews
- Collaborate with cross-functional teams in defining and developing new platform capabilities.
- Actively participate in architecture reviews, code reviews, and technical mentoring.
- Investigate and resolve complex performance and scalability issues.

Required Technical Qualifications

- Minimum - 4 years Computer Science engineering degree from a reputed institute. Masters in HPC (high performance computing) and Bigdata engineering areas preferred for Senior staff engineer and above roles
- Modern C++ Expertise: 12+ (Principal), 8+ (Senior Staff engineer), 6+(Staff engineer), years of production experience writing hands-on, low-level systems in C++ (C++17/20/23 preferred). Mastery of template metaprogramming, concepts, coroutines, and custom memory allocators is key.
- HPC Parallelism: Very good to expert-level knowledge of cluster-scale distributed programming using MPI (Message Passing Interface) paired with shared-memory thread-level programming via OpenMP and MapReduce frameworks.




- Hardware-Aware Engineering: Documented experience with microarchitectural optimization, loop vectorization, pointer-aliasing constraints, and manual SIMD intrinsics programming.
- Concurrency Mechanics: Deep understanding of OS internals, memory barriers, lock-free data structures, and atomics (std::atomic).
- Database Internals: Practical experience designing core database subsystems (e.g., vectorized query operators, LSM-trees, B-trees, custom buffer pools, columnar serialization frameworks and query plan optimizers).
- Strong debugging, performance tuning, and problem-solving skills.
- Proficiency in one or more leading memory and compute monitoring/performance tools (such as Valgrind, Purify, Intel/AMD/ARM diagnostics etc.)

Bonus Qualifications

- Track record contributing to open-source or work experience on proprietary distributed engines (e.g., ClickHouse, DuckDB, Doris, CockroachDB, RocksDB, ScyllaDB, or Velox).
- Experience with cloud-native distributed, elastic environments, object storage tiering (S3/GCS), or zero-copy networking.
- Experience integrating heterogeneous acceleration (NVIDIA CUDA / ROCm) into big data workloads and hybrid CPU/GPU setups.
- Vector & Graph Algorithmic experience: Experience building or modifying high-performance vector index/search engines, graph-based indices, or localized quantization frameworks.
- Functional domain awareness – CPG, Retail, Pharma verticals

What We Offer

- Work on the bleeding edge of data infrastructure, tackling next gen big data engineering problems to support real time querying and analytics against some of the largest data sets at high concurrency
- Challenging engineering problems - all the way from functional/data science algorithm implementations on a HPC setup to bare metal/close to OS kernel level programming involving handling of hundreds of billions of records to support real-time analytics at scale.
- Plenty of valuable exposure to broader BI/Analytics/Bigdata ecosystem including great opportunity to functionally better understand CPG, Retail and related domains.
- Highly competitive compensation package (Base, Bonus)
- Stimulating work environment and a outstanding Circana work culture that allows members to bring their best daily.

📌 Distributed MPP Database Engine (HPC/C++) (Bengaluru)
🏢 Circana
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: distributed mpp database engine (hpc/c++) (bengaluru) / bengaluru