Principal Data Engineer (Pune)

Principal Data Engineer (Pune)

08 Sep
|
HG Insights
|
Pune

08 Sep

HG Insights

Pune

About HG

HG Insights is the pioneer of Revenue Growth Intelligence. For more than a decade, we have delivered comprehensive, AI- driven datasets on B2B buyers, technology adoption, IT spend, and buyer intent, sourced from billions of data points.

Today, we are a trusted partner to Fortune 500 technology companies, hyperscalers, and innovative B2B vendors seeking precise go-to-market analytics and decision-making. Through an evolving suite of AI agents that incorporate first-party data and buyer signals, HG Insights enables AI-powered GTM automation across sales, marketing, RevOps, and data analytics teams, modernizing GTM execution from strategy through activation.

About the Role

This is a senior individual contributor role on the team that builds and runs HG's data platform: everything between a raw vendor file landing in our lake and a finished dataset arriving in a customer's hands. It includes the transformation stack that turns messy external signals into product-grade company attributes, the matching systems that decide which company a piece of evidence belongs to and the internal applications our analysts use to curate and correct the result.

You will be based in Pune, working with engineers at our Pune, US & Brazil locations.

What you’ll do

- Own the architecture of the data platform end to end, and make the calls on build vs. buy, batch vs. incremental, and where each workload belongs.
- Build the quality layer that catches silent failure: contracts at every handoff, freshness and completeness monitoring, and pipelines that quarantine bad data rather than publish it.




- Improve entity matching and attribution — and make match decisions explainable to the customers who depend on them.
- Make our release cycle boring. Recurring deliveries should not depend on people watching them.
- Treat compute cost as a first-class engineering metric across Spark, orchestration, and the serving tier.

What we are looking for

- 15+ years building production grade data engineering systems, with a minimum of 5 years in a staff or principal scope.
- Deep Spark and Databricks expertise at multi-terabyte scale, including the operational side of lakehouse table formats — merges, compaction, small files, schema evolution.
- Solid SQL and dimensional modelling, and the judgment to know when to break the rules.
- Experience with performant and scalable OLTP setups (MySQL/Postgres).
- Airflow orchestration (DAGs, operators, sensors) and integration with Spark/Databricks.
- Proven experience in AWS ecosystems (EC2, S3, EMR).
- Hands-on entity resolution, fuzzy matching, or record linkage at scale. This is central to what we do.
- Python/Scala/Java, with real software discipline: testing, CI/CD, code review, infrastructure as code is a plus.
- Experience in Docker - Kubernetes, Terraform is a plus.
- Experience in integrating AI first implementations in traditional data engineering setups will greatly help in shaping our future designs.
- Experience with machine learning pipelines (Spark MLlib, Databricks ML) for predictive analytics.
- Knowledge of data governance frameworks and compliance standards (GDPR, CCPA).

Preferred Qualifications

Helpful, not required: B2B firmographic, technographic or intent data; streaming and CDC; warehouse-to-lakehouse migration experience.

📌 Principal Data Engineer (Pune)
🏢 HG Insights
📍 Pune

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: principal data engineer (pune) / pune