11 Sep
|
HG Insights
|
Pune
About HG /n HG Insights is the pioneer of Revenue Growth Intelligence. For more than a decade, we have delivered comprehensive, AI- driven datasets on B2B buyers, technology adoption, IT spend, and buyer intent, sourced from billions of data points. /n Today, we are a trusted partner to Fortune 500 technology companies, hyperscalers, and cutting-edge B2B vendors seeking precise go-to-market analytics and decision-making. Through an evolving suite of AI agents that incorporate first-party data and buyer signals, HG Insights enables AI-powered GTM automation across sales, marketing, RevOps, and data analytics teams, modernizing GTM execution from strategy through activation. /n About the Role /n This is a senior individual contributor role on the team that builds and runs HG's data platform: everything between a raw vendor file landing in our lake and a finished dataset arriving in a customer's hands.
It includes the transformation stack that turns messy external signals into product-grade company attributes, the matching systems that decide which company a piece of evidence belongs to and the internal applications our analysts use to curate and correct the result. /n You will be based in Pune, working with engineers at our Pune, US & Brazil locations. /n What you’ll do /n /n
- Own the architecture of the data platform end to end, and make the calls on build vs. buy, batch vs. incremental, and where each workload belongs.
/n
- Build the quality layer that catches silent failure: contracts at every handoff, freshness and completeness monitoring, and pipelines that quarantine bad data rather than publish it.
/n
- Improve entity matching and attribution — and make match decisions explainable to the customers who depend on them.
/n
- Make our release cycle boring. Recurring deliveries should not depend on people watching them.
/n
- Treat compute cost as a first-class engineering metric across Spark, orchestration, and the serving tier.
/n /n /n What we are looking for /n /n
- 15+ years building production grade data engineering systems, with a minimum of 5 years in a staff or principal scope.
/n
- Deep Spark and Databricks expertise at multi-terabyte scale, including the operational side of lakehouse table formats — merges, compaction, small files, schema evolution.
/n
- Strong SQL and dimensional modelling, and the judgment to know when to break the rules.
/n
- Experience with performant and scalable OLTP setups (MySQL/Postgres).
/n
- Airflow orchestration (DAGs, operators, sensors) and integration with Spark/Databricks .
/n
- Proven experience in AWS ecosystems (EC2, S3, EMR).
/n
- Hands-on entity resolution, fuzzy matching, or record linkage at scale. This is central to what we do.
/n
- Python/Scala/Java, with real software discipline: testing, CI/CD, code review, infrastructure as code is a plus.
/n
- Experience in Docker - Kubernetes, Terraform is a plus.
/n
- Experience in integrating AI first implementations in traditional data engineering setups will greatly help in shaping our future designs.
/n
- Experience with machine learning pipelines (Spark MLlib, Databricks ML) for predictive analytics.
/n
- Knowledge of data governance frameworks and compliance standards (GDPR, CCPA).
/n /n /n Preferred Qualifications /n Helpful, not required: B2B firmographic, technographic or intent data; streaming and CDC; warehouse-to-lakehouse migration experience. /n
📌 Principal Data Engineer (Pune)
🏢 HG Insights
📍 Pune