10 Sep
|
HG Insights
|
Amravati
10 Sep
HG Insights
Amravati
About HGHG Insights is the pioneer of Revenue Growth Intelligence. For more than a decade, we have delivered comprehensive, AI- driven datasets on B2B buyers, technology adoption, IT spend, and buyer intent, sourced from billions of data points.Today, we are a trusted partner to Fortune 500 technology companies, hyperscalers, and cutting-edge B2B vendors seeking precise go-to-market analytics and decision-making. Through an evolving suite of AI agents that incorporate first-party data and buyer signals, HG Insights enables AI-powered GTM automation across sales, marketing, RevOps, and data analytics teams, modernizing GTM execution from strategy through activation.About the RoleThis is a senior individual contributor role on the team that builds and runs HG's data platform: everything between a raw vendor file landing in our lake and a finished dataset arriving in a customer's hands. It includes the transformation stack that turns messy external signals into product-grade company attributes, the matching systems that decide which company a piece of evidence belongs to and the internal applications our analysts use to curate and correct the result.You will be based in Pune, working with engineers at our Pune, US & Brazil locations.What you'll do- Own the architecture of the data platform end to end, and make the calls on build vs. buy, batch vs. incremental, and where each workload belongs.- Build the quality layer that catches silent failure: contracts at every handoff, freshness and completeness monitoring,
and pipelines that quarantine bad data rather than publish it.- Improve entity matching and attribution — and make match decisions explainable to the customers who depend on them.- Make our release cycle boring. Recurring deliveries should not depend on people watching them.- Treat compute cost as a first-class engineering metric across Spark, orchestration, and the serving tier.What we are looking for- 15+ years building production grade data engineering systems, with a minimum of 5 years in a staff or principal scope.- Deep Spark and Databricks expertise at multi-terabyte scale, including the operational side of lakehouse table formats — merges, compaction, small files, schema evolution.- Strong SQL and dimensional modelling, and the judgment to know when to break the rules.- Experience with performant and scalable OLTP setups (MySQL/Postgres).- Airflow orchestration (DAGs, operators, sensors) and integration with Spark/Databricks.- Proven experience in AWS ecosystems (EC2, S3, EMR).- Hands-on entity resolution, fuzzy matching, or record linkage at scale. This is central to what we do.- Python/Scala/Java, with real software discipline: testing, CI/CD, code review, infrastructure as code is a plus.- Experience in Docker - Kubernetes, Terraform is a plus.- Experience in integrating AI first implementations in traditional data engineering setups will greatly help in shaping our future designs.- Experience with machine learning pipelines (Spark MLlib, Databricks ML) for predictive analytics.- Knowledge of data governance frameworks and compliance standards (GDPR, CCPA).Preferred QualificationsHelpful, not required: B2B firmographic, technographic or intent data; streaming and CDC; warehouse-to-lakehouse migration experience.
📌 Principal Data Engineer (Amravati)
🏢 HG Insights
📍 Amravati