Where you’ll work: Remote
About the Role
We're looking for a Principal Data Engineer to own our Databricks data platform. You'll be the go-to expert on Unity Catalog, Delta Lake, and Spark, and you'll set the technical direction for how our teams build, govern, and scale data pipelines. This is a hands-on role: you'll write code, review designs, and fix hard problems.
We're looking for a Principal Data Engineer who operates at architect level across our data platform and ML/AI infrastructure. You'll own the technical vision and multi-year roadmap for our Databricks Lakehouse, set the architectural standards that govern how Data Engineering teams build and ship, and be an active contributor to our ML and AI platform.
You won't be maintaining someone else's architecture. You'll be designing the next version of it, and you'll have the scope to make it matter.
Your Day to Day
- Own the architecture and roadmap for our Databricks platform: Unity Catalog, Delta Lake, cluster configuration, and workflow orchestration.
- Lead the architecture of our data platform, pipelines, and AI-ready datasets, so they're built to support GenAI and ML use cases, not just BI.
- Design and evolve our medallion (Bronze/Silver/Gold) lakehouse architecture so it scales with the business.
- Set standards for data quality, security, and access control across the platform, and make sure teams actually follow them.
- Tune Spark jobs and cluster configs for performance and cost. Find the waste, right-size the clusters, kill the idle spend.
- Build and maintain the core ETL frameworks and reusable components other teams build on top of.
- Lead migrations from legacy platforms onto Databricks, including data modeling and pipeline redesign.
- Embed infrastructure-as-code and CI/CD practices into how we build, deploy, and run pipelines and platform infrastructure, so deployments stay consistent and repeatable.
- Troubleshoot production issues, do root cause analysis, and put fixes in place that prevent repeat incidents.
- Mentor senior engineers and act as the technical escalation point across the Data teams.
- You'll set the standard for spec-driven development and agentic AI tooling, then drive adoption for pipeline generation, testing, documentation, and incident response across the Data teams.
- Stay current on Databricks releases and bring in recent capabilities when they solve a real problem, not just because they're new.
- What We’re Looking For
- 12+ years in data engineering, with at least 3+ years working hands-on with Databricks in production.
- Deep expertise in Unity Catalog, Delta Lake, and Spark performance tuning at scale.
- Real experience designing and running medallion/lakehouse architectures.
- Strong SQL and Python.
- Experience with at least one major cloud platform (AWS, Azure, or GCP) and infrastructure-as-code tooling.
- Experience with CI/CD for data pipelines, including Databricks Asset Bundles.
- A track record of leading technical decisions and mentoring other engineers.
- Experience rolling out AI agents or AI-assisted tooling across an engineering team. You've set standards and gotten a team to adopt them.
- Clear communicator who can work across engineering, analytics, and business stakeholders.
Nice to Have
- Experience with orchestration tools like Airflow or dbt.
- Experience with streaming/messaging systems like Kafka.
- Exposure to GenAI, RAG, or LLM-based data applications.
- Experience leading a legacy data warehouse to Databricks migration.
📌 Principal Data Engineer- Databricks & AI platform (India)
🏢 GoTo
📍 India