30 Sep
|
Mastercard
|
Pune
Job Description
Overview
n
n
The Mastercard Services Technology team is looking for a Principal Data Engineer to unlock the potential of our data assets by innovating continuously, removing friction from how large-scale data is stored, managed, governed, and accessed, and establishing standards across public-cloud and on-premises environments.
n
n
We are looking for a hands-on, passionate engineer and architect with strong PySpark, Python, cloud and modern data-architecture expertise. This role will design scalable distributed and lakehouse platforms across regions and execution environments; establish federated governance, metadata, access, and query patterns; build reliable data solutions; shape engineering culture; and mentor others. It is a role for builders and collaborators who value clean pipelines, cloud-native design, sound architecture decisions, and helping teammates succeed.
n
n
n
n
Role
n
n
Design and build scalable, secure, cloud-native and hybrid data platforms and control-plane services using Java and/or Python, PySpark, REST and gRPC APIs, and modern data and software engineering practices, supporting business-critical batch, streaming, API, and secure data-sharing use cases.
n
n
Architect large-scale distributed data and lakehouse platforms spanning regions, public clouds, on-premises, and private or sovereign execution environments.
n
n
Define control-plane and data-plane responsibilities and clear boundaries among global governance, local enforcement, metadata and orchestration, and local workload execution.
n
n
Design control-plane services and interfaces for platform configuration, metadata, policy, identity and access, provisioning, orchestration, lifecycle management, and observability using secure API-first and event-driven patterns.
n
n
Design federated metadata, catalog, identity, access, query, and data-sharing patterns; evaluate federation, replication, and compute-to-data using latency, cost, regulation, residency, freshness, reliability, and operational complexity.
n
n
Define data-product and data-contract standards covering ownership, interfaces, schema evolution, compatibility, lineage, quality, observability, and SLA/SLO commitments.
n
n
Design processing and query architectures using Databricks, Snowflake, Spark, and Trino, with open table and file formats such as Iceberg, Delta Lake, and Parquet.
n
n
Define catalog and metadata architectures using Unity Catalog, AWS Glue Data Catalog,
and open catalog patterns such as Apache Polaris.
n
n
Design AWS data-platform architecture using S3, IAM, EKS, networking, and relevant cloud-native data services, and Azure architecture using ADLS, Microsoft Entra ID, Azure RBAC, Key Vault, AKS, and relevant analytics services.
n
n
Design Kubernetes platforms for portable compute, workload isolation, security, scaling, observability, and deployment across cloud, private, sovereign, and on-premises environments.
n
n
Own and optimize infrastructure for data storage, processing, orchestration, networking, identity, secrets, security, reliability, and cost.
n
n
Create architecture decision records, reference architectures, engineering standards, reusable components, and golden paths; decompose complex problems into modular, scalable, and maintainable solutions aligned with platform and product goals.
n
n
Lead by example through high-quality, secure, testable code, architecture and design discussions, code reviews, testing, CI/CD, version control, documentation, observability, and performance tuning.
n
n
Build governance, privacy, quality, lineage, cataloging, retention, and access management into the platform by design.
n
n
Mentor engineers, share knowledge, and foster curiosity, growth, and continuous improvement while collaborating with product managers, data scientists, application engineers, architects, security partners, and stakeholders.
n
n
Participate in architectural discussions, iteration planning, feature sizing, Agile ceremonies, and delivery-risk assessment; communicate trade-offs clearly and influence technical direction across teams.
n
n
Continuously evaluate and apply relevant technologies and patterns to improve productivity, interoperability, reliability, and total cost of ownership.
n
n
n
n
All About You
n
n
14+ years of hands-on data and software engineering experience, with strong Java and/or Python, PySpark, API and backend service development skills, and a track record of delivering production-grade data platforms, control-plane services, data models,
pipelines, and batch or streaming systems.
n
n
Proven experience designing large-scale distributed or lakehouse platforms across multiple regions, clouds, clusters, or execution environments.
n
n
Strong understanding of control-plane versus data-plane architecture and the boundaries among global governance, local enforcement, and local execution.
n
n
Experience designing distributed control-plane applications using Java and/or Python, REST and gRPC APIs, asynchronous messaging, workflow and state management, authentication and authorization, multi-tenancy, idempotency, resiliency, auditability, and operational observability.
n
n
Experience with federated metadata, catalog, access, query, and secure sharing patterns, including evaluating federation, replication, and compute-to-data using latency, cost, regulatory, freshness, and operational criteria.
n
n
Strong knowledge of data products, data contracts, schema evolution, compatibility, lineage, quality, observability, and SLA/SLO design.
n
n
Ability to produce architecture decision records, reference architectures, technical standards, reusable patterns, and golden paths adopted across teams.
n
n
Solid cloud data-platform experience across AWS, Azure, or GCP, including AWS S3, IAM, Glue, EKS, networking, and cloud-native data services; and Azure ADLS, Microsoft Entra ID, Azure RBAC, Key Vault, Data Factory, Databricks, AKS, and analytics services.
n
n
Hands-on architecture experience with Databricks, Snowflake, Spark, Kubernetes, Iceberg, Delta Lake, Parquet, Unity Catalog, AWS Glue Data Catalog, open catalog patterns such as Apache Polaris, and distributed-query concepts using Trino or Spark.
n
n
Experience designing batch, streaming, API, and secure data-sharing architectures and integrating heterogeneous systems across cloud environments.
n
n
Strong foundations in modern data architectures such as lakehouse and medallion, data lifecycle management, data modeling, database design, distributed systems, and performance optimization.
n
n
Comfort with Git, CI/CD, automated testing, infrastructure as code, documentation, Agile or Scrum delivery, and production operational practices.
n
n
Curious, adaptable, and committed to continuous learning and improvement.
n
n
A bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent hands-on experience.
📌 Principal Data Engineer (Pune)
🏢 Mastercard
📍 Pune