02 Sep
|
TransUnion
|
Maharashtra
02 Sep
TransUnion
Maharashtra
Team Overview
We are looking for a Data Engineering Operations Lead to join our growing Data Engineering and Analytics Practice who will drive building next generation suite of products and platforms by designing, coding, building, and deploying highly scalable and robust data solutions. The team reports to the Head of Data Engineering & Analytics and is part of the wider Global Technology function.
Role Overview and Core Responsibilities The role exists to own and manage the internal function responsible for development, support and operation of our ingestion pipelines as well as the assets created and owned by the team across our cloud (GCP) and on-premises environments.
As Data Engineering
Operations lead, you will be responsible to ensure the quality and robustness of the developed pipelines, operational resilience, SLA compliance for all of our main data assets created by our team. As a lead engineer most of your time will be spent on hands-on engineering work with expected ~25% spent on operational activities.
Responsibilities include:
- Own and manage daily Data Engineering operations, including data ingestion workflows, pipeline monitoring, and incident resolution.
- Oversee end-to-end pipeline health - proactively monitor, triage, and resolve failures, bottlenecks, and data quality issues across all ingestion layers.
- Organize and prioritize the engineering operations function’s daily workload - assign tasks, track progress, and remove blockers to ensure smooth operational delivery.
- Maintain and improve ingestion pipelines from various data sources into data lakes, warehouses, and real-time streaming systems.
- Act as the first point of escalation for pipeline failures, SLA breaches, and data anomalies - driving root cause analysis and permanent fixes.
- Define and enforce operational standards - runbooks, alerting thresholds, on-call procedures, and incident management practices.
- Coordinate with upstream data providers and downstream consumers to manage dependencies and communicate pipeline status.
- Lead daily stand-ups and work organization for the function under the wider Data Engineering practice.
- Mentor and guide junior engineers on operational best practices, debugging, and pipeline development.
- Collaborate with data scientists, analysts, and product partners to onboard new data sources and meet ingestion SLAs.
- Drive continuous improvement in pipeline reliability, observability, and efficiency.
- Implement best practices in data governance, data quality monitoring, and compliance across all pipelines.
- Design, build, test, and deploy Data solutions at scale, including data lakes, data warehouses, and real-time analytics.
- Lead technical delivery on use cases, plan and delegate tasks to junior team members, and oversee work from inception to final product.
Required Knowledge and Experiences
- Bachelor’s degree in Computer Science, Engineering, Statistics or a related field
- 10 years of data engineering experience with at least 3 years in lead roles.
- 6+ years of experience in Big Data technologies (e.g., Spark, Hive, Hadoop, Databricks).
- Excellent knowledge of data engineering concepts and best practices.
- Proven experience leading, mentoring and supporting junior team members.
- Ability to lead technical deliverables autonomously and guide junior data engineers.
- Ability to organize and manage daily engineering workload - task assignment, prioritization, and delivery tracking.
- Strong attention to detail and adherence to best practices and defined policies.
Essential Technical Skills
- Advanced proficiency with Apache Spark (PySpark) including tuning and performance optimisation experience.
- Proficiency in Python, Pandas (Scala/Java knowledge is desirable).
- Working knowledge of Apache Hive.
- Strong SQL knowledge and experience (T-SQL, working with SQL Server, SSMS, GCP BigQuery).
- Expertise in designing and implementing scalable data pipelines and ETL processes using the GCP data stack, including BigQuery, Dataflow, Pub/Sub, Cloud Storage, Cloud Composer, Cloud Functions, Dataproc (Spark).
- Experience with batch, real-time streaming, and ETL processes, including incident resolution and pipeline recovery.
- Experience building and managing ETL workflows using Apache Airflow, including DAG creation, scheduling, and error handling.
- Source control with Git.
- Knowledge of CI/CD concepts and experience designing CI/CD for data pipelines.
- Knowledge of Delta Lake concepts and common data formats, Lakehouse architecture.
- Software engineering principles including OOP, design patterns, SDLC, Agile, TDD, and performance optimization.
Desirable Technical Skills
- Experience designing logical data models and physical data models, including data warehouse and data mart designs.
- Relevant certifications (e.g. Google Cloud Qualified Data Engineer).
- Experience with streaming services such as Kafka is a plus.
- R & Sparklyr experience is a plus.
- Knowledge of MLOps concepts, AI/ML lifecycle management, and MLflow.
- Harness experience is a plus.
We’re also looking for the preferred skills below. Whether you are proficient or could use some brushing up, we’re happy to support your career development and growth in:
- AI engineering fundamentals.
- GCP Certifications
📌 Data Engineering Operations Lead (Maharashtra)
🏢 TransUnion
📍 Maharashtra