16 Sep
|
Tanisha Systems
|
Bengaluru
16 Sep
Tanisha Systems
Bengaluru
Mandatory : One round will be F2F
Exp 7 + years
Location : PAN INDIA
Role Summary:
Primary focus is on data pipeline monitoring, establishing CI/CD processes, and resolving pipeline failures. Leads debugging, structured logging, and the design of robust error-handling frameworks. Additionally supports data ingestion, transformations, and Redshift data pipeline development as a secondary responsibility.
- Deep understanding of data pipeline monitoring, proactive alerting, and observability (e.g., AWS CloudWatch)
- Solid experience building CI/CD processes and deployment automation
- Advanced pipeline debugging, root-cause analysis, and exception handling
- Apache Airflow (DAG writing, scheduling, task dependencies, and failure management)
- Strong experience in SQL, Python (pandas, json, logging) and PySpark for troubleshooting and automation
- Version control with Git and branching strategies
- Working knowledge of AWS Glue, Redshift, and S3
- Familiarity with JSON parsing, flattening, and automated data validation checks
- SQL proficiency for investigating data issues and validating tables
- EMR exposure
- GitLab, Airflow, and CloudWatch
Key Skills to Look For:
- Monitoring & Debugging: Proven ability to track job failures, parse driver/executor logs, and set up automated alerts for pipeline SLA breaches.
- CI/CD & DevOps: Hands-on experience configuring GitLab CI to automate the testing, validation, and deployment of data models and pipeline code.
- Error Handling: Demonstrated experience writing robust, fault-tolerant Python/PySpark code with structured logging and automated retry mechanisms.
- Data Engineering Foundation:
Solid competency in maintaining, optimizing, and supporting existing AWS data workflows (S3, Glue, Redshift) alongside the core reliability tasks.
Below is the summary of the client expectation.
1. Production Support / Troubleshooting
- Ability to troubleshoot web services and data pipelines.
- Understanding of incident diagnosis and step-by-step problem resolution.
- Some candidates demonstrated strong troubleshooting skills, while others struggled.
1. Data Engineering & Pipeline Design
- Designing end-to-end data pipelines.
- Spark architecture, optimization, and troubleshooting.
- Data modelling and warehouse design.
Key Observations
- Most candidates performed well on Spark-related topics.
- Several candidates lacked depth in data pipeline design and operational best practices.
Major Gaps Identified
- Backfill design and execution: Most candidates could not explain how to safely and scalably backfill historical data, a common real-world requirement.
- Handling duplicate data and idempotency: Limited understanding of preventing duplicates when pipelines are rerun.
- Data warehouse design: Weaknesses around Iceberg/Delta tables, partitioning strategies, and query optimization.
- Production readiness: Some candidates struggled with practical troubleshooting and operational ownership scenarios.
Overall Candidates generally had solid Spark knowledge but were weaker in real-world data engineering practices such as backfills, idempotent pipelines, warehouse optimization, partitioning, and production support. These areas were the main differentiators between stronger and weaker candidates.
📌 AWS Data Engineer_PAN INDIA (Bengaluru)
🏢 Tanisha Systems
📍 Bengaluru