28 Sep
|
Infosys
|
Bengaluru
Key Responsibilities
- Lead the architecture and implementation of lakehouse and analytics solutions using Iceberg, Doris, and Trino for scalable querying and reporting.
- Design and maintain Iceberg table layouts, partitioning strategies, schema evolution patterns, and data lifecycle management (compaction, snapshots, retention).
- Build and optimize distributed query workflows in Trino, including connector configuration, query tuning, resource governance, and workload management.
- Develop and optimize analytical data models and ingestion patterns leveraging Doris for high-performance OLAP workloads.
- Implement robust batch/stream processing pipelines using Spark, ensuring correctness, scalability, and cost efficiency.
- Establish performance benchmarks, monitor SLAs, and troubleshoot production issues across compute, storage, and query layers.
- Drive best practices for data quality, reliability, and operational excellence through automation, documentation, and runbooks.
- Mentor engineers, conduct design reviews, and lead technical decision-making aligned with long-term platform goals.
Minimum
Qualifications:
- Bachelor’s or Master’s degree in BTECH, MTECH, MCA, MSC or a related field.
- 6–8 years of experience in data engineering, analytics engineering, or building distributed data platforms.
- Strong hands-on expertise with Iceberg,
including table design, partitioning, schema evolution, and maintenance operations.
- Strong hands-on expertise with Trino for federated/distributed querying, performance tuning, and operational troubleshooting.
- Strong hands-on expertise with Doris for OLAP use cases, data modeling, and query performance optimization.
- Proven experience building data pipelines using Spark in production environments.
- Solid understanding of distributed systems, query execution concepts, and data storage formats for analytics workloads.
Preferred
Qualifications:
- Experience designing end-to-end lakehouse architectures integrating Iceberg with multiple compute engines and downstream consumers.
- Advanced expertise in query optimization techniques (statistics, partition pruning, file sizing, caching strategies) across Trino and OLAP systems.
- Experience with Spark optimization (shuffle tuning, join strategies, adaptive execution) and building reusable pipeline frameworks.
- Robust operational ownership: monitoring, alerting, incident management, and capacity planning for analytics platforms.
- Ability to lead cross-team technical initiatives, influence standards, and improve platform adoption through enablement and documentation.
📌 Iceberg, Doris, Trino (Bengaluru)
🏢 Infosys
📍 Bengaluru