31 Jul
|
Trinity Mobility
|
Bengaluru
31 Jul
Trinity Mobility
Bengaluru
Role & responsibilities
- Lead the design, development, and implementation of scalable enterprise Data Platform solutions supporting analytics, AI/ML, and business intelligence initiatives.
- Architect and develop robust workflow orchestration pipelines using Apache Airflow for scheduling, monitoring, dependency management, and workflow automation.
- Design, build, and optimize distributed data processing applications using Apache Spark for large-scale batch and real-time data processing.
- Develop scalable ETL/ELT frameworks for ingesting structured, semi-structured, and unstructured data from multiple enterprise data sources.
- Design and implement modern Data Lake, Data Warehouse, and Lakehouse architectures using technologies such as Apache Iceberg, Delta Lake, or Apache Hudi.
- Build high-performance analytical data models and OLAP solutions using Apache Kylin to enable low-latency multidimensional analytics.
- Develop and optimize real-time analytics platforms using Apache Druid for interactive dashboards, time-series analytics, and high-speed query performance.
- Integrate streaming data pipelines using Apache Kafka or similar messaging platforms for real-time data ingestion and processing.
- Design scalable data models that support reporting, business intelligence, advanced analytics, and machine learning workloads.
- Optimize Spark applications through partitioning, caching, memory tuning, resource allocation, and query optimization to maximize performance.
- Configure, monitor, and optimize Apache Airflow environments, including DAG development, scheduling strategies, retries, alerting, logging, and failure recovery mechanisms.
- Develop reusable data engineering frameworks, metadata-driven pipelines, and common libraries to improve development efficiency.
- Implement data governance, metadata management, lineage, quality validation, and monitoring across the enterprise data platform.
- Develop APIs and data services to enable secure and efficient data consumption across applications and business platforms.
- Integrate data from databases, cloud storage, APIs, IoT platforms, enterprise applications,
and third-party systems into centralized data platforms.
- Design and optimize SQL queries, distributed processing workflows, and data storage strategies for maximum scalability and performance.
- Deploy and manage data platform components using Docker and Kubernetes, ensuring scalability, reliability, and high availability.
- Work with cloud platforms such as Microsoft Azure, AWS, or Google Cloud Platform to design, deploy, and manage cloud-native data engineering solutions.
- Implement CI/CD pipelines using Git, Jenkins, GitHub Actions, or GitLab CI to automate application deployment and infrastructure changes.
- Collaborate closely with Data Scientists, BI Developers, Product Managers, DevOps Engineers, and Business stakeholders to deliver enterprise data solutions.
- Lead architecture discussions, establish coding standards, conduct code reviews, and mentor Data Engineers on best engineering practices.
- Monitor platform performance, troubleshoot production issues, perform root cause analysis, and drive continuous improvements in platform reliability and efficiency.
- Ensure platform security, access controls, compliance, backup strategies, and disaster recovery planning.
- Prepare technical documentation, architecture diagrams, deployment guides, and operational runbooks.
- Evaluate emerging technologies and recommend enhancements to modernize the organization's data platform ecosystem.
- Drive technical strategy, innovation, and platform roadmap initiatives while ensuring alignment with organizational business objectives.
Preferred candidate profile
- 58 years of hands-on experience in Data Engineering, Big Data, or Data Platform development.
- Strong expertise in Apache Airflow for workflow orchestration, scheduling, monitoring, and pipeline automation.
- Hands-on experience with Apache Spark (Spark Core, Spark SQL, DataFrames, and performance tuning) for large-scale distributed data processing.
- Proficiency in Python, SQL, and Shell scripting for developing scalable data pipelines and automation.
- Strong understanding of ETL/ELT, Data Warehousing, Data Lakes, and modern Lakehouse architectures.
- Experience designing and implementing scalable, high-performance data platforms capable of handling large volumes of structured and unstructured data.
- Hands-on experience with Apache Kylin for OLAP analytics and multidimensional data modeling is highly preferred.
- Experience with Apache Druid for real-time analytics, time-series data processing, and low-latency query performance is an added advantage.
- Knowledge of Apache Kafka or similar streaming technologies for real-time data ingestion and processing.
- Familiarity with distributed storage technologies such as Apache Iceberg, Delta Lake, or Apache Hudi.
- Experience working with relational and NoSQL databases such as PostgreSQL, MySQL, MongoDB, Cassandra, or similar databases.
- Exposure to cloud platforms such as Microsoft Azure, AWS, or Google Cloud Platform.
- Experience with Docker, Kubernetes, Git, and CI/CD tools such as Jenkins or GitHub Actions is preferred.
- Strong understanding of data governance, data quality, metadata management, and security best practices.
- Proven ability to optimize Spark workloads, SQL queries, and large-scale data processing pipelines for performance and scalability.
- Experience in mentoring junior engineers, participating in architecture discussions, conducting code reviews, and contributing to technical decision-making.
- Solid analytical, troubleshooting, and problem-solving skills with the ability to resolve complex production issues.
- Excellent communication, stakeholder management, and collaboration skills, with the ability to work effectively in cross-functional Agile teams.
- Self-motivated, proactive, and passionate about building scalable enterprise data platforms and adopting modern data engineering technologies.
📌 Data Platform Lead (Bengaluru)
🏢 Trinity Mobility
📍 Bengaluru