07 Aug
|
Innova Solutions
|
Hyderabad
07 Aug
Innova Solutions
Hyderabad
Job Title: Senior Open Source Data Engineer
Job Summary
We are looking for a highly skilled Senior Open Source Data Engineer to design, build, and scale a modern enterprise data platform using open-source technologies. The ideal candidate will have extensive experience in distributed data architectures, real-time data processing, data governance, and large-scale data platform engineering. This role will drive the design and implementation of secure, scalable, and high-performance data solutions supporting analytics, business intelligence, and operational systems across global environments.
Key Responsibilities
Data Platform Architecture & Engineering
- Design and implement scalable Data Lake architectures using Medallion Architecture (Bronze, Silver, Gold layers).
- Build and maintain end-to-end batch and real-time data ingestion pipelines.
- Architect scalable, distributed data platforms supporting enterprise-wide data requirements.
- Ensure seamless data flow from source systems to analytics, APIs, and downstream applications.
- Support multi-region and globally distributed data environments.
Open Source Technology Leadership
- Lead implementation and optimization of open-source data platforms including:
- Data Ingestion: Debezium, Apache Airflow
- Streaming & Processing: Apache Kafka (KRaft), Apache Flink
- Storage: Hadoop, Apache Iceberg
- Query & Data Access: Trino, Apache Gravitino
- Metadata & Governance: Apache Atlas
- Visualization: Apache Superset
- Drive containerized deployments using Docker and Kubernetes.
- Promote open standards and extensible architectures while minimizing dependence on proprietary technologies.
Data Modeling & Governance
- Design robust and scalable data models for complex business domains.
- Implement enterprise-wide data governance frameworks.
- Establish data quality, metadata management, lineage tracking, and access control mechanisms.
- Ensure compliance with organizational security and data protection standards.
- Maintain consistency, integrity, and quality of data across multiple systems and regions.
Data Processing & Integration
- Design and optimize real-time and near real-time processing pipelines.
- Integrate data platforms with APIs, B2B systems, and analytics platforms.
- Support high-volume transactional data processing across multiple geographies.
- Enable reliable and scalable data movement across enterprise platforms.
Analytics Enablement
- Build data foundations that support self-service analytics.
- Enable reporting and visualization through Apache Superset and other analytics platforms.
- Ensure data readiness, performance, governance, and secure access for analytics users.
Required Skills & Qualifications
- Bachelor's Degree (Any Discipline).
- 8+ years of experience in Data Engineering and Data Platform development.
- Proven expertise in designing and implementing large-scale distributed data platforms.
- Strong hands-on experience with:
- Apache Kafka
- Apache Flink
- Hadoop Ecosystem
- Apache Iceberg
- Apache Airflow
- Debezium
- Apache Atlas
- Apache Superset
- Robust understanding of:
- Data Warehousing Concepts
- Medallion Architecture
- Data Modeling
- Data Governance
- Data Quality Management
- Security & Access Control
- Experience with Change Data Capture (CDC) and Real-Time Streaming Architectures.
- Strong programming skills in Python, Java, or Scala.
- Experience working in globally distributed teams and environments.
Preferred Skills
- Experience in Logistics, Supply Chain, or Transportation domain.
- Knowledge of multi-region compliance and regulatory requirements.
- Exposure to AI/ML data pipelines and feature engineering.
- Experience integrating with platforms such as Snowflake, Databricks, or similar analytics ecosystems.
- Contributions to or experience working with open-source communities/projects.
Desired Competencies
- Strong leadership and mentoring capabilities.
- Excellent stakeholder management and communication skills.
- Strategic thinking with a hands-on execution mindset.
- Strong analytical and problem-solving abilities.
- High ownership and accountability.
- Focus on engineering excellence and business value delivery.
Employment Details
- Role: Senior Open Source Data Engineer
- Experience: 8+ Years
- Industry: IT Services / Product Development / Data Engineering
- Employment Type: Full-Time
- Education: Any Graduate
📌 Senior Data Engineer (Hyderabad)
🏢 Innova Solutions
📍 Hyderabad