07 Aug
|
ChainBrain
|
India
Job Summary
We are looking for an experienced Senior Data Platform Engineer to design, build, and scale the core data platform powering AI-driven decision systems. The ideal candidate should have strong expertise in data lakehouse architecture, real-time streaming, large-scale data processing, cloud-native data platforms, and cost optimization. This is a hands-on technical role requiring ownership of the end-to-end data infrastructure that supports analytics, machine learning, and business intelligence workloads.
Key Responsibilities
- Design, build, and own the end-to-end Data Lakehouse architecture, ensuring scalability, reliability, and performance.
- Develop and maintain data ingestion pipelines from CDC sources such as PostgreSQL and DynamoDB.
- Build and operate real-time streaming frameworks to support use cases including anomaly detection, customer analytics, and supply chain intelligence.
- Design and optimize OLAP data stores for real-time and batch analytics across AI/ML and business intelligence platforms.
- Develop self-service ETL frameworks and query platforms to enable productive data access for internal teams.
- Build and maintain Reverse ETL pipelines and data movement APIs for downstream applications and business systems.
- Implement scalable workflow orchestration using tools such as Apache Airflow, YARN, and AWS EMR.
- Monitor and optimize infrastructure costs by analyzing compute, storage, and query utilization.
- Ensure high availability, performance,
security, and reliability of the enterprise data platform.
- Collaborate with engineering, analytics, AI/ML, and product teams to deliver scalable data solutions.
- Troubleshoot platform issues and provide technical leadership for critical production incidents.
Desired Candidate Profile
- 7- 15 years of experience in Data Engineering, with at least 3- 7 years building and managing enterprise-scale data platforms.
- Strong hands-on experience with Apache Spark, Kafka, Airflow, Debezium, Delta Lake/Hudi, Presto/Trino, DBT, and Airbyte.
- Expertise in AWS data services including EMR, S3, Athena, Glue, and CloudWatch.
- Experience processing terabytes of data daily and managing high-volume, event-driven architectures.
- Proven experience optimizing infrastructure costs and improving platform efficiency.
- Strong programming skills in Java, Python, and/or Scala.
- Experience leading technical teams or serving as a technical lead is preferred.
- Excellent analytical, problem-solving, and communication skills.
Preferred Skills
- Experience with OLAP technologies such as Apache Pinot, Apache Druid, or ClickHouse.
- Knowledge of Reverse ETL frameworks and data movement APIs.
- Familiarity with Feature Stores such as Feast or Feathr.
- Experience with Data Catalog and Metadata Management tools such as DataHub.
- Exposure to AI/ML data platforms and modern Lakehouse architectures.
📌 Data Platform Engineer (India)
🏢 ChainBrain
📍 India