10 Aug
|
Smart Analytica
|
Pune
10 Aug
Smart Analytica
Pune
Principal Architect
Experience: 12–15 years
Location: Pune (onsite)
About the Role
We are looking for a Principal Architect to lead the design and development of a Lakehouse solution end to end. This is an opportunity to architect and build a reusable data platform framework from the ground up.
You will own the technical architecture, write code for the core framework, mentor a growing engineering team, and work closely with leadership on the overall technology direction. The ideal candidate is a hands-on architect with a proven track record of building complex frameworks and platforms using open-source big-data technologies such as Spark, Hadoop, and modern Lakehouse stacks (e.g., Iceberg, Trino, Databricks, etc).
Key Responsibilities
Architecture (40%)
- Own the end-to-end architecture of the Lakehouse platform and its reusable framework components
- Design a metadata-driven data engineering framework built on Spark SQL and PySpark
- Architect a pipeline engine supporting dynamic SQL generation, incremental loads, CDC, SCD implementations, and schema evolution
- Design Kubernetes-native deployment with multi-tenancy, resource isolation, scaling strategy, high availability, and disaster recovery
- Define API contracts, plugin/extension architecture, and integration patterns
- Drive and document key architecture decisions (ADRs)
Hands-on Engineering (40%)
- Build core framework components: pipeline engine, dynamic SQL generation, data quality checks, error handling and retry mechanisms
- Implement observability across the platform
- Optimize Spark job performance (partitioning, shuffle tuning, memory management, query plans)
- Establish and implement CI/CD pipelines, automated testing (unit and integration), and release practices
- Embed data governance and security into the platform (access control, auditability, cataloging)
Engineering Leadership (20%)
- Mentor and guide junior engineers; conduct design and code reviews
- Define coding standards, engineering best practices, and documentation norms
- Contribute to broader technical direction and team growth
- Continuously raise the bar on engineering quality and delivery practices
Required Experience
- 12–15 years of software engineering experience overall
- 8+ years in data platform engineering
- 5+ years as an Architect on Spark/Hadoop-based platforms
- Proven success building complex, reusable frameworks or platforms (not just applications) using open-source technologies
- Expert-level proficiency in Python, PySpark, Spark SQL, and Spark internals
- Deep knowledge of Apache Iceberg , distributed systems, Kubernetes, Docker , REST APIs, microservices, Linux, Git, and CI/CD
Technical Stack
- Data Platform: Apache Spark, Apache Iceberg, Apache Ozone, Apache Ranger, Trino, Kafka, Airflow,
- Platform Engineering: Kubernetes-native deployments, Helm, Terraform, Prometheus, Grafana, containerization, multi-tenant architecture, storage management, HA/DR
- Framework Capabilities You'll Build: JSON-driven pipelines, metadata-driven ETL, dynamic SQL generation, incremental loads & CDC, SCD handling, schema evolution, data quality, lineage, plugin architecture, unit testing, performance optimization
What We're Looking For
- A hands-on architect who designs systems and writes the code that proves the design
- Strong architecture and design instincts with the judgment to keep frameworks simple and extensible
- Complex problem-solving ability and speed in picking up recent technologies
- Experience with Agile and spec-driven development
- Clear communication — able to explain architecture decisions to engineers and leadership alike
Nice to Have
- Contributions to open-source data projects (Spark, Iceberg, Trino, etc.)
- Experience with Databricks or comparable managed Lakehouse platforms
- Prior experience taking a platform or framework from concept to production
📌 Principal Architect (Pune)
🏢 Smart Analytica
📍 Pune