30 Jul
|
Taskus India
|
India
30 Jul
Taskus India
India
The Data Platform Architect is the lead designer and technical authority for our modern data ecosystem. Architect is responsible for ensuring the platform provides a high-performance, cost-effective, and secure foundation that evolves alongside business needs. This is a "design-and-direction" role where you set the standards, review engineering output, and serve as the primary bridge between complex technical systems and business value.
Work Mode : Remote
What the Architect will own:
- Platform Design & Research: Lead architectural decisions regarding compute engine selection, open-table format implementation, and tiered storage design.
- Agnostic Infrastructure: Architect a decoupled data environment that ensures interoperability across multiple engines and prevents proprietary vendor lock-in.
- Governance & Compliance: Design and oversee the implementation of automated data governance, including PII discovery, row/column-level security, and auditability.
- Standards & Frameworks: Define the "Definition of Done" for data pipelines, establishing coding standards, CI/CD patterns, and technical documentation requirements.
- Architectural Decision Records (ADR): Maintain a version-controlled repository of all consequential technical decisions, documenting the rationale, trade-offs, and long-term implications.
- Performance & FinOps: Monitor and optimize platform performance and spend, ensuring sub-second query speeds for massive user bases while maintaining a lean cloud footprint.
- Technical Stewardship:
Conduct deep-dive code and design reviews for all data models and orchestration workflows; mentor and unblock senior engineering staff.
Technical Skills & Experience:
- 8+ Years in Data Engineering / Architecture: Proven experience delivering production-grade Lakehouse environments for high-concurrency (1,000+ user) organizations.
- Up-to-date Data Stack Fluency: Extensive experience with Databricks (Lakehouse/Unity Catalog) and high-performance warehouses like Amazon Redshift or Snowflake.
- Open-Table Formats: Deep hands-on expertise with Apache Iceberg or Delta Lake, including optimization strategies for partitioning and schema evolution.
- Transformation & Modeling: Mastery of dbt (Core) for complex SQL-based modeling and PySpark or Python for sophisticated data processing.
- High-Efficiency Compute: Familiarity with vectorized/embedded engines like DuckDB for specialized or cost-sensitive processing tasks.
- Orchestration Mastery: Advanced experience with Apache Airflow, specifically in designing resilient, dependency-aware DAGs in resource-constrained environments.
- Cloud Ecosystems: Expert-level knowledge of AWS (S3, EC2, IAM) or equivalent services in Azure/GCP, with a focus on storage-compute separation.
- Experience with open-source catalog implementations like Apache Polaris.
- Knowledge of Data Ops principles and automated data quality testing frameworks.
- Experience translating technical debt and architectural roadmaps for C-level executives.
- Background in managing fixed-resource infrastructure (e.g., EC2/VM-based processing) vs. elastic serverless models.
- Kubernetes (K8s). Deep understanding of Pods, Deployments, Services, ConfigMaps, and Secrets management.
📌 Data Platform Architect (India)
🏢 Taskus India
📍 India