02 Sep
|
Pashtek • Salesforce Partner | Data u0026 AI
|
India
02 Sep
Pashtek • Salesforce Partner | Data u0026 AI
India
About the Role
As our Data Architect, you’ll define the end-to-end architecture for a modern, hybrid on-premises and AWS cloud data platform. You’ll standardize our data lake/lakehouse patterns, choose the right compute engines, table formats, and catalogs, and partner with security, platform, analytics, and app teams to deliver governed, performant data products at scale.
What You’ll Do
- Own the reference architecture for our data platform across on-prem (e.g., Hadoop/Cloudera, Spark on K8s, SQL Server/Oracle) and AWS (S3, Glue/EMR, Lake Formation, Redshift, ECS/EKS).
- Design lakehouse patterns using Apache Iceberg (and/or Delta Lake/Hudi) with the right catalog (e.g., AWS Glue, Hive Metastore, Polaris/REST catalogs) and optimize for ACID, time travel, schema evolution, and multi-engine reads.
- Select and integrate compute engines: Apache Spark (Databricks/EMR), Trino/Starburst, Dremio, Snowflake, and pushdown-aware query acceleration.
- Model data (conceptual/logical/physical) and define domain-driven data products, semantic layers, and dimensional/ELT designs.
- Establish governance: data cataloging, lineage, PII classification, Lake Formation permissions, IAM roles, row/column-level security, and auditing.
- Drive performance & cost: partitioning, Z-order/cluster, predicate pushdown, file sizes/compaction, caching, and workload isolation.
- Lead migration roadmaps from on-prem EDW/DW to lakehouse (e.g., Informatica/SSIS/SAP BW → Spark/dbt/ELT), including dual-run cutovers and data validation.
- Partner with Data Engineering to standardize pipeline frameworks (Airflow/Glue Workflows/dbt/Kedro), CI/CD, and reliability (SLAs, SLOs, observability, quality rules).
- Guide streaming (Kafka/Kinesis/MSK, Spark Structured Streaming) and batch patterns; define CDC approaches (Debezium, DMS,
Fivetran).
- Create architecture docs and guardrails (naming, zones, schemas, S3 layout, encryption, backup/DR, multi-region).
- Mentor engineers and review designs for scalability, security, and operability.
Required Experience
- 7+ years in data architecture / data engineering with at least 2 large-scale platforms in production.
- Hands-on with AWS data services: S3, Glue/EMR, Lake Formation, IAM, and at least one warehouse (Snowflake or Redshift).
- Deep experience with Apache Spark and at least one of: Databricks, EMR Spark, Snowflake Snowpark, Dremio, or Starburst/Trino.
- Production experience with open table formats: Apache Iceberg (preferred), Delta Lake, or Apache Hudi; solid grasp of manifest/metadata, compaction, and schema evolution.
- Solid understanding of on-prem stacks: Hadoop/Hive, Spark on Kubernetes, SQL Server/SSIS and/or Oracle/Exadata, Netezza/Teradata helpful.
- Proven data modeling (3NF, dimensional, Data Vault), ELT/ETL design, and SQL performance tuning.
- Security and governance: RBAC/ABAC, row/column-level security, tokenization/masking, key management, KMS.
- Infrastructure-as-Code (Terraform/CloudFormation) and CI/CD for data (GitHub Actions/GitLab CI/Azure DevOps).
- Excellent communication; able to align execs, engineers, and analysts around clear roadmaps.
Nice to Have
- dbt for ELT, Airflow orchestration, Great Expectations/Deequ for data quality, OpenLineage/Marquez.
- Streaming: Kafka/MSK, Kinesis, Flink.
- Catalogs/semantic: AWS Glue Data Catalog, Unity Catalog, Amundsen/DataHub, Atlan/Collibra.
- BI/serving: DuckDB, Athena, QuickSight, Tableau/Power BI/Looker.
- Security/Compliance: SOC 2, HIPAA/PHI, GDPR, PCI;
experience with Okta/OIDC and SSO to data tools.
- Multi-tenant data platform or federated governance experience
📌 Data Architect (India)
🏢 Pashtek • Salesforce Partner | Data u0026 AI
📍 India