03 Oct
|
RandomTrees
|
Hyderabad
03 Oct
RandomTrees
Hyderabad
bout the Role
We are seeking an experienced Data Engineer with 5+ years of hands-on expertise across the modern data stack to design, build, and maintain high-performance ELT pipelines. In this role, you will be the core technical owner of our analytics foundationleveraging Snowflake, dbt, and Apache Airflow hosted on cloud infrastructure (AWS, Azure, or GCP). You will partner closely with data analysts, analytics engineers, and business stakeholders to ensure our data models are scalable, reliable, and governed.
Key Responsibilities
- Data Pipeline Architecture & ELT: Design, develop, and optimize robust, idempotent ELT data pipelines moving data from diverse operational sources into Snowflake.
- Transformation & Modeling (dbt): Architect modular, production-grade dbt projects. Implement dimension and fact tables (Kimball/medallion architecture), write reusable Jinja macros, and establish automated testing (data freshness, uniqueness, referential integrity).
- Workflow Orchestration (Airflow): Build, monitor, and maintain production Airflow DAGs with proper error handling, SLA alerts, retries, and dependency management.
- Snowflake Administration & Optimization: Configure and optimize virtual warehouses, storage, clustering keys, zero-copy cloning, and Time Travel. Enforce cost-governance practices and resource monitors.
- Cloud Infrastructure & Ingestion: Integrate cloud storage services (e.g., AWS S3, Azure Blob/ADLS, or Google Cloud Storage) with Snowflake using Snowpipe, external stages, and event-driven architectures.
- Data Quality & Governance: Establish schema migrations, data testing frameworks, and role-based access control (RBAC), object tagging, and data masking policies.
- CI/CD & DataOps: Implement version control best practices,
automated deployment pipelines (GitHub Actions, GitLab CI, or Bitbucket Pipelines), and environment management for data assets.
Qualifications & Requirements
Must-Have:
- Experience: 5+ years of dedicated professional experience in data engineering, data warehousing, or business intelligence backend systems.
- Snowflake: Deep expertise in Snowflake architecture, query profiling, warehouse right-sizing, Snowpipe, and security implementations (RBAC, row-level security).
- dbt (Data Build Tool): Proven track record modeling complex datasets in dbt Core or dbt Cloud, using incremental models, snapshots, tests, documentation, and Jinja/SQL macros.
- Apache Airflow: Strong experience authoring complex Python DAGs, utilizing custom operators/sensors, managing pools/connections, and debugging pipeline bottlenecks.
- Any Major Cloud Platform: Solid working experience with at least one major cloud provider (AWS, Azure, or GCP), particularly with IAM roles, object storage, and serverless compute.
- Programming & SQL: Mastery of advanced SQL (window functions, CTEs, query plan optimization) and intermediate-to-advanced proficiency in Python for data processing and automation.
- Version Control & CI/CD: Hands-on experience with Git branching workflows and automated testing/deployment pipelines.
Nice-to-Have:
- Experience with streaming technologies (Apache Kafka, AWS Kinesis) or Snowflake Streaming Snowpipe.
- Exposure to Snowflake Agile Tables, Streams & Tasks, or Snowpark.
- Relevant certifications: SnowPro Core / Advanced, dbt Certified Developer, or Associate/Professional Cloud Architect/Data Engineer (AWS/Azure/GCP).
- Familiarity with data observability tools (e.g., Monte Carlo, Datafold, Elementary).
📌 Data Engineer (Hyderabad)
🏢 RandomTrees
📍 Hyderabad