31 Aug
|
Crest Infosystems
|
India
31 Aug
Crest Infosystems
India
5-10 year(s)
Apache KafkaAWSCI/CDClient CommunicationETLSQL
Description
We're hiring a Data Engineer to design, build, and operate the data pipelines and platforms behind our client projects — from ingestion through transformation to warehouse/lake delivery. This is a client-facing delivery role: you'll work directly with international stakeholders to understand data requirements and translate them into production-grade pipelines, not one-off scripts. We're looking for someone who brings DataOps discipline to the work — CI/CD for pipelines, automated testing, monitoring, and version control — rather than treating operational reliability as an afterthought. AWS experience is a plus, since our US partner firm is an AWS Advanced Partner, but it is not mandatory — we care more about your ability to reason about data architecture across any cloud stack.
Responsibilities
- Design, build, and maintain scalable ETL/ELT pipelines from diverse source systems into client data warehouses or lakes, with robust error handling and alerting.
- Build and manage orchestration workflows (Airflow, Prefect, Dagster, or cloud-native equivalents) to schedule and monitor pipeline execution.
- Develop transformation layers using dbt, Spark, or Python-based frameworks to produce clean, well-modeled, queryable datasets.
- Apply DataOps practices to the pipeline lifecycle — CI/CD for data pipelines, automated testing, version control discipline, and environment promotion — to bring engineering rigor and speed to data delivery.
- Implement data quality checks, validation rules, and pipeline observability/monitoring so failures and data anomalies are caught before they reach downstream consumers.
- Design and optimize data warehouse/lake architecture (Snowflake, Databricks, BigQuery, Redshift, Synapse, or similar), balancing performance, cost, and scalability.
- Work directly with international clients to understand data requirements and translate them into pipeline and platform designs; contribute to scoping and estimation.
- Collaborate with data scientists, AI engineers, analysts, and client stakeholders to ensure the data platform supports downstream analytics, reporting, and ML/AI use cases.
- Build streaming/real-time pipelines (Kafka, Flink, Spark Streaming, or similar) where client use cases require low-latency data.
- Document data lineage, pipeline logic, and architecture decisions for maintainability and audit purposes.
- Review code and architecture, mentor junior data engineers, and contribute to reusable data engineering accelerators and technical standards for the practice.
Requirements
- 5+ years of production data engineering experience building and operating pipelines at scale.
- Strong SQL — complex query writing, performance tuning, and schema/data modeling (dimensional modeling, normalization).
- Strong Python for pipeline and transformation development.
- Hands-on experience with a cloud data warehouse or lakehouse platform (Snowflake, Databricks, BigQuery, Redshift, or Synapse).
- Production experience with a workflow orchestration tool (Airflow, Prefect, Dagster, or a cloud-native equivalent).
- Working knowledge of DataOps practices: CI/CD for data pipelines, automated pipeline testing, version control discipline, and environment promotion workflows.
- Experience with at least one major cloud platform (AWS, GCP, or Azure) for data infrastructure.
- Experience implementing data quality checks and pipeline monitoring/observability.
- Comfort working with both SQL and NoSQL data stores.
- Strong verbal and written communication — comfortable working directly with international clients across cultures and time zones.
- Bachelor's or Master's degree in Computer Science, Data Engineering, or a related field, or equivalent practical experience.
Natural Abilities
- Systems thinking — traces data end-to-end, anticipates failure points
- Attention to detail — catches quality issues and anomalies early
- Ownership mindset — treats reliability and monitoring as core, not an afterthought
- Comfort with ambiguity — turns vague client requirements into solid designs
- Clear, cross-cultural communication — works smoothly with international stakeholders
- Sound architectural judgment — balances performance, cost, and scalability
- Collaborative — works well with data scientists, analysts, and junior engineers
Preferences
- AWS experience specifically (Glue, Redshift, EMR, Kinesis, Lambda, etc.) — a plus, since our US partner firm holds AWS Advanced Partner status, but not required.
- Production dbt experience, ideally managing a meaningful number of models across multiple environments.
- Streaming/real-time architecture experience (Kafka, Flink, Spark Streaming).
- Big data framework experience (Apache Spark, Hadoop) for large-scale distributed processing.
- Containerization and orchestration (Docker, Kubernetes) for data platform deployment.
- Basic ML/analytics literacy to collaborate effectively with data science and AI teams.
- Cloud or data platform certifications.
- Prior experience in an IT services or consulting environment delivering to overseas clients.
Benefits
- Excellent Salary
- Friendly Environment
- Work-life Balance
- 5 Days Working
- Versatile Office Timings
- Employee Friendly Leave Policies
📌 Data Engineer (India)
🏢 Crest Infosystems
📍 India