31 Jul
|
Neurodrift
|
India
Lead Data Engineer, AWS Data Platform
Hiring immediately . We are looking for someone who can start now or on a short notice period.
Experience: 5+ years in data engineering, 1 to 2 of those leading Type: Full-time Location: Remote (India) Working hours: 5:00 PM to 2:00 AM IST. Hard requirement, not a preference.
Stack: AWS (S3, Glue, Redshift, Athena, Lake Formation, SNS/SQS, RDS Postgres), Apache Airflow, a metadata-driven ingestion framework, SQL, Python, GitHub, VS Code, GitHub Copilot
About the role
We run a modern data platform on AWS that feeds analytics and AI/ML teams at an enterprise client. Data moves from a raw landing zone through curated layers into Redshift, driven by a metadata framework rather than hand-written one-off pipelines, and it runs across both non-production and production environments. We want one person accountable for the design, the standards, and the roadmap of that platform.
Senior enough to say no to a bad model before it ships, and calm enough to run a backfill without turning it into a war room.
This is an immediate opening on a live project, so we are moving fast on interviews and offers. If you are available on a short notice period, say so in your application.
The evening shift is real and it is the reason this role is hard to fill. Your day overlaps almost entirely with a US-based stakeholder group, and this role talks to them directly rather than through a proxy. If that overlap does not work for you long term, please do not apply. We would rather lose you now than in month three.
What you will own
- The metadata-driven ingestion framework. Extending it, hardening it, and making it the default path for new sources instead of another bespoke pipeline.
- The raw to land to Redshift flow. Layer contracts, partitioning, file formats, compaction, and load patterns that stay reliable as source count grows.
- The Redshift warehouse. Dimensional models, distribution and sort keys, WLM behaviour, and a real answer for cost per query.
- Orchestration in Airflow. DAG standards, retries, SLAs, alerting, dependency design, and backfills that do not become incidents.
- Non-prod and prod parity. Promotion process, environment configuration, and the discipline that keeps the two from drifting apart.
- Governance in Lake Formation. Fine-grained access, row and column controls, and a defensible answer to "who can see what and who approved it".
- Event-driven flows on SNS and SQS, including the parts people skip: DLQs, replay, ordering, idempotency.
- RDS Postgres. Schema design, indexing, and query performance for the operational and metadata layers.
- The contract with analytics and ML. Curated datasets and models those teams can build on without messaging you first.
- The engineering bar. Code review, modelling and naming standards, documentation, and mentoring 2 to 4 engineers.
Must have
- 5+ years building production data platforms, including 1 to 2 years as a lead or tech lead accountable for other engineers' output, not just your own.
- Hands-on AWS data stack: S3, Glue (jobs,
crawlers, catalog), Redshift, Athena, Lake Formation, and RDS Postgres.
Ownership, not exposure.
- SNS and SQS used inside real data or event-driven pipelines, including failure handling and replay.
- Experience with metadata-driven or config-driven ingestion frameworks, and a transparent view of why they beat one-off pipelines at scale.
- Airflow in production. You have authored DAGs, run backfills, and been the person paged when they failed.
- SQL you can defend line by line, plus genuine dimensional modelling background (Kimball or equivalent).
- Redshift performance work: distribution and sort key decisions, load strategy, and diagnosing a slow query rather than throwing compute at it.
- Comfortable working across separate non-production and production environments with a controlled promotion process.
- Python for pipeline and framework development, with Git-based collaboration and code review.
- Proven remote leadership: async written communication, running reviews, and holding a technical position with stakeholders you have never met in person.
- Ability to work 5:00 PM to 2:00 AM IST consistently.
Nice to have
- Curated data layers or a feature store built for ML teams.
- Infrastructure-as-code for data infra: Terraform, CDK, or CloudFormation.
- Cost governance: Redshift concurrency scaling and WLM tuning, Glue DPU sizing, S3 lifecycle policies.
- dbt or a comparable transformation framework.
- Working fluently with AI coding assistants such as GitHub Copilot.
We review applications daily and respond within 48 hours. If you fit the brief and can start soon, apply now.
📌 Senior AWS Data Engineer (Remote) (India)
🏢 Neurodrift
📍 India