04 Sep
|
Neurodrift
|
Madurai
04 Sep
Neurodrift
Madurai
Senior Data Engineer, AWS Data Platform
Hiring immediately. We are looking for someone who can start now or on a short notice period.
Experience: 5+ years in data engineering
Type: Full-time
Location: Remote (India)
Working hours: 5:30 PM to 2:30 AM IST. Hard requirement, not a preference.
Stack: AWS (S3, Glue, Redshift, Athena, Lambda, SNS/SQS, SES, RDS Postgres), Delta Lake, Apache Airflow, a metadata-driven ingestion framework, SQL, Python, GitHub, GitHub Copilot
About the role
We run a modern data platform on AWS for an enterprise client in the payments space. Data moves from a raw landing zone through curated Delta layers into Redshift, driven by a metadata framework rather than hand-written one-off pipelines, across separate non-production, UAT and production environments.
Alongside the platform sits a live customer-facing service: SLA breach detection built on SNS, SQS, Lambda and Athena, with configuration-driven checks and notifications going to real customers. New customers are onboarded onto it continuously.
You would own both halves. This is a hands-on engineering role, and it is also the role that talks to the client's stakeholders directly rather than through a proxy.
This is an immediate opening on a live project, so we are moving fast on interviews and offers. If you are available on a short notice period, say so in your application.
A note on the hours. The shift is a genuine part of this role rather than a formality. Your day overlaps almost entirely with a US-based stakeholder group. It suits some people well and it is worth being sure it works for you before applying.
What you will own
The SLA alerting service. Config-driven breach checks, the detection engine on Lambda and Athena, customer notifications,
onboarding new customers onto the platform, and the parts people skip: DLQs, replay, idempotency, and not sending four thousand emails when a config goes wrong.
Source onboarding into Redshift. Taking a new source from SFTP, SAP, Salesforce, an API or a file feed through the land and trusted layers into Redshift, including schema design, primary keys, and the data validation that catches duplicates before anyone downstream does.
Production deployments and UAT. Promoting changes through non-prod, UAT and prod, running UAT rounds with client stakeholders, triaging defects, and getting sign-off.
The metadata-driven ingestion framework. Extending it and making it the default path for current sources instead of another bespoke pipeline, with control tables in RDS Postgres.
The Delta layers on S3. Table maintenance, MERGE patterns, compaction, VACUUM and retention, partitioning and file formats that stay reliable as source count grows.
The Redshift warehouse. Dimensional models, distribution and sort keys, load and upsert strategy, WLM behaviour, and a real answer for cost per query.
Orchestration in Airflow. DAG standards, retries, SLAs, alerting, dependency design, and backfills that do not become incidents.
Data correctness. Reconciliation between source and target, row count and freshness assertions,
and pipelines that fail loudly rather than shipping something wrong quietly.
The engineering bar. Code review, modelling and naming standards, documentation, and mentoring two to four engineers you do not formally manage.
Requirements
- 5+ years data engineering on production AWS
- AWS Lambda, SNS, SQS, EventBridge: event-driven pipelines, DLQs, retries, idempotency
- AWS Glue (PySpark jobs), S3, Athena
- Amazon Redshift: COPY, staging upserts, distribution keys, sort keys, WLM, query tuning
- Delta Lake on S3: MERGE, OPTIMIZE, VACUUM, retention
- Metadata-driven or config-driven ingestion frameworks; control tables; incremental watermarks
- Apache Airflow: DAG authoring, scheduling, backfills, SLA alerting
- Python, SQL
- Dimensional modelling: fact and dimension design, star schema, SCD
- Source onboarding: SFTP, SAP, Salesforce, REST APIs, file feeds
- Non-prod, UAT and production promotion; deployment coordination
- Data validation and reconciliation; row count checks; primary key and duplicate resolution
- Git, code review, CI/CD
- UAT support, defect triage, direct client stakeholder communication
- Able to work 5:30 PM to 2:30 AM IST consistently
Preferred
- RDS PostgreSQL for metadata and control tables
- AWS SES; MS Teams or Slack webhook notifications
- Amazon EMR or EMR Serverless
- AWS AppFlow, AWS DMS
- AWS Secrets Manager
- S3 cost optimisation: lifecycle policies, storage class analysis, resource tagging
- Terraform, CloudFormation or CDK
- dbt
- Alation, Lake Formation or comparable catalog and governance tooling
- Bitbucket
- QuerySurge or comparable data testing frameworks
- Payments, fintech or BFSI domain
- GitHub Copilot, Claude Code or comparable AI coding assistants
📌 Senior AWS Data Engineer (Remote) (Madurai)
🏢 Neurodrift
📍 Madurai