About the Company Aistra is an AI transformation and deployment company operating across India, MENA, the UK and Australia. We take clients through all four stages of AI transformation: strategy and roadmap, solution build, production deployment, and operations at scale. Most firms pick one of those and call it a practice.
Delivery runs on 80 assets of registered IP, spanning accelerators, toolkits and codified domain playbooks, carried into engagements by embedded pods rather than sold as a slide pack.
Aistra
Talent builds and calibrates that workforce against defined role archetypes and places them directly into client teams. The AI and Data Engineer is one of those archetypes. On most programmes it is the one everything else ends up waiting for.
About the Role Almost every AI programme stalls in the same place. Not the model, not the interface, the data. It sits fragmented across systems nobody has documented, of uncertain quality, loosely governed, and shaped for reporting rather than for the systems that now need to consume it.
Somebody has to go in and fix that before anything else in the programme is real. That is this job. The AI and Data Engineer owns the data problem end to end, from the first look at a client's estate through to pipelines running reliably in their production environment.
This is a forward deployed role, based in India and deployed across India, MENA, the UK and Australia, for engineers with four to eight years of building behind them. You work inside the client's environment, on their systems, alongside their people, onsite or remote as the engagement requires. When a data question is asked in the room, you are the Aistra person who answers it.
There is nobody behind you to check with first. Expect to rotate across engagements, sectors and geographies, and to spend time on client sites. Every data estate is twenty years of somebody's accumulated decisions, and you will inherit a great many of them.
The role is AI-native in both directions. You build the data foundations AI systems depend on, including retrieval layers, context pipelines and evaluation datasets, and you build them using AI, because compressing turnaround is how one engineer now covers ground that used to take a team. The role is not tiered.
There is no version of this job where you build a slice of the pipeline while somebody more senior owns the architecture. You own the assessment, the design, the build, the quality, the deployment and the stability of what you shipped.
Responsibilities Assess a client's data estate quickly and honestly, and turn what you find into a feasibility position the team can commit to, including when that position is unwelcome.
Design and build the data architecture, pipelines, integrations and retrieval layers that AI solutions depend on,
and take them to production.
Make data AI-ready, meaning clean, structured, labelled, traceable and retrievable, and prove it through evaluation rather than assertion.
Own quality, lineage, governance and cost for everything you build, and instrument it so failures are found by monitoring rather than by the client.
Work directly with client stakeholders across the range, from data owners and platform teams to the business users who live with the output.
Use AI tooling as the default mode of work across schema exploration, transformation, testing, documentation and debugging, and stay accountable for the correctness of what it produces.
Stabilise and tune what you built in the weeks after go-live. You built it, so it is yours when it wobbles.
Document and hand over well enough that the client can run and extend the system without you, and without a standing call.
Qualifications A degree in computer science, engineering, statistics or a related quantitative discipline, or equivalent demonstrated capability. We weigh what you have built more heavily than what you studied or where.
Four to eight years building data systems, with explicit evidence of production ownership rather than contribution to someone else's platform.
A body of work we can look at. A GitHub profile or portfolio is the straightforward route, and for this archetype we expect one. If your work sits behind client NDAs, as much of the best work in this field does, a technical write-up, an architecture you can walk us through, or an open-source contribution serves the same purpose.
Required Skills Production ownership of data systems. At least one substantial pipeline or platform you took from first look at the data estate through to running reliably in production, and stayed with afterwards.
Core data engineering depth. Strong SQL including analytical functions and query optimisation, and data modelling across dimensional, normalised and wide-table approaches with the judgement to pick between them.
Pipeline build at production standard. ETL and ELT, batch and streaming, in Python or equivalent, as version-controlled, tested, reviewable code rather than scripts that happen to run.
Warehouse or lakehouse delivery. Hands-on production work on at least one of BigQuery, Snowflake, Databricks, Redshift or Synapse, with a view on the trade-offs between them.
Quality as a gate. Testing,
lineage and governance applied so failures are caught before the client sees them, and so you can answer where a number came from without a two-day investigation.
AI-native engineering. Fluent use of coding agents and LLM tooling across schema exploration, transformation, testing, documentation and debugging, with accountability for verifying output rather than shipping it unread. This is a core requirement of the role, not a differentiator.
Ownership and client composure. You are the Aistra person answering the data question in the room. Deliver an unwelcome feasibility position early, hold it when pushed, and act when the system is undocumented and nobody is coming to help.
Preferred Skills 4 to 8 years building data systems, with a GitHub profile, portfolio or equivalent body of work we can look at. For this archetype we expect one.
Change data capture, incremental loads, backfills, late-arriving data and slowly changing dimensions treated as routine rather than as edge cases.
Performance and cost tuning through partitioning, clustering, file formats and compaction.
Lake and lakehouse architecture with open table formats such as Delta, Iceberg or Hudi, and distributed processing on Spark or the Hadoop ecosystem.
Genuine high-volume experience at billions of rows or terabytes, including what broke when you got there, plus streaming infrastructure such as Kafka or Pub/Sub.
Enterprise source integration across ERP and finance, CRM, ticketing, core banking or policy administration, HRIS and document management, with sensible handling of pagination, rate limits, retries and idempotency.
Comfort with legacy and on-premise estates: SFTP and flat files, database replicas, systems with no supported API, and unstructured data such as documents, transcripts and logs.
Data contracts, schema evolution, master data and entity resolution across systems that disagree, and governance in practice covering access models, audit trails, retention and erasure.
Retrieval architecture across chunking strategy, embedding model selection, hybrid and semantic search, reranking and context assembly, with vector and search infrastructure such as pgvector, Pinecone, Weaviate, Qdrant or OpenSearch, and a view on when a vector store is not the answer.
Evaluation datasets and retrieval evaluation, measuring recall and precision and diagnosing whether a failure sits in the data, the retrieval or the model.
Production experience on AWS, Azure or GCP covering storage, compute, networking, identity and cost, with orchestration in Airflow, Dagster or Prefect and Git-based CI/CD.
Willingness to rotate across engagements, sectors and geographies, and to work on client sites as engagements require.
📌 AI & Data Engineer (AIDE) (Delhi)
🏢 Aistra
📍 Delhi