30 Jul
|
Axiom Global Technologies
|
New Delhi
30 Jul
Axiom Global Technologies
New Delhi
AI/ML & Knowledge Platform Engineer
Retrieval, knowledge modeling, and grounded generation Full-time
About the role
We are hiring a founding-level platform engineer to build the trusted data and retrieval plane behind our AI products. You will own how unstructured enterprise content and structured records are ingested, modeled, resolved into a canonical knowledge layer, and served to LLM-driven services with full provenance. The central engineering problem is making generated outputs traceable to the exact evidence that supports them, at production quality.
This is a builder role, not a research-only or dashboard role. You will write production pipelines, design canonical schemas, implement entity resolution, and stand up hybrid retrieval and grounded-generation evaluation that other engineers and AI services depend on.
What you will work on
Ingestion and pipeline engineering
- Build connectors for APIs, databases, event feeds, secure file transfer, customer exports, and document repositories.
- Ingest and normalize PDF, DOCX, PPTX, XLSX, email, and structured records while preserving source metadata, version, access restrictions, and licensing boundaries.
- Design idempotent, restartable pipelines with schema evolution, backfills, retries, dead-letter handling, reconciliation, and observability.
Canonical models, ontology, and entity resolution
- Own canonical schemas and data contracts across organization, workforce, requirement, solution, pricing, and evidence domains.
- Resolve organizations, people, roles, titles, skills, technologies, and identifiers across inconsistent sources while preserving match confidence and reviewability.
- Model time and change: current versus historical attributes, effective dates, revisions, and superseded records using durable IDs and crosswalk tables.
Retrieval, knowledge, and provenance services
- Build the document/record processing layer for chunking, metadata enrichment, embeddings, and retrieval.
- Support hybrid retrieval across structured filters, full text, vectors, and relationships while enforcing tenant and document-level permissions.
- Expose facts, evidence, confidence, and lineage to downstream AI services without coupling them to raw source schemas.
Applied LLM quality and evaluation
- Build grounded-generation datasets, retrieval evaluation (recall/precision, citation coverage), and hallucination/error analysis.
- Stand up replayable evaluation corpora and regression tests so model and pipeline changes are measured, not guessed at.
- Work with self-hosted and API LLMs behind a routing/proxy layer; reason about cost, latency, and where a smaller local model is sufficient.
- Expose data quality, freshness, failed ingestion, duplicate/match confidence, and provenance coverage as operational metrics.
What we are looking for
Required
- Production data/backend/search systems: 3+ years building data, backend, search, or knowledge systems with senior ownership of architecture and operations, not ticket-level ETL.
- Solid Python and SQL: expert relational modeling plus practical experience with object storage, queues/events, pipeline orchestration, APIs, and cloud deployment.
- Entity resolution and modeling: hands-on deduplication, taxonomy/ontology design, slowly-changing/temporal data, schema evolution, lineage, and source reconciliation.
- Retrieval systems in production: full text, vectors, metadata filters, ranking, and reranking; comfort with graph-shaped models where they help.
- Data-quality discipline: ability to instrument pipelines, debug silent corruption, perform safe backfills, and reject a convenient pipeline that creates an untraceable data product.
- Azure, Kubernetes, PostgreSQL/pgvector, OpenSearch/Elasticsearch, a graph database, dbt, and an orchestrator (Dagster/Airflow/Temporal); depth in a coherent subset matters more than breadth.
- Evaluation and observability for RAG or agentic systems: groundedness, retrieval recall/precision, citation coverage, and regression testing.
- Experience fine-tuning, distilling, or serving open-weight models (vLLM or similar) and routing between local and hosted models.
Preferred
- Early-stage instinct: ship a correct v1, document why, and improve it without waiting for a platform team to appear.
- Unstructured document processing: extracting and serving information from enterprise documents metadata, section structure, tables, revisions, and citations.
Not the right profile
- A warehouse/BI specialist whose primary output is dashboards and reports.
- A research-only knowledge-graph or NLP profile that has not operated pipelines in production.
- A prompt engineer who treats ingestion, identity, provenance, and quality as someone elses problem.
How we evaluate
Expect a working session: given a sample document set, a structured feed, and an export with inconsistent organizations and roles, design the canonical model, source lineage, entity resolution, ingestion/retry strategy, permission model, and a service that lets an AI layer generate a recommendation with citations with explicit quality metrics and a plan for a source that silently changes its schema.
📌 AI/ML & Knowledge Platform Engineer (New Delhi)
🏢 Axiom Global Technologies
📍 New Delhi