15 Sep
|
TalentDryft
|
Bengaluru
15 Sep
TalentDryft
Bengaluru
What you'll do
- Own agentic-system architecture. Define how agents reason, plan, use tools, maintain state, recover from failure, collaborate, and involve humans. Make autonomy boundaries and behavioural contracts explicit.
- Design harness behaviour. Specify and prototype the control loop around the model: context assembly, planning and execution, tool selection, memory, retries, reflection, escalation, checkpoints, and termination conditions.
- Make quality measurable. Build the evaluation strategy for non-deterministic systems across offline test sets, simulations, trajectory and tool-use evaluation, online experiments, human review, and production monitoring.
- Architect retrieval and context. Design retrieval, ranking, context assembly, grounding, citation, permissions, and freshness patterns that remain reliable across enterprise knowledge sources and changing data.
- Turn intent into buildable specifications. Produce architecture decisions, behavioural specifications, interface contracts, evaluation plans, and reference implementations precise enough for teams to implement consistently.
- Prototype and de-risk. Write production-quality proofs of concept for difficult agent behaviours, retrieval strategies, or evaluation methods before a pod commits to an approach. You still build.
- Guard execution fidelity. Lead design reviews, identify behavioural and quality drift early, make evidence-based trade-offs, and keep implementation aligned with the intended architecture.
- Raise the engineering bar. Mentor senior engineers, establish reusable patterns and review standards, and serve as the VP's depth partner on AI/ML feasibility, risk, and opportunity.
Core depth requirement (non-negotiable) This is an agentic-engineering architect role, not a generalist AI advisory position. The successful candidate must demonstrate substantial hands-on depth in the following areas:
- Agentic engineering and harness behaviour
- control loops, planning and execution, state and memory, tool use, structured outputs, error recovery, long-running tasks, human-in-the-loop patterns,
multi-agent coordination, and safe autonomy.
- Complex evaluations
- task-success and trajectory evaluation, tool-call correctness, groundedness, retrieval quality, safety, latency and cost, simulation, regression suites, online measurement, human calibration, and careful use of model-based judges.
- Information retrieval and context engineering
- ingestion and indexing, hybrid and semantic retrieval, metadata and permission filtering, ranking and reranking, graph or relationship-aware retrieval, context compression, grounding, citations, and retrieval evaluation.
- Production AI quality
- prompt and model selection, guardrails, observability and tracing, auditability, failure analysis, and the cost-latency-quality trade-offs required for systems that serve real users.
- Architecture through implementation
- the ability to move between research papers, experiments, Python code, APIs, design documents, and production constraints without losing technical precision.
Must-have qualifications
- Approximately 8+ years building production software or ML systems, with several years architecting complex systems that span multiple components and teams.
- Demonstrated hands-on delivery of production agentic systems used by real users; experience that is limited to prompt engineering, demos, or vendor configuration is not sufficient.
- Deep practical knowledge of agent harnesses or runtimes, including the behaviour of the execution loop rather than only application-level use of a framework.
- Evidence of designing sophisticated evaluation systems for non-deterministic, tool-using, or retrieval-augmented applications and using results to drive engineering decisions.
- Strong information-retrieval architecture experience, including ranking, grounding, context construction, security or permissions, and quantitative retrieval evaluation.
- Robust Python and software-engineering ability, with the credibility to review implementation details and create reference code when needed.
- Technical authority to guide senior engineers without relying on positional management; you influence through depth, evidence, and clear design.
- Strong written communication: architecture documents, behavioural specifications, and evaluation plans that engineers can implement unambiguously.
Preferred experience
- LLM inference and serving
- exposure to throughput and latency engineering, batching, caching, model routing, quantization, GPU economics, and observability for hosted or self-managed models.
- Fine-tuning and adaptation
- exposure to data curation, supervised fine-tuning, parameter-efficient methods, preference optimization, synthetic data, and the evaluation needed to justify adaptation over prompting or retrieval.
- Conversational and multimodal AI
- experience with dialogue state, turn-taking, streaming, voice or speech interfaces, interruption handling, personalization, and end-to-end conversational quality.
- Experience working across more than one major AI ecosystem such as AWS, Azure, Databricks, Google Cloud, or leading open-source stacks.
- Background in a services, systems-integration, or accelerator context where reusable engineering patterns must work across multiple clients and environments.
What this role is - and isn't
- Is: the senior technical IC for agent intelligence and behaviour; the person who sets the architecture for harness behaviour, retrieval, and evaluation and keeps delivery evidence-driven.
- Isn't: a platform-infrastructure architect, a research-only role, or a people manager. The Platform Architect owns runtime and substrate architecture; this role partners closely on the interfaces between the harness and the platform.
📌 AI Architect (Bengaluru)
🏢 TalentDryft
📍 Bengaluru