09 Oct
|
Tekwissen India
|
India
09 Oct
Tekwissen India
India
Overview :
Tek Wissen is a global workforce management provider throughout India and many other countries in the world. The job opportunity described below is for one of our clients, who has developed a core competence in creating and deploying cost-effective capabilities using an offshore-centric business model.
Position: Lead Backend AI Engineer
Location: Remote
Job Type: Contract
Duration: 6 Months
Work Type: Remote
Job Description:
- Seeking an elite Lead Backend AI Engineer to architect and build the core multi-agent execution runtime, domain-adapted model pipelines, and deterministic safety harnesses for our enterprise Agentic Ops platform.
- In this role, you will own the engine room of autonomous IT and Sec Ops: translating high-level operational intent into reliable, multi-agent reasoning DAGs, grounded retrieval mechanisms, and sandboxed infrastructure actions.
- You will engineer durable agent state machines, optimize low-latency inference pipelines, train and fine-tune open-weight models for domain-specific infrastructure tasks, and construct fail-safe policy boundaries that prevent destructive executions.
- The ideal candidate pairs high-performance systems programming in Go (Golang) and Python with source-level mastery of agentic frameworks, open-source model ecosystems, and rigorous AI-driven evaluation methodologies.
Core Responsibilities:
- Autonomous Multi-Agent Runtime Engineering: Architect and implement fault-tolerant, stateful agent runtimes using Go and Python. Build advanced execution patterns including hierarchical agent-to-agent delegation, planner-executor-critic loops, multi-step reflection, and dynamic schema-constrained tool selection.
- Domain Model Training, Fine-Tuning & Distillation: Design data synthesis and fine-tuning pipelines (LoRA, QLoRA, full fine-tuning, direct preference optimization) to adapt open weight models (Llama, Mistral, Qwen, Deep Seek) for specialized infrastructure tasks-such as CLI syntax parsing, telemetry anomaly correlation, and structured action generation.
- Deterministic Safety Harness & Execution Guardrails: Engineer execution sandboxes, preflight dry-run validators, blast-radius limiters, and automated state-rollback mechanisms. Ensure agents cannot commit hallucinated commands or execute unverified, destructive payloads against live networks or compute clusters.
- Open-Source Agent Tooling & Web-Execution Harnesses: Build extensible tool-calling layers integrating open-source agentic and computer/browser-use tooling (e.g., Open Claw, Browser-Use, Headless CDP) alongside traditional enterprise controller SDKs, gRPC streams, and SSH/NETCONF interfaces.
- Hybrid Contextual Memory & GraphRAG Substrate: Develop scalable, low-latency memory pipelines combining live topology graphs (Neo4j, AWS Neptune), vector search (pgvector, Qdrant, Milvus), and temporal session stores to provide agents with grounded infrastructure context.
- AI-Native SDLC & Continuous Eval Pipelines: Implement automated,
CI/CD-integrated evaluation suites (LLM-as-a-judge, synthetic scenario generation, regression benchmarks) to rigorously test agent trajectory accuracy, tool call precision, and safety compliance prior to production deployment.
Required Qualifications & Technical Depth:
- Experience & Track Record: 8+ years of production backend engineering in high-throughput distributed systems, with at least 3+ years specifically architecting LLM application runtimes, autonomous agent frameworks, or ML inference infrastructure.
- Dual-Language Systems Mastery (Go & Python): Expert proficiency in Go (Golang) for high-concurrency, low-latency microservices, worker pools, and memory-protected network programming (goroutines, channels, context cancellation). Advanced expertise in Python (asyncio, typing, Pydantic) for AI/LLM orchestration and model training.
- Agentic Frameworks & Runtime Internals: Source-code level fluency with agentic orchestration frameworks (Lang Chain, Lang Graph, Semantic Kernel, Llama Index) or custom actor-model state machines. Deep understanding of how to manage durable distributed execution, state persistence across restarts, and graceful failure recovery.
- Model Training, Fine-Tuning & Local Inference: Direct hands-on experience fine-tuning and distilling open-source models using Hugging Face PEFT, TRL, Axolotl, or Deep Speed. Production expertise deploying and tuning high-throughput inference engines (vLLM, TensorRT-LLM, Ollama) and quantization formats (AWQ, GPTQ, GGUF).
- Open-Source Tooling Ecosystem: Practical experience working with open-source autonomous agent tools and execution harnesses (e.g., Open Claw, browser automation agents, sandbox container environments like gVisor/Firecracker).
- Enterprise Infrastructure & Protocol Fluency: Hands-on experience building integrations against enterprise infrastructure interfaces: REST APIs (Cisco Intersight, Catalyst, Nexus, Splunk), streaming telemetry (gNMI/Open Telemetry), SNMP, and CLI over SSH.
- AI-Driven SDLC & Evals Rigor: Proven experience establishing automated evaluation frameworks (Deep Eval, Ragas, promptfoo) to continuously benchmark agent decision-making, track tool hallucination rates, and validate regression safety across model updates.
- Knowledge Retrieval & Data Systems: Deep experience architecting production GraphRAG and hybrid search pipelines: dense/sparse hybrid retrieval (BM25 + embeddings), rerankers(Cross-Encoders, Cohere), and transactional data stores with graph/vector capabilities.
Mandatory:
Experience
- 8+ years of production backend engineering in high-throughput distributed systems
- 3+ years architecting LLM application runtimes, autonomous agent frameworks, or ML inference infrastructure Programming Languages
- Go (Golang): high concurrency microservices, worker pools, goroutines, channels, context cancellation
- Python: asyncio, typing, Pydantic, for AI/LLM orchestration and model training Agentic Frameworks & Runtime
- Hands-on depth in at least one of Lang Chain, Lang Graph, Semantic Kernel or Llama Index, or custom actor-model state machines
- Durable distributed execution, state persistence across restarts, graceful failure recovery
- Multi-agent patterns: planner-executor-critic loops, agent-to-agent delegation, schema-constrained tool selection Model Fine-Tuning & Inference
- Fine-tuning open-weight models (LoRA/QLoRA, DPO) using Hugging Face PEFT, TRL, Axolotl or Deep Speed
- Deploying and tuning inference engines: vLLM, TensorRT-LLM or Ollama
- Quantization formats: AWQ, GPTQ or GGUF Safety & Execution Guardrails
- Sandboxed execution (gVisor, Firecracker or similar container isolation)
- Dry-run validation, blast-radius limiting, rollback mechanisms for destructive actions Enterprise Infrastructure Integration
- REST API integrations (Cisco Intersight, Catalyst, Nexus, Splunk or equivalent)
- Streaming telemetry (gNMI/Open Telemetry), SNMP, CLI over SSH Evals & AI-Driven SDLC
- Automated evaluation frameworks: Deep Eval, Ragas or promptfoo.
- Tracking tool hallucination rates and regression safety across model updates, integrated into CI/CD
Retrieval & Data Systems
- Production GraphRAG and hybrid search (BM25 + embeddings)
- Rerankers (Cross-Encoders, Cohere)
- Graph and vector stores: Neo4j or Neptune, plus pgvector, Qdrant or Milvus
Good to Have (Optional):
- Open Claw, Browser-Use, Headless CDP
- NETCONF, gRPC streams
- Full fine-tuning and model distillation at scale
- Temporal session stores
- Experience with Llama, Mistral, Qwen or Deep Seek specifically
- Synthetic scenario generation and LLM-as-a-judge pipelines
Required Competencies
- Must possess excellent communication skills oral and written
- Must possess knowledge of latest technology trends
- Must be a keen learner should be able to drive "Self Learning"
- Must practice principle of "First Time Right"
- Must have an Eye for Details
- Must have high Customer Orientation
- Must be adaptable to working in multiple / matrix work environment
- Must possess good systems thinking
- Must possess good negotiation, analytical and interpersonal skills. Good leadership & team player qualities.
- High on personal integrity with ability to establish relationships and work in teams and should be able to influence stakeholders. Should poses independence, robust ethics and resilience.
Educational Qualification:
- ME (IT, Computer), BE (IT, Computer), MCA, MSC-IT, BCA
Tek Wissen Group is an equal opportunity employer supporting workforce diversity
📌 Lead Backend AI Engineer (India)
🏢 Tekwissen India
📍 India