10 Oct
|
Tekwissen India
|
India
10 Oct
Tekwissen India
India
Overview :
TekWissen is a global workforce management provider throughout India and many other countries in the world. The job chance described below is for one of our clients, who has developed a core competence in creating and deploying cost-effective capabilities using an offshore-centric business model. Position: Engineering Manager Location: Remote Job Type: Contract
Duration: 6 Months Work Type: Remote
Job Description:
- Engineering Manager & Principal Systems Architect (AgenticOps Core)
- We are seeking a hands-on Engineering Manager & Principal Systems Architect to serve as the technical anchor, customer-facing architect, and engineering leader for our enterprise AgenticOps platform.
- In this player-coach role, you will lead a specialized engineering pod building production grade autonomous runtimes for enterprise IT and SecOps. You will own the architectural blueprint, write production code, fine-tune models, and work directly with enterprise customers and technical partners across NVIDIA and Cisco.
Core Responsibilities:
- Architecture & Hands-On Engineering: Lead, design, and personally code an event-driven, high-throughput multi-agent platform in Golang for closed-loop IT self-healing and SecOps containment.
- Customer & Partner Advisory: Serve as primary technical lead for enterprise clients and design partners (Cisco, NVIDIA). Lead deep-dive architectural workshops and translate complex operational workflows into multi-agent specifications.
- NVIDIA AI Stack & Inference Acceleration: Architect GPU-accelerated inference pipelines using NVIDIA NIM microservices, NeMo Framework/Guardrails, and TensorRT-LLM across DGX clusters, hybrid cloud, and edge.
- Model Adaptation & Fine-Tuning: Lead data curation and fine-tuning pipelines (PEFT, LoRA/QLoRA, distillation) to adapt open-weight models (Llama, Mistral, Qwen, DeepSeek) for IT/SecOps tasks.
- Agent Frameworks & Open-Source Ecosystem: Build durable agent runtimes using LangChain, LangGraph, or custom state machines. Leverage open-source agentic and execution tooling (e.g., OpenClaw, Nemoclaw) and local inference engines (vLLM, Ollama).
- Safety, Governance & AI-Native SDLC: Engineer deterministic safety harnesses (dry-run preflights, blast-radius containment, rollbacks) and establish an AI-augmented SDLC (agentic PR reviews, synthetic evals, LLM-as-a-judge).
- Team Leadership & Integration: Recruit, mentor, and lead an elite pod; enforce rigorous code reviews, evals, and OpenTelemetry observability while integrating with enterprise substrates (Cisco Cloud Control, Intersight, Nexus, Splunk).
Required Qualifications:
- Experience & Scope: 10+ years architecting scalable distributed systems and backend infrastructure, with 3+ years as a hands-on player-coach leading engineering teams.
- Customer-Facing Acumen: Proven ability to interface with enterprise architects and leaders on technical discovery, system architecture, and complex deployments.
- Core Languages: Deep production mastery of Go (Golang) for high-concurrency, low-latency microservices; strong proficiency in Python for AI/ML pipelines and orchestration.
- NVIDIA AI Stack: Hands-on experience with NVIDIA AI Enterprise, NIMs, NeMo Guardrails, TensorRT-LLM, and Triton Inference Server on modern GPU architectures.
- LLM Engineering & Fine-Tuning: Practical experience with fine-tuning toolchains (Hugging Face, PEFT/LoRA, Axolotl, DeepSpeed), alignment (DPO/RLHF), quantization, and agent frameworks (LangChain, LangGraph, DSPy).
- Open-Source Tooling: Experience self-hosting open models (vLLM, Ollama) and integrating open-source agentic/browser automation tools (OpenClaw, Browser-Use).
- Distributed Systems & Safety: Strong background in event-driven streaming (Kafka, NATS, gRPC), vector/graph databases (pgvector, Qdrant, Neo4j), and zero-trust safety execution guardrails.
- Domain Fluency: Working knowledge of enterprise IT fabrics (SDN, data center networking), observability platforms (Splunk, ThousandEyes), and SOC incident lifecycles.
Leadership & Role:
- Engineering Manager, Principal Systems Architect, Player-Coach
- Hands-on Technical Leadership, Team Mentoring, Code Reviews
- Customer-Facing Architect, Technical Discovery, Architectural Workshops
Core Languages & Backend
- Golang (Go), Python
- Distributed Systems, Event-Driven Architecture,
Microservices
- High-Concurrency, Low-Latency, High-Throughput
- gRPC, Kafka, NATS
Agentic AI
- AgenticOps, Multi-Agent Systems, Autonomous Agent
- LangChain, LangGraph, DSPy, Custom State Machines
- OpenClaw, Browser-Use
- Closed-Loop Automation, IT Self-Healing, SecOps Containment
NVIDIA AI Stack:
- NVIDIA AI Enterprise, NVIDIA NIM, NeMo Framework, NeMo Guardrails TensorRT-LLM, Triton Inference Server
- DGX, GPU-Accelerated Inference
LLM Engineering & Fine-Tuning
- Fine-Tuning, PEFT, LoRA/QLoRA, Distillation, Quantization
- DPO, RLHF, Alignment
- Hugging Face, Axolotl, DeepSpeed
- Open-Weight Models (Llama, Mistral, Qwen, DeepSeek)
- vLLM, Ollama, Self-Hosted LLMs
Data & Storage
- Vector Databases (pgvector, Qdrant), Graph Databases (Neo4j)
Safety & Governance
- Zero-Trust, Safety Guardrails, Dry-Run Preflight
- Blast-Radius Containment, Rollbacks
- LLM-as-a-Judge, Synthetic Evals, Agentic PR Reviews, AI-Native SDLC
Observability & Enterprise Integration
- OpenTelemetry, Splunk, ThousandEyes
- Cisco Cloud Control, Intersight, Nexus
- SDN, Data Center Networking, SOC Incident Lifecycle.
Experience Requirements
- 10+ Years in Distributed Systems / Backend Infrastructure
- 3+ Years Player-Coach / Team Lead
- Enterprise Customer Engagement (Cisco, NVIDIA partners)
Required Competencies
- Must possess excellent communication skills oral and written
- Must possess knowledge of latest technology trends
- Must be a keen learner should be able to drive "Self Learning"
- Must practice principle of "First Time Right"
- Must have an Eye for Details
- Must have high Customer Orientation
- Must be adaptable to working in multiple / matrix work environment
- Must possess good systems thinking
- Must possess good negotiation, analytical and interpersonal skills. Good leadership & team player qualities.
- High on personal integrity with ability to establish relationships and work in teams and should be
- able to influence stakeholders. Should poses independence, robust ethics and resilience.
Years of Experience:
- 10+ years of relevant work experience with a reputed organization.
Educational Qualification:
- ME (IT, Computer), BE (IT, Computer), MCA, MSC-IT, BCA
TekWissen Group is an equal opportunity employer supporting workforce diversity
📌 Engineering Manager (India)
🏢 Tekwissen India
📍 India