09 Aug
|
Hitya Global
|
Bengaluru
09 Aug
Hitya Global
Bengaluru
Responsibilities
- Design and implement asynchronous multi-agent orchestration.
- Own end-to-end latency from user message to AI response.
- Build resilient inference pipelines that gracefully degrade under load.
- Implement intelligent request routing and load balancing for AI workloads.
- Migrate critical AI conversation flow from monolith to dedicated services.
- Implement WebSocket/streaming infrastructure for real-time chat.
- Design circuit breakers and fallback strategies for AI model failures.
- Build comprehensive observability for AI system performance.
- Optimise credit data retrieval and caching strategies.
Requirements
- 3-5 years building production systems handling > 10k concurrent users.
- Proven experience with async/event-driven architectures (not just REST APIs).
- Hands-on experience scaling ML/AI inference in production.
- Deep understanding of caching strategies (Redis, in-memory, CDN).
- Experience with message queues and real-time communication protocols.
- Built systems integrating multiple LLM/AI models in production.
- Experience with AI model serving frameworks (TensorFlow Serving, Triton, etc.).
- Understanding of AI inference optimisation (batching, caching, model quantisation).
- Knowledge of conversation state management and context handling.
- Has debugged production issues under high AI inference load.
(ref:hirist.tech)
📌 AI Engineer - LLM/Agentic AI (Bengaluru)
🏢 Hitya Global
📍 Bengaluru