Note: Kindly attend the interview based on HR recruiter confirmation.
Job description
Key Responsibilities
- Deploy and manage LLM-based applications across cloud and hybrid environments.
• Build and maintain production-grade inference pipelines for LLM workloads.
• Design and implement LLMOps pipelines for model, prompt, and configuration lifecycle management.
• Implement monitoring for latency, throughput, cost, errors, drift, and hallucinations.
• Build dashboards and alerts to ensure SLA and reliability targets.
• Optimize inference cost and performance using caching, batching, routing, and autoscaling strategies.
• Enforce security, access control, audit logging, and governance controls for LLM usage.
• Collaborate with Responsible AI teams to operationalize safety and compliance guardrails.
• Enable AI Engineers and Fullstack teams through reusable LLMOps frameworks and platforms.
Required Skills & Experience
- Strong understanding of LLMs, Generative AI architectures, and inference workflows.
• Hands-on experience running LLM-based systems in production environments.
• Robust experience with cloud platforms (Azure, AWS, or GCP).
• Hands-on experience with Docker, Kubernetes, and container orchestration.
• Experience with CI/CD for ML or AI applications.
• Programming and automation skills in Python or similar languages.
• Familiarity with MLOps / LLMOps tooling, monitoring, and logging systems.
📌 Walk-in || Gen AI Engineer (Noida)
🏢 HCLTech
📍 Noida