29 Sep
|
Trine Infotech
|
Gurugram
29 Sep
Trine Infotech
Gurugram
About the Role
We are looking for an Infrastructure AI Engineer to join our SRE team and build, deploy, and optimize high-performance cloud and bare-metal infrastructure for AI/ML and LLM workloads.
Key Responsibilities
- Design and manage scalable infrastructure for AI model training and real-time inference.
- Manage GPU/TPU clusters, Kubernetes, Ray, Slurm, and distributed storage.
- Build AIOps capabilities using Prometheus, Datadog, OpenTelemetry and automated anomaly detection.
- Develop self-healing infrastructure and automated remediation workflows.
- Define and monitor SLOs, SLAs, and Error Budgets.
- Build CI/CD and MLOps infrastructure for continuous model deployment and validation.
- Optimize GPU/CPU utilization, infrastructure performance, and cloud costs.
- Perform capacity planning, load testing, and stress testing for large-scale inference workloads.
Required Skills
- 6+ years in SRE,
Systems Engineering, Infrastructure Engineering, or Cloud Platform Engineering, preferably with AI/ML workloads.
- Solid hands-on experience with Kubernetes, containers, and GPU provisioning.
- Advanced experience with Terraform, Pulumi, Ansible, or CloudFormation.
- Programming experience in Python, Go, or C++.
- Strong knowledge of Prometheus, Grafana, OpenTelemetry, and modern observability platforms.
Good to Have
- Experience with vLLM, TensorRT-LLM, or Triton Inference Server.
- Knowledge of DeepSpeed, Megatron-LM, or Ray Cluster.
- Understanding of NVLink, InfiniBand, RoCE, GPU interconnects, and memory optimization.
- Exposure to Pinecone, Milvus, Qdrant, or other vector databases.
📌 Infrastructure AI Engineer (Gurugram)
🏢 Trine Infotech
📍 Gurugram