29 Sep
|
Trine Infotech
|
Gurugram
29 Sep
Trine Infotech
Gurugram
We are looking for a Lead DevOps / AI Engineer to join our Site Reliability Engineering (SRE) team. The role combines DevOps, cloud infrastructure, automation, AIOps and MLOps to build highly reliable and scalable production environments.
Key Responsibilities:
- Design, build and manage scalable cloud infrastructure using Infrastructure as Code.
- Manage and optimize AWS, Kubernetes and Docker environments.
- Build and maintain CI/CD and MLOps pipelines for ML models and LLMs.
- Implement AI-driven anomaly detection, predictive alerting and automated incident response.
- Develop automation and self-healing workflows to improve system reliability and reduce MTTR.
- Manage observability using Datadog, Prometheus, Grafana and OpenTelemetry.
- Optimize AI/ML workloads, including GPU/CPU utilization, inference performance and cloud costs.
- Implement security, disaster recovery and reliability best practices.
- Lead technical decisions and provide guidance across DevOps, SRE, automation and MLOps.
Required Skills:
- 8+ years of experience in DevOps, SRE or Cloud Platform Engineering.
- Solid hands-on experience with AWS, Kubernetes and Docker.
- Proficiency in Terraform, Pulumi or CloudFormation.
- Strong scripting/programming skills in Python, Go or Bash.
- Experience with GitHub Actions, GitLab CI or ArgoCD.
- Hands-on experience with Prometheus, Grafana, Datadog or OpenTelemetry.
- Hands-on exposure to AI/ML deployment pipelines.
Good to Have:
- Experience with MLflow, Kubeflow, Ray, vLLM or Triton Inference Server.
- Experience deploying LLMs or Generative AI solutions in production.
- Knowledge of vector databases such as Pinecone, Milvus or Qdrant.
- Familiarity with LangChain or LlamaIndex.
📌 Lead DevOps AI Engineer (Gurugram)
🏢 Trine Infotech
📍 Gurugram