We are seeking experienced AI-OPS Engineers with a robust blend of Site Reliability Engineering (SRE), Cloud Operations, Automation, Observability, and Generative AI expertise. Candidates should be capable of designing and implementing intelligent operational solutions that improve incident detection, diagnosis, remediation, and operational efficiency.
Required Technical Skills
AI/ML & Generative AI
- OpenAI / GenAI solutions on GCP or AWS
- Machine Learning fundamentals
- MLOps
- AI Agents / Agentic AI
- Retrieval Augmented Generation (RAG)
- Enterprise Search
- AI Governance
- Prompt Engineering
SRE & IT Operations
- Reliability Engineering
- Incident Management
- Production Support
- ITSM platforms (Remedy preferred)
- Problem Management
- Operational Excellence
Preferred Experience
- 7+ years in Infrastructure Operations, SRE, Platform Engineering, DevOps, or AIOps
- 3+ years working with Cloud Platforms (AWS/GCP)
- Experience implementing automation and self-healing solutions
- Experience building AI-assisted operational workflows
- Experience with enterprise monitoring and observability platforms
Preferred Certifications
- AWS Certified Solutions Architect / DevOps Engineer
- Google Professional Cloud Architect
- Certified Kubernetes Administrator (CKA)
- ITIL Foundation
- AI/ML or Generative AI certifications
Key Responsibilities
- Design and implement self-healing operational workflows.
- Develop AI-assisted RCA and operational intelligence capabilities.
- Build and maintain knowledge and runbook co