We are seeking experienced AIoPS Engineers with a strong blend of Site Reliability Engineering (SRE), Cloud Operations, Automation, Observability, and Generative AI expertise. Candidates should be capable of designing and implementing intelligent operational solutions that improve incident detection, diagnosis, remediation, and operational efficiency.
Required Technical Skills
AI/ML & Generative AI
- OpenAI / GenAI solutions on GCP or AWS
- Machine Learning fundamentals
- MLOps
- AI Agents / Agentic AI
- Retrieval Augmented Generation (RAG)
- Enterprise Search
- AI Governance
- Prompt Engineering
- Reliability Engineering
- Incident Management
- Production Support
- ITSM platforms (Remedy preferred)
- Problem Management
- Operational Excellence
Preferred Experience
- 7+ years in Infrastructure Operations, SRE, Platform Engineering, DevOps, or AIOps
- 3+ years working with Cloud Platforms (AWS/GCP)
- Experience implementing automation and self-healing solutions
- Experience building AI-assisted operational workflows
- Experience with enterprise monitoring and observability platforms
Preferred Certifications
- AWS Certified Solutions Architect / DevOps Engineer
- Google Qualified Cloud Architect
- Certified Kubernetes Administrator (CKA)
- ITIL Foundation
- AI/ML or Generative AI certifications
Key Responsibilities
- Design and implement self-healing operational workflows.
- Develop AI-assisted RCA and operational intelligence capabilities.
- Build and maintain knowledge and runbook copilots.
- Improve monitoring, observability, and incident response processes.
- Automate operational tasks using Infrastructure as Code and orchestration tools.
- Collaborate with SRE, platform, and application teams to improve reliability and operational efficiency.
- Evaluate and implement Agentic AI solutions for autonomous operations.
Ideal Candidate Profile
Candidates should possess a strong combination of:
- SRE/Operations background
- Cloud Engineering expertise
- Automation and Infrastructure as Code experience
- Observability and Incident Management knowledge
- AI/ML and Generative AI capabilities
- Excellent troubleshooting, RCA, and problem-solving skills