oDeploy configure and manage AI models agentic systems and supporting infrastructure in cloud eg GCP and onpremise environments
oImplement and maintain CICD pipelines for AIML models and agentic applications MLOpsAgent Ops
oManage and optimize cloud resources ensuring costeffectiveness and scalability for AI workloads
oCollaborate with infrastructure teams to ensure network storage and compute resources meet the demands of AI systems
Monitoring Logging ing
oDevelop and implement comprehensive monitoring logging and ing solutions for AI agents and infrastructure to ensure high availability and performance
oProactively identify and address potential issues performance bottlenecks and anomalies in production AI systems
oTrack key operational metrics and create dashboards for system health and performance
Incident Response Troubleshooting
oProvide operational support for production AI systems including incident response root cause analysis and resolution of technical issues
oDevelop and maintain runbooks and standard operating procedures for common operational tasks and incident management
oParticipate in oncall rotations as needed to support critical AI services
Automation Operational Excellence
oAutomate routine operational tasks deployment processes and system maintenance activities using scripting eg Python Bash and automation tools
oContribute to the development and enforcement of operational best practices security standards and compliance requirements for AI systems
oWork with development teams to improve the deployability manageability and observability of AI applications
Collaboration Documentation
oCollaborate effectively with AI developers data scientists AI architects and other stakeholders to ensure smooth transitions from development to production
oMaintain clear and comprehensive documentation for system configurations operational procedures and troubleshooting guides
oProvide feedback to development teams on operational aspects and system performance
Required Qualifications Experience
Bachelors degree in Computer Science Information Technology Engineering or a related technical field
47 years of experience in a MLOps or Agent Ops role preferably supporting AIML or dataintensive applications
Handson experience with cloud computing platforms eg Google Cloud Platform especially Vertex AI and managing cloudbased infrastructure
Proficiency in scripting languages such as Python Bash or PowerShell for automation
Experience with CICD tools and practices eg Bitbucket GitLab CI GitHub Actions
Familiarity with containerization technologies eg Docker Kubernetes and orchestration
Experience with monitoring and logging tools eg Prometheus Grafana ELK Stack Datadog Google Cloud Monitoring Langfuse
Understanding of networking concepts security best practices and infrastructureascode IaC principles eg Terraform Ansible
Robust troubleshooting and problemsolving skills with an analytical mindset
Excellent communication skills and ability to work collaboratively in a team environment
A proactive approach to identifying and resolving issues and improving system reliability
Preferred Qualifications Experience
Masters degree in a relevant field
Specific experience in MLOps or Agent Ops including deploying and managing machine learning models or large language model applications in production
Familiarity with AIML frameworks and libraries eg TensorFlow PyTorch scikitlearn