We are seeking a proactive and technically skilled Senior AI Engineer to join our AI CoE. This role is crucial for ensuring the reliability, scalability, and effective operation of our cutting-edge AI and Agentic AI systems in production environments. You will be responsible for deploying, monitoring, maintaining, and troubleshooting our AI agents and the underlying infrastructure. Working closely with AI developers, architects, and data scientists, you will implement and manage Agentic Ops/MLOps practices, automate operational tasks, and contribute to building a robust and resilient operational framework for our AI initiatives.
Deployment & Infrastructure Management: o Deploy, configure, and manage AI models, agentic systems, and supporting infrastructure in cloud (e.g., GCP) and on-premise environments.
o Implement and maintain CI/CD pipelines for AI/ML models and agentic applications (MLOps/Agent Ops). o Manage and optimize cloud resources, ensuring cost-effectiveness and scalability for AI workloads. o Collaborate with infrastructure teams to ensure network, storage, and compute resources meet the demands of AI systems. Monitoring, Logging & Alerting: o Develop and implement comprehensive monitoring, logging, and alerting solutions for AI agents and infrastructure to ensure high availability and performance. o Proactively identify and address potential issues, performance bottlenecks, and anomalies in production AI system