10 Aug
|
Insight Global
|
Hyderabad
10 Aug
Insight Global
Hyderabad
Insight Global is seeking an AI DevOps / MLOps Engineer to support and scale the infrastructure powering enterprise AI applications and services. This role focuses on the operational side of AI systems, ensuring model-driven applications are reliable, observable, secure, and production-ready.
The ideal candidate is an infrastructure-focused engineer who has experience operating Kubernetes-based platforms, automating deployments, improving system reliability, and supporting AI workloads in production environments.
What You'll Do
- Deploy, operate, and optimize AI services running on Kubernetes environments.
- Support MLOps and AI platform tooling, including workflow orchestration, model evaluation pipelines, and AI-enabled automation services.
- Build and maintain CI/CD pipelines using modern GitOps practices.
- Improve system reliability through health checks, readiness probes, autoscaling, rollback strategies, and resiliency patterns.
- Develop observability solutions across logs, metrics, dashboards, alerts, and operational runbooks.
- Troubleshoot complex production issues across containers, services, networking, secrets management, and cloud infrastructure.
- Collaborate with AI, platform, and operations teams to establish scalable operating practices for AI systems.
- Participate in incident response activities and drive continuous improvements based on lessons learned from production events
Required Qualifications
- 5 years of experience in DevOps, MLOps, Site Reliability Engineering (SRE), Platform Engineering, or Cloud Infrastructure.
- Hands-on experience administering and troubleshooting production Kubernetes environments.
- Experience with machine learning platform technologies such as Kubeflow or similar workflow orchestration tools.
- Strong understanding of Docker, Helm, ArgoCD, GitOps methodologies, and up-to-date CI/CD practices.
- Experience supporting production systems with monitoring, alerting, incident management, and operational support processes.
- Familiarity with Azure cloud services and AI technologies, including Azure OpenAI, Azure AI Search, RAG architectures, or LLM-powered applications.
- Strong understanding of reliability engineering concepts including scaling, retries, failover, health monitoring, and service recovery.
- Experience creating technical documentation, operational procedures, and runbooks.
Preferred Qualifications
- Experience with service mesh technologies such as Istio.
- Experience with event-driven architectures and messaging technologies including Kafka, Event Hub, or Service Bus.
- Experience supporting Python-based application environments utilizing FastAPI, Uvicorn, or Gunicorn.
- Familiarity with SLO, SLI, incident response, and production support best practices.
- Experience supporting AI platforms, RAG pipelines, vector search applications, or enterprise LLM deployments.
Ideal Candidate Profile
You are an engineer who enjoys building stable, scalable platforms that enable AI applications to operate reliably in production. You have a strong operational mindset, thrive in cloud-native environments, and are comfortable troubleshooting complex infrastructure challenges. Your background combines DevOps excellence with exposure to modern AI platforms and services.
Technical Environment
- Kubernetes
- Kubeflow
- Docker
- Helm
- ArgoCD
- GitOps
- Azure
- Azure OpenAI
- Azure AI Search
- CI/CD Pipelines
- Observability & Monitoring
- RAG Architectures
- LLM-Based Services
- Service Reliability Engineering Practices
📌 AI DevOps/MLOps Engineer (Hyderabad)
🏢 Insight Global
📍 Hyderabad