10 Sep
|
Randstad
|
Hyderabad
10 Sep
Randstad
Hyderabad
Job Title: AI-MLOps SRE Lead Engineer
Location: Hyderabad. (Hybrid)
Discover your role
- Drive service reliability, availability, and performance across multi-cloud environments, establishing SLOs, SLIs, error budgets, and reliability standard methodologies.
- Design, build, and scale enterprise ML platform infrastructure using technologies such as Dataiku, Amazon SageMaker AI, Databricks, and Google Vertex AI.
- Develop AI-driven observability capabilities using anomaly detection, predictive analytics, and automated remediation solutions to proactively identify and resolve operational issues.
- Lead the implementation and monitoring of LLM, SLM, RAG, and AI Agent platforms, ensuring performance, governance, operational efficiency, and scalability.
- Design and implement Infrastructure as Code, CI/CD pipelines, self-healing systems, and platform automation capabilities to improve engineering productivity and operational resilience.
- Architect enterprise ChatOps solutions that integrate operational events, AI workflows, observability platforms, and automated remediation capabilities.
- Partner with Data Science, AI Engineering, and Platform teams to deliver secure, scalable, and production-ready AI/ML solutions.
- Evaluate emerging AI-native operational technologies and integrate innovative solutions that enhance platform reliability, engineering efficiency, and business value.
- Conduct technical debt assessments, identify architectural risks, and provide strategic recommendations to improve enterprise platform maturity.
- Serve as a technical leader and trusted advisor, mentoring engineers and influencing reliability engineering, MLOps, and cloud platform strategy across the organization.
- This role requires
- Bachelor's degree in Computer Science, Information Technology, Engineering, Data Science, Artificial Intelligence,
or a related subject area; Master's degree preferred.
- 6-8 years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or related technology fields with enterprise-scale delivery experience.
- Strong hands-on experience operating across two or more major cloud platforms, including AWS, GCP, or Azure.
- Deep expertise with ML platform technologies including Databricks, Amazon SageMaker AI, Dataiku, and Google Vertex AI.
- Proven experience implementing end-to-end ML workflows including model training, deployment, experiment tracking, monitoring, and pipeline orchestration.
- Advanced proficiency in Infrastructure as Code tools such as Terraform, Pulumi, AWS CDK, and modern CI/CD automation practices.
- Solid programming and scripting skills in Python, Go, Bash, or similar languages.
- Experience building enterprise observability solutions using Prometheus, Grafana, Datadog, OpenTelemetry, distributed tracing, metrics, and logging platforms.
- Demonstrated expertise in anomaly detection, predictive analytics, automated remediation, and AI-assisted operational capabilities.
- Proven experience designing and implementing enterprise ChatOps solutions and AI-enabled operational workflows.
- Strong ability to assess technical debt, influence technical strategy, and drive platform modernization initiatives.
- Experience using AI tools, LLM-powered assistants, and AI Agents to enhance engineering operations and productivity.
- Experience with Kubernetes and container orchestration platforms such as EKS, GKE, or AKS preferred.
- Familiarity with MLOps technologies including Kubeflow, Feast, MLflow, and model evaluation frameworks such as LangSmith, RAGAS, Evidently AI, or Weights & Biases preferred.
- Knowledge of cloud cost optimization, policy-as-code, compliance automation, and multi-cloud governance practices preferred.
📌 AI-MLOps SRE Lead Engineer (Hyderabad)
🏢 Randstad
📍 Hyderabad