DevOps Specialist Engineer – SRE, Cloud & Applied AI (Hyderabad)

DevOps Specialist Engineer – SRE, Cloud & Applied AI (Hyderabad)

29 Aug
|
Clarus Advisers
|
Hyderabad

29 Aug

Clarus Advisers

Hyderabad

Company Overview

Our client is a leading technology organization focused on building scalable, cloud-native platforms and intelligent software solutions. The organization combines software engineering, cloud technologies, Site Reliability Engineering, and Applied AI to deliver highly resilient and production-ready products at scale.

Position Overview

We are looking for a DevOps Specialist Engineer with strong experience in Site Reliability Engineering, cloud platform engineering, software development, and production operations. The ideal candidate will have hands-on experience operating large-scale distributed systems across Azure, AWS, or GCP, along with exposure to AI/ML, GenAI, LLMOps/MLOps, observability, performance engineering, and cloud cost optimization. The role requires a strong engineering mindset and the ability to build reliable, secure, scalable, and highly automated production platforms.

Responsibilities

- Design, build, operate, and continuously improve large-scale cloud-native production systems.
- Define and own SLIs, SLOs, SLAs, error budgets, and reliability objectives.
- Lead production incident response, on-call operations, root-cause analysis, and reliability improvements.
- Build and manage CI/CD pipelines, Kubernetes platforms, infrastructure automation, and multi-environment deployments.
- Implement Infrastructure as Code using Terraform and deployment automation using tools such as ArgoCD.
- Develop production-grade observability across metrics, logging, and distributed tracing.
- Work with technologies such as Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace, Splunk, Azure Monitor,



and AWS CloudWatch.
- Operate AI/ML and GenAI workloads in production, addressing reliability, performance, model drift, output variance, and train/serve skew.
- Support MLOps/LLMOps platforms and AI control-plane capabilities such as model gateways and guardrails.
- Implement Kubernetes/Docker-based solutions and cloud-native networking across multiple environments.
- Conduct load and performance testing using tools such as LoadRunner, k6, or JMeter.
- Implement chaos engineering, capacity planning, autoscaling, and resilience testing.
- Drive cloud and AI FinOps, including GPU, inference, and token-cost attribution and optimization.
- Implement security controls including RBAC, least privilege, secrets management, deployment approvals, and segregation of duties.
- Collaborate with software engineering, security, data/AI, and product teams to improve platform reliability and operational excellence.

Skills & Experience

- 6–9 years of experience in DevOps, SRE, Site Reliability Engineering, Platform Engineering, or Software Engineering.
- Bachelor's degree in Computer Science, Software Engineering, Data Science, Machine Learning, or a related discipline.
- Strong programming experience in one or more of Python,



C#/.NET, Go, Java, or Bash.
- Strong hands-on experience with Azure, AWS, or GCP; Azure/AWS preferred.
- Experience with Kubernetes, Docker, Terraform, and cloud-native architectures.
- Robust understanding of CI/CD, GitHub, Azure DevOps (ADO), and ArgoCD.
- Experience with production observability and monitoring using tools such as Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace, Splunk, CloudWatch, or Azure Monitor.
- Strong understanding of SRE principles, SLIs, SLOs, SLAs, error budgets, incident management, and production on-call operations.
- 3+ years of experience operating or supporting large-scale production systems.
- Experience with AI/ML or GenAI workloads in production and familiarity with Azure OpenAI, AWS Bedrock, or Vertex AI.
- Knowledge of MLOps/LLMOps, MLflow, LangFuse, LangSmith, or equivalent AI/agent orchestration platforms.
- Experience with load/performance testing, capacity planning, autoscaling, and chaos engineering.
- Understanding of cloud cost optimization/FinOps, including AI/GPU/inference and token-cost management.
- Knowledge of DevSecOps, RBAC, secrets management, security controls, and environment integrity.
- Strong understanding of OOP/OOD, data structures, algorithms, code instrumentation, and software engineering practices.
- Ability to understand and work with business context, sequence, activity, state, entity-relationship, and data-flow diagrams.
- Exposure to XP, Lean, SRE, and AI-augmented/spec-driven software development is an advantage.

📌 DevOps Specialist Engineer – SRE, Cloud & Applied AI (Hyderabad)
🏢 Clarus Advisers
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: devops specialist engineer – sre, cloud & applied ai (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: devops specialist engineer – sre, cloud & applied ai (hyderabad) / hyderabad