Site reliability Engineer Devops (Hyderabad)

Site reliability Engineer Devops (Hyderabad)

06 Sep
|
Arminus
|
Hyderabad

06 Sep

Arminus

Hyderabad

Observability & Monitoring • Architect and operate enterprise-scale observability platforms on AWS/Azure, covering microservices, Lambda/serverless functions, and Kubernetes-based workloads.

• Build full-stack observability solutions using Prometheus, Grafana, CloudWatch, Datadog, Splunk, and Fluentd — covering metrics, logs, and distributed traces.

• Define and implement SLIs, SLOs, and error budgets; build actionable dashboards and alerting policies aligned to business and technical objectives.

• Standardize application logging across engineering codebases by designing and championing reusable Python logging modules and log aggregation pipelines.

• Instrument applications and services with distributed tracing to identify performance bottlenecks and reduce MTTR. Automation & SRE • Design and build reusable automation frameworks standardized as team-wide baselines for all operational automation development.

• Automate high-volume operational workflows (ServiceNow tickets, VDI lifecycle, user provisioning) using Python, PowerShell, and REST APIs.

• Implement zero-touch automation with approval and rollback mechanisms, Managed Identity security controls, and audit logging.

• Build and own CI/CD pipelines using Jenkins, AWS CodePipeline, and Azure DevOps for rapid, safe software delivery.

• Provision and manage infrastructure using Terraform and CloudFormation, enforcing IaC standards across the organization.

Cloud Platform

Engineering • Design and operate Kubernetes-based workloads on AWS EKS and Azure AKS, supporting high-transaction production environments.

• Optimize cloud costs and platform performance — including PostgreSQL query tuning, Redis-based retry workflows, and Lambda function efficiency.

• Integrate automation services securely using Azure API Management (APIM) and AWS API Gateway as governed entry points.

• Ensure enterprise security compliance using IAM, RBAC, Azure Managed Identity, and secure API design patterns reviewed by Security and Architecture teams.

Incident

Management & Reliability • Lead root-cause analysis (RCA) for critical production incidents, ensuring 100% uptime SLAs for mission-critical workloads.

• Develop and maintain incident response runbooks, post-mortem processes,



and reliability improvement roadmaps.

• Proactively identify and resolve performance degradation, capacity risks, and reliability gaps before they impact production. GenAI & Innovation • Evaluate and integrate GenAI tooling to enhance operational efficiency, including LLM-based automation using AWS Bedrock for ticket classification and event triage.

• Drive innovation initiatives through PoC development and internal engineering engagement programmes. Leadership & Collaboration • Mentor junior and mid-level engineers; establish team standards, best practices, and engineering documentation.

• Collaborate with architects, managers, directors, and security teams on platform and observability architecture decisions.

• Champion a culture of reliability, automation-first thinking, and continuous improvement across the engineering organization.

Required Skills &

Experience Core Requirements • Minimum 7+ years of hands-on experience in SRE, DevOps, or Platform Engineering roles.

• Cloud Platforms (Advanced): AWS (EKS, Lambda, DynamoDB, RDS, EC2, SQS, CloudWatch, Bedrock) and/or Azure (Azure Functions, Azure AD, Managed Identity, Azure Monitor, APIM).

• Python (Advanced): production-grade scripting for automation, tooling, backend services, and logging frameworks.

• Observability Tooling (Advanced): Prometheus, Grafana, Datadog, Splunk, CloudWatch, Fluentd, or equivalent.

• Kubernetes (Advanced): EKS or AKS in production; Helm, container lifecycle, and workload troubleshooting.

• Infrastructure as Code (Advanced): Terraform and/or CloudFormation for consistent, auditable infrastructure.

• CI/CD (Advanced): Jenkins, Azure DevOps, AWS CodePipeline — building and maintaining production pipelines.

• Security & Governance (Advanced): IAM/RBAC, Managed Identity, secure API design, and audit logging in enterprise environments.





• Proven track record of leading RCA processes and delivering reliability improvements in high-transaction-volume environments.

• Excellent communication skills with the ability to collaborate across engineering, architecture, and business stakeholders. Automation & Scripting • Proficiency in Bash and PowerShell for operational automation.

• Experience with REST API integration and enterprise workflow automation (e.g., ServiceNow, ticketing platforms).

• Familiarity with database performance tuning, particularly PostgreSQL long-running query optimization. Nice to Have • Hands-on experience with GenAI / LLM-based automation (AWS Bedrock, Azure OpenAI, or similar).

• Experience with ServiceNow automation workflows using JavaScript and REST APIs.

• Familiarity with AIOps or ML-based anomaly detection and monitoring solutions.

• Exposure to chaos engineering practices (Chaos Monkey, Gremlin, LitmusChaos).

• Prior experience in a consulting or client-facing environment (Big 4 / System Integrator a plus).

• Knowledge of FinOps principles and cloud cost optimization strategies.

• Experience in multi-cloud or large-scale enterprise environments spanning Fortune 500 organizations.

Preferred

Certifications • AWS Certified Solutions Architect – Associate or Professional • Microsoft Azure Administrator / Azure Developer / Azure AI Engineer • Certified Kubernetes Administrator (CKA) or CKAD • HashiCorp Certified: Terraform Associate • Datadog Fundamentals / Splunk Core Certified User Behavioral & Soft Skills • Strong problem-solving mindset with the ability to diagnose complex distributed systems issues under pressure • Excellent written and verbal communication — able to translate technical findings to non-technical stakeholders, including directors and architects.

• Proactive self-starter with a bias for automation and eliminating toil at scale.

• Collaborative team player who thrives in fast-paced, cross-functional enterprise environments.

• Ownership mentality — takes end-to-end responsibility for platform quality, reliability, and operational excellence.

• Demonstrates initiative in driving innovation, including PoC development and internal hackathon participation.

📌 Site reliability Engineer Devops (Hyderabad)
🏢 Arminus
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer devops (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer devops (hyderabad) / hyderabad