06 Sep
|
Arminus
|
Hyderabad
Observability & Monitoring • Architect and operate enterprise-scale observability platforms on AWS/Azure, covering microservices, Lambda/serverless functions, and Kubernetes-based workloads.
• Build full-stack observability solutions using Prometheus, Grafana, CloudWatch, Datadog, Splunk, and Fluentd — covering metrics, logs, and distributed traces.
• Define and implement SLIs, SLOs, and error budgets; build actionable dashboards and alerting policies aligned to business and technical objectives.
• Standardize application logging across engineering codebases by designing and championing reusable Python logging modules and log aggregation pipelines.
• Instrument applications and services with distributed tracing to identify performance bottlenecks and reduce MTTR. Automation & SRE • Design and build reusable automation frameworks standardized as team-wide baselines for all operational automation development.
• Automate high-volume operational workflows (ServiceNow tickets, VDI lifecycle, user provisioning) using Python, PowerShell, and REST APIs.
• Implement zero-touch automation with approval and rollback mechanisms, Managed Identity security controls, and audit logging.
• Build and own CI/CD pipelines using Jenkins, AWS CodePipeline, and Azure DevOps for rapid, safe software delivery.
• Provision and manage infrastructure using Terraform and CloudFormation, enforcing IaC standards across the organization.
Cloud Platform
Engineering • Design and operate Kubernetes-based workloads on AWS EKS and Azure AKS, supporting high-transaction production environments.
• Optimize cloud costs and platform performance — including PostgreSQL query tuning, Redis-based retry workflows, and Lambda function efficiency.
• Integrate automation services securely using Azure API Management (APIM) and AWS API Gateway as governed entry points.
• Ensure enterprise security compliance using IAM, RBAC, Azure Managed Identity, and secure API design patterns reviewed by Security and Architecture teams.
Incident
Management & Reliability • Lead root-cause analysis (RCA) for critical production incidents, ensuring 100% uptime SLAs for mission-critical workloads.
• Develop and maintain incident response runbooks, post-mortem processes,
and reliability improvement roadmaps.
• Proactively identify and resolve performance degradation, capacity risks, and reliability gaps before they impact production. GenAI & Innovation • Evaluate and integrate GenAI tooling to enhance operational efficiency, including LLM-based automation using AWS Bedrock for ticket classification and event triage.
• Drive innovation initiatives through PoC development and internal engineering engagement programmes. Leadership & Collaboration • Mentor junior and mid-level engineers; establish team standards, best practices, and engineering documentation.
• Collaborate with architects, managers, directors, and security teams on platform and observability architecture decisions.
• Champion a culture of reliability, automation-first thinking, and continuous improvement across the engineering organization.
Required Skills &
Experience Core Requirements • Minimum 7+ years of hands-on experience in SRE, DevOps, or Platform Engineering roles.
• Cloud Platforms (Advanced): AWS (EKS, Lambda, DynamoDB, RDS, EC2, SQS, CloudWatch, Bedrock) and/or Azure (Azure Functions, Azure AD, Managed Identity, Azure Monitor, APIM).
• Python (Advanced): production-grade scripting for automation, tooling, backend services, and logging frameworks.
• Observability Tooling (Advanced): Prometheus, Grafana, Datadog, Splunk, CloudWatch, Fluentd, or equivalent.
• Kubernetes (Advanced): EKS or AKS in production; Helm, container lifecycle, and workload troubleshooting.
• Infrastructure as Code (Advanced): Terraform and/or CloudFormation for consistent, auditable infrastructure.
• CI/CD (Advanced): Jenkins, Azure DevOps, AWS CodePipeline — building and maintaining production pipelines.
• Security & Governance (Advanced): IAM/RBAC, Managed Identity, secure API design, and audit logging in enterprise environments.
• Proven track record of leading RCA processes and delivering reliability improvements in high-transaction-volume environments.
• Excellent communication skills with the ability to collaborate across engineering, architecture, and business stakeholders. Automation & Scripting • Proficiency in Bash and PowerShell for operational automation.
• Experience with REST API integration and enterprise workflow automation (e.g., ServiceNow, ticketing platforms).
• Familiarity with database performance tuning, particularly PostgreSQL long-running query optimization. Nice to Have • Hands-on experience with GenAI / LLM-based automation (AWS Bedrock, Azure OpenAI, or similar).
• Experience with ServiceNow automation workflows using JavaScript and REST APIs.
• Familiarity with AIOps or ML-based anomaly detection and monitoring solutions.
• Exposure to chaos engineering practices (Chaos Monkey, Gremlin, LitmusChaos).
• Prior experience in a consulting or client-facing environment (Big 4 / System Integrator a plus).
• Knowledge of FinOps principles and cloud cost optimization strategies.
• Experience in multi-cloud or large-scale enterprise environments spanning Fortune 500 organizations.
Preferred
Certifications • AWS Certified Solutions Architect – Associate or Professional • Microsoft Azure Administrator / Azure Developer / Azure AI Engineer • Certified Kubernetes Administrator (CKA) or CKAD • HashiCorp Certified: Terraform Associate • Datadog Fundamentals / Splunk Core Certified User Behavioral & Soft Skills • Strong problem-solving mindset with the ability to diagnose complex distributed systems issues under pressure • Excellent written and verbal communication — able to translate technical findings to non-technical stakeholders, including directors and architects.
• Proactive self-starter with a bias for automation and eliminating toil at scale.
• Collaborative team player who thrives in fast-paced, cross-functional enterprise environments.
• Ownership mentality — takes end-to-end responsibility for platform quality, reliability, and operational excellence.
• Demonstrates initiative in driving innovation, including PoC development and internal hackathon participation.
📌 Site reliability Engineer Devops (Hyderabad)
🏢 Arminus
📍 Hyderabad