06 Aug
|
AltiusHub
|
Hyderabad
06 Aug
AltiusHub
Hyderabad
Role Overview
We are looking for a seasoned DevOps SRE Specialist who brings deep, hands-on expertise across both cloud infrastructure engineering and site reliability. This is a hybrid role that requires someone equally comfortable designing Kubernetes-based platforms on AWS and Azure, as they are building observability pipelines and responding to production incidents.
You will own the full lifecycle of platform infrastructure — from cluster provisioning and auto-scaling to distributed tracing, on-call incident management, and continuous delivery pipelines. You will work closely with product engineering teams to raise the reliability and performance bar across all services.
What You Will Do
Platform & Infrastructure
- Design, deploy, and maintain Kubernetes clusters on both AWS EKS and Azure AKS across multi-region production environments.
- Manage ingress and traffic routing using Traefik Gateway and Nginx Ingress Controllers, including TLS termination and routing rules.
- Implement and maintain Karpenter for intelligent, cost-aware node auto-provisioning on EKS.
- Configure KEDA (Kubernetes Event-Driven Autoscaler) for workload-driven horizontal scaling based on queue depth, custom metrics, and event sources.
- Manage TLS certificate lifecycles across all services using Cert Manager with automated renewal.
- Own Helm chart development, versioning, and dependency management for all platform and application deployments.
Networking & Security
- Design and manage VPC/VNet architectures including NAT Gateways, public/private subnets, route tables, and security groups.
- Configure and maintain Azure Bastion and Hub-Spoke network topologies for secure administrative access.
- Manage secrets, credentials,
and certificates using Azure Key Vault and AWS Secrets Manager.
- Administer and secure SFTP server infrastructure for external data exchange partners.
Storage & Messaging
- Manage cloud-native storage solutions including AWS S3 and Azure Blob Storage — lifecycle policies, access controls, versioning, and cross-region replication.
- Administer AWS RDS (PostgreSQL) clusters — parameter tuning, replication, backup policies, and failover configuration.
- Operate and monitor messaging infrastructure on AWS SQS and Azure Service Bus, including DLQ management and consumer scaling.
CI/CD & Automation
- Build and maintain GitHub Actions pipelines for build, test, security scan, and multi-workplace deployment workflows.
- Configure and manage GitHub Actions Runner Controller (ARC) for self-hosted, scalable runner fleets on Kubernetes.
- Write automation tooling and internal platform utilities in Python and/or Go.
Observability & Reliability (SRE)
- Build and maintain the full observability stack: metrics with Prometheus and PromQL, logs with Loki, traces with Tempo, and continuous profiling with Pyroscope.
- Instrument services and infrastructure using OpenTelemetry — SDKs, collectors, and exporters.
- Design and maintain Grafana dashboards, alerting rules, and SLO/SLI tracking across all production systems.
- Manage on-call rotations and incident response workflows using Grafana Cloud IRM (Incident Response & Management).
- Drive post-incident reviews, blameless retrospectives, and reliability improvements from production learnings.
- Define and enforce error budgets, SLOs, and reliability targets in collaboration with product and engineering teams.
📌 Senior Devops Engineer (Hyderabad)
🏢 AltiusHub
📍 Hyderabad