02 Sep
|
Vagaro
|
Ahmedabad
About the Company :
Vagaro is a global leader in cloud-based business management software, serving the salon, spa, beauty, fitness, and wellness industries. Our platform enables businesses to manage scheduling, payments, marketing, and operations from one centralized system.
We are a rapidly growing, product-led company transforming how wellness businesses operate and connect with their customers.
Why Join Us :
- 5-day work week & flexible schedule
- Fun Fridays, Tech sessions & cultural events
- Learning resources & gaming zones
- Annual bonus & recognition programs
- Leave encashment & maternity benefits
- 15 Paid Leaves + 11 Holidays
- Comprehensive family medical insurance
Role: Director of Devops Location: Ahmedabad (Work from Office)
Position Summary:
We are looking for an experienced Director of DevOps, SRE & Infrastructure to lead the strategy, architecture, and operations of our Microsoft Azure Cloud Platform and engineering infrastructure.
This role will own DevOps, SRE, Infrastructure, Platform Engineering, Observability, Security, Disaster Recovery, Scalability, and Cloud Cost Optimization. The ideal candidate combines strong technical expertise with proven leadership experience in designing and operating highly available, secure, scalable, and mission-critical SaaS platforms on Microsoft Azure.
Strong hands-on and leadership experience with Microsoft Azure Cloud Platform is mandatory.
Key Responsibilities:
DevOps & Platform Engineering:
- Define and execute the DevOps and Platform Engineering strategy.
- Establish standardized CI/CD, release, deployment, and environment management practices.
- Drive automation, self-service infrastructure, and developer productivity.
- Establish engineering standards for Infrastructure as Code and configuration management.
- Promote safe deployment practices including Blue/Green, Canary, and Rolling deployments.
Site Reliability Engineering:
- Build and mature the SRE organization and culture.
- Define and manage SLIs, SLOs, SLAs, and Error Budgets .
- Improve system availability, reliability, performance, and scalability.
- Establish production readiness and service ownership standards.
- Reduce operational toil and recurring production issues.
Infrastructure & Architecture:
- Own the architecture and operation of Azure infrastructure .
- Design highly available, scalable, fault-tolerant, and multi-region systems.
- Establish standards for compute, networking, storage, databases, caching, messaging, and traffic management.
- Drive infrastructure modernization and cloud migration initiatives.
Observability & Operations:
- Establish enterprise observability across metrics, logs, traces, APM, and infrastructure monitoring .
- Define monitoring, alerting, and operational dashboards.
- Establish effective incident management, escalation, and post-incident review processes.
- Drive reduction of MTTD, MTTR, and recurring incidents.
Security & Compliance:
- Partner with Security and Engineering to establish secure infrastructure.
- Implement standards for IAM, secrets, certificates, encryption,
network security, vulnerability management, and supply-chain security.
- Support SOC 2, ISO 27001, PCI DSS , and other applicable compliance requirements.
Disaster Recovery & Business Continuity:
- Own infrastructure-level DR and Business Continuity strategy.
- Define and maintain RTO/RPO objectives.
- Establish backup, recovery, regional failover, and disaster recovery standards.
- Conduct regular DR and failover testing.
Performance, Scalability & Cost:
- Establish infrastructure capacity and performance planning.
- Identify and eliminate infrastructure bottlenecks.
- Ensure platforms can scale with business growth.
- Own Azure cost optimization and FinOps initiatives.
- Balance reliability, performance, scalability, and infrastructure cost.
Leadership:
- Lead DevOps, SRE, Infrastructure, and Platform Engineering teams.
- Recruit, mentor, and develop high-performing technical teams.
- Establish engineering objectives, standards, and KPIs.
- Partner with CTO/CPTO, Engineering, Architecture, Security, Product, QA, and Finance.
- Manage strategic technology vendors and infrastructure budgets.
Core Technical Skill Set: Azure — Required:
Solid production experience with
- Azure App Services
- Azure Kubernetes Service (AKS)
- Azure Functions
- Azure Virtual Machines
- Azure Storage
- Azure SQL
- Cosmos DB
- Azure Cache for Redis
- Azure Service Bus
- Azure Event Hubs
- Azure Front Door
- Azure Application Gateway
- Azure Load Balancer
- Azure CDN
- Azure DNS
- Azure API Management
- Azure Virtual Network
- Azure Monitor
- Application Insights
- Log Analytics
- Azure Key Vault
- Microsoft Entra ID
- Azure Firewall
- Azure WAF
- Azure Backup / Site Recovery
DevOps & CI/CD:
- GitHub / GitHub Enterprise
- Azure DevOps
- GitHub Actions
- CI/CD architecture
- Release automation
- Deployment automation
- GitOps
- Feature flags
- Blue/Green deployment
- Canary deployment
- Rolling deployment
Infrastructure as Code:
- Terraform
- Bicep / ARM
- Ansible
- Helm
- Kubernetes manifests
- Policy as Code
Containers & Kubernetes:
- Docker
- Kubernetes
- AKS
- Helm
- Kubernetes networking
- Autoscaling
- Container security
- Cluster management
- Kubernetes observability
SRE:
- SLI / SLO / SLA
- Error Budgets
- Incident Management
- Production Readiness
- Reliability Engineering
- Capacity Planning
- Toil Reduction
- Chaos Engineering
- Fault Tolerance
- Disaster Recovery
Observability:
- Azure Monitor
- Application Insights
- Log Analytics
- Prometheus
- Grafana
- Elastic
- OpenTelemetry
- Distributed Tracing
- APM
- Centralized Logging
- Alerting
Networking :
- TCP/IP
- HTTP/HTTPS
- DNS
- TLS/SSL
- CDN
- WAF
- DDoS Protection
- Load Balancing
- API Gateway
- Reverse Proxy
- VNet
- Firewall
- VPN
- Routing
- Private Networking
Security:
- Microsoft Entra ID
- IAM / RBAC
- OAuth 2.0
- OpenID Connect
- Secrets Management
- Certificate Management
- Encryption / KMS
- Zero Trust
- Vulnerability Management
- Container Security
- SAST / DAST
- Supply-Chain Security
Data & Messaging:
Working knowledge of
- SQL Server
- PostgreSQL
- MongoDB
- Redis
- Elasticsearch
- Azure Service Bus
- Azure Event Hubs
- Kafka
- RabbitMQ
- Distributed systems
- Replication
- Partitioning
- Sharding
- Backup & Recovery
Leadership & Management Skills:
- Strategic planning
- Technical vision
- Engineering leadership
- Team building
- Hiring and talent development
- Executive communication
- Incident leadership
- Architecture decision-making
- Risk management
- Vendor management
- Budget management
- Cross-functional collaboration
- Stakeholder management
- Performance management
- Change management
Experience & Qualifications:
Required:
- 12+ years of experience in DevOps, SRE, Infrastructure, Cloud, Platform Engineering, or related fields.
- 5+ years of engineering leadership/management experience .
- Proven experience managing DevOps/SRE/Infrastructure teams.
- Strong hands-on experience with Microsoft Azure .
- Strong experience with Kubernetes/AKS and modern CI/CD.
- Experience operating large-scale production SaaS systems.
- Experience designing highly available and distributed systems.
- Experience with production incident management and disaster recovery.
- Strong understanding of security, observability, scalability, and cloud cost management.
Preferred:
- Experience operating 24×7 mission-critical SaaS platforms .
- Experience with multi-region / active-active architectures .
- Experience with SOC 2, ISO 27001, PCI DSS, or similar compliance environments.
- Experience building or transforming DevOps/SRE organizations.
- Experience managing significant infrastructure/cloud budgets.
Key Success Metrics: Reliability: Availability, SLO Achievement
Incidents: MTTD, MTTR, Incident Frequency
Deployment: Deployment Frequency, Lead Time
Quality: Change Failure Rate
Scalability: Capacity & Performance
Recovery: RTO / RPO
Automation: Reduction in Operational Toil
Security: Critical Vulnerability Remediation
Cost: Azure Cost Optimization
DR: Successful DR/Failover Tests
Observability: Monitoring & Alert Coverage
Developer Experience: CI/CD Reliability & Deployment Efficiency
Ideal Candidate:
The ideal candidate is not just a DevOps manager . They should be able to operate at three levels:
Strategic: Define the long-term DevOps, SRE, infrastructure, reliability, and platform strategy.
Architectural: Make critical decisions around Azure, Kubernetes, networking, distributed systems, observability, security, scalability, and DR.
Operational: Lead critical production incidents, improve reliability, and build systems that prevent recurring failures.
Lead. Architect. Transform. Your journey as Vagaro’s Director of DevOps starts here—apply now!
📌 Director of DevOps | 12+ Years | Ahmedabad Location
🏢 Vagaro
📍 Ahmedabad