09 Aug
|
Azuga
|
Bangalore Urban
09 Aug
Azuga
Bangalore Urban
Principal Site Reliability Engineer (SRE)
Position Summary
We are seeking an experienced Site Reliability Engineer (SRE) with 12+ years of hands-on experience designing, building, automating, and operating large-scale cloud infrastructure and mission-critical production environments.
The ideal candidate is a highly technical engineer with deep expertise in AWS cloud technologies, Kubernetes, Infrastructure as Code (IaC), automation, observability, and production operations . This individual will drive platform reliability, operational excellence, automation, and continuous improvement while partnering closely with Engineering, Architecture, Security, and Product teams.
As part of this role, the candidate will also provide technical leadership for enterprise database platforms, including AWS RDS (MySQL & PostgreSQL), Amazon Aurora, and MongoDB , ensuring they remain secure, highly available, performant, and operationally resilient.
This is a hands-on technical leadership role for an engineer who enjoys solving complex production challenges, mentoring teams, and building reliable cloud platforms at scale.
Key Responsibilities
Site Reliability Engineering
- Lead the adoption and continuous improvement of Site Reliability Engineering (SRE) practices across the organization.
- Design, build, and operate highly available, scalable, and resilient production platforms.
- Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
- Lead production incident response, Root Cause Analysis (RCA), and post-incident reviews.
- Drive operational excellence through automation, standardization, and engineering best practices.
- Develop operational runbooks, playbooks, and reliability standards.
- Perform capacity planning, scalability assessments, and performance optimization.
- Design and validate High Availability (HA) and Disaster Recovery (DR) strategies.
- Partner with software engineering teams to improve production readiness and platform reliability.
Cloud Platform Engineering
- Design, implement, and manage enterprise AWS cloud infrastructure.
- Lead Kubernetes platform engineering using Amazon EKS.
- Build and maintain Infrastructure as Code (IaC) solutions using Terraform.
- Support cloud modernization and platform engineering initiatives.
- Optimize cloud infrastructure for availability, performance, scalability, security, and cost efficiency.
- Collaborate with development teams to improve deployment strategies and platform reliability.
Automation & DevOps
- Design and implement automation solutions using Python, Shell scripting, Terraform,
and Ansible.
- Build reusable operational tooling and self-service automation capabilities.
- Improve CI/CD pipelines and deployment reliability.
- Eliminate repetitive operational tasks through automation.
- Promote Infrastructure as Code (IaC) and GitOps best practices.
Observability & Monitoring
- Design enterprise observability solutions for infrastructure, applications, and databases.
- Implement monitoring and alerting using Prometheus, Grafana, Alertmanager, Amazon CloudWatch, and related technologies.
- Build meaningful dashboards and operational metrics.
- Improve system visibility and reduce alert fatigue through effective monitoring strategies.
- Drive proactive monitoring to identify and resolve issues before customer impact.
Database Platform Engineering Provide Technical Leadership For Enterprise Database Platforms, Including
- AWS RDS MySQL
- AWS RDS PostgreSQL
- Amazon Aurora (MySQL & PostgreSQL)
- MongoDB (Atlas and Self-Managed)
Responsibilities Include
- Manage database provisioning, upgrades, backups, patching, and lifecycle management.
- Perform database performance tuning, query optimization, indexing, and capacity planning.
- Design and maintain High Availability (HA), backup, recovery, and Disaster Recovery (DR) solutions.
- Troubleshoot complex production database issues and drive long-term improvements.
- Ensure database security through encryption, IAM integration, auditing, and access controls.
Leadership & Collaboration
- Serve as a technical leader within the Site Reliability Engineering organization.
- Mentor SREs, Platform Engineers, and Cloud Engineers.
- Participate in architecture reviews and technical design discussions.
- Collaborate with Engineering, Architecture, Security, Product, and Infrastructure teams.
- Drive engineering standards, operational maturity, and continuous improvement initiatives.
Required Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline.
- 5 years of hands-on experience in Site Reliability Engineering, Platform Engineering, Cloud Engineering, DevOps, or Infrastructure Engineering.
- Strong expertise in AWS cloud services and production operations.
- Hands-on experience with Kubernetes and Amazon EKS.
- Strong Linux administration and troubleshooting skills.
- Experience implementing Infrastructure as Code (IaC) using Terraform.
- Strong automation skills using Python and Shell scripting.
- Experience with GitOps practices and tools such as Argo CD.
- Hands-on experience with CI/CD platforms such as Jenkins or GitHub Actions.
- Experience implementing observability solutions using Prometheus, Grafana, Alertmanager, and Amazon CloudWatch.
- Strong understanding of networking concepts, including:
- VPC
- DNS
- Load Balancers
- SSL/TLS
- IAM
- Cloud security best practices
- Experience leading production incidents, Root Cause Analysis (RCA), and operational improvements.
- Solid experience administering and optimizing:
- AWS RDS MySQL
- AWS RDS PostgreSQL
- Amazon Aurora
- MongoDB
- Excellent analytical, communication, collaboration, and problem-solving skills.
Preferred Qualifications
- AWS Certified Solutions Architect – Professional (or equivalent).
- Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD).
- Experience with Helm and Kubernetes package management.
- Experience with Redis/Valkey, Kafka, Elasticsearch/OpenSearch, or ClickHouse.
- Experience supporting enterprise SaaS or cloud-native production platforms.
- Experience mentoring engineering teams and leading cross-functional technical initiatives.
Technical Skills
Cloud Platforms
- Amazon Web Services (AWS)
- Amazon EKS
- EC2
- VPC
- IAM
- Route 53
- Application Load Balancer (ALB)
- Network Load Balancer (NLB)
- Auto Scaling
- Amazon CloudWatch
Containers & Platform Engineering
- Kubernetes
- Docker
- Helm
- Argo CD
- GitOps
Infrastructure Automation
- Terraform
- Python
- Shell Scripting
- Ansible
Monitoring & Observability
- Prometheus
- Grafana
- Alertmanager
- Loki
- Amazon CloudWatch
Database Technologies
- AWS RDS MySQL
- AWS RDS PostgreSQL
- Amazon Aurora
- MongoDB
DevOps & CI/CD
- Git
- Jenkins
- GitHub Actions
- CI/CD Pipelines
Core Competencies
- Site Reliability Engineering (SRE)
- Cloud Platform Engineering
- AWS Architecture
- Kubernetes Administration
- Infrastructure Automation
- Production Operations
- Observability & Monitoring
- Performance Engineering
- High Availability & Disaster Recovery
- Database Platform Engineering
- Security & Compliance
- Incident & Problem Management
- Technical Leadership
- Architecture & Design
- Continuous Improvement
- Cross-functional Collaboration
📌 Site Reliability Engineer (Bangalore Urban)
🏢 Azuga
📍 Bangalore Urban