We are looking for Site Reliability Engineer (SRE) with project opportunities.
Job Role: Site Reliability Engineer (SRE)
Location: Bangalore
Interview Mode:Face to Face interview
Experience: 5–10 Years
Relevant Experience: Minimum 5 Years in Site Reliability Engineering, Production Support, Operations, DevOps, or Software Development
Mandatory: SRE + DevOps + Production Support + Microservices + Cloud (Azure/GCP) + Kubernetes (AKS/GKE) + Terraform/Ansible + CI/CD + Python/Java/C#/Go/Ruby + GitHub Actions + Monitoring/Observability + Splunk/Grafana + ITIL/ServiceNow + SLI/SLO/Error Budgets
REQUIREMENT FOR SITE RELIABILITY ENGINEER (SRE):
Key Responsibilities:
Work within cross-functional product teams as the reliability expert for assigned products or product areas.
Apply Site Reliability Engineering practices and standards in collaboration with SRE governance teams.
Ensure high-quality service delivery and provide operational KPI reporting.
Collaborate closely with product teams to maintain predictable operations and minimize production disruptions.
Drive continuous improvement initiatives by sharing best practices and enhancing operational processes.
Monitor, manage, troubleshoot, and resolve application and infrastructure issues across production environments.
Perform technical analysis and Root Cause Analysis (RCA) for complex production incidents.
Improve system reliability through proactive monitoring, alerting, and preventive measures.
Analyze application code and logs to identify opportunities for product and operational improvements.
Develop automation solutions for monitoring, housekeeping activities, and incident prevention.
Ensure application and workplace stability, availability, scalability, and performance.
Automate development and operational processes using scripting and infrastructure automation tools.
Participate in on-call support rotations and resolve business-critical incidents within SLA targets.