Hello,
Greetings from ZettaMine Labs Pvt Ltd!! We are looking for Site Reliability Engineer (SRE) with project opportunities.
Job Role: Site Reliability Engineer (SRE)
Location: Bangalore
Interview Mode:Face to Face interview
Experience: 5–10 Years
Relevant Experience: Minimum 5 Years in Site Reliability Engineering, Production Support, Operations, DevOps, or Software Development
Mandatory: SRE + DevOps + Production Support + Microservices + Cloud (Azure/GCP) + Kubernetes (AKS/GKE) + Terraform/Ansible + CI/CD + Python/Java/C#/Go/Ruby + GitHub Actions + Monitoring/Observability + Splunk/Grafana + ITIL/ServiceNow + SLI/SLO/Error Budgets
REQUIREMENT FOR SITE RELIABILITY ENGINEER (SRE):
Key Responsibilities
Work within cross-functional product teams as the reliability expert for assigned products or product areas.
Apply Site Reliability Engineering practices and standards in collaboration with SRE governance teams.
Ensure high-quality service delivery and provide operational KPI reporting.
Collaborate closely with product teams to maintain predictable operations and minimize production disruptions.
Drive continuous improvement initiatives by sharing best practices and enhancing operational processes.
Monitor, manage, troubleshoot, and resolve application and infrastructure issues across production environments.
Perform technical analysis and Root Cause Analysis (RCA) for complex production incidents.
Improve system reliability through proactive monitoring, alerting, and preventive measures.
Analyze application code and logs to identify opportunities for product and operational improvements.
Develop automation solutions for monitoring, housekeeping activities, and incident prevention.
Ensure application and environment stability, availability, scalability, and performance.
Automate development and operational processes using scripting and infrastructure automation tools.
Participate in on-call support rotations and resolve business-critical incidents within SLA targets.
Define and track reliability metrics including SLIs, SLOs, and Error Budgets.
Contribute to performance engineering and application reliability initiatives.
Who Can Apply?
✔️ 5–10 years of experience in SRE, Production Support, Operations, DevOps, or Software Development.
✔️ Strong experience supporting and operating eCommerce platforms.
✔️ Hands-on experience with DevOps practices, CI/CD, automated testing, and release automation.
✔️ Experience troubleshooting complex distributed systems and microservices-based architectures.
✔️ Strong understanding of solution architecture and Root Cause Analysis techniques.
✔️ Experience with API-driven frameworks such as CommerceTools, Fabric, or similar platforms.
✔️ Experience with ITIL processes and ITSM tools such as ServiceNow.
✔️ Knowledge of application reliability and performance engineering principles.
✔️ Experience supporting web, desktop, and mobile applications.
Cloud & Infrastructure:
✔️ Hands-on experience with Microsoft Azure and/or Google Cloud Platform (GCP).
✔️ Experience with managed Kubernetes services such as AKS and/or GKE.
✔️ Experience provisioning and managing infrastructure using Terraform and/or Ansible.
✔️ Knowledge of cloud-native architecture, scalability, and reliability best practices.
Development & Automation:
✔️ Proficiency in at least one programming language: Python, Java, C#, Go, or Ruby.
✔️ Experience with GitHub Actions for CI/CD workflow development.
✔️ Familiarity with Azure DevOps and other deployment automation platforms.
✔️ Understanding of ReactJS, React Native, and Node.js is an advantage.
Monitoring & Reliability:
✔️ Hands-on experience with observability and monitoring tools such as Splunk, Grafana, or similar monitoring platforms.
✔️ Solid understanding of Service Level Indicators (SLIs), Service Level Objectives (SLOs), Error Budgets, Incident Management, and Reliability Engineering practices.
Please share the following details along with your updated resume and reach out to
[email protected] Full Name:
Contact Number
Current Location
Total Experience
Relevant Experience in SRE/DevOps
eCommerce Platform Experience:
Cloud Experience (Azure/GCP)
Kubernetes Experience (AKS/GKE)
Terraform/Ansible Experience
Programming Language
CI/CD & GitHub Actions Experience:
Monitoring Tools (Splunk/Grafana)
ServiceNow/ITIL Experience
Current Company
Notice Period
Current CTC
Expected CTC
For any further clarification, feel free to reach out at
[email protected] ,(phone hidden) .
We look forward to connecting with you! Thanks & Regards
TAG-Team
📌 Site Reliability Engineer-F Interview Bangalore (Bengaluru)
🏢 ZettaMine Labs
📍 Bengaluru