SRE Engineer
Position: SRE Engineer
Location: Mumbai – BKC
Budget: Up to 10 LPA
Primary Skill: Site Reliability Engineering (SRE)
Role Overview
We are looking for an SRE Engineer responsible for improving the reliability, scalability, observability, performance, and availability of critical applications and business services.
The candidate will work closely with Application Development, Infrastructure, Architecture, and Product teams to identify service-impacting issues, improve monitoring and alerting, automate operational processes, and ensure highly reliable production environments.
Key Responsibilities
- Define and monitor SLA, SLO, and SLI based on SRE best practices and the Four Golden Signals.
- Develop and maintain monitoring, alerting, and observability solutions.
- Drive proactive reliability practices such as Chaos Engineering, Game Days, and Synthetic Monitoring.
- Work with application and architecture teams to improve system scalability, performance, and reliability.
- Troubleshoot production issues and support incident management and root-cause analysis.
- Improve CI/CD, deployment, environment management, and operational processes.
- Design and implement solutions for high availability, resilience, and self-healing.
- Perform capacity planning and performance optimization while considering cost efficiency.
- Develop and execute performance, scalability, reliability, and automation testing.
- Ensure Dev/Test/Production environments are consistent, reproducible, and highly available.
- Collaborate with engineering teams to simplify real-time troubleshooting and operational response.
- Maintain documentation,
testing schedules, status reports, and reliability metrics.
- Follow Agile/Scrum methodologies and contribute to continuous improvement.
Required Technical Skills
- Strong experience in SRE / DevOps / Production Engineering.
- Strong knowledge of AWS Cloud and infrastructure provisioning using Terraform.
- Hands-on experience with CI/CD, Docker, Kubernetes, and Ansible/Salt.
- Strong programming/scripting knowledge in Python, Java, Golang, or similar languages.
- Experience with microservices, APIs, and event-driven architectures.
- Strong understanding of monitoring and observability.
- Experience with Prometheus, Telegraf, OpenTSDB or similar monitoring platforms.
- Hands-on experience with Grafana and Kibana.
- Knowledge of metrics collection and time-series queries.
- Experience with test automation and performance/load testing.
- Good understanding of networking, storage, security, and infrastructure architecture.
- Understanding of messaging protocols, caching strategies, and software design principles.
- Experience working in Agile/Scrum environments.
Preferred / Good to Have
- Experience in Trading Systems / Capital Markets / Financial Services.
- Experience with private cloud environments.
- Exposure to Chaos Engineering, Synthetic Monitoring, and Game Days.
- Experience handling high-volume, highly available production systems.
- Robust incident management and troubleshooting skills.
If the JD aligns with your roles and responsibility shares us your updated resume to
[email protected] or WhatsApp you resume on this number (phone hidden)
Looking for immediate joiner
📌 SRE - 3i Infotech - BKC (Goregaon)
🏢 3i Infotech
📍 Goregaon