20 Aug
|
VY SYSTEMS PRIVATE
|
Bengaluru
20 Aug
VY SYSTEMS PRIVATE
Bengaluru
We are looking for an experienced SRE / Automation Engineer with strong expertise in SLA, SLO, SLIs, observability, monitoring, automation, and production reliability. The ideal candidate should have hands-on experience designing reliability metrics, automating operational processes, improving system availability, and driving SRE best practices across enterprise environments.
Key Responsibilities
- Define, implement, and maintain SLIs, SLOs, and SLAs for critical applications and services.
- Monitor service reliability, availability, latency, and performance against defined SLOs.
- Develop automation solutions to reduce manual operational effort and eliminate repetitive tasks.
- Build and enhance monitoring, alerting, dashboards, and observability capabilities.
- Analyze reliability trends and proactively identify risks to service availability and performance.
- Participate in incident management, troubleshooting, RCA, and post-incident reviews.
- Implement error-budget practices and use reliability metrics to drive engineering decisions.
- Automate health checks, remediation, reporting, and operational workflows.
- Develop scripts/tools using Python, Shell, or equivalent automation technologies.
- Collaborate with development, infrastructure, DevOps, and operations teams to improve system reliability.
- Create and maintain runbooks, SOPs, and automation documentation.
- Identify opportunities for improving MTTD, MTTR, availability, and operational efficiency.
- Support capacity planning, performance optimization, and reliability improvements.
Must-Have Skills
Site Reliability Engineering (SRE)
SLA / SLO / SLI
SRE Automation
Observability & Monitoring
Incident Management
Root Cause Analysis
Error Budgets
Python / Shell Scripting
Alerting & Dashboards
Production Support
Required Skills & Tools
- 6+ years of experience in SRE, Production Engineering, DevOps, or Reliability Engineering.
- Strong understanding of SLA, SLO, SLI, and error-budget concepts.
- Hands-on experience building automation for production operations.
- Strong scripting/programming skills in Python, Shell, or similar languages.
- Experience with monitoring and observability platforms such as Prometheus, Grafana, Splunk, AppDynamics, Datadog, or equivalent.
- Strong understanding of incident, problem, and change management.
- Experience performing RCA for complex production issues.
- Positive understanding of cloud, infrastructure, applications, and distributed systems.
- Strong analytical and troubleshooting skills.
Preferred Skills
- Experience with Kubernetes / Docker.
- Exposure to CI/CD and Infrastructure as Code.
- Experience with cloud platforms such as AWS, Azure, or GCP.
- Knowledge of OpenTelemetry and modern observability practices.
- Experience implementing automated remediation and self-healing solutions.
- Exposure to ITIL and enterprise service management processes.
Key Competencies
- Strong reliability and automation mindset.
- Excellent problem-solving and analytical skills.
- Ability to work independently and take ownership of production reliability.
- Strong communication and stakeholder management skills.
- Ability to work effectively with cross-functional and global teams.
Qualifications
Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline.
Acknowledge your availability with your updated Resume.
Kindly attach your LWD (Last Working Day) screenshot along with your updated resume, formal photograph, and PAN card copy for profile submission.
Please do share this opportunity with your friends or references who may be looking for a change.
Skills:- SRE, SLA and SLO
📌 SRE SLA SLO (Bengaluru)
🏢 VY SYSTEMS PRIVATE
📍 Bengaluru