Bangalore, Karnataka
Job Summary
Looking for a highly skilled and motivated Site Reliability Engineer to join Verizon Project. As a Site Reliability Engineer, you will play a crucial role in ensuring the reliability, scalability, and performance of our systems. You will be responsible for incident management, release management, automation, infrastructure monitoring, and collaborating with cross-functional teams.
Key Responsibilities
- Incident Management: Act as a key resource in incident management, responding promptly and effectively to incidents to minimize impact. Lead incident resolution efforts, working closely with stakeholders and subject matter experts.
- Release Management: Manage the planning, coordination, and execution of releases across multiple environments. Ensure smooth release processes, including risk assessment, communication, and rollback strategies.
- Automation: Identify opportunities for automation and drive the development of tools and frameworks to improve system resiliency, efficiency, and performance. Collaborate with the development and operations teams to implement automation solutions.
- Infrastructure Monitoring: Establish and maintain comprehensive monitoring systems to ensure high availability and performance of applications and services. Proactively identify potential issues and bottlenecks, and work towards their resolution.
- Collaboration: Work closely with cross-functional teams, including development, operations, and support, to understand requirements, address issues, and drive continuous improvement. Foster a collaborative and proactive culture within the organization.
- Incident Post-Mortems: Conduct post-incident analysis and root cause investigations. Identify opportunities for process improvements and work with stakeholders to implement preventive measures.
- Documentation: Maintain accurate documentation of system configurations, processes, and procedures. Contribute to the knowledge base and provide training and support to team members.
Skill Requirements
- Proven experience as a Site Reliability Engineer or in a similar role, with a focus on high-availability production environments.
- Strong understanding of cloud computing platforms, such as Amazon Web Services (AWS) or Microsoft Azure.
- Solid understanding of Linux/Unix systems administration and troubleshooting.
- Familiarity with monitoring and observability tools like Prometheus, Grafana, Elasticsearch, or Splunk.
- Splunk: Expert-level — dashboards, complex queries, production log analysis under pressure
- Troubleshooting under pressure: Non-negotiable core competency — diagnosing multi-service production failures in real time
- Strong analytical and problem-solving skills, with the ability to diagnose and resolve complex technical issues.
- Excellent communication and collaboration skills,
with the ability to work effectively in cross-functional teams.
- Knowledge of DevOps principles and practices, including CI/CD pipelines and version control systems (e.g., Git).
- Non-negotiable core competency — diagnosing multi-service production failures in real time
Other Requirements
Soft Skills — Equally Important
- Leadership: Confidently leads triage calls with engineers, DBAs, and senior stakeholders — not a passive participant
- Communication: Transparent, concise verbal and written communication; comfortable on recorded bridge calls
- Ownership & Accountability: Follows through to completion — open items do not sit unattended, incidents do not close until truly resolved
- Analytical Thinking: Strong problem-solving skills; drives root cause analysis, not surface-level fixes
- Composure: Calm, focused, and decisive during high-pressure P1 incidents
Preferred Skills:
- Experience with Openshift in addition to Kubernetes
- Familiarity with configuration management tools like Stash or GitHub.
- Familiarity with networking concepts, REST APIs and security best practices.
- Prior experience in IoT, telecom, or large-scale connected device platforms
- Confluence and JIRA proficiency for documentation and ticket management
- Certification in relevant technologies is a plus.
#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-
📌 Project Lead (India)
🏢 HCLTech
📍 India