06 Aug
|
Vanguard Technologies
|
Hyderabad
06 Aug
Vanguard Technologies
Hyderabad
Role Summary
This role provides subject matter expertise and coordination to site reliability efforts across the subdivision. This role also includes ensuring system reliability by meeting service-level objectives (SLOs), driving automation of operational tasks, defining and tracking key performance indicators (KPIs), designing scalable systems, managing incident responses, and collaborating with development teams to ensure software reliability and scalability.
Responsibilites
- Coordinates cross-product chaos experimentation.
- Maintains the centralized incident response playbook for the subdivision to document standards for managing communication and escalation during an incident. Aggregates quantifiable data about availability to report back to senior leadership. Makes contributions to centrally managed (IT-wide) inner source libraries for reliability.
- Facilitates blameless post-incident reviews for high severity incidents or incidents involving more than one product family.
- Regularly attends Reliability Engineering and Resilience communities of practice. Remains informed about site reliability engineering activities happening within the subdivision.
- Communicates new standards and newly available tools and frameworks across subdivisions. Enforces reliability standards.
- Participates in special projects and performs other duties as assigned.
- Drives automation of routine operational tasks to improve system efficiency, reduce manual intervention, and enhance deployment and monitoring workflows.
- Leads incident response efforts by diagnosing root causes rapidly, applying timely fixes,
and establishing preventive measures to avoid recurrence.
- Defines and tracks key system performance metrics such as availability, latency, and error rates to evaluate and optimize system health and reliability.
- Collaborates with development teams to align software architecture with reliability and scalability goals, ensuring seamless operations across the deployment lifecycle.
Qualifications and Skills
- Minimum 8 years of experience in software engineering or site reliability roles, with at least 2 years in development or operational support functions
- Bachelor s degree (B.E./B.Tech) in Computer Science, IT, Software Engineering, or a related field; Master s degree preferred
- Cloud Platforms: AWS (EC2, ECS, Lambda, S3, CloudWatch)
- Programming Languages: Java, Node.js
- Front-End Frameworks: Angular
- Containerization: Docker (ECS-based container patterns)
- CI/CD DevOps: Git, GitHub, JIRA, Confluence
- Observability Monitoring: Splunk, CloudWatch, Honeycomb
- Automation Testing: Cucumber, JUnit, Playwright, Puppeteer
- Robust experience in incident management, system performance optimization, automation of operational tasks, reliability engineering principles, and cross-functional collaboration
Location
This role is based in Hyderabad, Telangana at Vanguard s India office. Only qualified external applicants will be considered.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Site Reliability Champion, Specialist (Hyderabad)
🏢 Vanguard Technologies
📍 Hyderabad