20 Sep
|
Vikash Technologies
|
Hyderabad
20 Sep
Vikash Technologies
Hyderabad
Role Overview :As the Director of Site Reliability Engineering, you will spearhead the vision, strategy, and operational excellence of our global IT infrastructure. You will lead high-performing engineering teams to ensure our platforms remain resilient, scalable, and highly available in a fast-paced, cloud-native environment. Working closely with executive leadership, product managers, and cross-functional engineering heads, you will bridge the gap between development and operations to drive architectural improvements. Your leadership will directly influence the reliability of our services, ensuring a seamless experience for our customers while optimizing infrastructure costs and performance to support the company's long-term business objectives. This role is based in Hyderabad and requires a seasoned leader capable of navigating complex technical ecosystems.Key Responsibilities :- Define and execute the long-term SRE strategy to enhance system reliability and performance, ensuring that our infrastructure meets the rigorous demands of our growing user base.- Lead and mentor large, distributed engineering teams, fostering a culture of continuous improvement, automation, and operational excellence across the organization.- Oversee the design and maintenance of robust AWS-based cloud architectures, ensuring that our infrastructure is secure, cost-effective,
and capable of handling massive scale.- Drive the adoption of Kubernetes and container orchestration best practices to streamline deployment cycles and improve the efficiency of our microservices architecture.- Optimize CI/CD pipelines to accelerate release velocity while maintaining strict quality gates, ensuring that product teams can ship features rapidly without compromising system stability.- Manage complex incident response protocols and post-mortem processes, transforming technical failures into actionable insights that prevent future outages and improve system resilience.Required Skillset :- Demonstrated expertise in architecting and managing large-scale, mission-critical IT infrastructure within AWS environments, with a deep understanding of cloud-native design patterns.- Proven ability to lead and scale high-performing engineering organizations, with a focus on talent development, performance management, and building inclusive, high-output teams.- Extensive experience in implementing and managing Kubernetes clusters at scale, including deep knowledge of service meshes, observability, and infrastructure-as-code principles.- Exceptional communication and stakeholder management skills, with the ability to translate complex technical challenges into strategic business outcomes for executive leadership.- Strong background in CI/CD automation and DevOps methodologies, with a track record of reducing manual toil and improving deployment frequency through advanced tooling.- A minimum of 18 - 25 years of experience in infrastructure engineering and reliability roles, ideally within high-growth technology companies or large-scale enterprise environments.- Ability to thrive in a hybrid work setting in Hyderabad, demonstrating flexibility and a proactive approach to leading teams across different time zones and locations. (ref:hirist.tech)
📌 Director - Site Reliability Engineering (Hyderabad)
🏢 Vikash Technologies
📍 Hyderabad