31 Jul
|
Guidewire Software
|
Bengaluru
31 Jul
Guidewire Software
Bengaluru
- We are seeking a Site Reliability Engineer III who is eager to contribute to the transformation of the insurance industry with our leading cloud platform. As a member of the SRE-Application team, youll play a critical role in ensuring the reliability, performance, and scalability of applications running on our Guidewire Cloud Platform. This position offers a unique opportunity to apply your skills in automation, software engineering, and operational discipline to support our cloud-based solutions.
What Youll Do
- Work with development teams to troubleshoot and resolve issues, minimizing customer impact.
- Develop and maintain automated runbooks to manage issues proactively.
- Apply engineering principles and automation to enhance our operating environments.
- Monitor and improve the reliability and performance of applications on the Guidewire Cloud Platform.
- Use your software engineering expertise to optimize systems and reduce manual toil.
- Document incidents and develop processes to prevent future occurrences.
- Stay current with industry trends, tools, and best practices in site reliability engineering.
- Foster a culture of innovation, learning, and continuous improvement.
- Participate in on-call rotations to ensure the availability and reliability of our services.
What Youll Bring
- Experience as an SRE or similar role, with a focus on improving system reliability.
- Strong problem-solving skills and the ability to analyze complex systems and devise effective solutions.
- Effective collaboration and communication skills to work cross-functionally and document processes clearly.
- Experience with automation, monitoring, and performance optimization tools and techniques.
- Commitment to maximizing uptime, scalability, and delivering an exceptional end-user experience.
- Passion for technology and a desire to continuously learn and grow your skills.
- Alignment with Guidewires mission to leverage technology to help protect and support others.
Required Skills:
- 5 years of relevant work experience
- Software engineering background with experience in Python, Go, or Java, following best practices (SOLID, DRY, KISS) and writing clean, testable code
- Experience with designing and implementing SLIs, SLOs, and Error Budgets
- Familiarity with application performance monitoring (APM) and telemetry tools to maintain expected service levels for applications
- Experience troubleshooting and debugging distributed systems on cloud infrastructure
- Experience with CICD pipelines within K8S and legacy ecosystems
- Experience creating monitors, dashboards, and synthetic transactions in monitoring tools like Datadog
- Experience deploying and managing scalable infrastructure within AWS and Kubernetes ecosystems using Terraform and other cloud-native approaches
- Experience with infrastructure configuration management using tools such as GitOps, Puppet, or Ansible
- Valuable understanding of cloud networking, security, and vulnerability management, with the ability to programmatically remediate infrastructure issues
Preferred Skills
- SRE Certification in one or more categories
- AWS Certification in one or more categories
- Experience with SQL, database administration, data pipelines, performance tuning, and schema design
- Familiarity with pipelining tools such as Team City, Bitbucket Pipelines, Jenkins, or GitHub Actions
- Exposure to open-source distributed data processing frameworks such as Hadoop, Apache Spark, AWS RedShift, etc.
- Experience with distributed systems, including microservices and event-driven architectures
📌 Site Reliability Engineer III (SRE) - Guidewire Cloud Platform (Bengaluru)
🏢 Guidewire Software
📍 Bengaluru