Site Reliability Engineer IV (Hyderabad)

Site Reliability Engineer IV (Hyderabad)

25 Aug
|
Candescent (Digital First Holdings
|
Hyderabad

25 Aug

Candescent (Digital First Holdings

Hyderabad

Candescent is a forward-thinking technology company transforming how financial institutions deliver Intelligent Banking experiences. We unite digital banking, account opening, and branch solutions that power and connect digital banking, account opening, and branch solutions—creating seamless engagement across digital, remote, and in-person channels.

Our Experience-Led, Intelligence-Driven approach combines human-centered design with data, automation, and cloud-based innovation. Built on an API-first architecture, our extensible ecosystem enables institutions to adapt quickly, integrate easily, and unlock new opportunities for growth—turning every customer interaction into a moment of clarity, confidence, and connection.

Position: Site Reliability Engineer IV

Experience: 9-12 Years

Location: HYDERABD

Role Overview

We are looking for a strong Application Site Reliability Engineer (SRE) to support and improve the reliability of Java-based production systems running on Kubernetes in cloud environments.

This role focuses on application-level reliability, JVM deep troubleshooting, production incident management, and close collaboration with development teams — not infrastructure provisioning or CloudOps.

The ideal candidate understands how Java applications behave in production and can proactively improve performance, scalability, and operational maturity.

As a senior member of the SRE organization, the candidate will also provide technical leadership, mentorship, and people management support to the SRE team, helping develop engineers, improve team effectiveness, and strengthen the overall reliability culture.

Key Responsibilities

- Support and operate production Java applications running on Kubernetes (GKE).
- Troubleshoot complex application issues using logs, metrics, traces, heap dumps, and thread dumps.
- Participate in incident response, root cause analysis, and blameless postmortems.
- Collaborate closely with development teams to understand application architecture, dependencies, and failure patterns.
- Analyse JVM behavior (heap, GC, memory leaks, OOM, thread contention) and recommend performance improvements.




- Define and improve SLIs, SLOs, alerts, and dashboards.
- Support application deployments, rollbacks, and runtime configuration changes.
- Identify reliability, performance, and scalability gaps in application behaviour.
- Automate repetitive operational tasks to reduce toil.
- Drive improvements in runbooks, operational readiness, and on-call effectiveness.
- Advocate and influence adoption of shift-left reliability practices.
- Provide technical guidance, mentorship, and coaching to SRE team members.
- Support team development through regular feedback, career development, and performance discussions.
- Help establish team goals, priorities, and development plans in alignment with organizational objectives.
- Promote a culture of ownership, collaboration, continuous improvement, and reliability within the SRE team.

Must-Have Skills & Experience

- Strong hands-on experience supporting Java applications in production.
- Deep understanding of JVM internals:
- Heap & memory management
- Garbage collection
- Tuning and OOM analysis
- Thread dump and performance analysis

- Proven experience in incident response and production troubleshooting.
- Experience operating applications on Kubernetes from an application/runtime perspective.
- Solid experience with application observability:

- Logs
- Metrics
- Monitoring tools
- Distributed tracing

- Solid understanding of SLIs, SLOs, and reliability-driven operations.
- Experience with deployment strategies (rolling, blue/green, canary).
- Ability to write scripts/automation (Python, Shell, or similar) to reduce operational toil.
- Strong understanding of application architecture and service dependencies (databases, messaging systems, external APIs).




- Ability to analyze and troubleshoot issues holistically across the entire application stack, rather than focusing on isolated components.
- Strong collaboration and communication skills.
- Demonstrates accountability and sound judgment during high-pressure production incidents.
- Experience mentoring, coaching, and developing engineers.
- Experience with people management responsibilities, including setting expectations, providing feedback, supporting career development, and managing performance.
- Ability to lead and support a high-performing SRE team while balancing operational responsibilities and engineering priorities.

Cloud & Platform Exposure

- Experience working with applications deployed on Kubernetes in GCP.
- Familiarity with GKE environments from an application operations perspective (not CloudOps or infrastructure engineering).
- Understanding of cloud constructs relevant to application behavior (networking, IAM, storage, compute).

Good-to-Have Skills

- CI/CD pipeline exposure (GitHub Actions, Jenkins).
- Familiarity with GitOps practices.
- Experience supporting cloud migrations or modernization initiatives.
- Exposure to platform or infrastructure concepts supporting application workloads.
- Experience managing or leading an SRE, DevOps, or application operations team.

What We Value

- Ownership mindset and reliability-first thinking.
- Curiosity to investigate deep production issues.
- Strong collaboration with development teams.
- Bias toward automation and continuous improvement.
- Clear communication during incidents and stakeholder updates.
- Strong people leadership, mentoring, and team development skills.
- Ability to build a culture of accountability, ownership, and continuous improvement.

Statement to Third Party Agencies
To ALL recruitment agencies: Candescent only accepts resumes from agencies on the preferred supplier list. Please do not forward resumes to our applicant tracking system, Candescent employees, or any Candescent facility. Candescent is not responsible for any fees or charges associated with unsolicited resumes.

📌 Site Reliability Engineer IV (Hyderabad)
🏢 Candescent (Digital First Holdings
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer iv (hyderabad) / hyderabad