Senior Site Reliability Engineer (Bengaluru)

Senior Site Reliability Engineer (Bengaluru)

07 Aug
|
Nexthink
|
Bengaluru

07 Aug

Nexthink

Bengaluru

Job Description

We are looking for an experienced, proactive and innovative skilled that is keen to join as a Senior Site Reliability Engineer! The mission of Nexthinks SRE team is to strengthen our infrastructure and enhance our ability to deploy, monitor, and scale systems effectively and reliably. They work closely with over 50 Product Engineering teams that develop our products and services, as well as with the Technical Platform Engineering, Security and Architecture teams to understand the reliability requirements, design and implement solutions, and promote them for adoption and usage.

As a Senior Site Reliability Engineer, you will:

- Implement and manage cloud-native systems (AWS) using best-in-class tools and automation.
- Operate and enhance Kubernetes clusters, deployment pipelines, and service meshes to support rapid delivery cycles.
- Design, build, and maintain the infrastructure powering our multi-tenant SaaS platform with reliability, security, and scalability in mind.
- Define and maintain SLOs, SLAs, and error budgets, and proactively address availability and performance issues.
- Develop infrastructure-as-code (Terraform or similar) for repeatable and auditable provisioning.
- Build internal platform tools and automation to support provisioning, monitoring, and operational efficiency.
- Monitor infrastructure and applications ensuring high-quality user experiences.
- Participate in a shared on-call rotation, responding to incidents, troubleshooting outages, and driving timely resolution and communication.
- Act as an Incident Commander during the on-call duty and coordinate cross-team responses effectively to maintain an SLA.
- Drive and refine incident response processes,



reducing Mean Time to Detect (MTTD) and Mean Time to Recovery (MTTR).
- Diagnose and resolve complex issues independently, minimizing the need for external escalation.
- Work closely with software engineers to embed observability, fault tolerance, and reliability principles into service design.
- Automate runbooks, health checks, and alerting to support reliable operations with minimal manual intervention.
- Support automated testing, canary deployments, and rollback strategies to ensure safe, fast, and reliable releases.
- Contribute to security best practices, compliance automation, and cost optimization.

- Bachelor's degree in Computer Science or equivalent practical experience.

- 5+ years of experience as a Site Reliability Engineer or Platform Engineer with strong knowledge of software development best practices.
- Strong hands-on experience with public cloud services (AWS, GCP, Azure) and supporting SaaS product.
- Strong programming or scripting skills (e.g., Python, Go, Bash...), and experience with infrastructure-as-code (e.g. Terraform).
- Proficiency with Kubernetes, container-based deployment (e.g., Docker) and related ecosystems (e.g., Helm).
- Experience supporting multi-tenant microservices architectures.
- Experience with CI/CD pipelines tools (e.g., Jenkins, GitHub Actions,



GitLab CI, FluxCD, Crossplane).
- Experience with managing monitoring solutions (e.g. Datadog).
- Comfortable participating in a rotating on-call schedule, managing critical incidents, and leading post-incident reviews.
- At ease with operating and managing production systems, striking the right balance between urgency and methodology.
- Strong system-level troubleshooting skills and a proactive mindset toward incident prevention.
- Deep understanding of Linux systems, networking, and common troubleshooting practices.
- Solid understanding of the network stack (e.g., TCP/IP, VPN, etc.), cloud architectures (VPC, subnets, firewalls, load balancers), service mesh (e.g., Istio) and storage (e.g., S3, EBS, etc).
- Knowledge of zero-downtime deployment strategies, blue/green and canary releases.
- Exposure to compliance standards such as SOC 2, ISO 27001, or HIPAA. FedRAMP experience is a big plus.
- Experience with chaos engineering or resilience testing practices.
- Excellent problem-solving skills, collaborative mindset, and a strong grasp of agile, iterative development.
- Self-driven, highly organised, and capable of independently managing priorities.
- Curiosity to learn recent things and discover recent technologies.
- Strong communication, presentation, and team collaboration skills.
- Excellent written and verbal skills in English.

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

📌 Senior Site Reliability Engineer (Bengaluru)
🏢 Nexthink
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineer (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineer (bengaluru) / bengaluru