About Searce
Searce is an AI outcome engineering company that helps businesses transform through deep technology expertise. With 3,500+ engineers spread across 12 countries, we partner with enterprises to design, build, and operate cloud-native, AI-first systems that generate measurable business outcomes. Searce is a Google Cloud Global MSP, AWS Advanced Consulting Partner, Databricks Partner, and Anthropic Claude Partner.
We are PCI DSS and ISO 27001 certified, ensuring enterprise-grade security and compliance.
Role overview
As a Lead SRE within the CSRE Delivery team, you will anchor reliability, scalability, and operational excellence for our client cloud environments — primarily on GCP, with hands-on exposure to at least one additional cloud platform (AWS or Azure). You will drive SRE practices, mentor engineers, and partner closely with DevOps and platform squads to engineer resilient systems that exceed SLA commitments.
Key responsibilities
SRE & Reliability Engineering
- Own SLOs, SLIs, and error budget policies; drive reliability roadmaps for client workloads.
- Lead incident response, post-mortems, and blameless RCA practices; eliminate toil through automation.
- Design and implement auto-healing, self-remediation, and chaos engineering frameworks.
- Ensure high availability, disaster recovery, and resilience across multi-cloud environments.
Cloud & Infrastructure
- Architect and operate production-grade GCP environments (GKE, Cloud Run, BigQuery, Pub/Sub, etc.).
- Extend multi-cloud coverage on AWS or Azure as a secondary platform (minimum 2 cloud platforms required).
- Manage infrastructure-as-code using Terraform / Deployment Manager; enforce GitOps workflows.
DevOps & CI/CD
- Own and optimise CI/CD pipelines (Cloud Build, Jenkins,
ArgoCD, GitHub Actions, or equivalent).
- Champion container-native delivery on Kubernetes / GKE; define deployment strategies (blue-green, canary).
- Integrate security and compliance gates (SAST, DAST, policy-as-code) within pipeline workflows.
Observability & Performance
- Build end-to-end observability stacks: Cloud Monitoring, Prometheus, Grafana, Datadog, or equivalent.
- Define alerting strategies, dashboards, and runbooks; conduct capacity planning and performance tuning.
- Drive AIOps and intelligent alerting adoption to reduce MTTR.
Leadership & Collaboration
- Mentor and guide a team of SRE/DevOps engineers; conduct code and design reviews.
- Collaborate with Delivery Managers, architects, and client stakeholders on SRE roadmaps.
- Contribute to CSRE capability building, tooling standards, and internal knowledge assets.
Skills & requirements
Mandatory
- 6 – 8 years of total experience in SRE, DevOps, or cloud infrastructure roles.
- Strong hands-on expertise in Google Cloud Platform (GCP) as primary cloud — GKE, VPC, IAM, Cloud Monitoring, Pub/Sub, BigQuery, Cloud SQL.
- Proficiency in at least one secondary cloud: AWS (EKS, EC2, RDS, CloudWatch) or Azure (AKS, Azure Monitor, ARM).
- Minimum 2 cloud platform certifications or equivalent production experience across 2 clouds.
- Strong DevOps background: CI/CD pipeline design,
container orchestration (Kubernetes), GitOps.
- Infrastructure-as-Code expertise: Terraform, Pulumi, or Cloud Deployment Manager.
- Demonstrated SRE practice: SLO/SLI definition, error budgets, incident management, toil reduction.
- Scripting / automation: Python, Go, Bash — for runbooks, tooling, and ops automation.
- Observability tooling: Prometheus, Grafana, Cloud Monitoring, Datadog, or PagerDuty.
Good to have
- Experience with service mesh (Istio, Anthos Service Mesh).
- Familiarity with FinOps practices and cloud cost optimisation.
- Exposure to AIOps platforms or ML-based anomaly detection tools.
- Google Professional Cloud DevOps Engineer or equivalent cloud certification.
- Experience in a managed services or cloud consultancy delivery setting.
What we offer
- Work on cutting-edge, large-scale GCP-first client environments across industries.
- Access to Searce's AI outcome engineering learning ecosystem and certifications.
- Cross-cloud exposure — GCP, AWS, and Azure — within a single delivery team.
- Transparent growth framework with clear paths to Senior Manager / AVP tracks.
- Collaborative, engineer-led culture with direct client impact and ownership.
About the CSRE team The Cloud Site Reliability Engineering (CSRE) Delivery team at Searce is responsible for designing, running, and evolving mission-critical cloud operations for enterprise clients globally. The team embeds SRE discipline into every delivery — from onboarding to steady-state ops — ensuring clients benefit from proactive reliability engineering, not just reactive support. Interested candidate mail the resume to
[email protected]
📌 Lead Site Reliability Engineer (Bengaluru)
🏢 Searce
📍 Bengaluru