21 Aug
|
Saanvi Nexus
|
India
21 Aug
Saanvi Nexus
India
Site Reliability Engineer (SRE) GCP Experience: 6+ Years Job Type: Full-Time Work Mode: Remote Shift: Australian Shift – 4:30 AM to 1:30 PM IST Joining: Immediate Joiners Only Interview Process: 3 Rounds Job Summary We are looking for an experienced Site Reliability Engineer (SRE) with 6+ years of experience to own the reliability, availability, scalability, and performance of production workloads and microservices. The ideal candidate will have strong hands-on experience with Google Cloud Platform (GCP), Cloud Run, Kubernetes, Terraform/Terragrunt, CI/CD, observability, and containerized microservices. The role requires strong production ownership, incident management skills, and the ability to work closely with engineering teams to build reliable, secure, and scalable systems. Key Responsibilities • • • • • • • • • • • • Own the reliability, availability, and performance of microservices and production workloads. Design, build, and improve resilient infrastructure on GCP, with strong focus on Cloud Run, Kubernetes, and containerized services. Build and maintain observability across logs, metrics, tracing, alerting, and service health. Improve CI/CD pipelines, release controls, rollback strategies, and environment consistency. Lead incident response, production readiness, postmortems, runbooks, on-call practices, and resilience testing. Automate repetitive operational tasks and reduce engineering toil through effective tooling. Partner with development teams to improve the operability, scalability, and fault tolerance of microservices.
Strengthen cloud security and infrastructure hygiene across IAM, secrets management, workload hardening, and production safeguards. Improve system performance, resource efficiency, and cloud cost optimization without compromising reliability. Participate in architecture and reliability reviews for critical services and high-traffic business events. Support capacity planning, disaster recovery, and production resilience initiatives. Identify reliability risks and recommend practical solutions based on business and engineering priorities. Required Qualifications & Skills • • 6+ years of experience in Site Reliability Engineering, DevOps, or closely related roles with meaningful production ownership. Solid hands-on experience managing production systems on Google Cloud Platform (GCP). 1 • • • • • • • • • • • Proven production experience with Cloud Run, Kubernetes, Docker, and containerized microservices. Strong experience with Terraform and Terragrunt for Infrastructure as Code. Strong understanding of observability and monitoring using OpenTelemetry, Google Cloud Monitoring, New Relic, or equivalent tools. Strong understanding of distributed systems, microservice architectures,
failure modes, reliability engineering, and production debugging. Experience building and improving CI/CD pipelines and release workflows, particularly using GitHub Actions. Programming/scripting experience in Python, Java, or another suitable language. Strong incident management and troubleshooting skills with a practical approach to reliability, recovery, and risk. Good understanding of cloud security, IAM, secrets management, and production safeguards. Strong communication and collaboration skills with the ability to work across engineering teams. Experience with AI tooling and agentic workflows in engineering or operational environments. Experience in retail, e-commerce, SaaS, or other customer-facing environments is an advantage. Preferred Skills • • • • • Experience with high-traffic, customer-facing production environments. Experience with reliability testing, chaos/resilience testing, and disaster recovery. Familiarity with SLOs, SLIs, SLAs, error budgets, and reliability metrics. Experience with cloud cost optimization and resource efficiency. Strong understanding of production readiness and engineering best practices. Important Hiring Details • • • • • • Experience: 6+ Years Work Mode: 100% Remote Shift: Australian Shift | 4:30 AM – 1:30 PM IST Joining: Immediate Joiners Only Interview Process: 3 Rounds Candidates should be comfortable working in the specified Australian shift and taking ownership of production systems. 2
📌 GCP - Site Reliability Engineer (India)
🏢 Saanvi Nexus
📍 India