21 Aug
|
Recrew AI
|
Bangalore Urban
21 Aug
Recrew AI
Bangalore Urban
Role: Site Reliability Engineer
Function: Engineering / Infrastructure / SRE
Location: Bangalore
Type: Full-time
Industry: Market Research / Data & Analytics / Marketing Technology
About Company The company is a global leader in data-driven marketing research. It serves over 4,000 brands across Asia-Pacific with actionable consumer insights.
With 25+ years of expertise and 130 million+ consumer panelists worldwide, it combines survey data, digital behavior, and purchase insights. The company is in an active growth-through-acquisition phase, backed by ~$848M in funding.
Its engineering teams build cloud-native platforms that power research at massive scale.
Position Overview
As a Site Reliability Engineer, you will own the end-to-end platform and infrastructure that powers the company's research systems at scale. You'll work closely with engineering teams to build reliable, secure, and cost-efficient systems on GCP — driving automation, observability, and developer productivity across the organisation.
Role & Responsibilities
- Own and manage end-to-end cloud infrastructure on GCP, including Compute Engine, GKE, Cloud SQL, Pub/Sub, and Cloud Storage
- Design, build, and maintain CI/CD pipelines using GitHub Actions to enable faster and safer deployments
- Implement and manage Infrastructure as Code using Terraform for all infrastructure provisioning and automation
- Build and enhance the observability stack (Datadog, OpenTelemetry) covering logging, metrics, and distributed tracing
- Lead incident management, root cause analysis, and post-mortem processes for production systems
- Define and maintain SLIs, SLOs, and error budgets to drive reliability decisions across services
- Automate operational processes, reduce toil,
and support service onboarding to modern platform architecture
Must Have Criteria
- 4+ years of experience building and operating production systems at scale
- Hands-on experience with GCP services (Compute Engine, GKE, Cloud SQL, Cloud Storage, Pub/Sub)
- Proficiency in Terraform for infrastructure provisioning and management in production environments
- Experience running containerised workloads with Docker and Kubernetes (GKE) in production
- Experience building and maintaining CI/CD pipelines (GitHub Actions or equivalent)
- Hands-on experience with observability tools — specifically Datadog and/or OpenTelemetry (metrics, logs, traces)
- Programming experience in Go and scripting experience in Bash for automation and tooling
Nice to Have
- Hands-on experience with SRE practices: SLO-driven operations, error budgets, and reliability reviews
- Experience building internal developer platforms or platform engineering initiatives
- Business-level Japanese proficiency (JLPT N3 or equivalent) for collaboration with Japan-based teams
- Experience applying AI/ML tools to enhance SRE automation or incident response
- Open-source contributions or experience mentoring engineers on SRE/DevOps practices
What We Offer
- Prospect to own and shape the entire platform infrastructure for a globally scaled research platform
- Work with a modern, cloud-native stack (GCP, Terraform, Datadog, Go) in an agile engineering culture
- Exposure to large-scale consumer data systems serving 4,000+ enterprise clients across Asia-Pacific
- Collaborative, transparent work culture with strong ownership and continuous learning
- Growth opportunities within a company in an active merger and acquisition phase
📌 Site Reliability Engineer (Bangalore Urban)
🏢 Recrew AI
📍 Bangalore Urban