Senior Site Reliability Engineer (Maharashtra)

Senior Site Reliability Engineer (Maharashtra)

04 Sep
|
THG Ingenuity
|
Maharashtra

04 Sep

THG Ingenuity

Maharashtra

About THG Ingenuity

THG Ingenuity is a fully integrated digital commerce ecosystem, designed to power brands without limits. Our global end-to-end tech platform is comprised of three products: THG Commerce, THG Studios, and THG Fulfilment. Each represents a single, unified solution, overcoming challenges and taking brands direct-to-consumer.

Our client portfolio includes globally recognised brands such as Coca-Cola, Nestle, Elemis, Homebase, and Proctor & Gamble. Technology is the driving force behind THG Ingenuity, and it starts with our people. We are ambitious with our goals and challenge conventional thinking.

THG Ingenuity is different because we support every single person to make a massive impact and drive their own work. Our people are always learning, and we work every day to ensure our technology, from our software platforms to our hosting services, to our AI capability and beyond, is world class. This enables us to keep powering THG Ingenuity and our partners on a global scale.

Tech at THG Ingenuity

Technology is the driving force behind THG Ingenuity, and it starts with our people. We are ambitious with our goals and challenge conventional thinking. THG Ingenuity is different because we support every single person to make a massive impact and drive their own work.

Our people are always learning, and we work every day to ensure our technology, from our software platforms to our hosting services, to our AI capability and beyond, is world class. This enables us to keep powering THG Ingenuity and our partners on a global scale.

About the Team and Your Role

K8saaS on GCP

Our primary Kubernetes-as-a-Service platform, built end-to-end on Google Cloud Platform, provides THG Ingenuity's engineering teams with scalable, cloud-native infrastructure. Leveraging GCP's managed services, including GKE, we deliver comprehensive Kubernetes capabilities including multi-tenant clusters, dedicated clusters and App Runtime integration. GCP is our strategic home for container orchestration and cloud infrastructure, delivering scalability, reliability and reduced operational overhead through GCP's managed offerings.

App Runtime

We maintain seamless integration with App Runtime, THG Ingenuity's internal developer platform owned by a separate team. This partnership enables our GCP-based clusters to support App Runtime's simplified deployment interface, allowing developers to deploy applications without complex YAML configurations or approval workflows. Our K8saaS on GCP platform provides the underlying infrastructure and cluster management that powers App Runtime's developer experience.

What will I be doing?

As a Senior SRE, you will own services and problem domains end-to-end: designing, building, operating and improving them. You will multiply the effectiveness of the team through technical leadership as well as your own delivery.

In a developmental capacity, you will be responsible for:

- Leading the design of automation and platform components and setting best practice for how infrastructure is built and operated.
- Owning problem domains end-to-end, from design through build to operation.
- Identifying sources of toil across the platform and driving projects that engineer them away.




- Evaluating new tooling and approaches and making well-reasoned adopt/build/buy recommendations.
- Contributing to upstream tooling where appropriate and representing the team in those communities.

In an operational capacity, you will be responsible for:
- Acting as incident commander for major incidents, leading blameless post-mortems, and ensuring corrective actions of land.
- Defining and refining SLIs and SLOs with tenant teams, and using error budgets to guide the balance between reliability work and feature work.
- Leading complex upgrades and migrations across the GKE estate using zero-downtime strategies.
- Capacity planning and performance engineering for the platform.
- Designing and testing backup, failover and disaster recovery procedures against agreed RTO/RPO targets.

In a wider capacity, you will be part of the team and helping to:
- Mentor and grow mid-level and junior engineers.
- Plan and communicate significant pieces of work to stakeholders, including timelines and risk.
- Evangelise App Runtime, GCP and Kubernetes within THG Ingenuity, and feed developer requirements into the platform roadmap.

What would success in this role look like?

Success is a platform whose reliability is designed rather than firefought. Services have meaningful SLOs, error budgets inform prioritisation, and the areas you own visibly reduce their operational load quarter on quarter.

Success is a new GCP region being brought online and the entire stack deploying to it in a short period of time, because the automation and GitOps pipelines you helped design make it routine.

Success is major incidents being rarer, shorter and better-learned-from because of the incident and post-mortem practice you drive.

Success is engineers around you getting measurably better because you mentor them, review their designs, and raise the bar on how the team works.

Role Requirements The ideal candidate has a proven track record of owning production systems at scale, encompasses the DevOps and SRE mindset of wanting to automate, maintain a good service, practise what they preach, address technical debt, and is ready to lead as well as build.

A Software Engineering background is required, with focus on availability, performance, monitoring, capacity planning and change management.

Desired Technologies and Skills

- Strong programming experience in Golang (primary) and Python (secondary), including designing maintainable tooling and services, not just scripts.
- Deep public cloud expertise with a strong preference towards GCP: designing solutions around GCP managed services, IAM, and understanding their cost, performance and reliability trade-offs.
- Infrastructure as Code at scale with Terraform: module design, state management and multi-environment patterns.
- Experience with configuration management tools and concepts, for example: Ansible.
- Deep Kubernetes (GKE) and Container expertise, including Helm chart authoring.
- Debugging, troubleshooting,



and extending CRDs and controllers.
- Designing and operating Ingress controllers such as Nginx or HAP Roxy, and the Gateway API as their successor, including migration between the two.
- Designing CI/CD pipelines in GitHub Actions and selecting appropriate zero-downtime deployment strategies (blue-green, canary, rolling) for a given service.
- GitOps at scale with Flux CD and Argo CD, including multi-cluster patterns.
- Observability architecture: operating Prometheus, Alert Manager and Grafana at scale, GCP's Cloud Operations suite, and designing SLO-based alerting that keeps pages actionable.
- Operating, tuning and recovering PostgreSQL and etc..
- Object storage design, including lifecycle and cost management.
- Messaging systems: operating ActiveMQ in production and designing with GCP Pub/Sub.
- End-to-end networking design: load balancing, firewall rules, DNS, CDN, VPC design, and automating the TLS certificate lifecycle.
- Familiarity with different open-source licenses, and their implications.
- Effective and appropriate use of AI tooling: using AI assistants and agents productively in day-to-day engineering, and building skills, agents and sub-agents with proper verification, validation and guardrails in place. Establishing patterns and guardrails for how the team uses these tools is expected at this level.

Nice to See

- Cloud-native workload migrations.
- Multi-tenant application architecture.
- Experience running services subject to compliance regimes such as PCI DSS.

Operational Skills

- Defining SLIs, SLOs and SLAs with stakeholders, and operating an error budget in practice.
- Incident command, on-call leadership, and driving a strong blameless post-mortem culture.
- Security best practices: IAM, secrets management, compliance (PCI DSS for ecommerce).
- Disaster recovery ownership: backup strategies, failover procedures, RTO/RPO planning and regular testing.
- Performance optimization: load testing, capacity planning, autoscaling design.
- Cost optimization: resource rights-sizing, committed use discounts, budget monitoring.
- Documentation: setting the standard for runbooks, architecture diagrams, and migration plans.

Soft Skills

- Mentoring: The ability to guide and grow junior and mid-level engineers.
- Project management: migration timeline planning, stakeholder communication.
- Risk assessment.
- Cross-team collaboration: working with development, QA, and business teams.
- Customer-facing platform management.

What’s in it for me?

- Build solutions using the latest technology.
- Work alongside genuine industry experts.
- Continuous development through THG Academy, our in-house L&D; team.

Equal Prospect Statement

THG Ingenuity is an equal opportunity employer. We are committed to creating an inclusive, respectful and merit-based workplace where all employment decisions are made without discrimination on the basis of race, color, religion, caste, gender, gender identity or expression, sexual orientation, disability, age, marital status, pregnancy, nationality, veteran status, or any other status protected under applicable Indian laws. We encourage applications from candidates of all backgrounds and are committed to providing reasonable accommodation throughout the recruitment process, where required.

📌 Senior Site Reliability Engineer (Maharashtra)
🏢 THG Ingenuity
📍 Maharashtra

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineer (maharashtra) / maharashtra

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineer (maharashtra) / maharashtra