SRE Guild Lead (Kolkata)

SRE Guild Lead (Kolkata)

06 Aug
|
tcg digital solutions
|
Kolkata

06 Aug

tcg digital solutions

Kolkata

TCG Digital is looking for SRE Guild Lead for Kolkata Location.

Justification for the vacancy:

We are looking for an experienced SRE Guild Lead to own the practice, standards, and craft of Site Reliability Engineering across our Factory pods. As guild lead, the person will be responsible for the technical and cultural anchor for reliability at TCG Digital setting the bar for how we define, measure, and defend service reliability, and ensuring every pod applies consistent SRE practices regardless of which product or client engagement they sit under. The role will work closely with the DevOps & Infra guild, Architecture & Design guild, and pod leads to embed observability, automation, and operational excellence into how we build and run systems.

To succeed in this role, the person should have hands-on production engineering depth combined with the ability to mentor and set standards across a distributed group of engineers who do not report to the role directly. This is a guild leadership role, not a pure people-management role: influence, technical credibility, documentation, and cross-pod coordination are your primary tools. The person will also be the escalation point for major incidents and a key voice in capacity planning for reliability roles across the delivery organization.

Essential Skills:

- Excellent communication and influencing skills able to set and land standards across pods without formal reporting authority.
- Deep hands-on experience with observability stacks (e.g., Prometheus, Grafana, ELK/OpenSearch, Datadog) designing dashboards, alerts, and golden signals.




- Proven experience defining and operationalizing SLIs, SLOs, and error budgets, and using them to drive engineering and release decisions.
- Strong background in incident management on-call design, escalation paths, severity classification, and blameless postmortems.
- Hands-on expertise with Kubernetes, container orchestration, and cloud-native infrastructure (AWS, Azure, or GCP).
- Infrastructure as Code and automation experience (Terraform, Ansible, Helm, or equivalent) with a strong toil-reduction mindset.
- Proficient in scripting/programming for automation and tooling (Python, Go, or Shell).
- Experience with CI/CD platforms and release engineering (Jenkins, GitLab CI, ArgoCD, or similar).
- Working knowledge of capacity planning, performance tuning, and cost optimization for production systems.
- Familiarity with chaos engineering, disaster recovery, and multi-region resilience practices.
- Ability to mentor engineers across multiple pods, run guild syncs, and maintain shared runbooks, standards, and playbooks.
- Proven ability to work cross-functionally with Architecture, DevOps & Infra, and Delivery leadership

Roles & Responsibilities:

- Own the SRE guild charter define reliability standards, tooling,



and best practices applied consistently across all Dev Factory pods.
- Define and govern SLIs, SLOs, and error budgets for key services in partnership with pod leads and architects.
- Lead the guild's incident management practice – on-call rotations, escalation paths, and blameless postmortems; act as senior escalation point for critical production incidents.
- Drive observability strategy across the organization, ensuring pods have consistent monitoring, alerting, and logging coverage.
- Champion automation and toil reduction, identifying repetitive operational work across pods and driving it toward self-service or automated remediation.
- Partner with the DevOps & Infra guild to align infrastructure, deployment, and reliability practices, and flag capacity or staffing gaps to guild and delivery leadership.
- Mentor and upskill SREs and DevOps engineers embedded in pods, running guild-wide knowledge-sharing sessions and maintaining shared runbooks.
- Contribute to hiring, onboarding, and career pathing for the SRE discipline across the delivery organization.
- Collaborate with the Architecture & Design guild on non-functional requirements – scalability, fault tolerance, and disaster recovery – during solution design.
- Track and report guild health metrics (incident trends, MTTR, SLO attainment, automation coverage) to Dev Factory leadership.
- Stay current on SRE tooling and industry practices, evaluating and introducing recent technologies where they improve reliability or reduce operational cost

📌 SRE Guild Lead (Kolkata)
🏢 tcg digital solutions
📍 Kolkata

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: sre guild lead (kolkata) / kolkata

Subscribe to this job alert:

Get the latest job offers by email for: sre guild lead (kolkata) / kolkata