Life at TCG Digital
At TCG Digital, we’re more than technologists—we’re visionaries shaping the future of business with AI-driven solutions that transform speed into measurable value. Our global team is a dynamic blend of engineers, designers, and strategists spanning six countries and over 15 languages, all working together to create creative solutions for complex challenges. From mathematicians and ethical hackers to musicians and comedians, our team thrives on diversity. Each unique background contributes fresh perspectives and creativity, powering the innovative solutions that define our success.
Join us to make an impact, not just on technology but on the world of business itself.
Role
SRE Guild Lead
Location
Kolkata
Experience
10-15 Years
Education
Bachelor’s or Master’s degree in Computer Science, Engineering, or
related field
Number of Positions
1
Skills Required
- Excellent communication and influencing skills –able to set and land standards across pods without formal reporting authority.
- Deep hands-on experience with observability stacks (e.g., Prometheus, Grafana, ELK/OpenSearch, Datadog) – designing dashboards, alerts, and golden signals.
- Proven experience defining and operationalizing SLIs, SLOs, and error budgets, and using them to drive engineering and release decisions.
- Strong background in incident management – oncall design, escalation paths, severity classification, and blameless postmortems.
- Hands-on expertise with Kubernetes, container orchestration, and cloud-native infrastructure (AWS, Azure, or GCP).
- Infrastructure as Code and automation experience (Terraform, Ansible, Helm, or equivalent) with a strong toil-reduction mindset.
- Proficient in scripting/programming for automation and tooling (Python, Go, or Shell).
- Experience with CI/CD platforms and release engineering (Jenkins, GitLab CI, ArgoCD, or similar).
- Working knowledge of capacity planning, performance tuning,
and cost optimization for production systems.
- Familiarity with chaos engineering, disaster recovery, and multiregion resilience practices.
- Ability to mentor engineers across multiple pods, run guild syncs, and maintain shared runbooks, standards, and playbooks.
- Proven ability to work cross-functionally with Architecture, DevOps & Infra, and Delivery leadership
Roles & Responsibilities
- Own the SRE guild charter – define reliability standards, tooling, and best practices applied consistently across all Dev Factory pods.
- Define and govern SLIs, SLOs, and error budgets for key services in partnership with pod leads and architects.
- Lead the guild’s incident management practice – on-call rotations, escalation paths, and blameless postmortems; act as senior escalation point for critical production incidents.
- Drive observability strategy across the organization, ensuring pods have consistent monitoring, alerting, and logging coverage.
- Champion automation and toil reduction, identifying repetitive operational work across pods and driving it toward self-service or automated remediation.
- Partner with the DevOps & Infra guild to align infrastructure, deployment, and reliability practices, and flag capacity or staffing gaps to guild and delivery leadership.
- Mentor and upskill SREs and DevOps engineers embedded in pods, running guild-wide knowledge-sharing sessions and maintaining shared runbooks.
- Contribute to hiring, onboarding, and career pathing for the SRE discipline across the delivery organization.
- Collaborate with the Architecture & Design guild on non-functional requirements – scalability, fault tolerance, and disaster recovery – during solution design.
- Track and report guild health metrics (incident trends, MTTR, SLO attainment, automation coverage) to Dev Factory leadership.
- Stay current on SRE tooling and industry practices, evaluating and introducing new technologies where they improve reliability or reduce operational cost
SPOC
Human Resources
Mail to
[email protected]
📌 SRE Guild Lead (India)
🏢 TCG Digital
📍 India