Company Profile:
Founded in 1976, CGI is among the largest independent IT and business consulting services firms in the world. With 94,000 consultants and professionals across the globe, CGI delivers an end-to-end portfolio of capabilities, from strategic IT and business consulting to systems integration, managed IT and business process services and intellectual property solutions. CGI works with clients through a local relationship model complemented by a global delivery network that helps clients digitally transform their organizations and accelerate results.
CGI Fiscal 2024 reported revenue is CA$14.68 billion and CGI shares are listed on the TSX (GIB.A) and the NYSE (GIB). Learn more at cgi.com.
Job Title: Production support Lead
Position: Senior Software Engineer/Lead
Experience: 7 -12 Years
Category: Software Development/ Engineering
Shift: General
Main location: Bangalore/Hyderabad/Chennai
Position ID:
Employment Type: Full Time
Education Qualification: Bachelors degree in computer science or related field or higher with minimum 7 years of relevant experience.
Position Description:
We are looking for an experienced Production Support / SRE Lead to lead L2/L3 production support operations for a highly available, microservices-based environment. The role requires strong technical expertise across cloud, container platforms, application technologies, databases, monitoring, automation, and incident management.
The ideal candidate will be responsible for ensuring high availability, reliability, performance, and operational excellence of production systems while leading and mentoring a team of production support engineers.
The candidate should be comfortable navigating multiple technologies and working closely with Development, DevOps, Architecture, QA, and business stakeholders to resolve critical production issues and continuously improve platform reliability.
Responsibilities:
- Lead the triage, investigation, and resolution of complex L2/L3 production incidents in a microservices environment while ensuring adherence to SLA/SLO requirements.
- Lead and mentor a team of production support engineers, manage shift rosters, prioritize workloads, and foster a culture of technical excellence, ownership, and accountability.
- Design, implement, and maintain comprehensive monitoring, logging, alerting, and observability strategies to proactively identify bottlenecks, anomalies, performance issues, and potential system failures.
- Drive detailed Root Cause Analysis (RCA) for critical P1/P2 incidents and work closely with Development, Architecture, DevOps, and other technical teams to identify permanent fixes and prevent recurrence.
- Promote SRE practices and identify operational toil and repetitive manual activities,
driving automation initiatives such as auto-remediation scripts, automated health checks, monitoring improvements, and operational tooling.
- Act as the primary technical point of contact during major incidents and provide clear, timely, and accurate updates to stakeholders, Delivery Managers, and leadership.
- Collaborate with DevOps and CI/CD teams to support seamless production deployments, blue/green deployments, canary releases, rolling deployments, and zero-downtime transitions.
- Monitor production systems proactively and identify application, infrastructure, database, and integration-related issues before they impact customers.
- Analyze system performance, availability, capacity, logs, metrics, and traces to identify and resolve production issues.
- Drive continuous improvement of production support processes, monitoring capabilities, incident response, automation, and system reliability.
- Work across multiple technology stacks and effectively troubleshoot issues involving AWS, OpenShift, Node.js, Python, Java, MongoDB, APIs, microservices, and integrations.
- Ensure proper incident documentation, RCA completion, corrective actions, preventive actions, and knowledge sharing across the support team.
- Participate in production readiness reviews, release validation, capacity planning, disaster recovery activities, and reliability improvement initiatives.
Must Have Skills
- Strong experience in L2/L3 Production Support within a large-scale, high-availability production environment.
- Strong understanding of microservices architecture and distributed systems.
- Hands-on experience with AWS and production troubleshooting in cloud environments.
- Strong experience with OpenShift/Kubernetes and containerized applications.
- Good working knowledge of Node.js, Python, and Java for application troubleshooting and automation.
- Strong experience with MongoDB, including troubleshooting, performance analysis, and production support.
- Strong understanding of REST APIs, integrations, service-to-service communication, and distributed applications.
- Robust knowledge of Incident Management, Problem Management, Major Incident Management, and Change Management.
- Proven experience handling P1/P2 incidents and driving complex production issues to resolution.
- Strong Root Cause Analysis and problem-solving capabilities.
- Good understanding of SLA, SLO, SLI, availability, reliability, and performance metrics.
- Experience with monitoring, logging,
alerting, and observability platforms.
- Experience with CI/CD pipelines and DevOps practices.
- Experience with production deployments, rollback strategies, blue/green deployments, and canary releases.
- Strong scripting and automation skills using Python, Node.js, Shell, or similar technologies.
- Strong leadership, communication, stakeholder-management, and team-management skills.
- Ability to work effectively under pressure during critical production incidents.
- Strong ability to navigate and troubleshoot across multiple technologies rather than being limited to a single technology stack.
Good to Have Skills
- Experience with Prometheus, Grafana, ELK/Elastic Stack, Splunk, OpenTelemetry, CloudWatch, or similar observability platforms.
- Experience with Terraform or other Infrastructure as Code (IaC) tools.
- Experience with Kafka or other messaging technologies.
- Knowledge of API Gateway, service mesh, and distributed tracing.
- Understanding of SRE principles, error budgets, and reliability engineering practices.
- Experience implementing automated health checks and auto-remediation solutions.
- Knowledge of disaster recovery, high availability, business continuity, and resilience engineering.
- Experience with chaos engineering and resilience testing.
- AWS or Kubernetes/OpenShift certifications.
- Experience working in Agile/DevOps environments.
- Knowledge of application and cloud security practices.
CGI is an equal opportunity employer. In addition, CGI is committed to providing accommodation for people with disabilities in accordance with provincial legislation. Please let us know if you require reasonable accommodation due to a disability during any aspect of the recruitment process and we will work with you to address your needs.
Life at CGI
It is rooted in ownership, teamwork, respect and belonging. Here, you’ll reach your full potential because
You are invited to be an owner from day 1 as we work together to bring our Dream to life. That’s why we call ourselves CGI Partners rather than employees. We benefit from our collective success and actively shape our company’s strategy and direction
Your work creates value. You’ll develop innovative solutions and build relationships with teammates and clients while accessing global capabilities to scale your ideas, embrace recent opportunities, and benefit from expansive industry and technology expertise
You’ll shape your career by joining a company built to grow and last. You’ll be supported by leaders who care about your health and well-being and provide you with opportunities to deepen your skills and broaden your horizons
Come join our team, one of the largest IT and business consulting services firms in the world
📌 Production support Lead/SSE (Bengaluru)
🏢 CGI
📍 Bengaluru