02 Aug
|
Capgemini
|
Telangana
02 Aug
Capgemini
Telangana
Role Overview:
Ensures reliability, scalability, and performance of systems across any cloud platform. Focuses on observability, incident management, and automation.
Key Responsibilities:
Manage infrastructure on AWS, Azure, or GCP.
Implement IaC using Terraform and Deployment Manager.
Build CI/CD pipelines using GitHub Actions and other tools.
Containerize and orchestrate workloads using Docker and Kubernetes.
Automate tasks using Linux, Bash, and Python scripting.
Monitor systems using Kibana, Splunk, Recent Relic, and APM tools like Grafana, Dynatrace, AppDynamics, Datadog.
Define and manage SLIs/SLOs, error budgets.
Work with ITSM/ITIL processes, ticketing tools like JIRA, ServiceNow.
Handle production support, release management, and chaos engineering.
Use tools like LitmusChaos or Gremlin for resilience testing.
Integrate PagerDuty or Opsgenie for automated escalation.
Automate deployment gates based on SLO compliance.
Use predictive analytics for scaling decisions.
Implement zero-trust principles and vulnerability management.
Must-Have Skills:
Multi-cloud experience (AWS, Azure, GCP).
Terraform and Deployment Manager.
GitHub Actions and CI/CD tools.
Docker and Kubernetes.
Linux, Bash, Python scripting.
Monitoring and APM tools (Splunk, Grafana, Dynatrace, Datadog, etc.).
SLI/SLOs, error budgets, ITSM/ITIL, ticketing tools.
Valuable-to-Have Skills:
Ansible, Helm, OpenShift.
OpenTelemetry, Jaeger for distributed tracing.
Service mesh and global load balancing strategies.
Integrate CI/CD with ITSM for automated change records.
📌 Sre Telangana
🏢 Capgemini
📍 Telangana