08 Aug
|
Tata Consultancy Services
|
India
08 Aug
Tata Consultancy Services
India
Reliability & Availability: Designing, building, and maintaining systems that are highly available, performant, and resilient, ensuring minimal downtime.
Automation (Toil Reduction): Automating repetitive operational tasks (like deployments, scaling, provisioning) using scripting (Python, Go) and Infrastructure as Code (IaC) to increase efficiency and reduce human error.
Monitoring & Alerting: Implementing comprehensive observability (logs, metrics, traces) using tools like Cloud Logging/Monitoring (Stackdriver) to detect issues proactively and define actionable alerts.
Incident Management: Leading the response to production incidents, performing root cause analysis (postmortems), and implementing preventative measures.
Capacity Planning & Performance Tuning: Forecasting resource needs, optimizing resource allocation (GCP Compute, Storage, Networking),
and tuning systems for optimal speed and efficiency.
SLO/SLI Management: Defining, measuring, and enforcing Service Level Objectives (SLOs) and Indicators (SLIs) using Error Budgets to balance reliability and feature velocity.
Collaboration & Consultation: Working with development teams (DevOps culture) to embed reliability into software design and CI/CD pipelines (Cloud Build, Cloud Deploy) from the start.
Security & Compliance: Ensuring systems adhere to security best practices, managing vulnerabilities, and maintaining compliance within the GCP workplace.
CI/CD & Deployment: Managing and improving continuous integration and continuous delivery pipelines for protected, reliable software releases.
📌 Site Reliability Engineer Hyderabad (India)
🏢 Tata Consultancy Services
📍 India