29 Sep
|
Suvidha Sewa
|
India
29 Sep
Suvidha Sewa
India
Roles & Responsibilities:
- Ensure production reliability and performance by monitoring service availability, latency, capacity, and overall system health while managing SLOs, SLIs, and error budgets.
- Automate operational processes and reduce toil by developing tools and scripts for deployment, infrastructure management, incident response, capacity planning, and other repetitive operational tasks.
- Manage incident response and root-cause analysis, participate in on-call rotations, troubleshoot production issues, and conduct blameless postmortems to implement long-term corrective actions.
- Design and improve scalable infrastructure by collaborating with software engineering teams on system architecture, distributed systems, reliability, scalability, and performance requirements.
- Manage cloud infrastructure and deployment processes using technologies such as Kubernetes, Terraform, Ansible, and CI/CD pipelines, including canary releases and automated deployment practices.
- Implement observability, capacity planning, and performance optimization across production systems, using monitoring and logging platforms such as Datadog or Splunk and optimizing SQL/NoSQL databases and distributed services.
Qualifications:
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related technical field, or equivalent practical experience.
- 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Cloud Infrastructure, or a closely related role, with strong exposure to production environments.
- Robust programming or scripting skills in at least one language such as Python, Go, Java, or C++, with the ability to develop automation tools and solve complex technical problems.
- Strong knowledge of Linux/Unix systems, networking, and distributed systems, including TCP/IP, DNS, load balancing, service reliability, and troubleshooting across distributed environments.
- Hands-on experience with cloud-native infrastructure and Infrastructure as Code, including Kubernetes, Terraform, Ansible, CI/CD, and automated deployment practices.
- Experience with observability and database technologies, including tools such as Datadog/Splunk and both relational and NoSQL databases, along with strong analytical, debugging, incident-management, and performance-tuning skills.
📌 Site Reliability Engineer (India)
🏢 Suvidha Sewa
📍 India