12 Sep
|
Tekskills
|
Chennai
Key Responsibilities
Reliability & Availability
- Ensure high availability of critical systems aligned to SLAs, SLOs, and error budgets
- Lead P1/P2 incident management, including root cause analysis and post-incident reviews
- Proactively identify and mitigate system risks
Engineering & Automation
- Build automation-first solutions for deployment, monitoring, and recovery
- Reduce operational toil using scripting (Python, Go, Java, Bash)
- Implement Infrastructure as Code (Terraform, ARM, CloudFormation)
Observability & Performance
- Enhance monitoring, logging, alerting, and tracing capabilities
- Define and track SLIs/SLOs based on business impact
- Perform capacity planning and performance testing
Cloud & Platform Engineering
- Manage systems across Azure, AWS, and hybrid environments
- Collaborate with DevOps and platform teams for secure and scalable design
- Advocate contemporary practices like Kubernetes and containerization
Governance, Risk & Compliance
- Ensure compliance with banking regulations and security standards
- Support audits, resilience testing, and risk assessments
- Maintain clear system and recovery documentation
Leadership & Collaboration
- Mentor junior engineers and provide technical guidance
- Drive adoption of SRE practices across teams
- Collaborate with stakeholders across Technology, Security, and Business
Key Skills & Experience
Essential
- Proven experience in SRE / DevOps / Production Engineering
- Strong programming skills in at least one language
- Hands-on expertise in cloud platforms (Azure/AWS)
- Experience handling high-availability systems in regulated environments
- Strong knowledge of Linux, networking, and distributed systems
- Excellent troubleshooting and incident management capabilities
Desirable
- Experience in banking or financial services
- Knowledge of Kubernetes, containers, and service meshes
- Familiarity with CI/CD pipelines
- Experience with SLOs, error budgets, and reliability frameworks
- Exposure to chaos engineering or resilience testing
Behavioural Competencies
- Strong ownership and accountability
- Data-driven mindset with focus on continuous improvement
- Effective communication with technical & non-technical stakeholders
- Ability to remain calm under pressure
- Proactive approach to identifying improvements
📌 Site Reliability Engineer (Chennai)
🏢 Tekskills
📍 Chennai