Responsible for ensuring the reliability scalability and performance of software systems through automation monitoring and incident management
Collaborate with development and operations teams to design and implement scalable and reliable systems Develop and maintain automation tools for deployment monitoring and incident response Monitor system performance and troubleshoot issues to minimize downtime Implement best practices for system security backup and disaster recovery Participate in oncall rotations to provide timely incident resolution and root cause analysis Continuously improve system reliability through proactive maintenance and capacity planning
Roles and Responsibilities
Design build and maintain infrastructure and tools to support largescale applications
Automate operational processes including deployment configuration and monitoring
Respond promptly to system s and outages performing root cause analysis and implementing fixes
Collaborate with software engineers to improve system architecture and performance
Manage and optimize cloud infrastructure and services
Ensure compliance with security policies and standards
Document system configurations procedures and incident reports
Drive continuous improvement initiatives to enhance system reliability and efficiency