Key Responsibilities
Reliability & Operations
• Maintain high availability and performance of production environments.
• Define and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).
• Perform root cause analysis (RCA) for production incidents.
• Participate in on-call support and incident response activities.
• Drive continuous improvements to system reliability and operational efficiency.
Automation & Infrastructure
• Automate infrastructure provisioning using Infrastructure as Code (IaC).
• Develop scripts and tools to eliminate repetitive operational tasks.
• Implement automated deployment and rollback mechanisms.
• Manage and optimize CI/CD pipelines.
Cloud & Platform Management
• Design, implement, and support cloud-native solutions.
• Manage Kubernetes clusters and containerized workloads.
• Optimize cloud resource utilization and cost.
• Ensure secure and scalable infrastructure architecture.
Monitoring & Observability
• Build and maintain monitoring, logging, and alerting solutions.
• Create operational dashboards and system health reports.
• Improve observability through metrics, logs, and tracing.
• Proactively identify and resolve performance bottlenecks.
Security & Compliance
• Implement security best practices across infrastructure.
• Support vulnerability remediation and patch management.
• Ensure compliance with organizational security standards.
• Collaborate with cybersecurity teams on infrastructure security initiatives.