- Design, implement, and maintain CI/CD pipelines using Azure DevOps to ensure effective software development lifecycle.
- Collaborate with cross-functional teams to identify areas for process improvement and implement changes using Site Reliability Engineering principles.
- Develop and manage infrastructure as code (IaC) using Terraform or CloudFormation to ensure consistency across environments.
- Troubleshoot issues related to pipeline failures or deployment errors, working closely with developers to resolve problems.
- Monitor system performance, availability, and security using tools such as Prometheus, Grafana, ELK Stack, and other observability platforms.
- Implement and maintain DevSecOps practices, including vulnerability management, compliance, and security automation. Collaborate with development, QA, and operations teams to improve release management and deployment efficiency.
- Monitor cluster health,
performance, scalability, and resource utilization. Optimize container orchestration and deployment strategies. Success Measure High application uptime and cluster stability. Improved application scalability and resource efficiency.
- Monitoring, Logging & Incident Management Implement and maintain monitoring and observability solutions using Prometheus, Grafana, ELK Stack, or equivalent tools.
- Proactively identify and resolve performance bottlenecks and infrastructure issues. Participate in root cause analysis and incident resolution.
- Success Measure Reduction in Mean Time to Detect MTTD. Reduction in Mean Time to Resolve MTTR. Increased system reliability and service availability.