Educational Requirements
Bachelor of Engineering, BCS, BBA, BCom, MCA, MSc
Service Line
Data Analytics Unit
Responsibilities
Reliability Engineering: Design, build, and maintain highly available and fault-tolerant production systems.
Define and monitor SLIs, SLOs, and SLAs for critical services.
Drive reliability improvements through automation and proactive engineering.
Conduct capacity planning and performance optimization activities.
Production Support Operations: Manage production environments and ensure service uptime.
Lead incident response, troubleshooting, and root cause analysis (RCA).
Develop runbooks, operational playbooks, and disaster recovery procedures.
Additional Responsibilities:
Automation DevOps: Automate deployments, infrastructure management, and operational workflows.
Improve CI/CD pipelines and release processes.
Implement self-healing, auto-scaling, and operational automation solutions.
Promote DevOps and SRE best practices across engineering teams.
Security Compliance:
Ensure production settings meet security and compliance requirements.
Manage secrets, access controls, and vulnerability remediation.
Partner with security teams to implement security best practices.
Technical and Qualified Requirements:
Cloud Infrastructure: Deploy and manage cloud-native infrastructure across AWS, Azure, or GCP.
Automate infrastructure provisioning using Infrastructure as Code (IaC).
Implement scalable and secure infrastructure solutions.
Support Kubernetes-based platforms and containerized workloads.
Observability/Monitoring: Build monitoring, logging, tracing, and alerting solutions.
Implement observability frameworks using industry-standard tools.
Monitor application health, performance metrics, and infrastructure utilization.
Drive continuous improvements in platform visibility and diagnostics.
Preferred Skills
Domain->Manufacturing->Production Planning
Technology->DevOps->Site Reliability Engineering(SRE)
📌 Senior Site Reliability Engineer Hyderabad (India)
🏢 Infosys
📍 India