28 Sep
|
Infosys
|
Bengaluru
SRE, Production engineering,Kubernetes, Terraform, CI/CD pipelines, Observability (Prometheus/Grafana)
Key Responsibilities: Reliability &
- Production Ownership
- Lead production engineering practices to ensure high availability, scalability, and performance across services and platforms.
- Define and drive SLOs/SLIs, error budgets, capacity planning, and reliability roadmaps aligned to business priorities.
- Partner with engineering teams to design resilient architectures and reduce operational risk through proactive improvements.
Incident
Management &
- Operational Excellence
- Own incident response processes (on-call readiness, triage, escalation, communication) and lead major incident bridges when needed.
- Drive blameless postmortems, root-cause analysis, and corrective/preventive actions to prevent recurrence.
- Establish operational runbooks, playbooks, and production readiness reviews for new releases and changes.
Cloud
Operations &
- Automation
- Lead cloud operations to ensure secure, cost-effective, and reliable environments across regions/accounts/subscriptions.
- Identify toil and implement automation to improve deployment safety, recovery time, and operational efficiency.
- Standardize operational tooling and workflows to improve service health, change success rate, and MTTR. Leadership &
- Collaboration
- Mentor engineers and influence cross-functional teams to adopt reliability engineering best practices.
- Provide technical leadership in prioritization, execution planning, and stakeholder communication for reliability initiatives.
Minimum
Qualifications:
- BTECH, MTECH, MCA, or MSC in Computer Science, IT, or a related field (or equivalent practical experience).
- 12–14 years of experience in SRE, Production Engineering, or Reliability Engineering roles supporting large-scale systems.
- Robust hands-on experience in cloud operations, incident management, and production support for critical services.
- Proven ability to drive automation initiatives that reduce manual effort and improve system reliability.
- Demonstrated experience leading operational processes such as on-call, postmortems, and production readiness practices.
Preferred
Qualifications:
- Experience designing and implementing SLO/SLI frameworks, error budgets, and reliability KPIs across multiple teams.
- Strong background in observability practices (monitoring, alerting, logging, tracing) and building actionable operational dashboards.
- Expertise in release/change management practices that improve deployment safety and reduce production incidents.
- Experience leading cross-team reliability programs, influencing stakeholders, and driving measurable improvements in uptime and MTTR.
- Track record of mentoring engineers and setting engineering standards for operational excellence and automation at scale.
📌 SRE / Production Engineering (Bengaluru)
🏢 Infosys
📍 Bengaluru