Ethos is looking for a Senior Engineering Manager to lead our SRE team in Bangalore. In this role, you will manage a team of engineers tasked with our core infrastructure, service reliability, and the development of our internal platforms including the toolchains supporting our AI initiatives. You will oversee the technical roadmap, drive hiring and career development, and ensure our production environments meet high standards for uptime, security, and cost-efficiency.
This is a hands-on management position. You will stay close to the technology participating in design reviews and troubleshooting complex issues while balancing the operational goals of the SRE charter with the emerging infrastructure needs of our AI platform.
Duties and Responsibilities:
Reliability Performance Engineering
- Incident Management: Oversee the incident response process, ensuring swift resolution of service disruptions and a high standard for post-mortem analysis.
- Service Standards: Help teams define and monitor SLIs and SLOs.
Implement practices like capacity planning and disaster recovery to ensure system endurance.
- Operational Health: Manage on-call rotations and develop forward-thinking monitoring and alerting systems to detect failures before they impact users.
Infrastructure Platform Tooling
- Scalable Systems: Lead the design and maintenance of our cloud infrastructure (AWS) using Infrastructure as Code (Terraform).
- AI Tooling Governance: Oversee the development of governed gateways for AI model routing. This includes managing model access, implementing audit trails, and securing data connectors (such as MCP) to prevent data exposure.
- Security Compliance: Ensure that all infrastructure components meet security and regulatory requirements. This includes AI prototyping environments. You will achieve this by using automated scanning and restricting data access.