Cloud Operations Engineers are responsible for building internal tools and process automation. Day-to-day duties are creating and monitoring systems alert dashboards, reviewing critical event and system logs, accessing customer instances that underpin their production databases, and performing server administration duties including performance troubleshooting. Applicants must be critical thinkers who are quick to detect, resolve, or escalate issues that are sometimes broad in scope and difficult to trace.
We are looking to speak to candidates who are based in Bengaluru for our hybrid working model.
Responsibilities
Help scale the Cloud Operations Engineering team with the strategic implementation and refinement of processes and tools
Provide career development feedback and advice to direct reports
Identify and measure team health indicators and performance metrics
Ensure proper team focus on priorities, objectives, and related deliverables
Collaborate with technical and non-technical teams across the company
Balance your time between leading your team, working on customer incidents and being involved in projects
Be a source of guidance and advice to your own team members and other teams within MongoDB
Build a relationship with your team around trust
Successfully coordinate with a global team of Cloud Operations Engineers who are tasked with ensuring our uptime guarantees to the MongoDB Atlas customer base
Participate in designing and building internal tools
Assist in scoping, designing and deploying systems that reduce Mean Time to Resolve for customer incidents
Monitor and detect emerging customer-facing incidents on the Atlas platform; assist in their proactive resolution
Automate internal processes, routine monitoring and troubleshooting tasks
Diagnose live incidents, differentiate between platform issues versus usage issues, and take the next steps toward resolution
Cooperate with our Product Management and Cloud