We are looking for a motivated and technically skilled operations responsible in the internal team leading the charge to resolve critical outages and stabilize production environments of the Online Remote Update Technology.
Key Responsibilities
- Rapid Response: During high-priority outages, coordinating engineers to restore service as quick as possible.
- Root Cause Analysis (RCA): Lead post-mortem investigations to find the "why" behind a crash and ensure it never happens again.
- Stakeholder Sync: Provide calm, clear updates to executives and customers while the technical team is busy fixing the problem.
- Focus on reducing the Mean Time To Repair, constantly looking for ways to speed up the recovery process.
- The role transforms into structured recovery, ensuring that every system failure becomes a blueprint for future resilience and improved reliability.
- Work together with Infra specialist to Design "self-healing" systems and backups so the app can recover instantly if a server fails.
- Optimize backend systems for high performance,
availability, and reliability.
- Collaborate closely with product managers, architects, QA, and operations to deliver solutions aligned with business goals.
- Create and maintain clear, comprehensive technical documentation for APIs and system components.
- Own service delivery from design through production, ensuring issues are addressed promptly.
- Mentor and support engineers, sharing knowledge and fostering professional growth.
- Identify and implement opportunities for continuous improvement in processes and technology.
- Ensure solutions adhere to internal quality standards and external compliance requirements.
- Communicate effectively with stakeholders, translating technical concepts into business value.