What You'll Do
As a Site Reliability Engineer (SRE), you will be responsible for building and maintaining robust, automated systems that reduce manual intervention and improve overall system efficiency. They play a key role in incident response, performance optimization, and infrastructure management, often working closely with development teams to integrate reliability into the software development lifecycle
Responsibilities
- Develop, test, and maintain high-quality software solutions, frameworks and automations.
- Collaborate with cross-functional teams to analyse requirements and design solutions around stability and reliability.
- Participate in code reviews to ensure code quality and shared knowledge.
- Identify, troubleshoot, and resolve various incidents, problems ensure DevOps/SRE best practices.
- Contribute to continuous improvement initiatives within the engineering team.
- Assist in mentoring junior engineers by providing guidance and feedback.
Requirements
- 3+ years of experience in Site Reliability Engineering, maintenance & operations and/or development.
- Strong working experience eCommerce.
- Hands-on experience with Firebase (Crashlytics, Performance Monitoring, Remote Config, inApp Messaging).
- Experience operating mobile applications in production: crash and ANR triage, stability rate ownership, release management through App Store Connect and Google Play Console.
- Working knowledge of the Apple and Google release model:
review guidelines and approval flow, phased rollouts, app signing, certificate and APNs key management, annual OS releases and target API level deadlines.
- Experience supporting multiple app versions live in production simultaneously, with backward-compatible API and a minimum supported version strategy.
- Experience mitigating production issues without shipping a new build: remote config, feature flags, kill switches, server-side fixes.
- Experience defining SLIs / SLOs for mobile (crash-free users and sessions, ANR rate, app start time, client-side API error and latency rates) and analysing them by app version, OS version, device model and market.
- Comfortable using store reviews and ratings as an operational signal alongside technical telemetry
- Experience within solutions architecture and how to rapid pinpoint causes of issues
Key Skills
- Site Reliability Engineering
- Firebase (Crashlytics, Performance Monitoring, Remote Config, inApp Messaging)
- App Store Connect
- Google Play Console
- Release management
- SLI/SLOs for mobile
- Solutions architecture
- Incident identification and troubleshooting
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Site Reliability Engineer (Bengaluru)
🏢 H&M
📍 Bengaluru