04 Aug
|
Adidas
|
Gurugram
Site Reliability Engineer (SRE)
Purpose & Overall Relevance for the Organization:
- Maintain and enhance monitoring framework (data collection, alert aggregation, dashboarding) and implement and enhance alerting logic (framework).
- Enable proactive incident alert and resolution leveraging knowledge scripts.
- Identify and detect repetitive incidents (stability, reliability) and develop solutions to fix problems.
- Work on technical resolution for incidents and identify technical root cause.
- Ensure tool standards, exploit tool capability to fine-tune product reliability.
- Integrate incident, release, monitoring, and alerting tools into the overall ecosystem.
- Measure and report SLI, MTTx in periodic reviews, analyze deviations, and take actions to closure.
- Update runbooks with changes to process/tools.
- Drive postmortems to arrive at remedial actions.
- Participate in On-Call Incident Technical Support.
- Ensure production release guidelines (entry/exit) and implementation are adhered to for changes to Production.
- Support CI/CD pipeline implementation and integration to quality, security.
- Scale systems sustainably through mechanisms like automation; evolve systems by pushing for changes that improve reliability and velocity.
Key Responsibilities:
- Maintain and enhance monitoring framework (data collection, alert aggregation, dashboarding).
- Implement and enhance alerting logic (framework).
- Enable proactive incident alert and resolution leveraging knowledge scripts.
- Identify and detect repetitive incidents (stability, reliability) and develop solutions to fix problems.
- Work on technical resolution for incidents and identify technical root cause.
- Ensure tool standards, exploit tool capability to fine-tune product reliability.
- Integrate incident, release, monitoring, alerting tools into the overall ecosystem.
- Measure and report SLI, MTTx in periodic reviews, analyze deviations,
and take actions to closure.
- Update runbooks with changes to process/tools.
- Drive postmortems to arrive at remedial actions.
- Participate in On-Call Incident Technical Support.
- Ensure production release guidelines (entry/exit) and implementation are adhered to for changes to Production.
- Support CI/CD pipeline implementation and integration to quality, security.
- Scale systems sustainably through mechanisms like automation; evolve systems by pushing for changes that improve reliability and velocity.
Required Skills:
- At least 5 years overall IT experience with 3 years in relevant area (DevOps / SRE).
- Robust awareness and experience of working with Site Reliability Engineering principles.
- Good understanding of public cloud offerings such as AWS components like EC2, IAM, RDS, Cloudwatch, Database (Redis, RDS, Dynamo DB).
- Hands-on experience on enterprise toolsets such as Grafana, Instana, Prometheus, ELK Stack, etc.
- Exposure to networking concepts (SSH, FTP, TCP/IP, DNS, Load balancing, CDN, etc.).
- Experience in any scripting language (bash / python / perl).
- Good experience with CI/CD pipelines including BitBucket, Jenkins.
- Experience operating high-availability, fault-tolerant, scalable, distributed software in production: building monitoring into your code, tweaking dashboards, defining alerts.
- Knowledge of Agile software development principles, including using JIRA.
- Experience in a 24/7 high availability production environment.
- Excellent organizational, verbal, and written communication skills.
- Aptitude to be a good team player and the desire to learn and implement new technologies.
- Knowledge of ITIL processes.
Nice to Have:
- Experience with building Rest APIs, API Integration, and Web Services is preferred.
- Knowledge in Messaging and Streaming frameworks like RabbitMQ / Kafka.
- Exposure to languages such as Typescript, Nodejs.
- ITIL V4 Foundation certified.
📌 Software Engineer (Gurugram)
🏢 Adidas
📍 Gurugram