- Responsibility for delivering on identifying, creating, and maintaining SLA's.
- Design, build and support automation and monitoring that improve system reliability
- Partner with R&D; and Operations teams to enhance telemetry and reliability
- Support services before they go live through system design consulting and launch reviews
- Maintain live services by measuring & monitoring availability, latency & overall system health
- Create monitoring to detect symptoms and pre-empt outages
- Debug production issues across services and levels of the stack
- Report problems and participate in related root cause analysis or incident Post Mortems
- Provide analysis of poor performance and instabilities identified in systems.
Experience & Skills :
- Container orchestration technologies like Docker
- Virtualization platforms, either on-prem or cloud-based (We use AWS)
- Understands Infrastructure as a code (Puppet, Ansible and Terraform) and containerization toolsets (Docker).
- Data-intensive applications and platforms like Kafka, Hadoop, Spark, Zookeeper, Cassandra, PostgreSQL OLAP, Druid
- Relational databases like MySQL, Oracle, PostgreSQL etc
- NoSQL databases like Redis, MongoDB, Cassandra, CouchDB etc
- One or more CI tools like Jenkins, Teamcity
- Centralized logging systems, metrics, and tooling frameworks such as ELK,
Prometheus, and Grafana.
- Web and Application servers like Apache, Nginx, Tomcat
- Versioning tools such as git.
- Ability to work independently and own problem statements end-to-end. - Outstanding communication, interpersonal and teamwork skills. - Adaptable to work in a fast-paced environment and alter priorities as per business needs
Qualifications :
- B.Tech/M.Tech or Equivalent in Computer Science, Information Technology or a related field
- 4-7 years of experience in handling services in a large-scale distributed system.
- Deep understanding of network stack (e.g. TCP/IP, routing, network topologies and hardware, SDN, etc)
- Deep understanding of modern software architectures, including load-balancing queueing, caching, distributed systems failure modes generally, microservices and big data technologies.
- Excellent programming (Asp.net, Node.js, Python, Go, Ruby or preferred scripting languages) and automation skills -
- AWS certification will give you an added advantage
100%Resize 100%
75%Resize 75%
50%Resize 50%
Auto size
Rotate left
Rotate right
Mirror, Horizontal
Mirror, Vertical
Align
- Basic
- Left
- Center
- Right
Insert description
Revert
Edit
Remove
📌 Site Reliability Engineer (India)
🏢 DronaHQ
📍 India