- Own responsibility for large scale/scope services/components that comprise our production infrastructure and systems
- Own responsibility of production support as well along with production infrastructure and systems
- Invent better ways to manage/automate the administration of new and existing systems across various services at Treebo via scripting and tools development.
- Consider, benchmark, propose, and implement new methods for monitoring and management of production servers. Collaborate with cross-functional organizations (engineering, QA, site operations, and security) on new product/feature design and/or diagnosis of problems with production systems.
- Perform expert-level debugging / troubleshooting / problem diagnosis and resolution.
- Monitor system health and performance.
- Engage in proactive communication and reporting.
- Should be ready to support 24x7 on a rotation basis,
active participation in war rooms on an as-needed basis and ability to drive root cause analysis (RCA) of failures.
What is expected out of you/What do we require from you:
- Educational Qualification: BE / B Tech /Mtech/MCA
- Build and maintain software modules for use and re-use in cloud systems automation
- Good programming (Python preferred) and automation skills
- String troubleshooting and system engineering exposure in UNIX/Linux environments. Positive experience with Linux
- Sound knowledge of cloud technologies such as AWS
- Good command on SQL
- Experience with Monitoring tools
- Performance evaluation and tuning of systems/application