As an SRE Engineer, you will be responsible for ensuring the seamless operation and optimal performance of large-scale distributed software applications. Your role revolves around maintaining a robust and high-performing environment, contributing to the reliability of our services, and innovating solutions to guarantee 24/7 availability. You will leverage your technical expertise to maintain a seamless experience for our users while upholding the highest standards of operational excellence.
Responsibilities:
Monitoring and Alerting:
- Review existing monitoring tools and set up new systems to track system performance and key metrics.
Incident Management:
- Monitor alerts and logs to promptly identify incidents or anomalies.
- Prioritize incidents based on severity and potential impact on stability and reliability.
- Engage in incident resolution, applying necessary fixes and mitigations to restore normal operations.
On-Call Responsibilities:
- Organize on-call schedules to ensure 24/7 coverage for incident response.
- Respond to alerts, troubleshoot issues, and coordinate with NOC and Engineering teams for incident resolution.
- Conduct post-incident reviews to identify root causes, learn from incidents, and implement preventive measures.
Automation and Tooling:
- Review and build new automation scripts and tools to streamline tasks, enhance efficiency, and reduce manual errors.
- Regularly update and maintain monitoring, deployment, and incident management tools.
Performance Optimization:
- Analyze application performance using profiling and monitoring tools to identify bottlenecks and areas for improvement.
- Work on optimizations, infrastructure upgrades, and architectural improvements to enhance system performance and efficiency.
Capacity Planning and Scaling:
- Monitor resource utilization and trends to predict capacity needs and plan for scaling.
- Scale resources (servers and databases)
based on usage patterns and anticipated growth.
- Automate the sizing process to ensure efficiency.
Disaster Recovery and Redundancy:
- Develop and maintain disaster recovery plans to ensure business continuity.
- Implement redundancy and failover strategies to minimize downtime and maintain service availability during failures.
Knowledge Sharing and Documentation:
- Create and maintain comprehensive documentation for configurations, procedures, incidents, and best practices.
- Foster a culture of knowledge sharing within the team through regular sessions and training programs.
Feedback Loop and Continuous Improvement:
- Collect feedback from incidents, post-mortems, and NOC/Dev team interactions to identify areas for improvement.
- Continuously iterate on processes, tools, and systems based on feedback to drive continuous improvement.
Collaboration and Communication:
- Collaborate closely with Engineering and DC/NOC teams to align goals and priorities.
- Ensure open communication within the team and with stakeholders, providing regular updates on incidents, progress, and initiatives.
Requirements:
- Bachelor's degree in Computer Science or related disciplines.
- 3+ years of experience in software application/product support.
- Proficiency in programming using Go, Shell, or Python scripting languages.
- Experience in technical engineering (preferred).
- A proactive approach to identifying problems, performance bottlenecks, and areas for improvement.
- Solid knowledge of Networking, Database (MySQL), Linux System concepts, and experience in debugging and analyzing core dumps.
- Hands-on experience with monitoring and observability tools like Grafana, Nagios, Influx, ELK, etc.
- Familiarity with orchestration tools like Docker, Grafana, and incident management systems like Zenduty.
- Excellent communication and collaboration skills with the ability to work effectively across teams.
- Self-motivated with a positive mindset to examine and solve incidents.
📌 Site Reliability Engineer (Pune)
🏢 PubMatic
📍 Pune