06 Sep
|
TalentOla
|
Chennai
- Monitoring and Automation: Proactively monitor software systems to prevent incidents and automate routine tasks.
2. Effective Monitoring: Build monitoring systems that alert based on symptoms rather than outages.
3. Application Performance Monitoring (APM): Implement and utilize APM tools such as Recent Relic or Dynatrace to monitor application performance, identify bottlenecks, and optimize resource usage.
4. Log Analysis with Splunk: Analyze logs using Splunk to troubleshoot issues, detect anomalies, and improve system reliability.
5. Dashboards Preparation: Create informative dashboards to visualize system health, performance, and key metrics.
6. Alerts Setup: Configure alerts based on thresholds and anomalies to promptly address issues.
7. Reports Scheduling: Set up regular reports to provide insights into system performance and reliability.
8. Reliability Metrics: Establish and track reliability metrics (e.g., SLOs, SLIs, error budgets) to measure system performance.
9. Observability Skills:
Proficiency in observability practices, including distributed tracing, logging, and metrics collection.
10. Collaboration: Partner with development, support teams to improve services through rigorous testing and release procedures.
11. Capacity Planning: Participate in system design consulting and capacity planning.
12. Debugging and Incident Response: Understand debugging information, handle incidents, and roll back faulty software pushes.
13. Mentoring L1/L2 Support Teams: Provide guidance, mentorship to establish best practices on monitoring and observability.
14. Infrastructure Management: Run and manage infrastructure using tools like Chef, Ansible, Terraform, GitLab CI/CD, and Kubernetes.
15. Documentation: Document processes and procedures to avoid redundancy.
16. Enthusiastic Attitude: Approach challenges with enthusiasm and a proactive mindset.
📌 Site Reliability Engineer (Chennai)
🏢 TalentOla
📍 Chennai