1. Monitor platform health, performance, utilization, and customer KPIs using tools like Dynatrace, SolarWinds, Grafana.
2. Ensure end-to-end incident management, RCA documentation, and service request fulfillment.
3. Implement and configure smart alerts and monitoring tools across application and infrastructure layers.
4. Drive continuous process improvements, build SOPs, and maintain adherence to SLAs.
Role Responsibilities:
1. Serve 24x7 with the team for proactive problem resolution and escalation handling.
2. Triage and log incidents, initiate bridge calls, and notify relevant stakeholders for major issues.
3. Publish platform health dashboards and maintain incident lifecycle documentation.
4. Design and deploy recent monitoring solutions and write automation scripts using Python, Shell, or PowerShell.
📌 Lead Engineer (Chennai)
🏢 Tanla Platforms
📍 Chennai
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.