08 Aug
|
Radiant Digital
|
Secunderabad
08 Aug
Radiant Digital
Secunderabad
Key Responsibilities
- Provide 24x7 support for incidents, alerts, and operational issues with a strong bias towards independent identification and resolution.
- Triage and diagnose incidents using logs, monitoring dashboards, and platform knowledge; resolve issues directly without defaulting to escalation.
- Engage Tier 2 only when incidents involve architectural complexity, infrastructure-level failures, or changes beyond Tier 1 resolution authority.
- Support client-submitted iTrack incident tickets and maintain end-to-end ticket ownership including resolution and closure.
- Respond to PagerDuty and automated alerts; validate, investigate, and remediate before escalating.
- Monitor production and non-production environment health; proactively identify anomalies and take corrective action.
- Monitor application support mailboxes and manage operational follow-through.
- Provide C2W emergency support as the first responder; independently handle and resolve wherever possible.
- Share client profile data and usage reports on request.
- Maintain clear incident communication to clients even when Tier 2 is engaged.
- Track and report operational metrics including MTTR and ticket resolution trends.
- Develop and maintain Tier 1 SOPs and operational runbooks based on real resolution patterns.
Required Qualifications / Must-Have Skills
- 3+ years of experience in application production support with a strong track record of independently diagnosing and resolving incidents.
- Solid working knowledge of the full technology stack in scope including event streaming platforms, integration middleware, AKS-hosted microservices,
and observability tooling.
- Hands-on experience with incident lifecycle management in ticketing systems (iTrack or equivalent), including root cause identification and resolution documentation.
- Proficient with Splunk, PagerDuty, Prometheus, and Grafana for active troubleshooting and issue resolution, not just monitoring.
- Hands-on operational experience with Kubernetes, especially Azure Kubernetes Service (AKS), including pod-level diagnostics, restarts, and health investigation.
- Practical working knowledge of Confluent Kafka and Azure Event Hub: consumer lag analysis, topic health checks, and message flow troubleshooting.
- Solid SQL/Postgres skills for data-level investigation and validation during incidents.
- Working ability to read and interpret Java, Spring Boot, and React application logs for issue identification.
- Basic Python scripting capability for operational checks and quick-fix automation.
- Good Linux/Unix command-line skills for real-time log analysis and system diagnostics.
- Robust written and verbal communication skills for incident updates, resolution documentation, and client coordination.
- Willingness to work in rotational 24x7 shifts.
Good-to-Have / Nice-to-Have
- Awareness of hybrid streaming ecosystems including Confluent Cloud, AWS-MSK, and Apache Flink.
- Exposure to IBM Sterling Integrator integration flows for context during incident triage.
- Telecom or high-availability enterprise support experience.
- Experience with CI/CD-driven deployment pipelines in a support context.
Experience Level
Associate to Mid-Level (typically 3 to 5 years)
📌 Production Support Analyst (Secunderabad)
🏢 Radiant Digital
📍 Secunderabad