16 Sep
|
Radiant Digital
|
Hyderabad
16 Sep
Radiant Digital
Hyderabad
Provide 24x7 support for incidents, alerts, and operational issues with a strong bias towards independent identification and resolution.
Triage and diagnose incidents using logs, monitoring dashboards, and platform knowledge; resolve issues directly without defaulting to escalation.
Engage
Tier 2 only when incidents involve architectural complexity, infrastructure-level failures, or changes beyond Tier 1 resolution authority. Support client-submitted i Track incident tickets and maintain end-to-end ticket ownership including resolution and closure. Respond to Pager Duty and automated alerts; validate, investigate, and remediate before escalating. Monitor production and non-production setting health; proactively identify anomalies and take corrective action. Monitor application support mailboxes and manage operational follow-through. Provide C2 W emergency support as the first responder; independently handle and resolve wherever possible. Share client profile data and usage reports on request. Maintain clear incident communication to clients even when Tier 2 is engaged. Track and report operational metrics including MTTR and ticket resolution trends. Develop and maintain Tier 1 SOPs and operational runbooks based on real resolution patterns.
Required Qualifications / Must-Have Skills
3+ years of experience in application production support with a strong track record of independently diagnosing and resolving incidents. Solid working knowledge of the full technology stack in scope including event streaming platforms, integration middleware, AKS-hosted microservices, and observability tooling.
Hands-on experience with incident lifecycle management in ticketing systems (i Track or equivalent), including root cause identification and resolution documentation. Proficient with Splunk, Pager Duty, Prometheus, and Grafana for active troubleshooting and issue resolution, not just monitoring. Hands-on operational experience with Kubernetes, especially Azure Kubernetes Service (AKS), including pod-level diagnostics, restarts, and health investigation. Practical working knowledge of Confluent Kafka and Azure Event Hub: consumer lag analysis, topic health checks, and message flow troubleshooting. Solid SQL/Postgres skills for data-level investigation and validation during incidents. Working ability to read and interpret Java, Spring Boot, and React application logs for issue identification.
Basic
Python scripting capability for operational checks and quick-fix automation. Good Linux/Unix command-line skills for real-time log analysis and system diagnostics. Strong written and verbal communication skills for incident updates, resolution documentation, and client coordination. Willingness to work in rotational 24x7 shifts. Good-to-Have / Nice-to-Have
Awareness of hybrid streaming ecosystems including Confluent Cloud, AWS-MSK, and Apache Flink. Exposure to IBM Sterling Integrator integration flows for context during incident triage. Telecom or high-availability enterprise support experience. Experience with CI/CD-driven deployment pipelines in a support context. Experience Level
Associate to Mid-Level (typically 3 to 5 years)
📌 Kafka Adminstrator (Hyderabad)
🏢 Radiant Digital
📍 Hyderabad