Production and Support Engineer ( Java, Node.js) (Hyderabad)

Production and Support Engineer ( Java, Node.js) (Hyderabad)

09 Sep
|
Solugenix
|
Hyderabad

09 Sep

Solugenix

Hyderabad

Solugenix is a leader in IT services, delivering cutting-edge technology solutions, exceptional talent, and managed services to global enterprises. With extensive experse in highly regulated and complex industries, we are a trusted partner for integrating advanced technologies with streamlined processes. Our solutions drive growth, foster innovation, and ensure complianceproviding clients with reliability and a strong competitive edge.

Recognized as a 2024 Top Workplace, Solugenix is proud of its inclusive culture and unwavering commitment to excellence. Our recent expansion, with new offices in the Dominican Republic, Jakarta, and the Philippines, underscores our growing global presence and ability to offer world-class technology solutions. Partnering with Solugenix means more than just business—having a dedicated our financial client focused on your success in today's fast-evolving digital world

Job Title: Tier 3 Support Engineer ( Java)

Job Location: Hyderabad or Indore

Experiences: 5+ Years

Job Type: Full time

Shift Timings: 11:30 am – 8:30 pm

:

We are seeking an experienced Tier 3 Support Engineer with 5+ years of experience in supporting mission-critical applications within the Insurance and Banking domain. The ideal candidate will be responsible for advanced incident management, root cause analysis, production stability, problem management, and continuous service improvement.

The candidate is expected to leverage Microsoft Copilot, GenAI, Datadog Watchdog AI, AI-driven incident management, and runbook automation to improve operational efficiency, accelerate incident resolution, and enhance application reliability.

Key Responsibilities

- Provide Tier 3 support for business-critical banking and insurance applications.
- Investigate and resolve complex production incidents, application issues, and performance bottlenecks.
- Perform detailed Root Cause Analysis (RCA) and drive permanent corrective actions.
- Collaborate with development, infrastructure, DevOps, and business teams to resolve high-priority incidents.
- Monitor application health, system performance, and service availability.




- Lead problem management activities and identify recurring issues for proactive remediation.
- Support application deployments, production releases, and environment management.
- Maintain operational documentation, support procedures, and knowledge repositories.
- Participate in on-call rotations and major incident management processes.
- Drive automation initiatives to reduce manual operational effort and improve service reliability.

AI-Powered Incident Management

- Utilize Microsoft Copilot to generate comprehensive RCA reports, post-incident summaries, and executive communications.
- Leverage Datadog Watchdog AI for anomaly detection, root cause insights, performance monitoring, and proactive issue identification.
- Implement AI-driven incident routing to automatically assign incidents to the appropriate support teams.
- Use AI-assisted knowledge management for faster troubleshooting and resolution recommendations.
- Develop and maintain automated operational runbooks to reduce Mean Time to Resolution (MTTR).
- Utilize GenAI tools to create incident summaries, trend analysis reports, remediation recommendations, and support documentation.
- Apply predictive analytics and AI models to identify potential service disruptions before customer impact.

Required Technical Skills: Production Support & Incident Management

- Strong experience in:
- Incident Management
- Problem Managemen
- Change Management
- Release Support
- Root Cause Analysis (RCA)
- Service Availability Management
- Experience supporting customer-facing enterprise applications.

Monitoring & Observability

- Hands-on experience with:
- Datadog
- Splunk
- Dynatrace
- AppDynamics
- Azure Monitor
- Grafana
- Expertise in log analysis, alerting, monitoring dashboards, and performance troubleshooting.

Cloud & Infrastructure





- Experience supporting applications hosted on:
- Microsoft Azure
- AWS
- GCP

Understanding of:
- Linux/Unix Administration
- Containers (Docker)
- Kubernetes
- Microservices Architecture

Database & Middleware

- Experience with:
- SQL Server
- Oracle
- PostgreSQL
- MongoDB
- Ability to troubleshoot database performance and application integration issues.

Automation & Scripting

- Experience with:
- PowerShell
- Python
- Shell Scripting
- Knowledge of workflow automation and operational tooling.

Domain Experience Experience in Banking, Financial Services, Insurance (BFSI) environments.

- Working knowledge of:
- Policy Administration Systems
- Claims Processing
- Underwriting
- Lending Platforms
- Payments Processing
- Customer Onboarding
- Regulatory and Compliance Requirements

Preferred Qualifications

- ITIL Foundation or ITSM certifications.
- Experience with Site Reliability Engineering (SRE) practices.
- Knowledge of AIOps platforms and intelligent monitoring solutions.
- Experience supporting cloud-native and distributed applications.
- Exposure to ServiceNow Incident and Problem Management workflows.

Soft Skills

- Strong analytical and troubleshooting skills.
- Excellent communication and stakeholder management capabilities.
- Ability to work under pressure during critical incidents.
- Robust documentation and reporting skills.
- Customer-focused mindset with a commitment to operational excellence.
- Ability to mentor junior support engineers and lead incident response activities.

Key Success Metrics

- Reduction in Mean Time to Resolution (MTTR).
- Improved service availability and application stability.
- Increased incident automation through AI and runbooks.
- Faster root cause identification using Datadog Watchdog AI and Copilot.
- Improved first-time resolution rates and operational efficiency.
- Reduced recurring incidents through proactive problem management.

Desired Skills:

- Incident Managemen
- Problem Management
- Change Management
- Release Support
- Root Cause Analysis (RCA)
- Service Availability Management

Education:

- BE/ BTech / BSC / MCA

📌 Production and Support Engineer ( Java, Node.js) (Hyderabad)
🏢 Solugenix
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: production and support engineer ( java, node.js) (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: production and support engineer ( java, node.js) (hyderabad) / hyderabad