06 Aug
|
Bajaj Life Insurance
|
India
06 Aug
Bajaj Life Insurance
India
Role : Bajaj Life Insurance - Chief Manager - Application Management
Experience : 8 to 12 years
Role Title : Chief Manager Application Management (Service Reliability & Observability)
Reports To : Vice President Service Reliability & Observability
Company : BAJAJ LIFE Insurance
Function/ Department : Technology
JOB PURPOSE :
This techno functional role is responsible for ensuring the availability, stability, and performance of business-critical insurance applications (policy administration, claims, payouts, renewals, and policy servicing platforms) through effective production support, DevOps practices, and team leadership. The role owns end-to-end incident management, release and change management, and continuous improvement of application reliability, while building and mentoring a team of junior production support engineers to deliver consistent, SLA-compliant service to business and policyholders.
Any Life Insurance
Background would be helpful.
PRINCIPAL ACCOUNTABILITIES :
Application Availability & Reliability Management :
- Own end-to-end availability of policy administration, claims, and servicing applications against defined SLA/OLA targets (e.g. 99.8%+ uptime).
- Implement and continuously improve monitoring, alerting, and observability (APM, logs, synthetic checks) to detect issues proactively before business impact.
- Analyze recurring incidents to identify and remediate root causes, driving down repeat failures and reducing Mean Time to Detect (MTTD) and Mean Time to Restore (MTTR).
- Drive capacity planning and performance tuning (application, database, infrastructure) to prevent availability degradation during peak policy-servicing cycles (e.g. renewal season, month-end/quarter-end).
- Own and periodically test disaster recovery (DR) and business continuity procedures for critical applications, ensuring rapid restoration with minimal data loss.
Production & Product Support Management :
- Ensure timely triage, prioritization, and resolution of production incidents and service requests as per the P1-P4 severity framework and agreed SLA/OLA timelines.
- Act as escalation point for critical (P1/P2) incidents, coordinating war-room bridges, vendor/IT teams, and communicating status to stakeholders until closure.
- Institutionalize a Root Cause Analysis (RCA) process for all major incidents, ensuring corrective and preventive actions (CAPA) are tracked to closure.
- Oversee ticket/queue management (via ITSM tools such as JIRA/ServiceNow) to ensure ageing, pendency, and SLA breaches are actively managed and reported.
- Manage patching, version upgrades, and security vulnerability remediation across application environments without disrupting business operations.
Policy Servicing Functional & Technical Issue Resolution :
- Ensure functional and technical issues affecting policy servicing journeys New Business, Renewals, Endorsements, Claims, Payouts, Free-Look Cancellation, Policy Loan, Grievance, and Revival are resolved accurately and within SLA.
- Coordinate with business teams (Operations, Claims, Customer Service) to understand servicing-related pain points and translate them into technical fixes or enhancements.
- Validate that fixes/releases affecting policy servicing modules are tested (functional + regression)
before go-live to avoid recurrence of customer-impacting defects.
- Track and report on policy-servicing-related technical issues (SDC%, TAT%, NOP, pendency) to demonstrate service quality improvements over time.
Change, Release & DevOps Process Management :
- Own the change management and release process for production deployments, ensuring controlled, low-risk releases with rollback plans.
- Champion DevOps and CI/CD practices (build/deploy automation, infrastructure-as-code, environment consistency) to improve release velocity and quality.
- Verify completion of development, testing, and approvals prior to production release; ensure post-deployment monitoring and sign-off.
- Assess and manage the impact of proposed changes on live applications, maintaining a change control log and risk register.
- Manage API integrations and data flows between core applications (policy admin, payment gateway, eKYC, IGMS, etc.) to ensure end-to-end process integrity.
Stakeholder & Vendor Management :
- Establish and run governance cadences (daily stand-ups, weekly/monthly reviews) with business and technical stakeholders on incidents, changes, and service health.
- Prepare and present Weekly/Monthly Senior Management Communication covering availability, incident trends, SLA compliance, and improvement actions.
- Liaise with application vendors/system integrators (e.g. core policy administration platform partners) on defect fixes, patches, and enhancement delivery.
- Coordinate with Business Analysts, Project Managers, Infosec, and Compliance teams to ensure production changes meet regulatory (IRDAI) and internal audit requirements.
- Maintain transparent, timely communication with stakeholders during incident resolution to avoid information gaps.
Team Leadership & Capability Building :
- Lead, mentor, and manage day-to-day performance of a team of junior production support engineers, including shift/roster planning for 24x7 coverage where applicable.
- Define individual and team KPIs aligned to availability, TAT, and quality goals; conduct regular performance reviews and feedback.
- Build team capability through structured upskilling on ITIL practices, monitoring/observability tools, cloud platforms, scripting/automation, and insurance domain knowledge.
- Ensure adequate documentation, knowledge transfer, and runbooks exist so that support quality does not depend on individual tribal knowledge.
- Foster a culture of ownership, proactive problem-solving, and continuous improvement within the support team.
MAJOR CHALLENGES :
Balancing the need for robust, consistent governance rhythms (SLA/OLA reviews, RCA closures, senior management reporting) against the team's operational bandwidth, particularly during high-load periods such as month-end/quarter-end policy servicing cycles. Correlating user experience issues with underlying system logs, alerts, and infrastructure signals to derive a precise,
prioritized action plan for application stability improvements, while managing a lean team and multiple concurrent production issues.
DECISIONS :
- Prioritization and severity classification (P1-P4) of incoming production and support issues.
- Go/no-go decisions on production releases and emergency changes, within defined change-control authority.
- Escalation timing and routing for major incidents to senior stakeholders and vendors.
- Documenting, tracking, and prioritizing RCA-driven technical improvement actions and change requests.
- Allocation of team workload/shift coverage across incidents, changes, and BAU support activities.
INTERACTIONS :
Internal Clients :
- IT Leadership, Application Development Teams, Infrastructure/Cloud Teams, Business Operations (New Business, Claims, Servicing), Business Analysts, Project Managers, Information Security, Compliance/Audit, Customer Service.
External Clients :
- Application/Platform Vendors and System Integrators, Cloud Service Providers, Third-Party API Partners (Payment Gateway, eKYC, IGMS), External Auditors/Regulatory Bodies as required.
DIMENSIONS :
Financial Dimensions (FY 24) : NA.
Other Dimensions (FY 24) : Total Team Size : 2/3, Number of Direct Reports : 2.
SKILLS AND KNOWLEDGE :
Educational Qualifications :
- Bachelor's degree in computer science, information technology, or a related field (or equivalent work experience).
- Proven experience in log analysis, DB performance fine-tuning measures, application support and management roles.
Work Experience :
- 8+ years of overall IT experience, including at least 3-4 years in a production support, application management, or DevOps leadership role.
- Proven experience managing a team, ideally within BFSI/insurance or another regulated industry.
- Track record of improving application availability, reducing MTTR, and driving incident/problem management maturity.
- Experience with core insurance systems (policy administration platforms, claims systems) preferred.
Technical Skills :
- Strong knowledge of ITIL processes : Incident, Problem, Change, and Release Management.
- Hands-on exposure to monitoring/observability tools (APM, log management, dashboards) and ITSM platforms (JIRA, ServiceNow, or equivalent).
- Working knowledge of CI/CD pipelines, scripting/automation, containerization (Kubernetes/Docker), and cloud platforms (GCP/AWS/Azure).
- Familiarity with database performance tuning (Oracle/PostgreSQL) and application log analysis. Understanding of API-based integrations and microservices architectures.
Domain & Regulatory Knowledge :
- Understanding of Indian life insurance business processes New Business, Renewals, Claims, Payouts, Endorsements, Free-Look, Revival, Grievance.
- Awareness of IRDAI regulatory requirements and data protection norms (DPDP Act 2023) as applicable to production systems and customer data.
Behavioral Competencies :
- Strong analytical and problem-solving skills with a structured, root-cause-driven mindset.
- Excellent stakeholder communication skills, including experience presenting to senior/CXO audiences.
- Ability to lead and develop a team under pressure, balancing governance rigor with operational bandwidth.
- High ownership and accountability for outcomes, with a continuous-improvement orientation.
📌 Bajaj Life Insurance - Chief Manager - Application Management/Service Reliability & Observability (India)
🏢 Bajaj Life Insurance
📍 India