Fault Management Engineer (Bengaluru)

Fault Management Engineer (Bengaluru)

02 Sep
|
HCLTech
|
Bengaluru

02 Sep

HCLTech

Bengaluru

IBM MAINFRAME L2 SUPPORT

Fault Analyzer & File Manager

Lead Software Engineer I • Level 3.2 • Senior / Technical Lead

Product Scope

IBM Fault Analyzer (FA) • IBM File Manager (FM) • IBM ADFz Common Components • IBM DDRIVEN Agent

Qualification

Bachelor's degree in Computer Science, IT, Engineering, or related field

Position Type

Full-Time

Role Category

Mainframe Systems Engineering, DevOps Tooling & Middleware Support Technical Lead

Role Summary The Lead Support Engineer – Fault Analyzer & File Manager is the primary subject matter expert (SME) and technical anchor for HCL Software's FA/FM support capability worldwide. Operating at the intersection of L2 case ownership and L3 collaboration, this role drives root cause analysis on the most complex customer escalations, defines data gathering standards, mentors junior engineers, and serves as the escalation path for incidents involving FA/FM integration with CICS, IMS, DB2, WebSphere MQ, Language Environment, and Eclipse-based development tooling. This is not a management role — it is the deepest technical individual contributor position within the FA/FM support track, requiring mastery of both products' diagnostic frameworks, configuration architecture, and integration patterns across the full z/OS ecosystem.

About HCL Software

HCL Software was launched as a new division of HCL Technologies in 2016 with the mission to develop and deliver a next-generation portfolio of enterprise-grade software-based offerings with flexible consumption models, spanning on-premises software, SaaS, and managed services. We bring speed, insights, and innovations to create value for our customers in DevOps, Automation, Data, Security, and Application Modernization.

Key Responsibilities

Technical Leadership & Case Escalation

- Own and drive resolution of the most complex FA/FM escalations — cases that have exceeded standard L2 resolution capacity or involve multi-product integration failures.
- Act as the final L2 escalation point before engagement of L3 development engineering, ensuring all diagnostic artefacts are complete, structured, and actionable.
- Perform deep-level root cause analysis across FA abend diagnostics, FM data corruption scenarios, batch FASTREXX failures, and DB2/IMS subsystem integration issues.
- Lead technical war-room sessions for Severity 1 customer incidents involving FA/FM, coordinating with L3, IBM support, and customer technical teams.
- Triage and validate reported defects before L3 submission — distinguish configuration issues, customer environment problems, and genuine product defects.
- Review and approve all L3 escalations raised by junior engineers on the FA/FM product stream.

Product Subject Matter Expertise

- Maintain authoritative knowledge of IBM Fault Analyzer configuration: IDFOPTxx members, site defaults, fault history file architecture (SYS-level and user-level), dump capture policies, and LE integration.
- Maintain authoritative knowledge of all four IBM File Manager tracks: Base 3270 ISPF (Labs 1–13), Batch/FASTREXX, DB2 subsystem interface, IMS hierarchical DB interface, and Eclipse graphical workbench (Labs 1–10).
- Maintain current awareness of FA/FM PTF streams, APARs, and IBM support bulletins — identify HIPERs and service-affecting defects before they surface as customer incidents.
- Serve as the team's expert on IBM ADFz Common Components (IPV Server) connectivity: RSE, FTP, DTCN, Eclipse tool connections (IDz, IBM Explorer for z/OS, Zowe Explorer, IBM Z Open Debug, VS Code).
- Maintain working knowledge of DDRIVEN (Distributed Agent for z/OS) support patterns, including the critical DDRIVEN vs. z-Centric agent triage distinction and HTTPOPTS/TWSOPTS configuration diagnostics.

Data Gathering & Diagnostic Standards

- Define and maintain the FA/FM data gathering procedures — the structured six-category diagnostic checklist used by the team for every case type (FA abend, FM ISPF, FM Batch, FM DB2, FM IMS, FM Eclipse).
- Develop and publish internal technotes and data-gather procedures to fill gaps where no official IBM documentation exists (e.g., DDRIVEN data gather procedure, SYSOUT DD missing from SEELSAMP(EELAGT)).
- Ensure all team members correctly apply the FA/FM agent-type identification checklist before committing to a diagnostic path — preventing misdirected investigations on z-Centric vs. DDRIVEN cases.
- Define SYSOUT DD verification as a mandatory pre-trace step in all FM Eclipse/DDRIVEN SDIAHTT (HTTPFLAGS) trace requests.

Mentoring & Knowledge Enablement

- Mentor Software Engineer I and Senior Software Engineer II colleagues through structured case shadowing, technical briefings, and self-learning training programme delivery.
- Develop and maintain the FA/FM self-learning HTML training modules — covering each track's lab sessions, workbook references, week-by-week study plans, and real-world troubleshooting scenarios.
- Deliver internal Kepner-Tregoe (KT) coaching sessions to reinforce IS/IS NOT problem specification methodology on complex FA abend and FM data management cases.




- Conduct post-incident technical debriefs and contribute lessons learned to the team knowledge base in KCS format.
- Review and approve knowledge base articles, internal runbooks, and technotes authored by junior engineers before publication.

Customer & Stakeholder Engagement

- Engage directly with enterprise customers during high-visibility escalations — including technical calls, screen-share diagnostic sessions, and written RCA delivery.
- Deliver technical presentations to customer technical teams on FA/FM diagnostic best practices, preventive configuration, and upgrade readiness.
- Represent HCL Software in joint support sessions with IBM when product defects require escalation through the IBM support chain.
- Participate in customer governance reviews, SLA performance reviews, and support roadmap discussions as the FA/FM technical voice.

On-Call & Shift Coverage

- Lead and coordinate the on-call rotation for FA/FM Severity 1 incidents — weekend and after-hours availability on a scheduled rotational basis.
- Function as the shift technical lead during critical incident periods, directing junior engineers' diagnostic efforts.

Required Technical Skills

IBM Fault Analyzer (FA) — Expert Level

- Complete diagnostic mastery of FA abend analysis: CICS abends, IMS abends, batch job abends, z/OS system abends, and LE condition handling.
- IDFOPTxx configuration: site defaults, fault history file sizing and allocation, IDILOOK parameters, exclusion rules, and interactive analysis options.
- FA integration with z/OS dump facilities: SVC dumps, SDUMP, SYSMDUMP, IEATDUMP — understanding which dump type FA captures and under what conditions.
- FA ISPF interface: full navigation of fault history reports, interactive reanalysis, source-level mapping, and Up-to-date Compiler Support.
- FA exit customisation: understanding of FA user exits for site-specific fault handling, notification, and data capture policies.
- FA and Language Environment (LE): CEE3DMP, CEEDUMP, CEETRACE interaction with FA capture and formatting.
- FA real-time analysis, batch reanalysis (IDILOOK), and Fault Analyzer API usage for programmatic fault retrieval.
- Source-level problem diagnostics: mapping optimised object code back to source statements using FA's compiler listing integration.

IBM File Manager (FM) — Expert Level

- Base 3270 ISPF Interface: all 13 lab exercises — panel navigation, copy routing (option 3.3), print variables, catalog exploration (panel 3.6), batch find/change, delta comparison (Lab 9), configuration switches (Lab 10), dataset allocation (Lab 12), validation scripts (Lab 13).
- FM Batch / FASTREXX: conditional routing (DDOUT/NORTHAM/OTHER), layout restructuring (SET_OLEN, FLD_OUT, OVLY_OUT), nested loop processing, PGM=IGYCRCTL scanning, and automated JCL member validation.
- FM DB2: subsystem attach, catalog space browsing, custom table column formatting, relational constraint management across joined layers, object drops (Labs 5–6), and migration validation.
- FM IMS: IAF1 subsystem connection, PSBLIB/DBDLIB mapping, hierarchical navigation (ROOT/CHILD/TWIN switches), COBOL copybook template mapping, shadow line masking, and batch JCL extract generation (Labs 1–6).
- FM Eclipse (Advanced): Explorer view configuration, dataset query filtering, split layout editing, template selection criteria, variable record redefinition maps, dataset allocation tools, and zFS directory synchronisation (Labs 1–10).
- FASTREXX scripting: authoring, debugging, and optimising FASTREXX procedures for batch data management automation.
- FM storage layer support: QSAM, VSAM (KSDS/ESDS/RRDS), IAM, OAM, PDS/PDSE, HFS/zFS, tape structures — hands-on diagnostic capability across all types.

IBM ADFz Common Components & Eclipse Tooling

- IPV Server administration: installation, configuration, SSL setup, and connectivity troubleshooting for RSE, FTP, and DTCN connections.
- Eclipse-based IDE connectivity: IDz, IBM Explorer for z/OS, Zowe Explorer, IBM Z Open Debug, VS Code — resolving RSE connection failures, debug profile issues, and data transfer errors.
- DTCN (Dynamic Transaction Control) and debug profile management for CICS debugging sessions.
- Cross-platform tooling: Windows, UNIX/Linux, z/OS USS — resolving connectivity issues across all operating environments.

DDRIVEN — Distributed Agent for z/OS

- DDRIVEN architecture mastery: understanding DDRIVEN as a Tracker connecting to MDM/DDM instead of a z/OS WS controller — the foundational distinction from z-Centric agents.
- DDRIVEN configuration: TWSOPTS (SSCMNAME, BUILDSSX, VARSUB, CODEPAGE), HTTPOPTS (TDWBHOSTNAME, TDWBPORTNUMBER, HOSTNAME, PORTNUMBER, CONNTIMEOUT, SSL,



SSLKEYRINGTYPE), EXITS (EELUX000/002/004), EWTROPTS.
- JES exit diagnostics: EXIT7 difference between DDRIVEN (TWSEXIT7/TWSENTR7) and z/OS WS (OPCAXIT7/OPCAENT7); EXIT51 identical for both (TWSXIT51/TWSENT51).
- Sysplex deployment rules: unique SSCMNAME per LPAR requirement — diagnosing subsystem name conflicts.
- DDRIVEN data gather procedure: six-category checklist (environment basics, JCL/config, log datasets, JES exit configuration, connectivity, dataset triggering).
- Critical triage: DDRIVEN vs. z-Centric agent identification — TWSOPTS HTTPOPTS presence in parmlib, EELMLOG vs. EQQMLOG log prefix, SEELLOAD vs. SEOPLOAD library check.
- Dataset triggering (DDRIVEN v9.4 ): IEFU83 SMF exit, EELJCLIB monitoring list, EELEVDS resource event diagnosis.
- Missing SYSOUT DD knowledge: SEELSAMP(EELAGT) diagnostic — required step before any SDIAHTT (HTTPFLAGS) trace request.

z/OS Platform & Diagnostic Tools

- z/OS system commands, operator console operations, IPL procedures, and system health monitoring.
- SDSF: job output analysis, ULOG, SYSLOG, held output management, and active address space monitoring.
- JES2/JES3: EXIT7 and EXIT51 parmlib definitions, internal reader management, EELBRDS dataset configuration.
- Dump analysis: IPCS (Interactive Problem Control System), SVC dump formatting, CEEDUMP/IEATDUMP interpretation.
- VSAM management: IDCAMS, DEFINE CLUSTER, REPRO, DELETE — relevant to FM storage layer diagnostics.
- USS (z/OS UNIX System Services): zFS file system management, permission issues, and path mapping for ADFz components.
- Language Environment (LE): CEE runtime options, condition handling, and debugging integration.

JCL, Languages & Scripting

- Advanced JCL: PROC libraries, symbolic parameters, generation data groups (GDGs), conditional execution (IF/THEN/ELSE, COND=), and overrides.
- COBOL: reading and tracing program logic to source-level mapping level — sufficient for FA abend analysis and FM COBOL copybook template creation.
- REXX and FASTREXX: authoring diagnostic utilities, batch automation scripts, and FM FASTREXX procedures.
- SQL/DB2: querying catalog tables, explaining execution plans, and diagnosing FM DB2 interface issues.
- CLIST: reading and modifying site-specific ISPF CLIST procedures that interact with FA/FM.

Support Operations

- ServiceNow: advanced case management, escalation workflows, SLA monitoring, and reporting.
- ITIL: incident, problem, change, and knowledge management — with particular depth in problem management and known error database (KEDB) maintenance.
- KCS (Knowledge-Centred Service): article creation, quality review, and knowledge loop closure.
- Kepner-Tregoe (KT) methodology: IS/IS NOT problem specification, root cause identification, and decision analysis — applied to complex FA/FM diagnostic scenarios.
- RCA documentation: structured root cause analysis reports suitable for executive customer stakeholders.

Preferred Qualifications

- 5 years of hands-on technical support or engineering experience specifically with IBM Fault Analyzer and IBM File Manager in enterprise customer environments.
- Demonstrated experience owning Severity 1 customer escalations end-to-end, including IBM PMR/case engagement and on-site customer calls.
- Completed Kepner-Tregoe (KT) Problem Solving & Decision Making training.
- IBM Certified System Administrator — z/OS or equivalent IBM mainframe certification.
- Experience developing and delivering internal technical training programmes — including self-learning HTML modules, PPTX decks, or structured lab exercises.
- Published knowledge base articles or technotes for IBM Fault Analyzer, IBM File Manager, or related IBM ADFz products.
- Experience with DDRIVEN (IWS Distributed Agent for z/OS) configuration and support — particularly HTTPOPTS/TWSOPTS parameters and JES exit diagnostics.
- Familiarity with HCL Workload Automation (HWA) product family context — understanding DDRIVEN's positioning within the IWS/HWA scheduling ecosystem.

What Success Looks Like in This Role

- Zero cases requiring L3 escalation without a complete, structured data package — all escalations arrive at L3 with full diagnostic artefacts attached.
- Team MTTR (Mean Time to Resolution) for FA/FM cases improves quarter-over-quarter through data gather standardisation and knowledge base growth.
- Junior engineers on the team can independently triage FA abend cases and FM ISPF/Batch/DB2/IMS cases within 6 months of onboarding — measurable outcome of mentoring effectiveness.
- DDRIVEN vs. z-Centric misidentification rate reaches zero — all cases correctly routed within the first 5 minutes of case review.
- At least one published technote or KB article per quarter that fills a documented gap in the public FA/FM support documentation.
- Customer satisfaction scores for FA/FM escalations consistently at or above team average — reflecting both technical depth and communication quality.

HCL Software • Mainframe L2 Support • Fault Analyzer & File Manager • Lead Software Engineer I (3.2) • Confidential — Internal Use

📌 Fault Management Engineer (Bengaluru)
🏢 HCLTech
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: fault management engineer (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: fault management engineer (bengaluru) / bengaluru