27 Sep
|
Airbus
|
Bengaluru
Job DescriptionQualification & Experience:
N8+ years of working experience in IT working on Linux environments (RHEL) with development
Nexperience. ● Solid Linux/RedHat system administration skills for Installing, managing Scientific
NComputing applications and Performance Analysis and debugging real time application execution.
NStorage Infrastructure & Data Lifecycle Management
N● Design & Engineering: Architectural planning, deployment, and administration of parallel file
Nsystems (e.G., Lustre, IBM Spectrum Scale, BeeGFS) and scalable object storage solutions.
N● Performance Engineering: Continuous monitoring of storage I/O performance;
identification and
nmitigation of I/O bottlenecks to ensure throughput requirements for compute workloads are met.
N● Data Lifecycle Strategy: Preparing for the future implementation of tiered storage concepts
N(HSM/ILM) to enable automated data migration across different performance tiers (hot-to-cold
Narchiving).
nNetwork Architecture & Troubleshooting
N● Advanced Diagnostics: Analyzing complex interdependencies between storage backends, file
Nsystems, and fabric architectures.
N● Protocol Analysis: Utilizing deep packet inspection tools (e.G., tcpdump, Wireshark) and
Nperformance profiling suites (e.G., fio, perf, iostat) to isolate packet loss, latency jitter, or metadata
Nlock contention.
n● Interconnect Optimization: Performance tuning of low-latency network configurations
N(InfiniBand/Ethernet) with a specific focus on RDMA protocols (RoCE/iWARP).
NAccess Control, Security & Compliance
N● Identity & Access Governance: Managing NFS (v3/v4) and SMB/CIFS exports, taking into account
Nmulti-tenant structures and cross-cluster access.
N● Hardening & Compliance: Enforcing network-level export policies (CIDR-based) and implementing
NPOSIX-compliant ACLs as well as POSIX group mappings via LDAP/NIS.
N● Data Security: Ensuring data-at-rest and data-in-transit encryption in compliance with regulatory
Nstandards (GDPR, Export Control) and institutional security policies.
NWorkload & User Enablement
N● Data Movement:
Orchestrating large-scale data transfers using specialized tools (e.G., XCP,
NAspera, Globus).
n● Consulting & Optimization: Advising users on I/O-efficient application development and best
Npractices for highly scalable workflows.
NCandidate Profile (Skills)
N● Expertise: Deep knowledge in administering HPC storage stacks and distributed file systems
N(POSIX-compliant).
n● Network Stack: Expert knowledge of TCP/IP networking, InfiniBand fabrics, routing & switching,
Nand RDMA technologies.
N● Automation: Experience in automating provisioning and configuration management (Ansible,
NPuppet, Bash/Python).
n● Containers & Orchestration: Solid understanding of integrating storage solutions into
Ncontainerized environments (Singularity/Apptainer, Kubernetes).
N● Good understanding of ITIL processes (Incident, Change, Problem, Service Management);
ncertification preferred.
N● Analytical mindset withstrong problem-solving skills, able to diagnose, reproduce, and resolve
Ncomplex issues.
n●Experience in stakeholder management and user support, with ability to adapt to evolving
Nbusiness processes.
nGood understanding of server architecture and Engineering applications.
N● Knowledge on AI tools (Copilot, Gemini Coding Assistant) for Quality improvement, Automation
Nand Value generation.
n● Understanding and experience in Cloud technologies (preferably AWS) and contributing to building
Ncloud ready solutions. (Bonus points if Certified).
N● Excellent communicationskills to work in a globally distributed team
NResponsibilities:
N● Operate, maintain, and optimize Scientific Computing environments (Linux/RedHat, HPC
Nclusters, scheduling systems).
N
n
- Develop and maintain legacy codebases, new codebases and automation scripts.
N
- Troubleshoot and resolve issues related to Scientific Computing Applications, its installation,
N
ndependency management and performance.
N
n
- Apply ITIL processes in daily operations (Incident/Change/Problem Management).
N
- Proactively challenge existing workflows and propose innovative improvements to enhance
N
nefficiency and user experience.
N● Collaborating with business users and relevant stakeholders to define project requirements,
Nscope and deliverables.
N● Provide day-to-day support for production processing by solving incidents and requests raised by
Nthe user community and anongoing event and alert management
N
n
- Analyze and propose/implement to Improve incident resolution quality.
N
- Perform incident categorisation/classification & find out critical issues/most popular incidents and
N
ndo root cause analysis
N● Perform root cause analysis for critical incidents & trend analysis for proactive
Nmeasures
n● Participate in Agile Development activities and deliver features for fix/improvement of the
Nservice Apply DevOps and AI tools, culture and mindset for all your activities on a daily
Nbasis
n
n
- Excellent communication skills to work in a globally distributed team
N
- Contribute to continuous improvement initiatives on demand.
N
- Engage and motivate thepeople (involved and stakeholders / customers) to promote Self-
N
nreflection, Self-management, Communication and Teamwork.
N
n
- Manage Conflicts and Negotiation to achieve optimized results for the business.
N
- Define and develop the objectives hierarchy of the project or product using appropriate strategies,
N
nmethods and tools.
N● Engage stakeholders in a governance that ensures effective use of time, budget and achievement
Nof quality.
n
n
- Integrate processes and people from other functions (matrix organizations).
N
- Ensure effective integration of lessons learnt. Ensure active risk and opportunity and contract
N
n(internal or external) management.
N● Ensure effective management of partners and suppliers.
📌 Hiring: Lead Technology Specialist (Bengaluru)
🏢 Airbus
📍 Bengaluru