01 Aug
|
HCLTech
|
Bengaluru
Bengaluru, Karnataka
Job Summary
Skill: Network (Enterprise / AI Infrastructure Context) L3 – Senior Engineer / SME Skill Requirement Strong hands-on expertise in enterprise networking and Linux systems Responsible for troubleshooting, performance tuning, and operational stability of network and OS layers supporting AI/HPC/Kubernetes workloads Networking (L3) Certifications (Preferred): CCNA (Mandatory baseline)CCNP (Strongly preferred) Experience: 5–10 years in data center / cloud networking operations Hands-on experience with: Routing & Switching (BGP, OSPF, VLANs)Load balancing, firewall basicsInfiniBand card and switchesNvidia UFMCCNA, CCNP preferred with working knowledge on InfiniBand and Mellanix Switches Skill Depth: Strong in incident troubleshooting and RCA Ability to diagnose: Latency, packet loss, connectivity failuresNetwork issues affecting GPU / distributed workloadsWorking knowledge of cloud networking (AWS/Azure/GCP)Working Knowledge of AI/HPC networking for distributed training for GPU-GPU communication.Define and maintain production readiness standards across platform, data, model, application, and security layers.Establish SLO/SLI frameworks for latency, availability, quality, safety, and drift implement error budget policies.Maintain compliance mappings (e.g., ISO 27001, SOC 2, GDPR/DPDP, HIPAA where applicable).Author PRR checklists, runbooks/playbooks, and DR/BCP blueprints (RTO/RPO, multi‑region/site failover). Drive enablement (trainings, brown-bags) and maintain knowledge repositories and decision records.Implement observability (tracing, metrics, logs), dashboards, and SLO burn and cost anomaly alerting.Execute safe releases (canary/shadow/blue green), prompt/model versioning, feature flags, and rollback plans.
Key Responsibilities
Skill Depth: Strong in incident troubleshooting and RCA Ability to diagnose: Latency, packet loss,
connectivity failuresNetwork issues affecting GPU / distributed workloadsWorking knowledge of cloud networking (AWS/Azure/GCP)Working Knowledge of AI/HPC networking for distributed training for GPU-GPU communication.Define and maintain production readiness standards across platform, data, model, application, and security layers.Establish SLO/SLI frameworks for latency, availability, quality, safety, and drift implement error budget policies.Maintain compliance mappings (e.g., ISO 27001, SOC 2, GDPR/DPDP, HIPAA where applicable).Author PRR checklists, runbooks/playbooks, and DR/BCP blueprints (RTO/RPO, multi‑region/site failover). Drive enablement (trainings, brown-bags) and maintain knowledge repositories and decision records.Implement observability (tracing, metrics, logs), dashboards, and SLO burn and cost anomaly alerting.Execute protected releases (canary/shadow/blue green), prompt/model versioning, feature flags, and rollback plans.
Skill Requirements
Skill Requirement Strong hands-on expertise in enterprise networking and Linux systems Responsible for troubleshooting, performance tuning, and operational stability of network and OS layers supporting AI/HPC/Kubernetes workloads
CCNA, CCNP preferred with working knowledge on InfiniBand and Mellanix Switches
Other Requirements
1. Optional But Valuable Certifications: Cisco Certified Network Associate (Ccna), Comptia Network+, Itil Foundation.
2. Excellent communication and presentation skills. Must be able to clearly communicate with the customer and be able to present solutions to customers at CIO, CXO level
3. Must be open for 24x7 environment
#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-
📌 Track Lead - Network LAN/WAN (Bengaluru)
🏢 HCLTech
📍 Bengaluru