24 Sep
|
Cisco
|
Bengaluru
Job Overview
We are a small, agile, and highly collaborative team at the forefront of AI Infrastructure Automation and Benchmarking & Certification. We partner closely with leading hardware and software vendors to design, validate, and deliver curated AI infrastructure solutions to our customers - all built on proven reference architectures that reduce risk and accelerate time-to-value.
Because were a lean team, every member has real ownership and visibility into outcomes - from automating complex infrastructure workflows to running rigorous benchmarking and certification processes that ensure our solutions perform reliably at scale. We move fast, communicate openly, and lean on each others expertise daily, making this a great setting for engineers who want to work across the full stack of AI infrastructure rather than being siloed into one narrow function.
Your Impact
As a Software Engineering Technical leader for AI Cluster Orchestrator & Automation, you will design and implementation of repeatable, end-to-end automation for AI cluster bring-up, configuration, validation, lifecycle management, and teardown across compute, network, and storage domains.
You will
- Build idempotent orchestration workflows for GPU nodes, service nodes, network fabrics, and storage.
- Automate PXE, NVIDIA BCM, DHCP, Redfish, BIOS, firmware, OS, Kubernetes/operators, and Slurm integration.
- Coordinate dependencies across compute, Cisco networking, storage/Vast, GPU platforms, and service nodes.
- Implement health checks, configuration drift detection, validation gates, rollback, failure recovery, and operational observability.
- Document runbooks, APIs, interfaces, and support handoffs.
Minimum Qualifications
- Bachelors + 12 years of related experience, or Masters + 8 years of related or equivalent related work experience.
- Experience with Linux systems and AI/GPU cluster architecture knowledge.
- Coding experience using Python and automation/API development
- Prior experience with PXE, DHCP, Kubernetes, Slurm, BIOS/firmware, networking, and storage integration.
- Experience troubleshooting distributed provisioning failures and system dependencies.
Preferred Qualifications
- Familiarity with REST/Redfish and infrastructure-as-code concepts.1024-GPU-class lab operations, Supermicro systems, Cisco UCS, NVIDIA platforms, Vast storage, and Cisco switching.
- NVIDIA BCM, Cisco network automation, storage automation, and GPU server platforms such as Supermicro and Cisco UCS.
- Experience automating multi-plane/ToR network designs and large-scale cluster lifecycle operations.
- Familiarity with CI/CD, configuration management, logging, and telemetry systems.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Software Engineering Technical Leader | Automation Engineer (Bengaluru)
🏢 Cisco
📍 Bengaluru