Key Responsibilities:
- Deploy, manage, and optimize AI/ML and LLM inference workloads across GPU clusters, HPC infrastructure, and cloud environments.
- Build and maintain scalable AI platform infrastructure using Kubernetes, containers, and enterprise orchestration platforms.
- Administer and optimize Linux servers including system configuration, patching, security hardening, performance tuning, and troubleshooting.
- Manage physical infrastructure including servers, storage, networking, and bare metal environments within enterprise data centers.
- Implement and maintain CI/CD and automation workflows for platform and infrastructure deployments.
- Optimize infrastructure performance, GPU utilization, resource allocation, and distributed workloads to meet operational requirements.
- Benchmark and evaluate AI workloads for scalability, latency, throughput, and resource efficiency.
- Collaborate with infrastructure, SRE, and platform engineering teams to provision compute resources and maintain enterprise-scale AI environments.
- Implement monitoring, logging, observability, and alerting solutions for platform reliability and operational visibility.
- Apply security patches, upgrades, compliance controls, and operational best practices for Linux and Kubernetes environments.
- Troubleshoot issues across hardware, networking, operating systems, Kubernetes clusters, and AI/ML workloads.
- Support enterprise operations through efficient incident, change, and ticket management processes.
- Automate infrastructure operations using scripting and infrastructure automation tools.
Required Qualifications:
- 8+ years of experience in Linux systems administration, cloud-native infrastructure, HPC environments, or platform engineering.
- At least 4 years of experience supporting AI/ML workloads or large-scale distributed compute environments in production.
- Comfortable leveraging AI-assisted tools for collaborative development, code generation, refactoring, and productivity enhancement.
- Strong hands-on expertise with Linux administration (RHEL, Ubuntu, or similar).
- Experience with Kubernetes administration, container orchestration, and cloud native infrastructure platforms.
- Robust understanding of GPU infrastructure, distributed computing, and HPC systems.
- Hands-on experience with bare-metal infrastructure, servers, storage systems, and enterprise networking.
- Strong understanding of networking fundamentals including TCP/IP, DNS, load balancing, and firewalls.
- Experience with scripting and infrastructure automation with Bash, Python etc.
- Experience with CI/CD, DevOps, or infrastructure deployment workflows.
- Experience with monitoring, observability, and logging platforms.
- Strong troubleshooting and performance optimization skills across Linux, Kubernetes, networking, and infrastructure stacks.
- Excellent problem-solving, communication, and collaboration skills.
- Ability to work effectively in fast-paced, mission-critical production environments.
📌 Infrastructure Engineer (Pune)
🏢 Emergys
📍 Pune