23 Aug
|
Alion
|
Anupgarh
We are looking for a Lead DevOps and Platform Engineer who combines strong software engineering expertise with deep cloud infrastructure and platform engineering experience. This role is ideal for someone who started as a Backend/Software Engineer (Python, Java, Go, or Node.js ) and evolved into leading DevOps, Platform Engineering, or Cloud Infrastructure initiatives. You'll architect scalable AI infrastructure, build cloud-native platforms, and enable high availability, reliability, and rapid product delivery.
Responsibilities
- Design and manage highly scalable cloud infrastructure on AWS (GCP/Azure exposure is a plus).
- Architect multi-region, high-availability (HA), and disaster recovery (DR) cloud platforms.
- Build and manage Kubernetes (EKS/GKE) clusters for large-scale AI and microservices deployments.
- Design networking, storage, compute, VPCs, NAT Gateways, PrivateLink, and secure cloud architectures.
- Automate infrastructure using Infrastructure as Code (Terraform, CloudFormation).
- Partner with engineering teams to build cloud-native, auto-scaling applications.
- Review backend architectures and contribute to platform tooling using Python, Java, Go, or Node.js .
- Implement observability using distributed tracing, logging, metrics, and monitoring.
- Lead production incident management, root cause analysis (RCA), and platform reliability improvements.
- Optimise CI/CD pipelines and integrate security best practices into deployment workflows.
- Mentor engineering teams on cloud architecture, Kubernetes, DevOps, and platform engineering best practices.
Requirements
- 8-12 years of experience in Software Engineering, DevOps, Platform Engineering, or Cloud Infrastructure.
- Strong backend development experience with Python, Java, Go, or Node.js .
- Deep expertise in AWS and cloud infrastructure architecture.
- Hands-on experience with Kubernetes (EKS/GKE) and Docker.
- Strong knowledge of Infrastructure as Code using Terraform or CloudFormation.
- Experience designing High Availability (HA) and Disaster Recovery (DR) architectures.
- Strong understanding of networking, VPCs, load balancing, storage, and cloud security.
- Experience building and supporting large-scale microservices and distributed systems.
- Strong troubleshooting, performance optimisation, and incident management skills.
- Experience working in AI, SaaS, or startup environments is preferred.
Good To Have
- CI/CD tools (Jenkins, GitHub Actions, GitLab CI).
- Service Mesh (Istio/Linkerd).
- Observability tools (Prometheus, Grafana, OpenTelemetry).
- Security scanning (SAST/DAST).
- CKA Certification.
- Experience with AI infrastructure, LLM inference, and vector databases.
- Robust technical leadership with a hands-on approach.
- Ownership mindset and ability to drive platform initiatives end-to-end.
- Excellent problem-solving and architectural thinking.
- Passion for building scalable cloud-native systems.
- Ability to mentor engineering teams and collaborate across functions.
Top Skills (Instahyre Skills Section)
- AWS and Cloud Architecture.
- Kubernetes and Platform Engineering.
- Backend Development (Python/Java/Go/Node.js ).
- Terraform and Infrastructure as Code.
📌 Engineering & DevOps Lead (Anupgarh)
🏢 Alion
📍 Anupgarh