We are looking for an outstanding DevOps and Site Reliability Engineer to join
the NVIDIA e-commerce team. You will be a key architect of our e-commerce
platform, ensuring that our systems are scalable, resilient, and automated. The
ideal candidate is a Terraform expert who views infrastructure as code (IaC) not
just as a tool, but as a philosophy. You will bridge the gap between development
and operations, focusing on system reliability, high availability, and the
performance of our global e-commerce platform.
What you’ll be doing:
Architect and refine automated deployment Jenkins pipelines to ensure
seamless, zero-downtime releases.
Design, build, and maintain enterprise-scale infrastructure using Terraform.
Establish modular, reusable patterns for AWS resources.
Optimize and manage sophisticated AWS settings with a focus on
cost-efficiency and security.
Transition our monitoring from reactive to proactive using AI-powered
observability tools (e.g., Datadog Watchdog) for automated root cause
analysis (RCA) and anomaly detection.
Define and monitor SLOs and SLAs. Lead incident response and conduct thorough
post-mortems to improve system resilience.
What we need to see:
8+ years or equivalent industry experience
Bachelor's/Master's Degree in Computer Science, Software Engineering, or
equivalent experience.
Exceptionally robust background in developing CI/CD processes and deployment
pipelines using Jenkins.
Extensive experience architecting on AWS Cloud and running services such as
API Gateway, Lambda, EKS/ECS, RDS, S3, and SQS.
Expert-level knowledge of Terraform (including state management, workspaces,
and complex module development).
Advanced experience with Kubernetes (EKS) and Docker, including
orchestration, service meshes, and Helm.
Robust proficiency in a scripting language, such as Python, for automation
and custom tooling.
Strong communication skills.
Ways to stand out from the crowd:
Deep understanding of DN
📌 Senior Devops Engineer Pune
🏢 NVIDIA
📍 Pune