We are looking for an outstanding DevOps and Site Reliability Engineer to join
the NVIDIA e-commerce team. You will be a key architect of our e-commerce
platform, ensuring that our systems are scalable, resilient, and automated. The
ideal candidate is a Terraform expert who views infrastructure as code (IaC) not
just as a tool, but as a philosophy. You will bridge the gap between development
and operations, focusing on system reliability, high availability, and the
performance of our global e-commerce platform.
What you’ll be doing:
* Architect and refine automated deployment Jenkins pipelines to ensure
seamless, zero-downtime releases.
* Design, build, and maintain enterprise-scale infrastructure using Terraform.
Establish modular, reusable patterns for AWS resources.
* Optimize and manage sophisticated AWS environments with a focus on
cost-efficiency and security.
* Transition our monitoring from reactive to proactive using AI-powered
observability tools (e.g., Datadog Watchdog) for automated root cause
analysis (RCA) and anomaly detection.
* Define and monitor SLOs and SLAs. Lead incident response and conduct thorough
post-mortems to improve system resilience.
What we need to see:
* 8+ years or equivalent industry experience
* Bachelor's/Master's Degree in Computer Science, Software Engineering, or
equivalent experience.
* Exceptionally robust background in developing CI/CD processes and deployment
pipelines using Jenkins.
* Extensive experience architecting on AWS Cloud and running services such as
API Gateway, Lambda, EKS/ECS, RDS, S3, and SQS.
* Expert-level knowledge of Terraform (including state management, workspaces,
and complex module development).
* Advanced experience with Kubernetes (EKS) and Docker, including
orchestration, service meshes, and Helm.
* Strong proficiency in a scripting language, such as Python, for automation
and custom tooling.
* Strong communication skills.
Ways to stand out from the crowd:
* Deep understanding of DN
📌 Senior DevOps Engineer (Pune)
🏢 NVIDIA
📍 Pune