Role- AI Product Tools and Frameworks - Cloud
Location- Noida & Hyderabad
Experience- 12-17 Years
Role Summary:
Own the cloud and infrastructure specification for the product. Build the deployment, workplace, and observability foundations that the rest of the org is validated against, including cost accountability and support for test/harness environments.
Key Responsibilities:
- Define infrastructure architecture, deployment topology, and environment standards (dev/test/staging/prod)
- Build and maintain CI/CD pipelines and Infrastructure-as-Code for the product org
- Establish end-to-end observability (metrics, tracing, logging) as the standard other teams validate against
- Own cloud cost visibility and reporting; drive cost-optimization recommendations
- Provision and support environments and test data infrastructure for harness/QA teams
Must-Have Requirements:
- 10+ years cloud/DevOps engineering, with production depth on at least one major cloud (AWS, Azure, or GCP)
- Terraform (or equivalent IaC) used in production, not just POC
- Kubernetes and container-based deployment at production scaleBuilt observability from scratch on at least one stack — either cloud-native (CloudWatch/Azure Monitor/GCP Operations) or a third-party tool (Datadog, Grafana/Prometheus) — must be able to describe the alerting/SLO design, not just tool names
- Has deployed and operated LLM-based or agentic AI workloads in production — specifically: model serving/inference infrastructure, GPU/compute scheduling, and cost-per-request tracking
- Has owned cloud cost visibility/reporting for a product or platform, with a specific example of a cost reduction driven
Preferred:
- Multi-cloud experience (any combination) — production integration across 3 specific providers is not required unless the product itself is multi-cloud today
- Databricks or Snowflake infrastructure management
- Security hardening (IAM, network policy, secrets management)
📌 AI Product Tools and Frameworks - Cloud (Noida)
🏢 HCLTech
📍 Noida