Sr Tower Lead (Support & Operations)
Gautam Buddha Nagar, Uttar Pradesh
Job Summary
Senior Forward Deployment Engineer – AWS Agentic AI
Job Type
Full-Time
Experience
8–12 years overall; at least 3 years in Generative AI / Agentic AI with significant enterprise production delivery on AWS
Locations
Noida, Hyderabad, Chennai, Bangalore, Pune
Primary Focus
Solution architecture, customer technical leadership, production strategy, technical governance and FDE leadership
Role Overview
We are looking for a Senior Forward Deployment Engineer (FDE) to lead the technical architecture and end-to-end deployment of enterprise Generative AI and Agentic AI solutions for strategic customers on AWS. The Senior FDE is a customer-facing solution architect who can code: they own the technical success of the customer engagement, make architecture and deployment decisions, lead customer workshops, review implementation, drive production readiness, and mentor FDEs.
- Customer-facing solution architect who can code
- Owns technical success of the customer engagement
- Makes architecture and production decisions
- Leads FDEs and influences the platform roadmap
Typical time allocation
- ~50% architecture and technical leadership
- ~30% customer leadership
- ~20% engineering oversight and mentoring
Key Responsibilities
- Own the overall technical delivery of enterprise Agentic AI deployments from discovery and architecture through production rollout, hypercare, and transition to BAU.
- Lead customer discovery, technical workshops, architecture reviews, executive technical discussions, solution demonstrations, POCs, pilots, and production planning.
- Design customer-specific AI and AWS architectures using Amazon Bedrock, Amazon Bedrock AgentCore, AWS services, enterprise data sources, applications, security controls, and operational services.
- Define agent architecture including multi-agent patterns, Planner/Critic/Supervisor/Orchestrator designs, domain agents, Skills, tools, RAG, human-in-the-loop controls, state, and enterprise integrations.
- Define the integration architecture for REST APIs, MCP tools, databases, ITSM platforms, monitoring systems, identity systems, messaging systems, and proprietary applications.
- Decide which requirements should be implemented through configuration, customer-specific extensions, or reusable platform capabilities; drive the appropriate engineering path.
- Define production deployment architecture across Amazon Bedrock AgentCore Runtime, Gateway, Memory, Identity, Policy and Evaluations, Lambda, ECS, EKS, API Gateway, EventBridge, Step Functions, S3, CloudWatch, IAM, networking, and other appropriate AWS services.
- Define production architecture requirements for availability, scalability, security, resiliency, observability, release management, rollback, disaster recovery, and operational support.
- Lead security and identity architecture discussions covering IAM,
role-based access, authentication, authorization, Secrets Manager, data protection, auditability, VPC controls, and network security.
- Define and review production readiness criteria covering functionality, security, performance, reliability, evaluation, observability, governance, supportability, and operational readiness.
- Define the customer AgentOps strategy including telemetry, tracing, logging, metrics, evaluation, quality monitoring, cost monitoring, and operational dashboards.
- Lead complex troubleshooting and root-cause analysis for production issues involving agents, models, RAG, tools, integrations, networking, authentication, and AWS infrastructure.
- Drive reliability, latency, scalability, model performance, and AI cost optimization across customer deployments.
- Define and review CI/CD, release management, environment promotion, versioning, rollback, and deployment processes.
- Review architecture, integration code, deployment plans, and technical deliverables produced by FDEs; establish engineering quality standards.
- Mentor and technically guide FDEs and Associate FDEs working on customer deployments.
- Establish reusable deployment patterns, integration patterns, reference architectures, runbooks, onboarding standards, and troubleshooting playbooks.
- Identify recurring customer requirements and drive their conversion into reusable agents, Skills, tools, connectors, frameworks, and platform capabilities with Agent Development / Platform Engineering teams.
- Act as the senior technical escalation point for complex customer issues and major production incidents.
- Provide technical feedback to Product, Agent Development, Platform Engineering, Security, Cloud Architecture, and Operations teams and influence platform roadmap priorities.
- Lead technical handover, customer enablement, operational readiness, and transition to support/BAU teams.
Skill Requirements
Must Have Skills
- Strong Python and software engineering skills with the ability to review and guide production code.
- Extensive hands-on experience with Generative AI, LLMs, RAG, prompt engineering, tool calling, evaluation, and multi-agent architectures.
- Strong hands-on experience with Amazon Bedrock and/or Amazon Bedrock AgentCore, with LangGraph/LangChain/Strands Agents or equivalent frameworks in production.
- Strong AWS architecture and engineering experience across Amazon Bedrock and core AWS services.
- Strong enterprise integration experience across REST APIs, databases, ITSM platforms, monitoring systems, identity platforms,
and customer-specific applications.
- Strong understanding of AWS IAM, role-based access, OAuth, secrets management, authentication, authorization, security, and enterprise networking.
- Strong experience with containerized applications, Docker, ECS/EKS, Lambda, AgentCore Runtime, and production deployment patterns.
- Strong experience with CI/CD, Git, release management, versioning, and production operations.
- Strong experience with observability, tracing, logging, monitoring, debugging, evaluation, and production incident management.
- Robust understanding of RAG architecture, knowledge integration, retrieval, evaluation, and agent quality measurement.
- Experience with Responsible AI, guardrails, AI security, data protection, and enterprise governance.
- Strong customer-facing communication skills and ability to lead architecture conversations with senior technical stakeholders.
Preferred Skills
- Deep experience with Amazon Bedrock, AgentCore Runtime, Gateway, Memory, Identity, Policy, Evaluations, and AgentCore Observability.
- Experience with OpenTelemetry/ADOT, enterprise AgentOps, and agent evaluation practices.
- Experience with MCP, A2A, AWS SDKs, and enterprise agent interoperability.
- Experience with Terraform / Infrastructure-as-Code and enterprise AWS landing zones.
- Strong experience with Lambda, ECS, EKS, API Gateway, EventBridge, Step Functions, SQS/SNS, S3, Secrets Manager, CloudWatch, networking, VPC endpoints, and private connectivity.
- Experience with ServiceNow, ITSM, CloudOps, SRE, AIOps, infrastructure automation, or enterprise operations.
- Experience designing highly available, scalable, secure production architectures.
- Experience leading POCs, pilots, MVPs, production rollouts, and strategic enterprise deployments.
- Exposure to multiple cloud platforms is advantageous.
Other Requirements
Qualifications
- Bachelor’s or Master’s degree in Computer Science, Engineering, Information Technology, or related field.
- 8–12 years of overall software engineering, cloud engineering, solution architecture, or equivalent technical experience.
- At least 3 years of practical Generative AI / Agentic AI experience.
- Demonstrated track record of delivering multiple enterprise technology or AI solutions into production.
- Demonstrated customer-facing technical leadership and architecture ownership.
Key Attributes
- Strong customer-facing presence and ability to build technical trust with enterprise customers.
- Strong ownership of outcomes rather than individual tasks.
- Excellent architecture and problem-solving skills.
- Comfortable operating across AI engineering, AWS cloud architecture, integration, security, and operations.
- Able to make pragmatic decisions between reusable platform capabilities and customer-specific implementation.
- Strong mentoring and technical leadership capability.
📌 Sr Tower Lead (Support & Operations) (India)
🏢 HCLTech
📍 India