SeniorAdministrator - Social Network Development, Netwok Automation
Hyderabad, Telangana
Job Summary
Key Responsibilities Agentic AI Development • Design and develop autonomous and assistive AI agents for NOC and fleet operations. • Build multi-agent workflows leveraging LLMs, RAG, planning, orchestration, and tool-calling frameworks. • Develop AI-driven operational assistants capable of diagnostics, root cause analysis, remediation recommendations, and autonomous execution. • Implement human-in-the-loop governance controls for production AI systems. Platform Engineering and Automation • Develop integrations with ServiceNow, Prometheus/Grafana, SolarWinds, PagerDuty, Ansible Automation Platform, Terraform, and cloud and virtualization platforms. • Build APIs, connectors, and automation workflows enabling agent actions across operational systems. • Create self-healing and auto-remediation capabilities for infrastructure and application incidents. Reliability Engineering • Partner with SRE teams to improve SLI/SLO-based monitoring and operational excellence. • Design event correlation and incident intelligence solutions. • Automate incident triage, escalation, and root cause analysis workflows. • Reduce Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR). AI and Data Engineering • Implement Retrieval Augmented Generation (RAG) solutions using operational knowledge bases, runbooks, and historical incident data. • Build evaluation frameworks for AI agent accuracy, safety, and effectiveness. • Develop telemetry, observability, and performance analytics for agent operations. • Ensure responsible and secure AI deployment practices. Collaboration and Leadership • Work closely with AMD Fleet Operations architects, infrastructure teams, SREs, and platform engineering teams. • Mentor junior developers and guide architecture decisions. • Participate in design reviews, technical roadmaps, and innovation initiatives. • Drive adoption of modern AI engineering and automation best practices. Required Qualifications • Bachelor's or Master's degree in Computer Science, Engineering, Artificial Intelligence, or a related field. • 8+ years of software engineering experience. • Strong hands-on expertise in Python; Java or Go; REST APIs and microservices; Git, CI/CD, and DevOps practices. • Experience building enterprise-scale automation solutions. • Strong understanding of networking, infrastructure, cloud, and operational support environments. • Experience with LangGraph, LangChain, Semantic Kernel, AutoGen, CrewAI, or equivalent multi-agent frameworks. • Experience integrating AI with enterprise tools and operational platforms. Preferred Qualifications • Experience in NOC, SRE, NetOps, CloudOps, or Infrastructure Operations. • Knowledge of VMware, Kubernetes, Azure, AWS, and Linux systems. • Experience with observability platforms such as Prometheus, Grafana, SolarWinds, Datadog, or Splunk. • Experience with ServiceNow development and integration. • Knowledge of routing, switching, DNS, firewalls, and load balancing. Desired Skills Agentic AI Architecture Generative AI and LLMs Multi-Agent Systems RAG and Knowledge Engineering Observability Engineering Infrastructure Automation Terraform and Ansible SRE Practices Event Correlation and AIOps Incident Management Automation Success Metrics • Reduction in manual operational effort. • Increase in automated incident resolution rates. • Improved NOC productivity and operational efficiency. • Reduction in MTTR and operational escalations. • Increased adoption of agent-driven operational workflows. • Delivery of production-ready autonomous operational capabilities. Program Alignment This position supports the AMD Fleet Operations direction of automation-first operations, SRE-led reliability, observability-first architecture, self-healing remediation, and autonomous NetOps, CloudOps, and virtualization operations agents.
Key Responsibilities
Key Responsibilities Agentic AI Development • Design and develop autonomous and assistive AI agents for NOC and fleet operations. • Build multi-agent workflows leveraging LLMs, RAG, planning, orchestration, and tool-calling frameworks. • Develop AI-driven operational assistants capable of diagnostics, root cause analysis, remediation recommendations, and autonomous execution. • Implement human-in-the-loop governance controls for production AI systems. Platform Engineering and Automation • Develop integrations with ServiceNow, Prometheus/Grafana, SolarWinds, PagerDuty, Ansible Automation Platform, Terraform, and cloud and virtualization platforms. • Build APIs, connectors, and automation workflows enabling agent actions across operational systems. • Create self-healing and auto-remediation capabilities for infrastructure and application incidents. Reliability Engineering • Partner with SRE teams to improve SLI/SLO-based monitoring and operational excellence. • Design event correlation and incident intelligence solutions. • Automate incident triage, escalation, and root cause analysis workflows. • Reduce Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR). AI and Data Engineering • Implement Retrieval Augmented Generation (RAG) solutions using operational knowledge bases, runbooks, and historical incident data.
• Build evaluation frameworks for AI agent accuracy, safety, and effectiveness. • Develop telemetry, observability, and performance analytics for agent operations. • Ensure responsible and secure AI deployment practices. Collaboration and Leadership • Work closely with AMD Fleet Operations architects, infrastructure teams, SREs, and platform engineering teams. • Mentor junior developers and guide architecture decisions. • Participate in design reviews, technical roadmaps, and innovation initiatives. • Drive adoption of modern AI engineering and automation best practices. Required Qualifications • Bachelor's or Master's degree in Computer Science, Engineering, Artificial Intelligence, or a related field. • 8+ years of software engineering experience. • Strong hands-on expertise in Python; Java or Go; REST APIs and microservices; Git, CI/CD, and DevOps practices. • Experience building enterprise-scale automation solutions. • Strong understanding of networking, infrastructure, cloud, and operational support environments. • Experience with LangGraph, LangChain, Semantic Kernel, AutoGen, CrewAI, or equivalent multi-agent frameworks. • Experience integrating AI with enterprise tools and operational platforms. Preferred Qualifications • Experience in NOC, SRE, NetOps, CloudOps, or Infrastructure Operations. • Knowledge of VMware, Kubernetes, Azure, AWS, and Linux systems. • Experience with observability platforms such as Prometheus, Grafana, SolarWinds, Datadog, or Splunk. • Experience with ServiceNow development and integration. • Knowledge of routing, switching, DNS, firewalls, and load balancing. Desired Skills Agentic AI Architecture Generative AI and LLMs Multi-Agent Systems RAG and Knowledge Engineering Observability Engineering Infrastructure Automation Terraform and Ansible SRE Practices Event Correlation and AIOps Incident Management Automation Success Metrics • Reduction in manual operational effort. • Increase in automated incident resolution rates. • Improved NOC productivity and operational efficiency. • Reduction in MTTR and operational escalations. • Increased adoption of agent-driven operational workflows. • Delivery of production-ready autonomous operational capabilities. Program Alignment This position supports the AMD Fleet Operations direction of automation-first operations, SRE-led reliability, observability-first architecture, self-healing remediation, and autonomous NetOps, CloudOps, and virtualization operations agents.
Skill Requirements
Key Responsibilities Agentic AI Development • Design and develop autonomous and assistive AI agents for NOC and fleet operations. • Build multi-agent workflows leveraging LLMs, RAG, planning, orchestration, and tool-calling frameworks. • Develop AI-driven operational assistants capable of diagnostics, root cause analysis, remediation recommendations, and autonomous execution. • Implement human-in-the-loop governance controls for production AI systems. Platform Engineering and Automation • Develop integrations with ServiceNow, Prometheus/Grafana, SolarWinds, PagerDuty, Ansible Automation Platform, Terraform, and cloud and virtualization platforms. • Build APIs, connectors, and automation workflows enabling agent actions across operational systems. • Create self-healing and auto-remediation capabilities for infrastructure and application incidents. Reliability Engineering • Partner with SRE teams to improve SLI/SLO-based monitoring and operational excellence. • Design event correlation and incident intelligence solutions. • Automate incident triage, escalation, and root cause analysis workflows. • Reduce Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR). AI and Data Engineering • Implement Retrieval Augmented Generation (RAG) solutions using operational knowledge bases, runbooks, and historical incident data. • Build evaluation frameworks for AI agent accuracy, safety, and effectiveness. • Develop telemetry, observability, and performance analytics for agent operations. • Ensure responsible and secure AI deployment practices. Collaboration and Leadership • Work closely with AMD Fleet Operations architects, infrastructure teams, SREs, and platform engineering teams. • Mentor junior developers and guide architecture decisions. • Participate in design reviews, technical roadmaps, and innovation initiatives. • Drive adoption of modern AI engineering and automation best practices. Required Qualifications • Bachelor's or Master's degree in Computer Science, Engineering, Artificial Intelligence, or a related field. • 8+ years of software engineering experience. • Strong hands-on expertise in Python; Java or Go; REST APIs and microservices; Git, CI/CD, and DevOps practices. • Experience building enterprise-scale automation solutions. • Strong understanding of networking, infrastructure, cloud, and operational support environments.
• Experience with LangGraph, LangChain, Semantic Kernel, AutoGen, CrewAI, or equivalent multi-agent frameworks. • Experience integrating AI with enterprise tools and operational platforms. Preferred Qualifications • Experience in NOC, SRE, NetOps, CloudOps, or Infrastructure Operations. • Knowledge of VMware, Kubernetes, Azure, AWS, and Linux systems. • Experience with observability platforms such as Prometheus, Grafana, SolarWinds, Datadog, or Splunk. • Experience with ServiceNow development and integration. • Knowledge of routing, switching, DNS, firewalls, and load balancing. Desired Skills Agentic AI Architecture Generative AI and LLMs Multi-Agent Systems RAG and Knowledge Engineering Observability Engineering Infrastructure Automation Terraform and Ansible SRE Practices Event Correlation and AIOps Incident Management Automation Success Metrics • Reduction in manual operational effort. • Increase in automated incident resolution rates. • Improved NOC productivity and operational efficiency. • Reduction in MTTR and operational escalations. • Increased adoption of agent-driven operational workflows. • Delivery of production-ready autonomous operational capabilities. Program Alignment This position supports the AMD Fleet Operations direction of automation-first operations, SRE-led reliability, observability-first architecture, self-healing remediation, and autonomous NetOps, CloudOps, and virtualization operations agents.
Other Requirements
Key Responsibilities Agentic AI Development • Design and develop autonomous and assistive AI agents for NOC and fleet operations. • Build multi-agent workflows leveraging LLMs, RAG, planning, orchestration, and tool-calling frameworks. • Develop AI-driven operational assistants capable of diagnostics, root cause analysis, remediation recommendations, and autonomous execution. • Implement human-in-the-loop governance controls for production AI systems. Platform Engineering and Automation • Develop integrations with ServiceNow, Prometheus/Grafana, SolarWinds, PagerDuty, Ansible Automation Platform, Terraform, and cloud and virtualization platforms. • Build APIs, connectors, and automation workflows enabling agent actions across operational systems. • Create self-healing and auto-remediation capabilities for infrastructure and application incidents. Reliability Engineering • Partner with SRE teams to improve SLI/SLO-based monitoring and operational excellence. • Design event correlation and incident intelligence solutions. • Automate incident triage, escalation, and root cause analysis workflows. • Reduce Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR). AI and Data Engineering • Implement Retrieval Augmented Generation (RAG) solutions using operational knowledge bases, runbooks, and historical incident data. • Build evaluation frameworks for AI agent accuracy, safety, and effectiveness. • Develop telemetry, observability, and performance analytics for agent operations. • Ensure responsible and secure AI deployment practices. Collaboration and Leadership • Work closely with AMD Fleet Operations architects, infrastructure teams, SREs, and platform engineering teams. • Mentor junior developers and guide architecture decisions. • Participate in design reviews, technical roadmaps, and innovation initiatives. • Drive adoption of modern AI engineering and automation best practices. Required Qualifications • Bachelor's or Master's degree in Computer Science, Engineering, Artificial Intelligence, or a related field. • 8+ years of software engineering experience. • Strong hands-on expertise in Python; Java or Go; REST APIs and microservices; Git, CI/CD, and DevOps practices. • Experience building enterprise-scale automation solutions. • Solid understanding of networking, infrastructure, cloud, and operational support environments. • Experience with LangGraph, LangChain, Semantic Kernel, AutoGen, CrewAI, or equivalent multi-agent frameworks. • Experience integrating AI with enterprise tools and operational platforms. Preferred Qualifications • Experience in NOC, SRE, NetOps, CloudOps, or Infrastructure Operations. • Knowledge of VMware, Kubernetes, Azure, AWS, and Linux systems. • Experience with observability platforms such as Prometheus, Grafana, SolarWinds, Datadog, or Splunk. • Experience with ServiceNow development and integration. • Knowledge of routing, switching, DNS, firewalls, and load balancing. Desired Skills Agentic AI Architecture Generative AI and LLMs Multi-Agent Systems RAG and Knowledge Engineering Observability Engineering Infrastructure Automation Terraform and Ansible SRE Practices Event Correlation and AIOps Incident Management Automation Success Metrics • Reduction in manual operational effort. • Increase in automated incident resolution rates. • Improved NOC productivity and operational efficiency. • Reduction in MTTR and operational escalations. • Increased adoption of agent-driven operational workflows. • Delivery of production-ready autonomous operational capabilities. Program Alignment This position supports the AMD Fleet Operations direction of automation-first operations, SRE-led reliability, observability-first architecture, self-healing remediation, and autonomous NetOps, CloudOps, and virtualization operations agents.
📌 SeniorAdministrator - Social Network Development, Netwok Automation (India)
🏢 HCLTech
📍 India