14 Sep
|
Nameless
|
Gurugram
Al Support Engineer is responsible to provide L3 support based on incident management system.
Own end-to-end resolution of high-priority (P1/ P2)
production incidents
• Perform deep troubleshooting and root-cause analysis across platform, data and Al layers
• Drive permanent fixes and platform improvements - not just incident recovery
• Build dashboards, monitoring and alerts for proactive issue management
• Maintain runbooks and knowledge base for continuous improvement
• Support platform operations - monitoring, observability, deployments and pipelines
• Coordinate and provide resolution with TD Product Owners and PGTLS
• Flexible to work in rotational shift across 24X7, 5 working days and 2 weekly off
Requirements
Root-cause analysis across distributed systems, APls, pipelines and data layers
• Identification and fixing of bugs in production workplace
• Data engineering fundamentals - data ingestion, storage, pipelines and consistency
• Azure cloud expertise with observability tools (eg; Dynatrace, LangSmith, OpenTelemetry)
• DevOps / MLOps and CI/CD pipelines; deployment and orchestration in Azure cloud
• Incident management with Jira / ServiceNow in SLA-driven environments
• 5-10+ years in production support, SRE or platform engineering
📌 TD AI Support EngineerAI Platform) (Gurugram)
🏢 Nameless
📍 Gurugram