27 Sep
|
Webscale
|
Bengaluru
27 Sep
Webscale
Bengaluru
Department : - Customer Support ECOM Infra Support
Location : - Remote (India)
Experience : - 2-5 years
Shift : - Rotational 247 shifts, including nights and weekends
About the Role
We are looking for an ECOM Infrastructure Support Engineer to own the full lifecycle of customer support: rapid response, end-to-end resolution, and deep root cause analysis. You will operate as part of a single, highly capable global team that balances speed, quality, and system reliability.
This role also covers technical support for our AI product line. You will work daily with AI agents and AI-assisted tooling, so we are looking for someone who is genuinely AI fluent, not just AI aware.
ECOM Support Engineers are full-spectrum operators. You will act as the front door for customer issues, resolve standard issues independently, and perform deep technical investigations as/when required.
This is a fully remote role based in India. The team runs 247 on a rotational shift model, which includes night and weekend shifts on a rotating basis. A reliable home internet connection and a quiet working setup are essential.
Key Responsibilities
1. Unified Incident Ownership (247)
- Monitor and respond across alerts, tickets, messages, and phone
- Classify severity and customer impact quickly
- Take immediate action to stabilise systems
2. End-to-End Troubleshooting and Resolution
- Resolve issues across Linux systems, HTTP/TLS/DNS/CDN/WAF, caching layers, and basic database and application behaviour
- Use runbooks, tooling, and experience to drive issues to completion
- Expand beyond runbooks as experience grows
3. Deep Investigation and Root Cause Analysis
- Perform advanced troubleshooting using logs, metrics, and traces
- Diagnose performance bottlenecks across database, cache, and infrastructure layers
- Identify true root cause, not just symptoms
- Own RCA documentation for significant incidents
4.
AI Product Support
- Provide technical support for our AI product, including configuration, and issue resolution
- Troubleshoot AI agent behaviour: prompt and instruction issues, tool and integration failures, retrieval and knowledge base gaps, latency, and inconsistent or incorrect outputs
- Investigate and explain agent decision paths, tool calls, and API/webhook integration failures
- Translate customer-reported AI behaviour into explicit, reproducible technical findings for engineering
- Contribute to AI product documentation, known-issue libraries, and support playbooks
5. Customer Communication
- Deliver clear, concise, and proactive updates
- Set expectations and communicate progress throughout an incident
- Provide strong closure summaries covering root cause and actions taken
6. Continuous Improvement
- Identify and act on repeat issues, alert noise, and documentation gaps
- Contribute to runbooks and SOPs
- Identify automation opportunities and drive AI-led ticket deflection
Day-to-Day Workflow
1. Acknowledge - meet the response SLO, confirm impact, share initial findings
2. Stabilise - apply immediate mitigations and log all actions
3. Investigate - determine root cause or the next best hypothesis using logs, metrics, and system knowledge
4. Resolve - drive the issue to closure with full documentation
5. Communicate and Close - maintain ownership of updates, document resolution, RCA, and improvements
What We Are Looking For
Technical Skills
- Linux administration, networking fundamentals, and the HTTP stack
- TLS, DNS, and CDN/WAF mechanics
- AWS infrastructure fundamentals (EC2, RDS, ElastiCache, CloudWatch)
- Database performance basics, ideally MySQL
- Magento, Shopware, or comparable e-commerce platform troubleshooting
- Observability tooling such as Grafana, ELK, or New Relic
AI Skills (Required)
- Hands-on working experience with AI agents and agentic workflows, not just conversational AI use
- Practical fluency with LLM-based tools and assistants in a day-to-day working context
- Ability to debug agent behaviour and explain what went wrong in plain language to a customer
- Comfort with APIs, JSON, and webhook-based integrations
- Basic scripting ability (Python preferred) for diagnostics and automation
Operational Excellence
- Strong triage instincts and a structured troubleshooting approach
- Consistent documentation discipline
- Sound judgement on when to own an issue and when to escalate
Communication
- Clear, concise, and structured written updates
- Strong asynchronous communication
- Ability to write effective escalation and RCA narratives
Tools and Systems You Will Use
Zendesk (system of record) Slack Grafana, ELK, New Relic AWS Console and CLI Traffic Viewer and CDN/WAF tooling AI tools
Your First 90 Days
- 30 days: tool proficiency established, meeting response expectations, handling scoped issues with guidance
- 60 days: resolving the majority of standard issues independently, contributing to documentation and runbooks
- 90 days: handling more complex issues, participating in RCA and improvement efforts
- Work Mode: Remote
- Shift: Rotational (247), including night shifts
- Experience: 2-5 years
- Education: B.Tech/B.E. or equivalent; graduates in any discipline with relevant experience considered
📌 Production Support Engineer - Ecommerce Infrastructure and
AI (Bengaluru)
🏢 Webscale
📍 Bengaluru