30 Sep
|
Mindsprint
|
Bengaluru
30 Sep
Mindsprint
Bengaluru
Skill - Production Support Engineer
Experience - 2 to 4 years
Location - Bengaluru
Mode of Working - Work from Office Only
Significant Note - Looking for candidate who can join us within 20 days.
Position Specific Requirements - Must have skills:
- 3+ years of experience in Level 2 Application support
- Expertise in application and backend issue triage & analysis. Should have analyzed failures, error logs, done basic analysis, gather inputs and coordinate with Level 3
- Own day-to-day product operations for cloud-hosted applications, including proactive monitoring, incident triaging, outage management, and resolution while ensuring adherence to defined SLAs.
- Demonstrate strong analytical capabilities to diagnose issues, drive incidents to closure, and adapt to evolving tools and techniques during real-time incident response.
- Having worked upon support layers L1, L2, L2.5, worked closely with the engineering L3 teams, Product teams and the Business Stakeholders
- Troubleshooting knowledge in AWS, containerized workloads, and distributed microservices architecture.
- Understanding the root causes, restoration of the services and final reporting
- Monitor application and infrastructure using
Datadog/Logicmonitor/Newrelic/Grafana or any monitoring tool
- Usage of multiple tools and techniques to analyse the production issues across Applications, backend, data and infrastructure layers
- Correlate logs, traces and metrics to isolate the failures across multiple services, understand the impact on business and generate high-quality artifacts aiding L3 investigations
- Support and lead release management, operational readiness,
deployment support and postproduction deployment validations
- Collaborate with the Stakeholders, Project teams and Vendors, handle the incident triages and meetings ensuring SLA adherences, escalation flows and operational excellence
- Good command in Microsoft Office suite
- Willing to work in rotational based 24X7 multiple Shifts
- Willing to learn and implement automation or AIOps Features
Monitoring and Observability: Datadog, ElasticSearch ELK / Open Search Dashboards, AWS Cloudwatch
Cloud & Infra: AWS, Lambda, API Gateway, Open Search, Athena, Web services, API Management, AWS Cognito, EKS, SQS, VPC/SG/ELB , DynamoDB , Kinesis (Data Stream, Firehose) , SNS
Incident Management ITSM tools: Jira, ServiceNow
Databases: PostgreSQL, Mongo No SQL, SQL
Good to have skills:
Familiarity with Gitlab/Github, Jenkins, CI/CD Pipelines or any source code version control.
Any Automation Skillset / Willing to learn Automation tools
Webservices and API Management
Experience in Change Management Process
Connected car features understanding
What will be the roles and responsibilities?
- Incident handling, Issue analysis, Investigate & Gather necessary data/insights
Incident escalation based on workflows.
- Own L2 incident handling: triage, isolate, restore,
and document; run bridges for P1/P2 and steer the right resolver group.
- Monitor & act on CloudWatch/Datadog/ELK dashboards & auto‑alerts using runbooks; tune noise vs. signal.
- Deep‑dive production issues across Lambda/API Gateway/OpenSearch/Athena; create actionable debug artifacts (timelines, queries, log snippets, correlation IDs). Partner with L3: supply high‑quality inputs (repro steps, traces, request samples, dashboards), verify fixes, and drive post‑incident actions.
- Operational hygiene: maintain runbooks, workflows, SOPs, and knowledge articles; contribute to SLA reporting and weekly ops reviews.
- Cross‑geo coordination: collaborate with business, SupportDesk, QA, SecOps, vendors; close tickets end‑to‑end.
- Continuous improvement: raise kaizens to reduce toil (alerts, auto‑remediations, dashboards, cost/perf optimizations in Athena/OpenSearch).
- In case of priority incidents, bridge the calls with multiple stakeholders and ensure resolution stakeholder is identified and assigned to incident.
- Monitoring Dashboard of cloud applications.
- Monitoring and acting based on the auto-alerts generated from system using runbooks.
- Monitoring and Handling of the slack Datadog alerts.
- Report creation and publishing to required stake holders.
- Responsible for maintaining the procedures/run books, workflows, work instructions used by the ops team. Fulfil any ad-hoc data or report request queries from different functional groups
Shifts
Rotational 24×7 (multiple shifts), including weekends oncall / holidays support as rostered.
📌 Production Support Engineer (Bengaluru)
🏢 Mindsprint
📍 Bengaluru