18 Aug
|
Blue Cloud Softech Solutions
|
Bengaluru
18 Aug
Blue Cloud Softech Solutions
Bengaluru
Title – Engineer, Production Support
Experience – 3 to 5 years
Location – Bengaluru ( Look profiles from Bangalore only )
Mode of Working – WFO Only
Time of Hire – 1 to 2 Weeks
ShiftsRotational 24×7 (multiple shifts), including weekends oncall / holidays support as rostered.
What are we looking for?
An L2 System & Application Support Engineer with solid production operations experience across applications hosted on cloud infrastructure. As a Technical Operations Engineer, you are responsible for the monitoring, incident management, reporting, release management and continuous operational improvement. You’ll partner with L3/engineering teams, stakeholders, product teams and clients to ensure highest availability of the services/features.
You will be responsible for root cause identification ensuring implementation of processes to prevent the recurrences. You’ll also help us continuously improve observability, runbooks, and operational readiness.
The candidate,
Must have skills
3+ years of experience in Level 2 Application support
Expertise in application and backend issue triage & analysis. Should have analyzed failures, error logs, done basic analysis, gather inputs and coordinate with Level 3
Own day-to-day product operations for cloud-hosted applications, including proactive monitoring, incident triaging, outage management, and resolution while ensuring adherence to defined SLAs.
Demonstrate strong analytical capabilities to diagnose issues, drive incidents to closure, and adapt to evolving tools and techniques during real-time incident response.
Having worked upon support layers – L1, L2, L2.5, worked closely with the engineering L3 teams, Product teams and the Business Stakeholders
Troubleshooting knowledge in AWS, containerized workloads, and distributed microservices architecture.
Understanding the root causes,
restoration of the services and final reporting
Monitor application and infrastructure using Datadog/Logicmonitor/Newrelic/Grafana or any monitoring tool
Usage of multiple tools and techniques to analyse the production issues across Applications, backend, data and infrastructure layers
Correlate logs, traces and metrics to isolate the failures across multiple services, understand the impact on business and generate high quality artifacts aiding L3 investigations
Support and lead release management, operational readiness, deployment support and postproduction deployment validations
Collaborate with the Stakeholders, Project teams and Vendors, handle the incident triages and meetings ensuring SLA adherences, escalation flows and operational excellence
Good command in Microsoft Office suite
Willing to work in rotational based 24X7 multiple Shifts
Willing to learn and implement automation or AIOps Features
Tools (Related should be fine too)
Monitoring and Observability:
Datadog, ElasticSearch ELK / Open Search Dashboards, AWS Cloudwatch
Cloud & Infra:
AWS, Lambda, API Gateway, Open Search, Athena, Web services, API Management, AWS Cognito, EKS, SQS
Incident Management ITSM tools
Jira, ServiceNow
Databases
PostgreSQL, Mongo No SQL, SQL
Good to have skills:
Familiarity with Gitlab/Github, Jenkins, CI/CD Pipelines or any source code version control.
Any Automation Skillset / Willing to learn Automation tools
Webservices and API Management
Experience in Change Management Process
Connected car features understanding
What will be the roles and responsibilities?
Incident handling, Issue analysis, Investigate & Gather necessary data/insights
Incident escalation based on workflows
Own L2 incident handling: triage, isolate, restore, and document; run bridges for P1/P2 and steer the right resolver group.
Monitor & act on CloudWatch/Datadog/ELK dashboards & auto‑alerts using runbooks; tune noise vs. signal.
Deep‑dive production issues across Lambda/API Gateway/OpenSearch/Athena; create actionable debug artifacts (timelines, queries, log snippets, correlation IDs).
Partner with L3: supply high‑quality inputs (repro steps, traces, request samples, dashboards), verify fixes, and drive post‑incident actions.
Operational hygiene: maintain runbooks, workflows, SOPs, and knowledge articles; contribute to SLA reporting and weekly ops reviews.
Cross‑geo coordination: collaborate with business, SupportDesk, QA, SecOps, vendors; close tickets end‑to‑end.
Continuous improvement: raise kaizens to reduce toil (alerts, auto‑remediations, dashboards, cost/perf optimizations in Athena/OpenSearch).
In case of priority incidents , bridge the calls with multiple stakeholders and ensure resolution stakeholder is identified and assigned to incident
Monitoring Dashboard of cloud applications
Monitoring and acting based on the auto-alerts generated from system using runbooks
Monitoring and Handling of the slack Datadog alerts
Report creation and publishing to required stake holders
Responsible for maintaining the procedures/run books, workflows, work instructions used by the ops team
Fulfil any ad-hoc data or report request queries from different functional groups.
📌 Production Support (Bengaluru)
🏢 Blue Cloud Softech Solutions
📍 Bengaluru