01 Oct
|
Mindsprint
|
Bengaluru
01 Oct
Mindsprint
Bengaluru
Skill Cloud Engineer
Experience – 3 to 4 years
Location – Bengaluru
Mode of Working – Hybrid
Important Note - Candidate who can join us within 15 days are preferred.
WHAT ARE WE LOOKING FOR?
We are looking for a hands-on Operations/SRE engineer who keeps production AWS workloads healthy and predictable, own alerts end to end, treat SLAs as a commitment you are accountable for, and know exactly when to escalate instead of holding on to an issue. Just as key, you will reduce the volume of repetitive work through automation, so the team spends its time on real problems instead of the same manual steps.
Must have qualification –
- Minimum 3 years in IT, with at least 2 years hands-on in AWS operations or production support, and availability for a 24x7 rotational shift.
- Hands-on triage across S3, Glue, EMR, Athena, Lake Formation, Lambda, Kinesis, SQS/SNS, Step Functions, ECS/EKS, ALB/NLB, API Gateway, RDS, DynamoDB, Redshift, and IAM/KMS/VPC.
- Daily working use of Datadog and CloudWatch - building dashboards, configuring monitors, searching logs, and following traces to isolate a failing component.
- Running the incident lifecycle in ServiceNow or Jira with paging via PagerDuty/DataDog including correct priority classification and escalation at the right SLA threshold.
- Authoring RCAs - accurate timeline,
trigger separated from contributing factors, corrective actions with owners and dates.
- Writing Python and Bash automation and modifying Terraform for routine infrastructure changes.
Good to have -
- ELK/OpenSearch for log analytics and AWS X-Ray or equivalent distributed tracing.
- Self-healing automation, such as EventBridge rules triggering Lambda remediation for known failure modes.
WHAT WILL BE THE ROLES AND RESPONSIBILITIES? 1.Act as first responder for production alerts on shift - assess impact, isolate the fault across application, infrastructure, dependency, and data layers, and drive it to mitigation.
2.Build and tune monitors, dashboards, and thresholds so alerts are actionable; systematically eliminate false positives and close observability gaps.
3.Track SLA/SLO adherence and escalate to engineering with evidence - symptoms, scope, timeline, and what you have ruled out.
4.Write RCAs for significant incidents and drive the corrective actions to closure.
5.Author and maintain runbooks so a known issue is handled identically by anyone on any shift.
6.Identify recurring manual work and eliminate it with Python, Bash, or Terraform automation.
7.Provide clear stakeholder updates during high-severity events and run disciplined shift handovers.
📌 Cloud Engineer (Bengaluru)
🏢 Mindsprint
📍 Bengaluru