06 Aug
|
Recognized
|
Navi Mumbai
06 Aug
Recognized
Navi Mumbai
Job Title: Site Reliability Engineer (SRE) / L1
Monitoring Engineer
Job Summary
We are seeking a proactive and technically driven SRE /
L1 Monitoring Engineer with 1 to 3 years of experience to join our core
digital infrastructure operations team. In this role, you will serve as the
first line of defense ensuring the high availability, security, and performance
of critical financial services and digital banking applications. You will be
responsible for real -time system monitoring, tracking alerts across modern
observability stacks, performing initial triage on infrastructure bottlenecks,
and managing API traffic performance. This is an excellent opportunity for an
early -career engineer looking to scale their skills in a high -volume, secure
cloud infrastructure environment. [1]
Key Responsibilities
L1 Infrastructure Monitoring & Alerts
- Real -time
Surveillance: Actively monitor production environments, enterprise
dashboards, and telemetry feeds using toolsets like Datadog, Dynatrace,
and Grafana to spot anomalies before they impact end -users. [1, 2, 3]
- Alert
Triage: Acknowledge, validate, and categorize incoming infrastructure,
database, and application alerts generated by Prometheus and
application performance monitoring (APM) agents using predefined Standard
Operating Procedures (SOPs). [1, 2, 3, 4, 5]
- Incident
Escalation: Document incident details clearly in the ticketing system
and swiftly escalate unresolved P1/P2 issues to L2 engineers or
specialized DevOps teams with complete log snippets and context.
Application Delivery & API Traffic Management
- Nginx
Operations: Monitor web server logs, verify reverse proxy
configurations, and troubleshoot basic traffic routing or SSL/TLS
certificate errors. [1, 2, 3]
- API
Gateways: Use Apigee to monitor API proxy performance, track
error rates (5xx/4xx codes), track latency spikes, and check developer
portal connectivity. [1, 2, 3, 4]
- Kubernetes
Support: Monitor cluster health, inspect pod statuses, view
application logs using kubectl, and track resource usage (CPU/Memory
limits). [1, 2]
Cloud Operations & Reliability
- GCP
Monitoring: Utilize Google Cloud logging, native monitoring tools, and
integrated observability dashboards to check the health of virtual
machines, storage, and networking layers. [1, 2, 3, 4]
- Health
Checks: Perform routine daily morning sanity checks and
post -deployment validation steps for critical banking services.
- Runbook
Execution: Execute automated or manual scripts to restart failed
services, clear disk space, or cycle pods safely in staging and production
environments.
- Experience: 1 to 3 years of hands -on experience in an L1 Support, Infrastructure
Monitoring, or Junior SRE role.
Dynatrace, Prometheus, and Grafana.
balancing, log analysis).
- Containerization: Foundational knowledge of Kubernetes (K8s) (understanding pods,
logs and kubectl get pods).
services and cloud monitoring concepts.
monitoring traffic flow and checking endpoint health.
for navigating directories and tailing logs.
Flexibility: Readiness to work in a 24/7 rotating shift model (including night shifts and weekends) to maintain uninterrupted banking