24 Aug
|
BridgeNext
|
Pune
We are looking for a 247 NOC Specialist to ensure the availability, performance, and reliability of production systems. This role is hands-on across monitoring/observability, incident response, troubleshooting, and automation, working closely with engineering and infrastructure teams to reduce downtime and improve operational excellence.
Shift / Support Model: 247 rotational shifts (including nights/weekends) with on-call participation as required.
Key Responsibilities
- Monitor applications and infrastructure using New Relic, Datadog, Grafana and related observability tooling; maintain dashboards and actionable alerting
- Provide L1/L2 incident response in a 247 environment; triage alerts, restore service quickly, and manage escalations
- Perform deep troubleshooting across Linux systems, Kubernetes workloads, infrastructure components, and network paths
- Conduct log analysis using Newrelic/ELK (and/or similar platforms) to identify patterns, correlate events, and support root cause analysis
- Manage infrastructure changes using Terraform and follow Infrastructure-as-Code practices (review, version control, rollback readiness)
- Support Kubernetes platform operations by assisting with deployments, performing cluster/service health checks, executing scaling and recycling activities, monitoring capacity and performance, and troubleshooting issues
- Maintain clear runbooks, SOPs, and shift handover notes; ensure knowledge is captured and reusable
- Partner with engineering and cloud/infrastructure teams to improve reliability through post-incident reviews, problem management, and continuous improvements to observability
Must Have Skills:
- Bachelors degree (B.Tech/B.E., MCA) or equivalent practical experience
- 3–6 years of experience in SRE / NOC / Production Support / DevOps / Infrastructure Operations
- Experience working in a shift-based operations environment with strong ownership and urgency
- Ability to document clearly (runbooks, post-incident notes) and collaborate effectively with cross-functional teams.Monitoring & Observability: Current Relic, Datadog, Grafana; strong alert triage and dashboarding skills
- Linux: administration fundamentals, process/service troubleshooting, permissions, performance basics
- Infrastructure as Code: Terraform (hands-on)
- Containers: Kubernetes (workload troubleshooting, cluster concepts)
- Networking: TCP/IP basics, DNS, HTTP/HTTPS, load balancing concepts, connectivity troubleshooting
- Log Analysis: ELK (or equivalent), querying/correlation for RCA support
Preferred Skills:
- Cloud infrastructure fundamentals (AWS/Azure/GCP)
- Good communication skills: clear incident updates, shift handovers, and stakeholder coordination
Professional Skills:
- Solid written, verbal, and presentation communication skills
- Strong team and individual player
- Maintains composure during all types of situations and is collaborative by nature
- High standards of professionalism, consistently producing high quality results
- Self-sufficient, independent requiring very little supervision or intervention
- Demonstrate flexibility and openness to bring creative solutions to address issues
Bridgenext is an Equal Opportunity Employer
📌 Noc Specialist / SRE engineer (Pune)
🏢 BridgeNext
📍 Pune