24 Aug
|
Net Connect Global
|
Bengaluru
24 Aug
Net Connect Global
Bengaluru
About the Role
Are you passionate about building highly available, scalable, and resilient systems?
We're looking for a Site Reliability Engineer (SRE) who enjoys solving production challenges, automating repetitive tasks, and improving platform reliability. You'll work alongside software engineers and infrastructure teams to ensure our mission-critical platforms remain reliable, performant, and resilient while driving automation across the production landscape.
If you're someone who enjoys debugging complex distributed systems, writing automation scripts, and improving operational excellence, we'd love to hear from you.
What You'll Be Doing
Own the reliability, stability, and availability of large-scale production platforms.
Monitor critical applications using tools like Prometheus, Grafana, Splunk, and Apica to proactively detect and resolve issues.
Troubleshoot production incidents across Linux servers, applications, databases, cloud infrastructure, and networking layers.
Participate in Incident, Problem, Change, Capacity, and Event Management activities while driving Root Cause Analysis (RCA).
Build automation using Python, Bash, or Shell scripting to eliminate manual tasks and improve operational efficiency.
Collaborate closely with Software Development and Infrastructure Engineering teams to improve system resilience and operational readiness.
Improve deployment, monitoring, alerting, and operational workflows through automation.
Participate in production releases,
operational readiness reviews, and on-call rotations supporting global business-critical applications.
Continuously identify opportunities to improve platform reliability, scalability, and performance.
Required Skills
We're looking for engineers with experience in:
Linux / Unix Administration
Site Reliability Engineering (SRE) or Production Support
Incident & Problem Management
Monitoring & Observability
- Prometheus
- Grafana
- Splunk
- Apica
Cloud Platforms (Azure preferred) Python / Bash / Shell Scripting
Automation using Ansible, GitHub, or similar tools
SQL troubleshooting (Sybase, DB2, Azure SQL or equivalent)
Distributed Systems, Microservices & Cloud-based Architectures
Nice to Have
Azure Networking
Azure Service Bus
Azure Virtual Machines
Azure SQL
Infrastructure Automation
CI/CD exposure
Reliability Engineering best practices
What Makes You Successful
Strong troubleshooting mindset
Passion for automation
Excellent debugging skills
Ability to work across infrastructure and application layers
Ownership mentality
Enjoy solving complex production problems
Comfortable working in a global support environment
Why Join Us?
Work on business-critical enterprise platforms
Exposure to large-scale distributed systems
Build automation that impacts production reliability
Collaborate with highly skilled engineering teams
Accelerate your career in Site Reliability Engineering
📌 SRE Engineer @ Investment Banking | Mumbai (Bengaluru)
🏢 Net Connect Global
📍 Bengaluru