02 Aug
|
Zybisys Consulting Services
|
Bengaluru
02 Aug
Zybisys Consulting Services
Bengaluru
Role & responsibilities
L1 Operations Engineer
Location: Bangalore
Experience Required: 0 - 2 Years
Employment Type: Full-Time | Shift-Based (24x7 Rotational)
About Zybisys
Zybisys is a technology company that helps banks, financial institutions, and FinTech businesses build and run secure, reliable, and high-performance technology platforms. We work closely with some of India's leading stock brokers to manage their cloud infrastructure, cybersecurity, platform operations, and observability. With deep expertise in the Capital Markets domain, we focus on simplifying complex technology, improving operational resilience, and helping our customers innovate with confidence.
The L1 Operations Engineer is the first line of defence in Zybisys's 24x7 managed cloud operations. This role is responsible for continuous monitoring of multi-region Azure cloud and hybrid on-premises infrastructure, timely triage of alerts, first-line incident response, and ensuring all issues are accurately logged, prioritised, and escalated within SLA thresholds.
This is a shift-based operational role requiring strong attention to detail, disciplined runbook execution, and clear communication during incidents. The ideal candidate is technically curious, process-oriented, and comfortable working across cloud monitoring tools, ITSM platforms, and infrastructure dashboards in a fast-paced managed services environment.
Key Responsibilities
Infrastructure Monitoring & Alert Management
- Perform continuous 24x7 monitoring of Azure cloud and hybrid datacenter infrastructure using Prometheus, Grafana, Azure Monitor, and related observability dashboards.
- Triage incoming alerts assess severity, validate against known patterns, and determine whether to resolve at L1 or escalate to L2 within defined SLA thresholds.
- Execute approved runbooks and SOPs for all known alert categories; document actions taken for every incident with accurate timestamps and observations.
- Monitor health and availability of compute (VMs), storage, network links, VPN tunnels, and platform services across cloud and on-premises environments.
- Track sFlow and NetFlow dashboards for network traffic anomalies; flag unusual patterns to the L2 team for deeper investigation.
Incident Logging & ITSM Management
- Log all incidents, service requests, and alerts in the ITSM platform with complete and accurate details symptoms, affected components, priority, and initial actions taken.
- Update ticket status throughout the incident lifecycle; ensure no incident is left without a current status update beyond the defined response window.
- Coordinate with L2 engineers during escalations provide explicit handover notes including timeline, alert context, initial diagnostics, and business impact assessment.
- Follow the priority matrix strictly: P1 (15 min), P2 (30 min), P3 (4 hr), P4 (8 hr) response and escalation thresholds.
Platform & Service Health Checks
- Execute scheduled shift health checks across all managed platforms Azure resources, on-premises servers, network devices, security appliances, and employee services.
- Verify availability and performance of core services: DNS, DHCP, NTP, Active Directory, and M365 platform components.
- Monitor security platform dashboards (firewalls, EDR, proxy services) for health status and alert flags; escalate anomalies per defined procedures.
- Review Azure Cost Management dashboards for unusual consumption spikes and flag to the lead for review.
Routine Operations & Maintenance Support
- Execute scheduled batch jobs, backup verifications, replication checks, and housekeeping tasks as per the operational calendar.
- Support L2 and Specialist engineers during planned maintenance windows, patch cycles, and change activities providing monitoring coverage and rollback readiness.
- Validate post-change infrastructure health after every approved change; raise a flag immediately if anomalies are detected post-implementation.
Documentation & Knowledge Management
- Maintain precise shift handover reports open tickets, ongoing incidents, recent changes, and watch-items for the next shift.
- Contribute to the knowledge base by documenting recurring alert patterns, resolution steps, and workarounds for L1-resolvable issues.
- Flag gaps in runbooks or SOPs to the operations lead so that documentation is continuously improved.
Required Skills & Experience Cloud & Infrastructure Fundamentals
- 2 - 5 years of experience in IT operations, infrastructure support, or cloud managed services.
- Working knowledge of Microsoft Azure: Azure Portal navigation, VM status checks, resource monitoring, and basic troubleshooting using Azure Monitor and Log Analytics.
- Hands-on familiarity with Windows Server (2016/2019/2022) and Linux (RHEL/Ubuntu) service management, log file reading, process monitoring,
and basic fault diagnosis.
- Understanding of hybrid infrastructure models on-premises datacenter integrated with Azure cloud via ExpressRoute or VPN.
- Basic familiarity with storage concepts: disk performance thresholds, capacity monitoring, backup job status, and replication health checks.
Monitoring & Observability Tools
- Experience reading and interpreting Prometheus metrics and Grafana dashboards understanding panel thresholds, alert states, and trend data.
- Familiarity with Azure Monitor alerts and Log Analytics at a basic level understanding alert rules, severity levels, and affected resources.
- Ability to read sFlow or NetFlow traffic dashboards for high-level network health assessment.
- APM dashboard familiarity understanding response time trends, error rates, and service health indicators from tools such as Dynatrace, AppDynamics, or Azure Application Insights.
Networking Fundamentals
- Good understanding of core networking concepts: TCP/IP, DNS, DHCP, NTP, VLANs, and basic routing.
- Practical diagnostic skills: ping, traceroute, nslookup, netstat to perform first-level connectivity checks and provide meaningful diagnostics to L2.
- Familiarity with VPN tunnel health monitoring and firewall status dashboards at an operational level.
ITSM & Process Discipline
- Experience working with ITSM platforms (ServiceNow, Zoho Desk, Freshservice, or equivalent) for incident logging, ticket updates, and escalation workflows.
- Understanding of ITIL incident management concepts priority, impact, urgency, escalation paths, and SLA tracking.
- Ability to follow runbooks and SOPs precisely and consistently, including under pressure during major incident scenarios.
- Clear and concise written communication for incident tickets, shift handover notes, and escalation summaries.
Soft Skills & Work Style
- Comfortable working in a 24x7 rotational shift environment including night shifts, weekends, and public holidays.
- High attention to detail accurate logging, precise documentation, and consistent process adherence.
- Team-oriented with a proactive attitude toward learning and improving operational knowledge.
- Ability to stay calm and systematic during high-pressure P1/P2 incident scenarios.
Certification (Preferred)
Microsoft Azure
AZ-900 (Azure Fundamentals) | AZ-104 (Administrator working towards)
Service Management
ITIL v4 Foundation
Networking
CompTIA Network+ | Cisco CCNA (advantageous)
Monitoring
Grafana Certified Associate (advantageous)
Preferred candidate profile
📌 Operation engineer (Bengaluru)
🏢 Zybisys Consulting Services
📍 Bengaluru