14 Aug
|
Datavail Infotech
|
Mumbai
14 Aug
Datavail Infotech
Mumbai
Description
Job Title: Lead Senior Production Support / Operations Engineer
Education: Any Graduate
Experience: 8+years
Job Location: Mumbai
Role Summary
We are seeking a highly Senior Production Support / Operations Engineer to represent and operationally support our in-house monitoring and observability platform across enterprise database environments.
This role acts as the critical operational bridge between:
- The in-house monitoring product development and QE team(s),
- Internal ITSM/ticketing teams using ServiceNow,
- Multiple enterprise DBA organizations across:
- Oracle Database
- Open-source database platforms (MySQL, PostgreSQL, MongoDB, Cassandra, etc.)
- Microsoft SQL Server
The ideal candidate combines strong operational troubleshooting skills, production support leadership, monitoring expertise, incident coordination capabilities, and excellent cross-team collaboration skills in large-scale enterprise environments.
This is a senior leadership position: the candidate must bring substantial experience leading Production Support Engineers, combined with exceptional communication skills for direct, confident engagement with customers and senior stakeholders of in-house management.
Key Responsibilities
Production Operations & Monitoring Support
- Provide operational ownership and production support for the in-house enterprise monitoring platform.
- Monitor health, performance, alerting quality, and operational stability of monitoring services.
- Analyze monitoring gaps, false positives, missed alerts, and operational inefficiencies.
- Ensure monitoring coverage across Oracle, Open-source, and SQL Server database environments.
Incident & Escalation Management
- Act as the operational point-of-contact during production incidents involving monitoring failures, alerting gaps, or infrastructure issues.
- Coordinate incident triage across:
- DBA teams,
- Monitoring development teams,
- Infrastructure teams,
- Service management teams.
- Drive bridge calls and ensure effective stakeholder communication during critical outages.
- Perform root cause analysis (RCA) and post-incident operational reviews.
ServiceNow & Ticket Workflow Coordination
- Work with ServiceNow for:
- Operational escalations,
- Service requests.
- Review ticket quality and ensure operational accuracy of issue classification and routing.
- Improve ticket workflows between DBA teams and monitoring platform support teams.
- Collaborate with internal support organizations to streamline escalation processes.
Cross-Functional DBA Collaboration
- Collaborate closely with enterprise DBA teams supporting:
- Oracle Database
- MySQL
- PostgreSQL
- MongoDB
- Apache Cassandra
- Microsoft SQL Server
- Cloud services (AWS, AZURE, GCP)
- Understand operational monitoring requirements specific to each database technology.
- Work with DBAs to validate alert thresholds, event correlation, and monitoring accuracy.
- Serve as the operational liaison between DBAs and monitoring team developers and QE.
Operational Excellence & Reliability Engineering
- Identify recurring operational pain points and recommend automation opportunities.
- Improve alert quality, event correlation, and monitoring reliability.
- Participate in operational readiness reviews for new monitoring features.
- Help define operational standards, playbooks, and escalation procedures.
Monitoring & Observability Engineering
- Support enterprise observability initiatives involving:
- Metrics,
- Events,
- Alerting,
- Dashboards,
- Health monitoring,
- Incident correlation.
- Work with both commercial and in-house monitoring systems.
- Analyse operational telemetry to identify systemic reliability concerns.
DevOps & CI/CD Enablement
- Collaborate with engineering teams to improve CI/CD pipelines.
- Implement deployment strategies (blue-green, canary, rolling updates).
- Advocate for reliability-focused design patterns.
Security & Compliance
- Ensure infrastructure adheres to security standards and compliance requirements.
- Participate in vulnerability assessments and remediation.
Required Technical Skills
- Solid production support and operations experience in enterprise environments.
- Substantial experience leading Production Support Engineers, in a senior/lead or team-management capacity.
- Strong experience with cloud platforms (AWS, Azure, and GCP).
- Substantial, practical hands-on experience with both Windows and Linux operating systems, and sound working familiarity with Databricks and Microsoft Fabric for analytics workloads.
- Expertise in monitoring & observability tools (e.g., Prometheus, Grafana, Datadog, or in-house tools).
- Working knowledge of:
- ServiceNow
- Incident workflows,
- Escalation management,
- Operational support models.
- Exposure to database technologies including:
- Oracle Database
- Microsoft SQL Server
- MySQL
- PostgreSQL
- NoSQL ecosystems preferred.
- Strong understanding of:
- Windows and Linux systems,
- Infrastructure monitoring,
- Alerting concepts,
- Production operations.
- Experience supporting 24x7 enterprise production environments.
Preferred Qualifications
- Experience working with in-house monitoring or observability product teams.
- Familiarity with SRE/DevOps operational practices.
- Exposure to enterprise event management systems.
- Knowledge of automation/scripting (Python, Shell, PowerShell).
- Experience handling high-severity production incidents.
Critical Non-Technical Skills An ideal candidate must demonstrate:
Operational Intuition
- Ability to detect operational anomalies early.
- Strong troubleshooting instinct and pattern recognition.
Fearless Communication
- Ability to speak confidently during incidents and escalations.
- Comfortable engaging customers and senior stakeholders of in-house management, along with multiple technical teams.
- Exceptional written and verbal communication skills, able to present operational status and risk directly to customers and senior leadership with clarity and confidence.
Cross-Team Collaboration
- Ability to coordinate effectively across DBA teams, support organizations, and development groups.
Calmness Under Pressure
- Structured decision-making during high-severity incidents.
Ownership Mindset
- Drives issues to closure rather than relying solely on assigned ownership boundaries.
Investigative Curiosity
- Continuously analyses why operational failures occur and how they can be prevented.
- Substantial track record leading, mentoring, and developing Production Support Engineers.
- Sets the standard for operational excellence and coaches junior/mid-level engineers toward it.
- Acts as an escalation point and mentor for less experienced engineers during high-severity incidents.
- Team Leadership & Mentorship
Preferred Qualifications
- Relevant Certifications in cloud platforms (AWS/Azure/GCP).
- Familiarity with SRE/DevOps operational practices.
- Familiarity with Databricks and Fabric domain.
- Exposure to enterprise event management systems.
- Knowledge of automation/scripting (Python, Shell, PowerShell).
📌 Lead Senior Production Support / Operations Engineer (Mumbai)
🏢 Datavail Infotech
📍 Mumbai