03 Aug
|
Datavail
|
India
Job Title: Engineer - AWS, Azure, GCP, AKS
Experience: 6+ Years
Education: Any Graduate
Location: Mumbai
Job Description
Role Overview
This role acts as the critical operational bridge between:
- The in-house monitoring product development and QE team(s),
- Internal ITSM/ticketing teams using ServiceNow,
- Multiple enterprise DBA organizations across:
- Oracle Database
- Open-source database platforms (MySQL, PostgreSQL, MongoDB, Cassandra, etc.)
- Microsoft SQL Server
Key Responsibilities
Production Operations & Monitoring Support
- Provide operational ownership and production support for the in-house enterprise monitoring platform.
- Monitor health, performance, alerting quality, and operational stability of monitoring services.
- Analyze monitoring gaps, false positives, missed alerts, and operational inefficiencies.
- Ensure monitoring coverage across Oracle, Open-source, and SQL Server database environments .
Incident & Escalation Management
- Act as the operational point-of-contact during production incidents involving monitoring failures, alerting gaps, or infrastructure issues.
- Coordinate incident triage across:
- DBA teams,
- Monitoring development teams,
- Infrastructure teams,
- Service management teams.
- Drive bridge calls and ensure effective stakeholder communication during critical outages.
- Perform root cause analysis (RCA) and post-incident operational reviews.
ServiceNow & Ticket Workflow Coordination
- Work with ServiceNow for:
- Operational escalations,
- Service requests.
- Review ticket quality and ensure operational accuracy of issue classification and routing.
- Improve ticket workflows between DBA teams and monitoring platform support teams.
- Collaborate with internal support organizations to streamline escalation processes.
Cross-Functional DBA Collaboration
- Collaborate closely with enterprise DBA teams supporting:
- Oracle Database
- MySQL
- PostgreSQL
- MongoDB
- Apache Cassandra
- Microsoft SQL Server
- Cloud services (AWS, AZURE, GCP)
- Understand operational monitoring requirements specific to each database technology.
- Work with DBAs to validate alert thresholds, event correlation, and monitoring accuracy.
- Serve as the operational liaison between DBAs and monitoring team developers and QE.
Operational Excellence & Reliability Engineering
- Identify recurring operational pain points and recommend automation opportunities.
- Improve alert quality, event correlation, and monitoring reliability.
- Participate in operational readiness reviews for new monitoring features.
- Help define operational standards, playbooks, and escalation procedures.
Monitoring & Observability Engineering
- Support enterprise observability initiatives involving:
- Metrics,
- Events,
- Alerting,
- Dashboards,
- Health monitoring,
- Incident correlation.
- Work with both commercial and in-house monitoring systems.
- Analyse operational telemetry to identify systemic reliability concerns.
DevOps & CI/CD Enablement
- Collaborate with engineering teams to improve CI/CD pipelines.
- Implement deployment strategies (blue-green, canary, rolling updates).
- Advocate for reliability-focused design patterns.
Security & Compliance
- Ensure infrastructure adheres to security standards and compliance requirements.
- Participate in vulnerability assessments and remediation.
Required Technical Skills
- Strong production support and operations experience in enterprise environments.
- Strong experience with cloud platforms (AWS, Azure, or GCP).
- Expertise in monitoring & observability tools (e.g., Prometheus, Grafana, Datadog, or in-house tools).
- Working knowledge of:
- ServiceNow
- Incident workflows,
- Escalation management,
- Operational support models.
- Exposure to database technologies including:
- Oracle Database
- Microsoft SQL Server
- MySQL
- PostgreSQL
- NoSQL ecosystems preferred.
- Solid understanding of:
- Linux systems,
- Infrastructure monitoring,
- Alerting concepts,
- Production operations.
- Experience supporting 24x7 enterprise production environments.
Preferred Qualifications
- Experience working with in-house monitoring or observability product teams.
- Familiarity with SRE/DevOps operational practices.
- Exposure to enterprise event management systems.
- Knowledge of automation/scripting (Python, Shell, PowerShell).
- Experience handling high-severity production incidents.
Critical Non-Technical Skills
An ideal candidate must demonstrate:
Operational Intuition
- Ability to detect operational anomalies early.
- Solid troubleshooting instinct and pattern recognition.
- Fearless Communication
- Ability to speak confidently during incidents and escalations.
- Comfortable engaging senior stakeholders and multiple technical teams.
- Cross-Team Collaboration
- Ability to coordinate effectively across DBA teams, support organizations, and development groups.
- Calmness Under Pressure
- Structured decision-making during high-severity incidents.
- Ownership Mindset
- Drives issues to closure rather than relying solely on assigned ownership boundaries.
- Investigative Curiosity
- Continuously analyses why operational failures occur and how they can be prevented.
Preferred Qualifications
- Certifications in cloud platforms (AWS/Azure/GCP).
- Familiarity with SRE/DevOps operational practices.
- Exposure to enterprise event management systems.
- Knowledge of automation/scripting (Python, Shell, PowerShell).
📌 Product Support Engineer (India)
🏢 Datavail
📍 India