24 Sep
|
OvalEdge
|
Hyderabad
24 Sep
OvalEdge
Hyderabad
Location: Hyderabad
Experience: 12+ years, including 5+ years in leadership roles
About OvalEdgeOvalEdge provides an enterprise Data Intelligence platform to global customers. The product is deployed through cloud-hosted SaaS, customer-managed cloud, and licensed on-premises models. These environments require strong availability, security, scalability, and compliance practices.
The RoleThe Senior DevOps Manager is responsible for infrastructure, platform engineering, reliability, and production operations across all supported deployment models.
The role combines team leadership with involvement in key technical decisions and operational issues. It also represents DevOps in customer discussions and works with executive and cross-functional stakeholders.
What You’ll DoDevOps Strategy &
- Governance
- Define and maintain the DevOps roadmap in alignment with engineering, product, security, customer, and business priorities.
- Establish standards for infrastructure, CI/CD, observability, incident management, disaster recovery, documentation, and operational readiness.
- Track and report availability, deployment performance, automation coverage, security posture, incident trends, and infrastructure costs.
- Identify operational risks and technical debt and track agreed remediation actions.
Cloud & • Infrastructure Management
- Manage infrastructure across AWS, Azure, customer-cloud, and licensed on-premises deployments.
- Standardize deployments across virtual machines, container platforms, and Kubernetes.
- Lead Infrastructure-as-Code and configuration automation using AWS CDK, Terraform, Ansible, or equivalent tools.
- Establish reusable infrastructure blueprints and self-service provisioning capabilities.
- Ensure infrastructure meets scalability, availability, security, capacity, performance, and cost requirements.
CI/CD & • Platform Engineering
- Define standards for CI/CD, artifact management, configuration management, and releases.
- Automate build, testing integration, security scanning, deployment, validation, and rollback.
- Implement rolling, blue-green, and canary deployment strategies where appropriate.
- Improve deployment frequency, lead time, change failure rate, and recovery time.
- Develop internal platform and self-service capabilities that reduce manual work and improve consistency.
Reliability, Observability & • Incident Management
- Establish an observability framework for infrastructure, applications, containers, APIs, and AI-enabled services.
- Standardize metrics, logs, distributed tracing, dashboards, synthetic monitoring, and alerting.
- Define SLIs, SLOs, service ownership, on-call processes, alert thresholds, and escalation paths.
- Implement centralized monitoring for service health and dependencies.
- Improve alert quality by reducing noise and introducing correlation and automated remediation where appropriate.
- Ensure critical incidents undergo root-cause analysis and that corrective actions are tracked.
- Maintain runbooks, operational procedures, handover documents, and incident communication templates.
AI Platform Operations & • AIOps
- Define operational standards for deploying, monitoring, scaling, securing, and supporting AI-enabled workloads.
- Ensure AI services follow approved CI/CD, observability, security, availability, rollback, and incident-management practices.
- Establish monitoring for availability, latency, failures, resource consumption, usage, and cost.
- Evaluate and implement approved AIOps use cases such as anomaly detection, alert correlation, capacity forecasting, and incident analysis.
- Establish review and access controls for AI-generated scripts and automated operational actions.
Disaster Recovery, Security & • Compliance
- Define backup, restoration, regional failover, and recovery requirements based on agreed RTO and RPO targets.
- Conduct periodic disaster recovery exercises and track identified gaps.
- Establish standards for IAM, SSO, privileged access, secrets management, encryption, network security, and infrastructure hardening.
- Include vulnerability scanning, container security, and compliance checks in CI/CD and infrastructure workflows.
- Work with Security and Compliance teams on SOC 2, ISO 27001, and customer requirements.
Cost Optimization
- Manage cloud cost reporting, allocation, forecasting, and optimization.
- Implement right-sizing, environment consolidation, scaling policies, commitment plans, and removal of unused resources.
- Review infrastructure and AI-platform decisions for cost impact without compromising reliability or security.
- Provide regular updates to leadership on costs, savings, anomalies, and optimization plans.
Customer Engagement & • Technical Communication
- Represent DevOps during pre-sales discussions, security reviews, architecture reviews, technical evaluations, and due diligence.
- Assess customer deployment requirements and explain the available deployment options.
- Communicate infrastructure prerequisites, responsibilities, dependencies, limitations, security controls, scalability, and cost implications.
- Lead DevOps communication during post-sales planning, onboarding, environment provisioning, deployment, validation, and production handover.
- Manage technical communication for upgrades, migrations, maintenance activities, infrastructure changes, and operational reviews.
- Participate in incident and escalation calls and communicate the impact, mitigation, recovery status, risks, and next steps.
- Present technical information clearly to architects, security teams, project teams, business stakeholders, and executives.
- Ensure customer commitments are technically validated and internally aligned.
- Work with Sales, Customer Success, Engineering, and Support to maintain consistent communication throughout the customer lifecycle.
- Use recurring customer feedback to improve automation, documentation, platform capabilities, and operating processes.
People Leadership & • Cross-Functional Collaboration
- Lead and develop DevOps engineers, leads, and managers working across cloud, platform, automation, and operations.
- Define team goals, ownership, performance measures, development plans, and succession plans.
- Promote automation, accountability, documentation, knowledge sharing, and continuous improvement.
- Work with Engineering and QA on deployment architecture, release readiness, scalability, and quality gates.
- Coordinate with Product, Security, Compliance, Sales, and Customer Success on roadmap initiatives, audits, customer commitments, and deployment requirements.
- Provide regular updates to leadership on operational risks, incidents, costs, investment needs, and improvement initiatives.
What You’ll BringRequired
- 10+ years of experience in DevOps, cloud infrastructure, platform engineering, automation, or related roles in a product development organization.
- At least 5 years of experience leading DevOps or platform engineering teams.
- Experience in hiring, goal setting, performance management, mentoring, and career development.
- Hands-on experience managing infrastructure across enterprise SaaS, customer-cloud, and licensed on-premises deployments.
- Strong AWS experience and working knowledge of Azure and hybrid environments.
- Production experience with Docker, ECS, EKS, AKS, or Kubernetes.
- Strong knowledge of CI/CD, Infrastructure-as-Code, configuration management, container orchestration, and deployment automation.
- Experience with observability, incident management, disaster recovery, high availability, and production operations.
- Experience defining operational standards for AI-enabled or data-intensive services.
- Strong knowledge of IAM, secrets management, DevSecOps, vulnerability management, SOC 2, and ISO 27001.
- Experience presenting infrastructure architecture and deployment approaches to enterprise customers.
- Ability to manage technical discussions during pre-sales, onboarding, production operations, maintenance, and escalations.
- Strong written communication, presentation, technical articulation, and stakeholder-management skills.
- Ability to explain technical risks, dependencies, trade-offs, and recommendations to technical and non-technical audiences.
Nice to Have
- Experience building internal developer platforms and self-service infrastructure capabilities.
- Experience with AIOps for anomaly detection, alert correlation, capacity forecasting, or automated remediation.
- Experience with AWS Control Tower or enterprise multi-account governance.
- Exposure to Data Intelligence, Data Governance, Data Lineage, Analytics, or Data Management platforms.
- Experience supporting large or globally distributed enterprise customers.
- Experience managing geographically distributed engineering or operations teams.
Education
- Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related field.
- Master’s degree is preferred.
- AWS Professional-level, CKA/CKS, ITIL, or equivalent certifications are an advantage.
Why OvalEdge
- Lead infrastructure, reliability, and platform operations for an enterprise product.
- Work across SaaS, customer-cloud, and on-premises deployment models.
- Work directly with enterprise customers and executive stakeholders.
- Work in a flexible, energetic environment based in Hyderabad.
📌 DevOps Manager (Hyderabad)
🏢 OvalEdge
📍 Hyderabad