17 Sep
|
Robotico Digital
|
Bengaluru
17 Sep
Robotico Digital
Bengaluru
Key Responsibilities
Operations Leadership & 24/7 Support Management
- Own and manage 24/7 support operations for the Atlas5 platform across all production environments
- Define and enforce SLAs, escalation paths, on-call rotations, and incident response processes
- Lead RCA (Root Cause Analysis) culture ensure all critical incidents have completed RCAs with permanent fixes before closure
- Drive continuous improvement in MTTR (Mean Time to Resolve) and MTTD (Mean Time to Detect) across the team
- Partner with engineering, product, and infrastructure teams to reduce recurring incidents and improve platform stability
People Management
- Directly manage a team of Dev Ops, SRE, and cloud operations engineers across multiple shifts
- Own hiring, onboarding, performance management, and career development for direct reports
- Foster a culture of accountability, ownership, and continuous learning within the operations team
- Balance workload and capacity planning manage staffing for 24/7 coverage without burnout
- Act as a mentor and escalation point for the team during critical production incidents
Azure & Cloud Operations
- Oversee Azure infrastructure operations including compute, networking, storage, and platform services
- Drive Azure cost governance work with the team to monitor spend, identify optimisation opportunities, and enforce cost controls
- Manage Savings Plans, Reserved Instances, and cloud financial operations (Fin Ops)
- Ensure production environments are secure, compliant, and aligned with Azure best practices
- Own Azure subscription governance, access controls, and RBAC policies across all environments
Dev Ops & Deployment Operations
- Oversee all deployment operations including release management, change management, and rollback procedures
- Drive deployment velocity improvement identify and eliminate bottlenecks in the release pipeline
- Ensure build artifacts, configurations, and environment-specific settings are correctly managed across Admin Portal, Client Portal, and all Atlas5 deployments
- Work with Engineering to resolve structural deployment issues (DB error handling, GFA configuration, build artifact management)
- Champion CI/CD best practices enforce branching strategies, pipeline governance, and automated testing gates
ETL & Data Feeds Operations
- Manage and oversee ETL pipeline operations ensure reliable data ingestion, transformation, and delivery across all data feeds
- Monitor data feed health proactively identify and resolve failures, latency issues, and data quality problems
- Work with data engineering and product teams to define SLAs for critical data feeds and ensure they are met
- Manage file exchange operations including FTP/SFTP, partner data feeds, and inbound file processing pipelines
- Ensure end-to-end data pipeline monitoring is in place and alerts are actionable
Monitoring & Observability
- Own the operational monitoring framework across all 5 layers: Infrastructure, Data Tier, Application, Integration (EDA), and File Exchange
- Ensure all alerts are configured, tuned, and routed correctly every alert triggers a Jira incident and Dev Ops channel notification
- Drive the implementation of proactive monitoring to move from reactive to predictive operations
- Review and approve monitoring parameters, alert thresholds, and escalation policies
Required Experience & Skills
Core Requirements
- 14 16 years of total experience in cloud / IT operations with a robust technical foundation
- Proven experience managing and leading 24/7 support and operations teams in enterprise environments
- Strong people management skills experience hiring, developing, and retaining engineering talent
- Demonstrated experience running operations at scale for SaaS or fintech platforms
Technical Skills
- Azure (Advanced): AKS, App Services, Azure Monitor, Azure Dev Ops, Log Analytics, Cost Management, RBAC, Networking
- Dev Ops & CI/CD: Azure Dev Ops Pipelines, YAML, Git, release management, branching strategies, artifact management
- ETL Operations: Experience managing ETL pipelines, data ingestion workflows, and transformation jobs in production environments
- Data Feeds: Hands-on experience managing FTP/SFTP file exchange, partner data feeds, inbound file monitoring, and SLA management
- Monitoring tools: Azure Monitor, Application Insights, Log Analytics, Prometheus, Grafana, Dynatrace, or equivalent
- Scripting: Power Shell, Python, or Bash for automation and operational tooling
- ITSM: Jira Service Management or equivalent incident, problem, and change management
Nice to Have
- Experience in fintech, wealth management, or regulated financial services environments
- Familiarity with Microsoft Fabric, Cosmos DB, or Azure SQL at an operational level
- Fin Ops certification or experience with Azure cost optimisation and reservation management
- Experience implementing Open Telemetry-based observability frameworks
📌 Operations Manager - DevOps/Azure & Engineering (Bengaluru)
🏢 Robotico Digital
📍 Bengaluru