Production Support Lead (Hyderabad)

Production Support Lead (Hyderabad)

27 Aug
|
Cloudxtreme
|
Hyderabad

27 Aug

Cloudxtreme

Hyderabad

Job Requirements

Technical Requirements

SRE & Production Engineering

- 8+ years of experience in Application Support, Site Reliability Engineering, Production Engineering, or Platform Operations.
- Minimum 2-4 years of experience managing technical teams in large-scale production environments.
- Hands-on experience supporting highly available, business-critical applications.

- Strong understanding of

- SRE principles

- Reliability Engineering

- Operational Excellence

- Error Budgets

- Service Level Indicators (SLIs)

- Service Level Objectives (SLOs)

- Availability Engineering

- Capacity Planning
- Expertise managing production incidents, service outages, and major incident bridges.

Authentication & Security Platforms Strong understanding and operational support experience in:

- OAuth 2.0
- OpenID Connect (OIDC)
- SAML
- Multi-Factor Authentication (MFA)
- Session Management
- Token Lifecycle Management
- API Authentication
- Okta
- Transmit Security
- Identity and Access Management (IAM)

Application Troubleshooting Hands-on troubleshooting expertise involving:

- Java applications
- .NET Core applications
- REST APIs
- Microservices
- Batch processing applications
- Application latency issues
- Memory leaks
- Thread contention
- Authentication failures
- Service-to-service communication issues

Cloud Platforms Experience managing applications hosted on:

- Google Cloud Platform (Preferred)
- Microsoft Azure
- Amazon Web Services
- PCF (Pivotal Cloud Foundry)

Areas of expertise:
- Application hosting
- Networking
- Cloud security
- Service reliability
- Monitoring
- Disaster recovery

Containers & Microservices

Strong experience with

- Kubernetes
- Docker
- Helm
- Service Mesh concepts
- Container troubleshooting
- Pod lifecycle management
- Ingress controllers
- Service discovery
- Scaling strategies

Observability & Monitoring Hands-on expertise in:

- Splunk
- AppDynamics
- Datadog
- Grafana
- ThousandEyes
- ITRS Geneos
- MoogSoft
- AppMetrics

Experience with:
- Log Analytics
- Distributed Tracing
- Metrics Monitoring
- Synthetic Monitoring
- Alert Tuning
- Dashboard Creation
- Event Correlation

Database & Messaging

Strong understanding of

- MongoDB
- PostgreSQL
- SQL Query Optimization
- Database Performance Tuning
- Replication Issues
- Database Connectivity Troubleshooting




- Kafka Monitoring and Troubleshooting

Automation & Scripting Strong coding and automation skills using:

- Shell Scripting
- Python
- PowerShell
- Go (Preferred)
- Java (Preferred)

Experience automating:
- Operational runbooks
- Deployment validation
- Monitoring
- Incident remediation
- Service recovery procedures

CI/CD & Release Management

Experience with

- Harness
- Bamboo
- Bitbucket Pipelines
- GitHub Actions
- Jenkins
- Azure DevOps

Strong understanding of:
- CI/CD
- Release Engineering
- Deployment Strategies
- Blue-Green Deployments
- Canary Deployments
- Rollback Procedures

ITIL & ITSM

Strong experience in

- Incident Management
- Problem Management
- Change Management
- Release Management
- Knowledge Management

Tools:
- ServiceNow
- Remedy
- JIRA

Key Responsibilities Delivery & Technical Contribution (80%)

Reliability Engineering

- Ensure overall platform reliability, availability, and performance.
- Drive continuous improvements to reduce incidents and operational risks.
- Design and implement SLI/SLO frameworks.
- Monitor service health and proactively address burn-rate violations.

Production Support

- Lead Sev1 and Sev2 incident investigations.
- Drive service restoration and stakeholder communication.
- Manage major incident bridges and technical war rooms.
- Perform Root Cause Analysis and post-mortem reviews.

Application Operations

- Troubleshoot complex application, infrastructure, database, and network issues.
- Support authentication and authorization services globally.
- Monitor business-critical transaction flows.
- Build advanced monitoring dashboards and synthetic health checks.

Deployments & Release Management

- Support production deployments and releases.
- Lead deployment readiness assessments.
- Ensure successful rollout and rollback execution.
- Drive release governance processes.

Automation & Engineering Excellence

- Automate operational processes and runbooks.




- Implement self-healing and auto-remediation capabilities.
- Identify toil reduction opportunities.
- Improve MTTD and MTTR metrics.

Agile Delivery

- Participate in sprint planning.
- Own and deliver assigned epics and user stories.
- Review technical solutions and implementation approaches.
- Ensure operational readiness for new platform capabilities.

Leadership & Team Management (20%)

Team Leadership

- Lead, mentor, and coach SRE engineers.
- Conduct technical reviews and guidance sessions.
- Support career development initiatives.
- Build a culture of operational excellence.

Delivery Governance

- Ensure SLA, SLO, and operational commitments are consistently achieved.
- Monitor service delivery metrics.
- Review team performance and workload distribution.
- Drive capacity and resource planning.

Stakeholder Management

- Act as primary escalation point for critical incidents.
- Communicate effectively with business, engineering, and executive leadership.
- Manage client expectations during outages and major events.

Continuous Improvement

- Drive service improvement initiatives.
- Lead automation programs.
- Improve reliability maturity across application portfolios.
- Contribute to organizational SRE best practices.

Soft Skills

- Excellent communication and stakeholder management
- Strong leadership and mentoring capabilities
- High ownership and accountability
- Strategic problem-solving mindset
- Ability to make decisions under pressure
- Customer-centric attitude
- Conflict resolution and collaboration skills
- Strong documentation practices
- Ability to lead globally distributed teams

Readiness & Work Conditions

- Ready to work from office 5 days a week.
- Ready to support 24x7 Production Support operations.
- Comfortable working in rotational shifts and on-call support.
- Ready to support mission-critical customer-facing platforms.
- Ready to upskill continuously on emerging technologies and cloud platforms.

Preferred Certifications

Cloud

- Google Cloud Associate Cloud Engineer (Preferred)
- Google Skilled Cloud Architect
- Microsoft Azure Administrator
- AWS Solutions Architect Associate

Reliability & Operations

- ITIL Foundation
- Certified Kubernetes Administrator (CKA)
- Splunk Certified Power User/Admin

📌 Production Support Lead (Hyderabad)
🏢 Cloudxtreme
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: production support lead (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: production support lead (hyderabad) / hyderabad