02 Oct
|
EdgeVerve Systems
|
Maharashtra
02 Oct
EdgeVerve Systems
Maharashtra
Experience:
- 5 -12 years
Key Responsibilities
- To work as a techno-functional SME to provide technical troubleshooting and product support for customers using our products.
- Take ownership of user problems and be proactive when dealing with user issues.
- Lead day-to-day production support for applications running in GCP, AWS, Azure (or any other cloud service) & Kubernetes environments.
- Monitor application health, respond to incidents, and ensure timely resolution within SLAs.
- Act as the primary escalation point for high-severity incidents and coordinate with engineering and cloud teams.
- Drive root cause analysis (RCA), post-incident reviews, and long-term problem resolution.
- Oversee implementation and maintenance of monitoring, alerting, and logging tools.
- Maintain operational documentation, runbooks, and knowledge base for support teams.
- Implementing automation tools to streamline operations and reduce the frequency of errors.
- Mentor and guide L1/L2 support teams and ensure effective knowledge transfer.
Technical Skills
- Java & Application Layer
Strong knowledge of Java/J2EE applications and microservices architecture
Familiarity with REST APIs
Experience with application performance troubleshooting and profiling
- Kubernetes (K8s)
Understanding of Kubernetes objects: Deployments, Pods, Services, ConfigMaps, Secrets, etc.
Hands-on with kubectl, Kubernetes troubleshooting, and log analysis
- GCP (Google Cloud Platform)
Experience with key GCP services: GKE (Google Kubernetes Engine), Cloud Logging, Cloud Monitoring, Cloud SQL
- Monitoring & Observability
Tools: Datadog or any other similar tool
Proficient in setting up alerts, dashboards, and log aggregation
Root cause analysis from logs and metrics
- Incident & Problem Management
Robust skills in triaging production incidents and leading RCA efforts
Familiar with ITIL practices (especially Incident, Change, Problem Management)
Tools: Jira or any other similar tool
Soft Skills & Leadership
- Team Leadership
Leading L1/L2/L3 support teams
Incident bridge management, escalations handling
Mentoring and upskilling junior team members
2.
Stakeholder
Communication
Effective communication with developers, QA, DevOps, and business stakeholders
Post-incident reporting and communication
3.
Operational
Excellence
SLAs, SLOs, error budgets, service reliability
Defining support processes, runbooks, and knowledge bases
4.
Change
Management
Coordinating deployments, patches, hotfixes
Managing go-live and rollback plans
Nice to Have
Knowledge of SRE practices
Familiarity with security & compliance in cloud environments (e.g., vulnerability scanning, IAM best practices)
Automation using Python, Bash, or other scripting languages
📌 Senior Member of Technical Staff (TM) (Maharashtra)
🏢 EdgeVerve Systems
📍 Maharashtra