20 Aug
|
HDFC Securities
|
Mumbai
20 Aug
HDFC Securities
Mumbai
As an Application Support Engineer, you will:
- Provide L1/L2 application support for production systems, ensuring timely detection, triage, and resolution of incidents to meet SLA/SDR commitments
- Design, build, and maintain Grafana dashboards to provide real-time visibility into application health, performance, and business-critical metrics
- Define and configure alerting thresholds and rules within Grafana (and integrated data sources) to enable proactive detection of anomalies and outages
- Work with data sources such as Prometheus, Elasticsearch, InfluxDB, MySQL/Oracle, or similar to build meaningful queries and visualizations
- Perform root cause analysis (RCA) for recurring incidents and problem records, and drive preventive/corrective actions with DEV, QA, and Infra teams
- Monitor batch jobs, cron schedules, and end-of-day processes, and escalate deviations promptly to relevant stakeholders
- Collaborate with cross-functional teams (DEV, QA, Infra, Product, and Customer Care) to troubleshoot production issues and coordinate fixes
- Maintain and enhance runbooks, SOPs, and knowledge base articles for common issues, dashboard usage, and escalation procedures
- Support incident, problem, and change management processes on ServiceNow, ensuring accurate categorization, documentation, and closure
- Participate in governance reporting by contributing dashboard-driven metrics (MTTR, SLA compliance, incident trends) for leadership reviews
Key Skills and Experience:
- 5-7 years of overall experience in Application Support, Production Support, or Site Reliability roles
- Hands-on experience in creating, customizing, and maintaining Grafana dashboards, including panels, variables, alerts, and data source integrations
- Working knowledge of at least one Grafana-compatible data source (Prometheus, Elasticsearch, InfluxDB, MySQL, PostgreSQL, or similar)
- Solid understanding of application support fundamentals - incident triage, log analysis, RCA, and escalation management
- Familiarity with monitoring and observability concepts such as SLIs, SLOs, alert thresholds, and dashboarding best practices
- Experience with log analysis tools (e.g., ELK stack, Splunk) for troubleshooting production issues
- Basic scripting skills (Python, Bash, or SQL) for automation, data extraction, and building custom reports
- Understanding of Linux/Unix environments and basic networking concepts (DNS, load balancers, ports, protocols)
- Exposure to ServiceNow or similar ITSM tools for incident, problem, and change management
- Ability to read and interpret application logs, thread dumps, and basic performance metrics
Additional Requirements:
- Bachelor's degree in Computer Science, IT, or a related field
- Robust analytical and problem-solving skills with a data-driven approach to troubleshooting
- Good communication skills, with the ability to explain technical issues to both technical and non-technical stakeholders
- Ability to work under pressure and manage multiple priorities in a fast-paced, production environment
- Willingness to work in rotational/on-call shifts to support production systems as needed
Preferred Qualifications:
- Experience with other APM/monitoring tools such as Dynatrace, New Relic, or AppDynamics
- Exposure to Kubernetes, Docker, or other container orchestration platforms
- Experience in FinTech, BFSI, or other highly regulated industries
- Familiarity with CI/CD pipelines and their integration with monitoring tools
- Grafana or related observability certifications
📌 Application Support Engineer (Monitoring & Grafana Dashboards) (Mumbai)
🏢 HDFC Securities
📍 Mumbai