10 Sep
|
Tech Mahindra
|
Hyderabad
10 Sep
Tech Mahindra
Hyderabad
Primary skills to focus: Shell Scripting, PL/SQL, Cassandra Queries, experience with ELK, New relic, Quantum Metrics, OpenSearch, and Dynatrace.
Shift timings: Rotational 24*7Role & responsibilities
Job Profile
Responsibilities:
- Monitoring - System Health Monitoring Continuously check system performance, server uptime, and network availability.
- Application Monitoring – Track application response times, errors, and crashes.
- Trend Analysis – Identify recurring issues and suggest long-term fixes.
- Log Analysis – Analyze logs for anomalies, failures, or suspicious activities.
- Performance Metrics Tracking – Monitor CPU, memory, disk usage, and database performance.
- Network Traffic Analysis – Observe bandwidth usage and detect unusual spikes or bottlenecks.
- Incident Detection & Logging – Identify and document incidents for further analysis.
- Triaging - Categorize incidents based on severity and business impact.
- Perform root cause analysis on production issues and work with cross-functional teams to implement effective solutions.
- Document performance findings and provide recommendations for addressing recurring performance issues.
- Use monitoring and logging tools (such as New Relic, ELK, or Dynatrace) to track system behavior and pinpoint potential issues.
- Assignment & Escalation – Assign incidents to appropriate teams or escalate critical issues.
- Collaboration & Communication – Update stakeholders about ongoing issues and expected resolution times.
- Work with internal teams, vendors, or service providers to resolve issues.
- Knowledge of ITIL, SRE principles, or Site Reliability Engineering practices.
Mandatory Skills:
- Proficient in Linux/Unix Shell Scripting
- Strong experience with PL/SQL and database management.
- Hands-on experience with Cassandra Queries.
- Experience: 3+ years of experience in New Relic APM, Infrastructure Monitoring, or related observability tools.
- Hands-on experience with New Relic One, APM, Logs, Synthetic Monitoring, and Distributed Tracing.
- Strong knowledge of Linux, Windows, and cloud-based environments (AWS, Azure, GCP).
- Proficiency in scripting languages (Python, Bash, PowerShell) for automation.
- Understanding of APIs, microservices, and containerized environments (Docker, Kubernetes).
- Familiarity with logging tools like ELK Stack
- Knowledge of CI/CD tools (Jenkins, GitLab, Azure DevOps) and how monitoring fits into DevOps pipelines.
- Robust problem-solving and analytical skills.
- Ability to work in high-pressure, production-critical environments.
- Excellent communication and documentation skills.
Preferred Skills:
- Experience in cloud infrastructure management.
- Knowledge of containerization technologies (e.g., Docker, Kubernetes).
- Understanding of networking concepts and protocols.
- Experience with monitoring and logging tools.
📌 Production and Support Engineer (Hyderabad)
🏢 Tech Mahindra
📍 Hyderabad