- Ensure high-quality telemetry data collection with minimal noise
- Configure intelligent alerting using baselines and anomaly detection and Reduce alert fatigue through alert tuning and deduplication strategies
- Implement event correlation to group related incidents
- Use AIOps features to identify patterns and predict failures and Continuously improve signal-to-noise ratio in monitoring systems
- Investigate production incidents using logs, metrics, and traces
- Perform root cause analysis (RCA) and document findings
- Collaborate with application teams to resolve critical issues
- Improve MTTR (Mean Time to Resolution) through better observability and maintain incident response runbooks
- Work with containerized environments (Docker, Kubernetes)
- Ensure observability coverage for APIs
- Apply basic ML/statistical techniques for anomaly detection and Query and analyze telemetry data using SQL-like languages (e.g., NRQL)
- Identify trends and capacity issues using time-series analysis
- Provide insights and recommendations to improve system reliability
- Maintain data retention and governance policies
Key Responsibilities*
- Observability tool -Current Relic
- SQL / query languages
- Python/Unix scripting/React JS
- Understanding of Cloud platforms like Azure
- Basic AIOps / anomaly detection concepts
- Experience reducing alert noise and Strong system thinking, not just tool usage
📌 Tcs Hiring Observability Engineer For Kolkata (West Bengal)
🏢 Tata Consultancy Services
📍 West Bengal
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.