02 Aug
|
TEKsystems
|
Pune
Job Summary
- Kubernetes
- Monitoring
- Shell scripting
Location: Pune, Maharashtra
Experience: 6+ years
About the Role
Join a high-impact Observability Engineering team responsible for maintaining the reliability, performance, and availability of technology platforms supporting over 100,000 servers and critical enterprise systems globally. As an SRE focused on OpenTelemetry and Observability, you will manage and optimize telemetry platforms, support large-scale monitoring initiatives, contribute to the migration from Griffin to Splunk, and work closely with engineering and infrastructure teams to ensure reliable data collection, automation, operational excellence, and platform resilience across a complex Kubernetes-based environment.
Required Skills
- SRE practices
- OpenTelemetry
- Kubernetes (IKP or equivalent)
- Linux/Unix Administration
- Python Scripting
- Bash Scripting
- Prometheus / Grafana / Splunk
Good to Have Skills
- ITIL (Incident, Problem &
- Change Management)
- Telemetry Pipeline Management
- Capacity Planning
- Observability Platform Support
- High Availability &
- Disaster Recovery
- AI &
- Operational Automation
- Banking / Financial Services Domain Experience
Roles &
- Responsibilities
- Monitor and maintain the health, performance, and availability of OpenTelemetry collectors and telemetry pipelines.
- Deploy, configure, and support OpenTelemetry Collectors, including receivers, processors, and exporters.
- Manage and troubleshoot OpenTelemetry workloads running on Kubernetes environments.
- Respond to incidents, perform root cause analysis (RCA), and implement preventive solutions.
- Define, monitor, and improve SLOs, SLIs, and platform reliability metrics.
- Ensure reliable telemetry data delivery to observability platforms such as Splunk, Prometheus, Grafana, and Jaeger.
- Support application instrumentation for traces, metrics, and logs.
- Investigate and resolve telemetry data loss, data quality, and pipeline performance issues.
- Automate operational and support activities using Python and Bash scripting.
- Participate in on-call support, major incident management, and problem management activities.
- Perform capacity planning and support telemetry platform scalability requirements.
- Create and maintain operational documentation, runbooks, and standard operating procedures.
- Collaborate with development, infrastructure, and platform teams to onboard recent services into the observability ecosystem.
- Drive continuous improvement initiatives focused on platform stability, resilience, and operational efficiency.
- Identify automation opportunities and contribute to AI-driven operational transformation initiatives.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 IKP Engineer (Pune)
🏢 TEKsystems
📍 Pune