We are looking for a skilled Site Reliability Engineering (SRE) professional with 5-15 years of experience to join our team. The ideal candidate will have a strong background in ensuring the reliability and performance of software systems.
JD:
- Design, implement, and maintain scalable and reliable software systems.
- Collaborate with cross-functional teams to identify and prioritize project requirements.
- Develop and execute test plans to ensure high-quality deliverables.
- Troubleshoot and resolve complex technical issues efficiently.
- Implement automation scripts to improve system efficiency and reliability.
- Monitor system performance and recommend improvements.
- Strong understanding of software development principles and methodologies.
- Experience with agile development methodologies and version control systems.
- Excellent problem-solving skills and attention to detail.
- Ability to work collaboratively in a fast-paced environment.
- Strong communication and interpersonal skills.
- Experience with monitoring tools and log management systems.
- Solid hands-on experience with ELK Stack (Elasticsearch, Logstash, Kibana)
- Proficient in log ingestion, parsing, indexing, and visualization
- Experience with distributed systems, Linux, and networking fundamentals
- Scripting skills in Python / Shell
- Hands-on with monitoring, alerting, and troubleshooting production issues
- Experience integrating ELK with cloud platforms (AWS)
- Experience with Beats (Filebeat, Metricbeat, Heartbeat)
- Exposure to Open Telemetry, APM tools, or Prometheus/Grafana
- Knowledge of container platforms (Docker, Kubernetes)
- CI/CD integration and DevOps/SRE practices
- Basic knowledge of security logging and SIEM concepts