02 Oct
|
HuntingCube
|
Hyderabad
02 Oct
HuntingCube
Hyderabad
WHAT YOU WILL DO DAY-TO-DAY: You will design and develop robust, scalable, high-performance tools and automation solutions for Linux environments using Go and Python, leveraging distributed systems and up-to-date orchestration platforms. You will also build and scale the firm's observability platform, including metrics (Grafana, Grafana Mimir), logging (ELK: Elasticsearch, Logstash, Kibana), and high-throughput telemetry stores (ClickHouse). You will develop and advocate for robust observability and reliability practices: logging, metrics, tracing, SLIs/SLOs, alerting, and capacity planning.
You will lead critical incident response and drive blameless postmortems that turn incidents into permanent fixes, with a standing mandate to engineer toil out of operations and support rather than absorb it. Additionally, you will develop innovative solutions and streamline day-to-day workflows, including enhancing the SDLC pipeline by leveraging the rapidly evolving GenAI landscape to improve efficiency, quality, and automation. You will also lead and collaborate within a cross-functional team, provide mentorship to junior engineers, and contribute to code and design reviews.
Furthermore, you will troubleshoot and resolve complex issues related to Linux infrastructure, core services, and the observability pipeline.
WHO WE ARE LOOKING FOR: The ideal candidate should hold- Basic qualifications:
- A bachelor's or master's degree in computer science, engineering, or a related technical discipline, with 4 to 8 years of software development and SRE experience
- Excellent proficiency in Go/Python for developing production-grade applications, services, and automation tooling
- Exceptional understanding of Linux internals; familiarity with Kerberos, TLS, Puppet, LDAP, DNS, DHCP, mail, and load balancers
- Solid grounding in observability and reliability principles (metrics, logs, traces, SLIs/SLOs, alerting, and capacity planning)
- Experience with observability tools (e.g., Grafana, Grafana Mimir, ELK)
- Familiarity with distributed systems, including experience designing and maintaining highly available systems
- Exceptional problem-solving skills and the ability to think critically
- Excellent communication and collaboration skills Preferred qualifications:
- Experience with GenAI and its application to observability and reliability (e.g., anomaly detection, alert-noise reduction, incident triage/summarization, log analysis, or AIOps)
- Experience with container orchestration and workflow systems (e.g., Kubernetes, Temporal)
- Hands-on experience operating ClickHouse or similar high-cardinality/high-throughput data stores
Required Skills
['Python', 'Linux', 'TLS', 'LDAP', 'DNS', 'DHCP']
Additional Information
NA
📌 Senior Software Engineer- Linux Engineering (Hyderabad)
🏢 HuntingCube
📍 Hyderabad