17 Sep
|
Artech Infosystems
|
Telangana
17 Sep
Artech Infosystems
Telangana
: Observability Contractor
About the Role
Join our Observability team to ensure the reliability and performance of the infrastructure powering our global gaming experiences. In this role, you will own the lifecycle of our monitoring, logging, and telemetry initiatives, working with high autonomy to scale our observability stack.
You will partner with engineering teams to tackle challenges ranging from Kubernetes deployments to deep-dive infrastructure troubleshooting, directly improving how we detect and resolve incidents at scale. If you have a deep technical background in the Grafana stack and are passionate about transforming complex system data into actionable insights for engineers,
we want to hear from you.
Key Responsibilities
Strategy &
- Architecture
● Architect and Evolve: Design and deploy comprehensive observability solutions that provide deep visibility into our global infrastructure.
● Scale the Stack: Drive scalability and performance of our core services, ensuring it scales with our traffic growth.
Operational Excellence
● Drive Reliability: Optimize the performance of containerized workloads in Kubernetes,
proactively identifying bottlenecks before they impact the player experience.
● Lead Incident Resolution: Serve as a technical authority during incident response,
driving faster root-cause analysis (RCA) and post-mortem improvements.
● Reduce Toil: Identify and automate repetitive operational workflows to reduce manual effort and system noise for the broader engineering team.
Consultation &
- Partnership
● Champion Observability Culture: Act as an internal consultant for development teams,
guiding them on best practices for application instrumentation and effective alerting strategies.
Required Skills &
- Qualifications
● Observability Expertise: Deep experience with the Grafana stack (Prometheus,
Loki, Grafana)
and modern observability agents (OTeL, Alloy).
● Container Orchestration: Strong, hands-on experience managing and troubleshooting containerized workloads in production-grade Kubernetes environments.
● Automation &
- Scripting: Proficiency in writing automation scripts (Golang preferred) to reduce operational toil and manage infrastructure.
● Cloud &
- Networking: Working knowledge of AWS cloud-native architectures, Linux/CLI environments, and core networking protocols (TCP/IP, DNS, Load Balancing).
● Collaboration &
- Documentation: Proven ability to document complex system architectures and communicate effectively with cross-functional engineering teams during incident response.
Core competencies
● Ability to advocate for observability best practices and influence development teams to adopt instrumentation standards, even without direct reporting lines.
● A ‘root-cause detective’ mindset - someone who naturally questions system behaviours and persists until the underlying issue is understood, rather than just treating symptoms.
● ability to communicate complex technical data or incident impacts to both engineering peers and non-technical stakeholders clearly and concisely.
● Proven ability to thrive in ambiguous environments, quickly auditing complex systems,
and delivering results with minimal oversight or hand-holding.
Preferred Skills
● Observability/Monitoring: Experience with monitoring/observability agents and tools
(Prometheus, Promtail, Grafana Alloy, OTeL, etc)
● Scripting: Experience writing automation with a robust preference for Golang.
● Infrastructure as Code (IaC): Familiarity with Terraform to manage and provision infrastructure.
● Configuration Management: Ideal candidates will have working knowledge of Chef cookbooks for configuration automation
📌 IND_Engineer (Telangana)
🏢 Artech Infosystems
📍 Telangana