09 Oct
|
Artech Infosystems
|
India
09 Oct
Artech Infosystems
India
Job Description: Observability Contractor
About the Role
Join our Observability team to ensure the reliability and performance of the infrastructure
powering our global gaming experiences. In this role, you will own the lifecycle of our
monitoring, logging, and telemetry initiatives, working with high autonomy to scale our
observability stack.
You will partner with engineering teams to tackle challenges ranging from Kubernetes
deployments to deep-dive infrastructure troubleshooting, directly improving how we detect and
resolve incidents at scale. If you have a deep technical background in the Grafana stack and
are passionate about transforming complex system data into actionable insights for engineers,
we want to hear from you.
Key Responsibilities
Strategy & Architecture
Architect and Evolve: Design and deploy comprehensive observability solutions that
provide deep visibility into our global infrastructure.
Scale the Stack: Drive scalability and performance of our core services, ensuring it
scales with our traffic growth.
Operational Excellence
Drive Reliability: Optimize the performance of containerized workloads in Kubernetes,
proactively identifying bottlenecks before they impact the player experience.
Lead Incident Resolution: Serve as a technical authority during incident response,
driving faster root-cause analysis (RCA) and post-mortem improvements.
Reduce Toil: Identify and automate repetitive operational workflows to reduce manual
effort and system noise for the broader engineering team.
Consultation & Partnership
Champion Observability Culture: Act as an internal consultant for development teams,
guiding them on best practices for application instrumentation and effective alerting
strategies.
Required Skills & Qualifications
Observability Expertise:
Deep experience with the Grafana stack (Prometheus, Loki, Grafana)
and modern observability agents (OTeL, Alloy).
Container Orchestration: Strong, hands-on experience managing and troubleshooting
containerized workloads in production-grade Kubernetes environments.
Automation & Scripting: Proficiency in writing automation scripts (Golang preferred) to
reduce operational toil and manage infrastructure.
Cloud & Networking: Working knowledge of AWS cloud-native architectures, Linux/CLI
environments, and core networking protocols (TCP/IP, DNS, Load Balancing).
Collaboration & Documentation: Proven ability to document complex system architectures
and communicate effectively with cross-functional engineering teams during incident
response.
Core competencies
Ability to advocate for observability best practices and influence development teams to
adopt instrumentation standards, even without direct reporting lines.
A ‘root-cause detective’ mindset - someone who naturally questions system behaviours
and persists until the underlying issue is understood, rather than just treating symptoms.
ability to communicate complex technical data or incident impacts to both engineering
peers and non-technical stakeholders clearly and concisely.
Proven ability to thrive in ambiguous environments, quickly auditing complex systems,
and delivering results with minimal oversight or hand-holding.
Preferred Skills
Observability/Monitoring: Experience with monitoring/observability agents and tools
(Prometheus, Promtail, Grafana Alloy, OTeL, etc)
Scripting: Experience writing automation with a solid preference for Golang.
Infrastructure as Code (IaC): Familiarity with Terraform to manage and provision
infrastructure.
Configuration Management: Ideal candidates will have working knowledge of Chef
cookbooks for configuration automation
📌 IND_Engineer (India)
🏢 Artech Infosystems
📍 India