22 Aug
|
AgileEngine
|
India
About the Role
We are looking for a Senior Software Engineer & Site Reliability Engineer to support platform reliability, monitoring, and modernization across Kubernetes-based microservices environments with a solid observability focus.
You will build and maintain Datadog solutions including dashboards, alerts, APM, metrics, logging, and tracing, integrate observability tooling into AWS and CI/CD pipelines, and automate monitoring and operational tasks using Python. The role blends software engineering (60–70%) with site reliability engineering (30–40%) and requires JST timezone overlap.
What you will do
- Support platform reliability, monitoring, and continuous improvement across internal systems.
- Work in Kubernetes-based environments.
- Build and maintain observability solutions, with a focus on Datadog .
- Configure dashboards, alerts, APM, metrics, logging, and tracing.
- Monitor containerized and microservices-based applications.
- Integrate observability tools into AWS environments.
- Integrate observability into CI/CD pipelines.
- Automate monitoring and operational tasks using scripting ( Python preferred).
- Install and configure Datadog agents and integrations.
- Manage API keys and secure configurations.
- Manage user roles and access controls within observability platforms.
- Lead maintenance efforts and platform improvements while driving reliability, scalability, and performance.
Must haves
- Strong proficiency in Python , JavaScript (Node.js), or Java.
- Hands-on experience with API integrations (designing, consuming, and integrating).
- Strong experience working in Kubernetes environments (deployment, operations, monitoring).
- Experience with Datadog (preferred) or similar tools (Prometheus, Grafana).
- Ability to configure dashboards, alerts, and APM (tracing, metrics, logging).
- Experience monitoring containerized/microservices architectures.
- Hands-on experience with AWS .
- Experience integrating observability tools into cloud environments.
- Experience integrating observability into CI/CD pipelines.
- Ability to automate monitoring and operational tasks using scripting ( Python preferred).
- Upper-intermediate English level.
Nice to haves
- Experience owning and operating an internal engineering platform.
- Demonstrated ownership of reliability, scalability, and performance.
- Proven ability to proactively lead maintenance efforts and platform improvements (not just reactive support).
- Familiarity with Golang .
- Experience with additional observability tools such as New Relic, Dynatrace, Elastic, or Splunk Observability
Perks and Benefits
- Remote work & Local connection: Work where you feel most productive and connect with your team in periodic meet-ups to strengthen your network and connect with other top experts.
- Legal presence in India: We ensure full local compliance with a structured, secure work environment tailored to Indian regulations.
- Competitive Compensation in INR: Fair compensation in INR with dedicated budgets for your personal growth, education, and wellness.
- Innovative Projects: Leverage the latest tech and create cutting-edge solutions for world-recognized clients and the hottest startups.
📌 Software Engineer (SRE) (India)
🏢 AgileEngine
📍 India