27 Aug
|
Remote_WFH_INDIA
|
Pune
27 Aug
Remote_WFH_INDIA
Pune
? Senior Site Reliability Engineer | Datadog | Kubernetes | AWS | Python | CI/CD | Remote India | JST Overlap
? We're AgileEngine!
AgileEngine is an award-winning software engineering and AI services company partnering with Fortune 500 companies to build world-class digital products..
You’ll work extensively with Datadog, Kubernetes, AWS, Python, APIs, and CI/CD pipelines, with a strong focus on building reliable and scalable observability solutions.
? Location: Remote / WFH — India
? Timezone: JST timezone overlap is mandatory
? Experience: 4+ years
? Client: Indeed
? Why You Should Apply
- Work on a high-impact platform reliability initiative for Indeed
- Build and maintain enterprise-grade Datadog observability solutions
- Work with Kubernetes-based microservices environments
- Design dashboards, alerts, APM, metrics, logging, and tracing
- Integrate observability into AWS and CI/CD pipelines
- Automate monitoring and operational tasks using Python
- Work across software engineering and site reliability engineering
- Drive platform reliability, scalability, performance, and continuous improvement
- Take ownership of platform modernization and maintenance initiatives
✅ Must-Have Skills
- 4+ years of professional software engineering / SRE experience
- Strong proficiency in Python, JavaScript (Node.js), or Java
- Strong hands-on experience with Kubernetes — deployment, operations, and monitoring
- Hands-on experience with Datadog or similar observability platforms such as Prometheus/Grafana
- Experience configuring Datadog dashboards, alerts, APM, metrics, logging, and tracing
- Experience monitoring containerized and microservices-based applications
- Hands-on experience with AWS
- Experience integrating observability tools into cloud environments
- Experience integrating observability into CI/CD pipelines
- Strong experience with API integrations — designing, consuming,
and integrating APIs
- Experience automating monitoring and operational tasks using scripting, preferably Python
- Experience installing and configuring Datadog agents and integrations
- Understanding of secure configuration, API keys, user roles, and access controls
- Upper-intermediate English communication skills
- JST timezone overlap — Mandatory
? What You'll Do
- Support platform reliability, monitoring, and continuous improvement across internal systems
- Work extensively in Kubernetes-based environments
- Build and maintain Datadog dashboards, alerts, APM, metrics, logs, and traces
- Monitor containerized and microservices-based applications
- Integrate observability solutions with AWS environments
- Integrate monitoring and observability into CI/CD pipelines
- Install and configure Datadog agents and integrations
- Manage API keys, secure configurations, user roles, and access controls
- Automate monitoring and operational activities using Python
- Lead maintenance initiatives and platform improvements
- Improve system reliability, scalability, and performance
- Proactively identify opportunities to modernize and improve platform operations
➕ Nice to Have
- Experience owning and operating an internal engineering platform
- Demonstrated ownership of reliability, scalability, and performance
- Experience proactively leading maintenance and platform improvement initiatives
- Familiarity with Golang
- Experience with New Relic, Dynatrace, Elastic, or Splunk Observability
- Strong experience with distributed and microservices architectures
? Priority Will Be Given To Candidates With
- Strong Python experience
- Strong Kubernetes experience
- Hands-on Datadog experience
- Strong AWS experience
- CI/CD integration experience
- Observability experience across APM, metrics, logging, tracing, dashboards, and alerts
- Experience monitoring containerized / microservices applications
- Solid API integration experience
- Experience automating operational tasks using Python
- Proven ownership of reliability, scalability, and performance
- JST timezone overlap
?️ Tech Stack
Datadog • Kubernetes • AWS • Python • CI/CD • Observability • APM • Monitoring • Metrics • Logging • Tracing • Microservices • APIs • Docker/Containers • Cloud • Site Reliability Engineering
⚠️ Important Timezone Requirement
? Candidates must be able to provide the required JST timezone overlap for this role.
Please apply only if you are comfortable working with the required JST overlap.
⚠️ Hiring Process
Step 1: Technical Assessment
Step 2: Video Interview
Step 3: Recruiter Discussion
Step 4: Technical Interview
Step 5: Offer ?
Please complete all LaunchPod steps promptly to avoid delays in the hiring process.
? To Apply, DM me with:
1. Email ID
2. Total Experience
3. SRE Experience
4. Python Experience
5. Kubernetes Experience
6. Datadog Experience
7. AWS Experience
8. CI/CD Experience
9. Observability Experience
10. API Integration Experience
11. Microservices / Container Experience
12. Other Observability Tools — Prometheus / Grafana / New Relic / Dynatrace / Elastic / Splunk
13. Current CTC
14. Expected CTC
15. Notice Period
? Please don't use LinkedIn Easy Apply. Send me a DM with the above details.
📌 SRE | Kubernetes | Datadog | AWS | Python | Remote India | Remote/WFH (Pune)
🏢 Remote_WFH_INDIA
📍 Pune