Interested?
Please share your updated resume with
[email protected]
Site Reliability Engineer (SRE) / DevOps Engineer Python | Kubernetes | Cloud
Location: Gurgaon / Gurugram Client Location
Experience: 36 Years
Employment Type: Permanent / Full-Time
Work Mode: Work from Office Client Location
Employment: Permanent rolls of our company
Work Location
Gurgaon / Gurugram Client Location( Tower-D, 39, Arya Samaj Rd, Durga Colony, Sector 39, Gurugram, Haryana 122003 )
The selected candidate will be on the permanent rolls of our company and will be required to work from the client location in Gurgaon.
About the Role
We are looking for a hands-on Site Reliability Engineer (SRE) / DevOps Engineer with strong Python development and automation skills.
This is an engineering-focused role where the initial focus will be on Python development, automation, APIs, utilities, and platform capabilities, enabling the engineer to develop a strong understanding of the applications and technology platform.
Over time, the role will expand into broader DevOps and SRE responsibilities, including CI/CD, cloud infrastructure, Kubernetes, observability, production reliability, incident management, and operational automation.
The ideal candidate should be comfortable working with both application code and production systems and should have a strong engineering and automation mindset focused on improving reliability, scalability, and operational efficiency.
Key Responsibilities
- Develop and enhance internal applications, automation tools, APIs, utilities, and platform capabilities using Python.
- Write clean, maintainable, testable, scalable, and production-ready code.
- Participate in code reviews, debugging, testing, and technical discussions.
- Build, maintain, and improve CI/CD pipelines and automated deployment processes.
- Work with Docker and Kubernetes for application deployment and operations.
- Support on-premises and cloud-based application and infrastructure deployments.
- Maintain reliable, scalable, secure,
and highly available production environments.
- Implement and manage monitoring, logging, alerting, and observability solutions.
- Contribute to defining and tracking SLIs, SLOs, availability targets, and error budgets.
- Troubleshoot application and production issues and perform Root Cause Analysis (RCA).
- Identify recurring operational problems and address them through automation and engineering improvements.
- Support incident response, change management, deployment governance, and disaster recovery practices.
- Develop and maintain runbooks, SOPs, incident documentation, and technical documentation.
- Collaborate closely with Engineering, Product, Platform, Security, Operations, and external teams.
Required Technical Skills
- Strong hands-on experience with Python development and automation.
- Experience developing scripts, APIs, integrations, utilities, or backend services using Python.
- Positive understanding of software engineering principles, debugging, logging, testing, exception handling, and code quality.
- Experience with REST APIs, JSON, Git, pull requests, and code reviews.
- Strong knowledge of Linux/Unix environments and basic Windows administration.
- Good understanding of networking concepts including DNS, TCP/IP, HTTP/HTTPS, load balancing, and firewalls.
- Experience with at least one major cloud platform: AWS, Azure, or GCP.
- Hands-on experience with Docker and Kubernetes.
- Experience with CI/CD tools such as GitHub Actions, GitLab CI, Jenkins, Azure DevOps, or equivalent.
- Familiarity with Infrastructure as Code (IaC) tools such as Terraform is preferred.
- Experience with monitoring and observability tools such as Prometheus, Grafana, CloudWatch, Azure Monitor, Datadog, Splunk, or equivalent.
- Ability to analyze logs, metrics, alerts, and traces for troubleshooting and performance analysis.
- Understanding of SRE concepts including SLIs, SLOs, availability, reliability, error budgets, and RCA.
- Experience with JIRA, ServiceNow, and Confluence is desirable.
Preferred Experience
- 3–6 years of experience in SRE, DevOps, Platform Engineering, Cloud Engineering, Software Engineering, or related roles.
- Strong Python development or automation experience.
- Experience supporting applications across development, deployment, and production environments.
- Exposure to cloud-native, distributed, microservices-based, or production-grade systems.
- Understanding of security, compliance, and operational best practices.
- Experience with infrastructure and application automation.
- Familiarity with AI-assisted engineering tools such as GitHub Copilot, Claude Code, or similar tools is a plus.
Soft Skills
- Robust analytical, problem-solving, and troubleshooting skills.
- Strong engineering and automation mindset.
- Good written and verbal communication skills.
- Effective cross-functional collaboration skills.
- Ownership-driven approach to problem solving.
- Ability to work independently and take responsibility for production systems.
- Ability to remain calm, structured, and methodical during production incidents.
- Continuous improvement mindset with a focus on reducing manual effort and improving reliability.
Important Note
This is a hands-on engineering role with a strong emphasis on Python programming, automation, DevOps, and SRE practices. Candidates
whose experience is primarily limited to monitoring, deployment coordination, or routine production support without hands-on programming and automation experience may not be suitable.
Interested?
Please share your updated resume with
[email protected]
📌 Site Reliability Engineer (Gurugram)
🏢 New Era Technology
📍 Gurugram