Job Title: Site Reliability Engineer (SRE)
Location: Gurgaon / Panchkula, India
Experience Required: 6–8 years
Notice Period: Immediate joiners or up to 1 month
Budget: 4 X (X=Total Number of Experience)
Position Summary
We are seeking a Site Reliability Engineer (SRE) with 6–8 years of hands-on experience in Production Support, Application Support, Tech Ops, Incident Management, and RCA. The role focuses on ensuring production reliability, monitoring, and operational ownership of business-critical applications hosted on AWS. You will coordinate closely with development, platform, and infrastructure teams to drive operational improvements and incident resolution.
Key Responsibilities
Own end-to-end incident lifecycle including triage, resolution, RCA, and post-incident documentation (ITIL practices).
Configure and optimize monitoring, alerting, and escalation policies using Datadog, Pager Duty, and Splunk.
Manage secrets and credentials lifecycle with AWS Secrets Manager, Hashi Corp Vault, and Beacon Vault.
Perform AWS cloud operations (S3, IAM, EC2, VPC, ELB,
Auto Scaling) for access provisioning and key management.
Automate operational tasks using Shell/Bash scripting for monitoring, health checks, and alert automation.
Onboard repositories to CI/CD workflows with Git Hub Actions and maintain operational runbooks/wikis.
Collaborate with cross-functional teams to ensure operational deliverables are completed on time.
Required Skills
6–8 years in SRE, Application Support, or Tech Ops roles.
Solid expertise in Incident Management, RCA, and Production Support.
Hands-on with AWS, Datadog, Pager Duty, Splunk.
Proficiency in Docker, Git Hub, CI/CD workflows, and scripting (Bash/Python).
Knowledge of Nginx, monitoring tools, and operational automation.
Positive to Have
Familiarity with Terraform for minor infra changes.
Exposure to Kubernetes/EKS for operational tasks.
Experience with Autosys job scheduling.
AWS certifications (Sys Ops Adminis
📌 Site Reliability Engineer Gurugram (India)
🏢 TekClap
📍 India