19 Aug
|
PwC Acceleration Centers
|
Hyderabad
19 Aug
PwC Acceleration Centers
Hyderabad
Associate – Forward Deployment Engineer (SRE)
Site Reliability Engineering | Forward Deployed Engineering
Location: Bangalore / Hyderabad
Experience Required
3–5 years.
Job Summary A hands-on, reliability-first engineer embedded within an enterprise client's team. You will help run and stabilise their production systems on AWS, respond to incidents, and use AI-assisted tooling to keep everything healthy day to day.
Key Responsibilities
- Keep production systems reliable for enterprise customers on AWS, and respond quickly when something isn't behaving.
- Own monitoring and alerting, and work to the SLIs, SLOs, and error budgets agreed with customers.
- Handle production incidents (P0–P3) within SLA and take part in the on-call rotation.
- Troubleshoot live issues under pressure, from networking and authentication to container problems.
- Operate and maintain Kubernetes clusters, containers, and core AWS services.
- Use AI SRE assistants and AIOps tooling to detect and resolve incidents faster.
- Automate repetitive operational tasks to reduce manual effort.
- Maintain runbooks and post-incident notes, and collaborate with client engineers over Teams, Slack, and email.
Required Qualifications
- Proven experience keeping production systems reliable on AWS in an SRE or similar role.
- Hands-on experience running Kubernetes and containers in production.
- Working knowledge of core AWS services and cloud-native operations.
- Experience owning monitoring and observability, and working to SLIs, SLOs, and error budgets.
- Experience managing production incidents end to end, including on-call.
- Scripting ability for automation (Python and/or Bash).
- A genuine understanding of how AI-assisted tools support operations, and comfort using them in incident triage and resolution.
- AWS Certified Solutions Architect – Associate (or actively working towards it).
Preferred Qualifications
- Building CI/CD pipelines.
- Infrastructure as Code with Terraform (or equivalent).
- GitOps workflows and Helm.
- Incident-management tooling such as PagerDuty.
- Previous customer-facing experience.
- Hands-on experience with AIOps or AI-driven monitoring and incident tooling.
- Exposure to AI/ML or LLM-based tooling applied to operations.
- Hands-on experience with AI-based work or projects (AI/ML or LLM-powered tools).
- AWS Certified Solutions Architect – Professional or AWS Certified DevOps Engineer – Qualified.
Technical Skills & Tools
- Cloud (AWS): EC2, EKS, ECS, Lambda, S3, RDS, VPC, IAM, CloudWatch, ELB/ALB, Route 53
- Containers & orchestration: Docker, Kubernetes (EKS)
- Observability & monitoring: Prometheus, Grafana, CloudWatch, Loki, OpenTelemetry, ELK/Elastic
- Reliability practices: SLIs/SLOs, error budgets, incident management, performance tuning
- Incident & on-call: PagerDuty, Opsgenie, runbooks, post-mortems
- Automation & scripting: Python, Bash, Git
- AI-assisted operations: AI SRE assistants, AIOps, anomaly detection, alert correlation
- DevOps (good to have): CI/CD (GitHub Actions, GitLab CI, Jenkins), Terraform, Ansible, GitOps (ArgoCD), Helm
📌 Forward Deployment Engineer (SRE) (Hyderabad)
🏢 PwC Acceleration Centers
📍 Hyderabad