29 Aug
|
CorroHealth
|
Noida
About Us:
Our purpose is to help clients exceed their financial health goals. Across the reimbursement cycle, our scalable solutions and clinical expertise help solve programmatic needs. Enabling our teams with leading technology allows analytics to guide our solutions and keeps us accountable achieving goals.
We build long-term careers by investing in YOU. We seek to create an environment that cultivates your professional development and personal growth, as we believe your success is our success.
ESSENTIAL DUTIES AND RESPONSIBILITIES:
Note:
We are seeking a skilled Site Reliability Engineer (SRE) with expertise in AWS, Azure, Kubernetes, Serverless (Lambda), Python, CI/CD, and observability tooling . The SRE will ensure system reliability, scalability, and performance across our cloud platforms, while driving automation, observability, and operational excellence.
Key Responsibilities
- Cloud Infrastructure Design, deploy, and optimize workloads on AWS and Azure .Manage Kubernetes clusters (AKS, EKS) and serverless solutions like AWS Lambda .Implement secure, scalable, and cost-optimized infrastructure.
- Automation Engineering Build automation frameworks and operational tooling using Python .Implement Infrastructure as Code (IaC) using Terraform, CloudFormation, or Azure Bicep .Automate deployments, scaling, monitoring, and remediation.
- CI/CD DevOps Practices Develop and maintain CI/CD pipelines (Jenkins, GitHub Actions, GitLab CI/CD, Azure DevOps).Partner with development teams to improve release velocity and reliability.
- Observability Monitoring Implement end-to-end observability practices : metrics, logging, tracing.Manage logging monitoring solutions (Loggly, ELK/EFK stack, CloudWatch, Azure Monitor, Prometheus, Grafana, Datadog, New Relic ).Establish alerting escalation workflows with PagerDuty, OpsGenie, or similar incident management platforms .
- Site Reliability Incident Management Participate in 24/7 on-call rotations for critical systems.Diagnose, resolve, and perform root cause analysis for incidents.Drive post-incident reviews and continuous improvement in reliability.
- Collaboration Mentorship Work closely with Dev, QA, and Security teams to ensure production readiness.Champion SRE best practices across teams.Mentor junior engineers on cloud-native operations and observability.
Required Skills Qualifications
- Experience: 3-6 years in SRE, DevOps, or Cloud Engineering.
- Cloud Platforms: Strong experience with AWS (EC2, S3, Lambda, CloudWatch, RDS, VPC, etc.) and Azure (VMs, Functions, Monitor, Networking, etc.) .
- Containers Orchestration: Hands-on with Kubernetes (EKS, AKS) and containerization (Docker).
- Programming: Proficient in Python for automation and system integrations.
- CI/CD Tools: Experience with Jenkins, GitHub Actions, GitLab, Azure DevOps .
- IaC: Proficiency with Terraform / CloudFormation / Bicep .
- Observability: Experience with Prometheus, Grafana, ELK/EFK, Datadog, New Relic, or similar .
- Incident Management: Familiar with PagerDuty, OpsGenie, or equivalent tools .
- Networking Security: Solid understanding of IAM, VPC, firewalls, encryption, and compliance .
- Soft Skills: Robust analytical, troubleshooting, and collaboration skills.
Preferred Qualifications
- Certifications: AWS Certified SysOps Administrator / Solutions Architect , Microsoft Azure Administrator Associate , CKA (Certified Kubernetes Administrator) .
- Experience with multi-cloud deployments .
- Familiarity with service mesh (Istio/Linkerd) and advanced observability (OpenTelemetry)
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Site Reliability Engineer - Level - 3 (Noida)
🏢 CorroHealth
📍 Noida