31 Jul
|
Model N
|
Telangana
As a Systems Development Engineer with an SRE focus, you ll build and maintain highly available, observable, and efficient systems that support mission critical services. You will focus on CI/CD automation, infrastructure and configuration as code, and deep monitoring/observability , while owning incident response and reliability for cloud native applications on AWS and Kubernetes.
Job Responsibilities
- Design, build, and maintain automated CI/CD pipelines using tools such as Harness, GitHub Actions, and ArgoCD.
- Develop and maintain infrastructure and configuration as code using CloudFormation, Terraform, Ansible , and related automation tools.
- Administer and optimize AWS environments , including core services, networking, security, and architecture for availability, performance, and cost.
- Manage and support Kubernetes clusters and containerized workloads, including configuration, scaling, and upgrades.
- Design, implement, and evolve end to end monitoring and observability frameworks using tools such as Open Telemetry, Groundcover, CloudWatch, Datadog, Prometheus, New Relic , or similar platforms.
- Create and maintain dashboards, logs, traces, SLIs/SLOs, and automated alerting systems to ensure reliability and rapid detection of anomalies.
- Embed observability, CI/CD best practices, and operational readiness into all stages of the software development lifecycle in partnership with engineering teams.
- Lead or participate in incident response, troubleshooting, and root cause analysis for production incidents, using observability data to drive fast resolution.
- Automate operational tasks, runbooks, and incident remediation workflows to reduce toil and improve service reliability.
- Contribute to risk mitigation, backup, and disaster recovery strategies, including periodic testing and continuous improvement.
- Participate in shared after hours support and project work as needed.
Job Qualifications
- 2-4 years of experience designing, implementing, and maintaining CI/CD pipelines (e.g., Harness, GitHub Actions, ArgoCD or similar tools).
- Hands on experience with automation tools and Infrastructure as Code / Configuration as Code (CloudFormation, Terraform, Ansible).
- Robust understanding of Infrastructure as Code and Configuration as Code principles and patterns.
- Solid grasp of the software development lifecycle and modern SRE/DevOps practices.
- AWS administration and architecture experience, including networking, security, IAM, and core services.
- Experience operating Kubernetes clusters (EKS or other distributions) and containerized workloads.
- Deep experience with monitoring and observability tools such as Open Telemetry, Groundcover, CloudWatch, Datadog, Prometheus, New Relic, or equivalent, including metrics, logs, and traces.
- Ability to define and track SLIs/SLOs and use them to guide reliability improvements.
- Proficiency in Linux administration, including system configuration, troubleshooting, and performance tuning.
- Programming/scripting skills in at least one language such as Python, Go, or Rust for automation, tooling, and observability integrations.
- Solid understanding of networking, load balancing, and performance tuning.
- Experience troubleshooting complex distributed systems, supporting incident response, and driving root cause analysis.
- Familiarity with risk mitigation, backup, and disaster recovery concepts.
Preferred
- Experience building unified observability platforms or standardized dashboards for multiple services/teams.
- Experience with GitOps workflows and tools for declarative infrastructure and application delivery.
- Background in incident command and post mortem frameworks.
- Experience integrating observability and reliability practices into microservices and/or serverless architectures.
- Experience integrating testing, security and compliance checks into CI/CD pipelines.
📌 Systems Development Engineer (SRE/DevOps) (Telangana)
🏢 Model N
📍 Telangana