02 Aug
|
Innodata
|
India
Site Reliability Engineer (SRE) / DevOps Engineer – AWS, Kubernetes (EKS), CI/CD Location: Remote / base location Noida Experience: 7+ years (flexible based on depth in AWS + Kubernetes + production ops) Role Summary: We are looking for an experienced Site Reliability Engineer (SRE) to own reliability, scalability, automation, and operational excellence for cloud-native platforms running on AWS and Kubernetes (EKS). You will build and operate CI/CD pipelines, standardize infrastructure provisioning, implement monitoring/alerting, drive incident response, and partner with architects and engineering teams to deliver secure, cost- efficient, highly available systems.
Key Responsibilities: Reliability & Operations (SRE Core) • Own availability, latency, performance, and capacity for production workloads.
- Define and track SLIs/SLOs, error budgets, and reliability KPIs.
- Run incident management (on-call, triage, RCA/postmortems, preventive actions).
- Improve MTTR through automation, runbooks, and self-healing patterns. Kubernetes & Platform Engineering (EKS) • Operate and evolve AWS EKS clusters (multi-namespace, multi-env).
- Manage deployments using Helm/Kustomize, and enable safe rollouts (blue/green, canary).
- Handle ingress (ALB/NLB), service discovery, autoscaling (HPA/VPA/Cluster Autoscaler).
- Implement policies using RBAC, NetworkPolicies, Pod Security, secrets management. CI/CD & Release Engineering • Build and maintain CI/CD pipelines using GitHub Actions / Jenkins / GitLab CI (based on org standard).
- Enforce best practices: pipeline templates, approvals,
artifact versioning, rollback strategy.
- Container build & security scanning workflows (SAST/DAST/image scanning) and SBOM where required.
- Promote everything in Git culture: infra/app configs + deployment manifests. Infrastructure as Code & Cloud Automation • Provision cloud infra using Terraform / CloudFormation / CDK (as applicable).
- Implement reusable modules for VPC, IAM, EKS, RDS, S3, CloudFront, WAF, Route53, etc.
- Manage multi-account AWS access patterns (dedicated Dev accounts, IAM roles, federation/SSO).
- Implement cost controls: tagging standards, budgeting alerts, right-sizing, savings plans guidance. Observability (Monitoring, Logging, Tracing) • Own observability stack: Prometheus, Grafana, CloudWatch, Alertmanager.
- Implement metrics and dashboards for K8s cluster health + application SLIs.
- Implement logging pipeline (e.g., Fluent Bit/Fluentd → CloudWatch/ELK/OpenSearch).
- Enable distributed tracing (OpenTelemetry / X-Ray / Jaeger) where needed. Security & Governance (DevSecOps) • Implement edge and app protection using AWS WAF / CloudFront, and coordinate rollout safely.
- Enforce least privilege IAM, secrets handling, KMS encryption,
secure networking.
- Support vulnerability remediation (CVEs), patching processes, and audit readiness. Collaboration & Process • Partner with Engineering + Architecture to review designs and keep solutions simple.
- Create/maintain runbooks, SOPs, onboarding guides, and deployment standards.
- Raise infra requests and coordinate with platform/GT teams to ensure timely provisioning. Must-Have Skills • Robust hands-on experience with AWS (EKS, EC2, IAM, VPC, ALB/NLB, CloudWatch, S3, RDS).
- Deep experience managing Kubernetes in production, preferably EKS.
- Strong CI/CD ownership: pipelines, environments, release controls, rollback strategies.
- Observability experience: Prometheus + Grafana, alerting, dashboards, incident response.
- Infrastructure as Code: Terraform (preferred) or CloudFormation/CDK.
- Solid Linux + networking fundamentals (DNS, TLS, load balancing, routing).
- Strong scripting skills: Python/Bash. Good-to-Have Skills • Service mesh (Istio/Linkerd), OPA/Gatekeeper, Kyverno.
- External SaaS monitoring (Grafana Cloud, Datadog, New Relic).
- GitOps tools: Argo CD / Flux.
- Experience with WAF tuning, bot rules, rate limiting, and safe production rollout.
- PostgreSQL/MySQL operations basics (performance, monitoring, connection pooling).
- Experience with high-volume ingestion/data pipelines is a plus. AWS,GitHub Actions , Jenkins , GitFlows, Terraforms, EKS If you have relevant exp share your resume to [Confidential Information]
📌 Site Reliability Engineer (India)
🏢 Innodata
📍 India