03 Sep
|
Important Business
|
Chennai
03 Sep
Important Business
Chennai
Key Responsibilities:
● Architect and manage the end-to-end SaaS platform infrastructure on AWS, including EKS cluster design, VPC networking, IAM, and multi-region availability.
● Build, maintain, and optimize Jenkins-based CI/CD pipelines and develop Python automation scripts for provisioning, deployments, and runbook automation.
● Define and enforce platform SLOs/SLAs; own the observability strategy across logging, metrics, and tracing.
● Manage and participate in the on-call rotation; act as escalation point for P1/P2 incidents and drive post-incident reviews.
● Drive Infrastructure-as-Code (IaC) practices with Terraform/CloudFormation and champions a culture of automation and operational excellence.
● Collaborate cross-functionally with product, security, and engineering teams to align infrastructure roadmap with business goals.
● Identify opportunities to leverage AI to automate operational and DevOps workflows. ● Design and implement AI-assisted solutions for incident triaging, root cause analysis, log analysis, and performance optimization.
● Drive the adoption of AI-powered tools for infrastructure management, deployment automation, monitoring, and troubleshooting.
● Build intelligent workflows that reduce manual effort in release management, capacity planning, and operational support.
● Integrate AI capabilities into CI/CD pipelines to improve code quality, deployment reliability, and operational efficiency.
● Collaborate with engineering teams to automate repetitive tasks and improve developer productivity.
● Define best practices and governance for the safe and effective use of AI across DevOps processes.
● Measure and report on productivity gains, operational improvements, and cost savings achieved through AI adoption.
Requirements:
● 10+ years of experience in DevOps, SRE, or cloud infrastructure engineering roles. ● Deep hands-on expertise with AWS services (EC2, EKS, RDS, S3, IAM, VPC, CloudFront, Route53, Lambda, etc).
● Robust Kubernetes experience: cluster management, Helm, autoscaling (HPA/KEDA). ● Proficiency with Jenkins for complex CI/CD pipeline design and maintenance.
Internal 3
JD | V.1.0 | RD:18-May-2026
● Solid Python scripting skills for automation, tooling, and infrastructure management tasks. ● Experience with Infrastructure-as-Code using Terraform and/or AWS Cloud Formation. ● Proven track record of architecting and managing end-to-end SaaS products in a cloud-native environment.
● Strong understanding of networking fundamentals, security best practices, and compliance frameworks (SOC 2, ISO 27001 a plus).
● Hands-on experience with on-call processes and incident management frameworks.
Preferred Qualifications:
● AWS certifications: Solutions Architect Professional, DevOps Engineer Professional, or equivalent.
● Familiarity with service mesh, secrets management (Vault, AWS Secrets Manager), and zero-trust security models.
● Experience with multi-tenant SaaS architectures and tenant isolation strategies. ● Knowledge of FinOps principles and AWS cost management tooling.
● Experience with database DevOps: RDS, Aurora schema migrations, and backup strategies.
Core Competencies:
● Strategic Thinking – ability to translate business goals into scalable technical architecture. ● Operational Excellence – strong bias for reliability, automation, and continuous improvement.
● Communication – ability to clearly articulate complex technical topics to non-technical stakeholders.
● Ownership Mindset – proactively identifies and resolves risks without waiting to be asked.
● Resilience Under Pressure – calm and decisive during incidents; leads by example in high-stress situations.
Requirements
Key Responsibilities:
● Architect and manage the end-to-end SaaS platform infrastructure on AWS, including EKS cluster design, VPC networking, IAM, and multi-region availability.
● Build, maintain, and optimize Jenkins-based CI/CD pipelines and develop Python automation scripts for provisioning, deployments, and runbook automation.
● Define and enforce platform SLOs/SLAs; own the observability strategy across logging, metrics, and tracing.
● Manage and participate in the on-call rotation; act as escalation point for P1/P2 incidents and drive post-incident reviews.
● Drive Infrastructure-as-Code (IaC) practices with Terraform/CloudFormation and champions a culture of automation and operational excellence.
● Collaborate cross-functionally with product, security, and engineering teams to align infrastructure roadmap with business goals.
● Identify opportunities to leverage AI to automate operational and DevOps workflows. ● Design and implement AI-assisted solutions for incident triaging, root cause analysis, log analysis, and performance optimization.
● Drive the adoption of AI-powered tools for infrastructure management, deployment automation, monitoring, and troubleshooting.
● Build intelligent workflows that reduce manual effort in release management, capacity planning, and operational support.
● Integrate AI capabilities into CI/CD pipelines to improve code quality, deployment reliability, and operational efficiency.
● Collaborate with engineering teams to automate repetitive tasks and improve developer productivity.
● Define best practices and governance for the safe and effective use of AI across DevOps processes.
● Measure and report on productivity gains, operational improvements, and cost savings achieved through AI adoption.
Requirements:
📌 Principal DevOps Engineer (Chennai)
🏢 Important Business
📍 Chennai