31 Jul
|
Five Data Products And Solutions
|
Telangana
31 Jul
Five Data Products And Solutions
Telangana
We are seeking a highly skilled DevOps Engineer AI & Cloud Infrastructure to lead cloud infrastructure, platform reliability, and AI-driven operational automation initiatives. This role requires deep expertise in AWS, modern DevOps practices, SaaS operations, and leveraging AI technologies to automate infrastructure management, deployment pipelines, monitoring, and incident response.
The ideal candidate has experience supporting production SaaS platforms running on AWS, building highly reliable cloud infrastructure, and using AI/LLM technologies to reduce operational overhead and improve engineering productivity.
This position is part of the core infrastructure team responsible for maintaining platform availability, performance, security, and scalability.
Key Responsibilities
Cloud Infrastructure & Platform Engineering
- Design, implement, and manage secure, scalable, and highly available AWS infrastructure.
- Build and maintain Infrastructure as Code (IaC) using Terraform, CloudFormation, or equivalent tools.
- Manage Kubernetes clusters (EKS preferred), containerized applications, and cloud-native services.
- Optimize infrastructure performance, reliability, security, and cloud spend.
- Architect solutions that support high-availability SaaS environments with minimal downtime.
AI-Powered DevOps Automation
- Leverage AI/Generative AI technologies to automate DevOps and infrastructure operations.
- Build AI-driven workflows for:
- Automated CI/CD pipeline creation and maintenance.
- Infrastructure provisioning and configuration management.
- Monitoring and alert generation.
- Incident triage and root-cause analysis.
- Automated remediation and self-healing workflows.
- Infrastructure documentation generation.
- Utilize tools such as OpenAI, Anthropic Claude, Amazon Bedrock, GitHub Copilot, Cursor, and similar AI platforms to improve operational efficiency.
- Evaluate and implement emerging AI technologies to streamline infrastructure management.
CI/CD & Release Engineering
- Design and maintain automated CI/CD pipelines for cloud-native applications.
- Automate deployments, testing, security scans, approvals, and rollback procedures.
- Create reusable deployment templates and deployment automation frameworks.
- Improve release reliability, deployment frequency, and developer productivity.
Monitoring,
Observability & Reliability
- Design and implement comprehensive monitoring, logging, and observability platforms.
- Build intelligent alerting systems that minimize noise and accelerate issue detection.
- Implement and manage tools such as:
- Datadog
- Grafana
- Prometheus
- CloudWatch
- OpenSearch / ELK
- New Relic
- Develop automated incident response and remediation workflows.
- Establish SRE best practices including SLIs, SLOs, error budgets, and reliability metrics.
SaaS Operations & Production Support
- Support mission-critical SaaS applications running in AWS production environments.
- Participate in infrastructure and production incident response.
- Perform maintenance, upgrades, deployments, and infrastructure changes outside of business hours when required to minimize customer impact.
- Be available for occasional evenings, weekends, and on-call support during critical releases, incidents, or infrastructure maintenance windows.
- Collaborate with engineering teams to ensure high availability and operational excellence.
Security & Compliance
- Implement DevSecOps best practices throughout the software delivery lifecycle.
- Automate security controls, compliance monitoring, and vulnerability management.
- Support SOC 2, ISO 27001, and other compliance initiatives.
- Maintain secure cloud architectures and access management policies.
Required Qualifications
- 4+ years of experience in DevOps, Site Reliability Engineering (SRE), Infrastructure Engineering, or Platform Engineering.
- Must be based in India.
- Proven experience supporting production SaaS applications hosted on AWS.
- Experience working in customer-facing production SaaS environments with uptime and SLA commitments.
- Robust expertise with AWS services including:
- EC2
- EKS/ECS
- Lambda
- VPC
- IAM
- RDS
- S3
- CloudWatch
- Route 53
- Secrets Manager
- Strong experience with Infrastructure as Code:
- Terraform (preferred)
- CloudFormation
- Experience building CI/CD pipelines using:
- GitHub Actions
- GitLab CI/CD
- Jenkins
- AWS CodePipeline
- Experience using AI tools and LLMs to automate DevOps workflows, CI/CD pipelines, infrastructure provisioning, monitoring, alerting, troubleshooting, and operational tasks.
- Experience with containerization and orchestration:
- Docker
- Kubernetes
- Strong scripting and automation skills using Python, Bash, or Go.
- Hands-on experience implementing monitoring, alerting, logging, and observability solutions.
- Strong troubleshooting and production incident management experience.
Preferred Qualifications
- Hands-on experience with Google Cloud Platform (GCP), including:
- GKE
- Cloud Run
- Cloud Functions
- Cloud Monitoring
- BigQuery
- Experience building AI-powered infrastructure automation solutions.
- Experience with AIOps platforms and automated remediation systems.
- Familiarity with:
- Amazon Bedrock
- OpenAI APIs
- Anthropic Claude
- LangChain
- AI Agent Frameworks
- Experience managing multi-cloud environments.
- AWS, Kubernetes, Terraform, or GCP certifications.
Additional Expectations
- Must be comfortable supporting production environments outside normal business hours when necessary.
- Ability to perform planned infrastructure maintenance during evenings or weekends to avoid customer downtime.
- Participate in on-call rotations and critical incident response.
- Demonstrated ownership mentality and ability to independently drive operational excellence initiatives.
- Strong communication skills and ability to collaborate across engineering, security, product, and customer-facing teams.
- Highly organized and methodical with the ability to manage multiple concurrent projects, incidents, and operational priorities.
- Comfortable operating in a fast-paced SaaS environment where priorities may shift based on customer, business, or operational needs.
- Demonstrated ability to balance strategic infrastructure initiatives with urgent production support requirements.
- Strong task management and execution skills, including tracking deliverables, managing dependencies, and driving initiatives to completion.
- Ability to effectively prioritize work based on business impact, customer commitments, system reliability, and security requirements.
📌 Devops AI Cloud Engineer (Telangana)
🏢 Five Data Products And Solutions
📍 Telangana