04 Aug
|
Anaptyss
|
Noida
We are seeking a highly skilled and adaptable Senior DevOps Engineer to join our core engineering team. In this role, you will be the backbone of our infrastructure, responsible for designing, deploying, and maintaining the cloud architecture that powers our cutting-edge AI-native products and intelligence platforms.
Because we build complex, agentic AI workflows and intelligent applications (often requiring strict security and compliance standards), this role goes beyond traditional infrastructure. You will bridge the gap between our AI/ML engineering team and production, ensuring our models and applications scale efficiently, securely, and seamlessly. While our primary ecosystem is Google Cloud Platform (GCP), we operate in a dynamic environment where familiarity with other enterprise platforms like Azure and Salesforce is highly valued.
Key Responsibilities
Cloud Infrastructure & AI Deployment (MLOps)
GCP Architecture: Architect, build, and manage highly available, secure, and scalable infrastructure on Google Cloud Platform (GCP) using services like GKE (Kubernetes Engine), Cloud Run, Compute Engine, VPCs, and Cloud Storage.
AI/ML Product Deployment: Partner closely with data scientists and AI architects to deploy machine learning models, LLMs, and agentic workflows into production.
GPU & Compute Management: Optimize and provision infrastructure for AI workloads, managing GPU quotas, scaling policies, and compute costs effectively.
Infrastructure as Code (IaC): Write, test, and deploy infrastructure using Terraform, ensuring all environments are reproducible and version-controlled.
CI/CD & Automation
Pipeline Engineering: Design and maintain robust CI/CD pipelines (via GitHub Actions, GitLab CI, or Cloud Build) to automate the testing, building, and deployment of both application code and AI models.
Containerization: Deep expertise in Docker and Kubernetes. Manage container registries, helm charts, and microservices orchestration.
Workflow Automation: Automate routine operational tasks and deployment procedures using Python and Bash scripting.
Cross-Platform Integration & Adaptability
Multi-Cloud/Enterprise Integrations: Support the integration of our core GCP products with external client environments, specifically involving Microsoft Azure and Salesforce.
Rapid Upskilling: Act as a technical chameleon. If you lack direct experience with Azure or Salesforce architectures, you must demonstrate a proven track record of rapidly acquiring new platform skills to support cross-functional product requirements.
Security, Monitoring, & Reliability
Observability: Implement comprehensive logging, monitoring, and alerting solutions using tools like Prometheus, Grafana, and Google Cloud Operations (formerly Stackdriver) to ensure high availability of our AI services.
Enterprise Security & Compliance:
Enforce cloud security best practices (IAM, network security, encryption at rest/transit) to ensure our infrastructure meets the stringent compliance and control-testing requirements of our enterprise clients.
Incident Response: Act as a key point of contact for system outages, participating in root-cause analysis and implementing preventative measures.
Required Qualifications
Experience: 4 to 6 years of hands-on experience in DevOps, Cloud Engineering, or Site Reliability Engineering (SRE).
GCP Mastery: Proven production experience with Google Cloud Platform, particularly GKE, IAM, Cloud Build, and networking components.
AI/ML Context: Prior experience deploying AI, ML, or data-heavy applications to production (familiarity with Vertex AI, Hugging Face deployments, or local LLM hosting is a massive plus).
Tooling: Strong proficiency in Kubernetes, Docker, Terraform, and Linux administration.
Scripting/Coding: Robust programming skills in Python (essential for AI pipelines) and Bash.
Agile Mindset: Experience working in fast-paced, product-focused Agile environments.
Preferred Qualifications (The "Nice-to-Haves")
Hands-on experience with Microsoft Azure (AKS, Azure DevOps, Azure Entra ID).
Familiarity with Salesforce API integrations and managed packages from an infrastructure perspective.
Google Cloud Professional Cloud DevOps Engineer Certification (or equivalent).
Background in developing infrastructure for fintech, banking controls, or highly regulated domains.
📌 Senior DevOps Engineer (Noida)
🏢 Anaptyss
📍 Noida