08 Sep
|
Live Connections
|
Chennai
08 Sep
Live Connections
Chennai
Location: Chennai
Experience: 7-15 Years
Mandatory Skill: GCP Cloud
Notice Period: Immediate to 30 Days
Knowledge, Skills, and Abilities
- Proficiency with at least one major cloud platform (Azure, AWS, GCP) and its core infrastructure services.
- Deep expertise with Infrastructure as Code (IaC) tools (e.g., Terraform, CloudFormation) to automate the deployment of complex application and AI environments.
- Solid experience with containerization and orchestration (such as Docker, Kubernetes).
- Expertise in provisioning, managing, and optimizing various compute instances, including GPU-accelerated compute for training and inference workloads.
- In-depth knowledge of cloud-native platforms for both application development and AI/ML (e.g., Azure App Service, AWS Lambda, Google Vertex AI, AWS SageMaker) and their underlying infrastructure.
- Familiarity with high-performance networking and storage configurations required for distributed systems, model training, and low-latency data access.
- Experience implementing and managing CI/CD pipelines for infrastructure, application, and ML model deployments (e.g., GitHub Actions, Azure DevOps).
- Strong skills in cloud cost management and optimization, with specific experience in managing the high costs associated with specialized resources like GPU/TPU.
- Knowledge of logging, monitoring, and observability tools, with a focus on monitoring the performance and utilization of all cloud infrastructure.
- Ability to architect for high availability, security, and scalability in the context of mission-critical application and AI services.
- Experience with cloud migrations.
Key Responsibilities / Day-to-Day Activities
- Design, build, and maintain scalable, automated cloud infrastructure for enterprise application, AI, and MLOps platforms.
- Automate the provisioning and management of storage and compute resources to support the entire software application development and machine learning lifecycles.
- Implement and manage the underlying containerized infrastructure that hosts application layers, model training, batch inference, and real-time serving endpoints.
- Collaborate with Software Engineers, AI Engineers, and MLOps Engineers to build robust CI/CD pipelines that automate the testing and deployment of infrastructure, applications, and machine learning models.
- Continuously monitor and optimize the performance of cloud workloads, implementing strategies for resource scheduling, auto-scaling, and cost-effective instance utilization.
- Implement and enforce cloud security best practices, including network isolation, IAM policies, and data encryption for all environments.
- Troubleshoot and resolve complex infrastructure issues that impact the performance or availability of production applications and models.
- Manage the infrastructure for core data services, including databases, vector databases, feature stores, and high-throughput data streaming platforms.
- Evaluate and benchmark emerging cloud technologies and services (e.g., new instance types, serverless technologies, custom AI chips) to ensure the enterprise remains on the cutting edge of cloud infrastructure.
- Serve as the subject matter expert on cloud infrastructure, providing guidance and best practices to engineering teams across the organization.
📌 GCP Cloud Engineer (Chennai)
🏢 Live Connections
📍 Chennai