11 Sep
|
EliteSquad.AI
|
India
11 Sep
EliteSquad.AI
India
Lead Cloud Platform Engineer – AWS & Databricks
Experience
10+ years overall IT experience
5+ years AWS / Cloud Platform Engineering
3+ years Databricks Administration / Platform Support Role
Work Mode: Remote
Job Summary
We are looking for an experienced Lead Cloud Platform Engineer / Cloud Architect with strong expertise in AWS infrastructure, Terraform, DevOps automation, Databricks platform administration, cloud security, monitoring and reliability engineering .
The role will be primarily focused on AWS Cloud Platform Engineering, Infrastructure as Code and DevOps automation , with approximately 20% focus on Databricks platform administration and governance .
The successful candidate will lead the design, implementation, governance, optimization and support of enterprise-scale AWS and Databricks environments. The candidate should be comfortable acting as a technical lead/architect , mentoring engineering teams and serving as an escalation point for complex cloud and platform issues.
Core Focus
- 80% – AWS / Terraform / DevOps / Cloud Platform Engineering
- 20% – Databricks Platform Administration & Governance
- Technical leadership and architecture
- Cloud security and governance
- Platform reliability and observability
- CI/CD and Infrastructure as Code
Key Responsibilities
1. Technical Leadership & Architecture
- Lead the design, implementation and support of enterprise-scale AWS and Databricks platforms .
- Define cloud platform architecture, engineering standards and best practices.
- Provide technical leadership and mentorship to Cloud, Platform and DevOps engineers.
- Act as the primary escalation point for complex infrastructure and platform issues.
- Drive cloud modernization, platform transformation and automation initiatives.
- Review existing architectures and recommend improvements for scalability, security, reliability and cost optimization.
- Collaborate with Data Engineering, Security, Application and Infrastructure teams.
1. AWS Infrastructure & Cloud Operations
Design, deploy, manage and optimize AWS environments across enterprise workloads. Strong hands-on experience with:
- EC2
- VPC
- IAM
- S3
- EKS
- Lambda
- RDS
- CloudWatch
- AWS Secrets Manager
- Networking and infrastructure architecture
- Security and governance
- High availability and disaster recovery
- Cloud resource utilization and cost optimization
Responsibilities include:
- Design secure and scalable AWS architectures.
- Manage compute, networking, storage and database infrastructure.
- Implement IAM, security controls and access policies.
- Review and optimize AWS resource utilization and performance.
- Troubleshoot complex cloud infrastructure issues.
- Establish cloud governance and operational standards.
- Support production environments and critical enterprise workloads.
1. Terraform & Infrastructure as Code
This is one of the most important areas of the role . The candidate should have strong hands-on Terraform experience, not just theoretical knowledge.
Responsibilities
- Design and implement Terraform-based Infrastructure as Code .
- Build reusable Terraform modules for AWS infrastructure.
- Manage infrastructure across Development, QA/UAT and Production environments.
- Implement Terraform state management and remote backends.
- Establish IaC standards, branching and deployment strategies.
- Perform infrastructure provisioning, changes and upgrades through Terraform.
- Implement automated infrastructure deployment through CI/CD.
- Troubleshoot Terraform plan/apply/state issues.
- Drive Terraform adoption across cloud environments.
- Follow IaC security and governance best practices.
Candidate should be able to explain: Terraform → Git → CI/CD → AWS and how infrastructure changes move safely from development to production.
1. DevOps & CI/CD
Design and implement enterprise DevOps automation and deployment frameworks. Experience with one or more of:
- GitHub Actions
- Azure DevOps
- Jenkins
- GitLab CI/CD
- Git
- CI/CD pipelines
- Release management
- Infrastructure automation
Responsibilities:
- Design enterprise CI/CD frameworks.
- Automate infrastructure and application deployments.
- Integrate Terraform with CI/CD pipelines.
- Implement automated validation, testing and deployment.
- Establish release management processes.
- Implement environment promotion strategies.
- Promote DevSecOps practices throughout the SDLC.
- Integrate security scanning and compliance checks into pipelines.
1. Databricks Platform Administration
Databricks represents approximately 20% of the role , but the candidate still needs genuine administration/platform experience.
Responsibilities
- Administer and manage Databricks workspaces .
- Configure and optimize Databricks environments.
- Manage clusters and cluster policies.
- Implement cluster sizing and optimization strategies.
- Manage Databricks jobs and scheduling.
- Support job orchestration and operational workflows.
- Implement and manage Unity Catalog .
- Configure access controls and permissions.
- Manage workspace governance.
- Support multiple Databricks environments.
- Troubleshoot Databricks platform and cluster issues.
- Work with Data Engineering teams to improve platform scalability and operational efficiency.
- Monitor Databricks workloads and identify performance bottlenecks.
Important Databricks areas
Candidates should understand
Workspace → Cluster → Cluster Policy → Jobs → Unity Catalog → Catalog/Schema/Table → Permissions → Governance
1. Databricks Governance & Security
- Implement Unity Catalog governance.
- Manage users, groups, roles and permissions.
- Configure access controls.
- Establish workspace governance policies.
- Implement secure data access patterns.
- Support data platform security and compliance requirements.
- Work with security teams on cloud and Databricks controls.
- Manage environment-level access and governance.
1. Monitoring & Observability
Implement monitoring and reliability strategies across AWS and Databricks.
Experience with
- AWS CloudWatch
- Datadog
- Grafana
- Splunk
- New Relic
Responsibilities:
- Define monitoring and observability strategies.
- Build operational dashboards.
- Configure alerts and notifications.
- Establish platform health metrics.
- Define and monitor SLAs / SLOs .
- Monitor AWS infrastructure and Databricks workloads.
- Perform incident response and troubleshooting.
- Conduct Root Cause Analysis (RCA).
- Drive continuous improvement initiatives.
- Improve platform availability, reliability and performance.
1. Automation & Scripting
Strong automation skills are expected.
- Python
- Shell scripting
- AWS automation
- Terraform automation
- CI/CD automation
- Operational tooling
The candidate should be able to automate repetitive cloud/platform administration tasks rather than manually performing them.
1. Security, Governance & Compliance
- Implement AWS security best practices.
- Manage IAM and least-privilege access.
- Implement secrets management using AWS Secrets Manager .
- Establish cloud governance standards.
- Support security and compliance requirements.
- Implement DevSecOps practices.
- Review infrastructure for security vulnerabilities.
- Ensure platform configurations follow enterprise security standards.
Required Technical Skills
Must Have
AreaRequired Skills
Cloud
AWS
AWS Infrastructure
EC2, VPC, IAM, S3, EKS, Lambda, RDS
Monitoring
CloudWatch
Security
IAM, Secrets Manager, Cloud Security
IaC
Terraform – strong hands-on
DevOps
CI/CD, Git, GitHub Actions / Jenkins / Azure DevOps / GitLab
Databricks
3+ years administration/platform support
Databricks Governance
Unity Catalog, workspace governance, access control
Databricks Operations
Cluster management, jobs, scheduling, troubleshooting
Scripting
Python, Shell
Observability
Datadog, Grafana, Splunk, New Relic, CloudWatch
Leadership
Technical leadership, architecture, mentoring
Architecture
Cloud/platform architecture and transformation
Required Experience
- 10+ years of overall IT experience
- 5+ years of AWS Cloud / Platform Engineering
- 3+ years of Databricks administration and platform support
- Strong hands-on Terraform / IaC experience
- Strong DevOps and CI/CD experience
- Experience designing enterprise cloud infrastructure
- Experience with cloud security and governance
- Experience supporting enterprise data/analytics platforms
- Experience leading engineering teams or platform transformation initiatives
- Robust troubleshooting and incident-management experience
- Excellent communication and stakeholder-management skills
📌 Lead Cloud Platform Engineer – AWS, Terraform & Databricks (India)
🏢 EliteSquad.AI
📍 India