04 Aug
|
Yash Technologies
|
Bengaluru
04 Aug
Yash Technologies
Bengaluru
Role: Platform SRE + AI
Experience: 7–10 yrs
Location: Bangalore
Senior Platform SRE:
- Owns reliability and performance of the platform end to end — leads incident response, drives cluster and runtime tuning standards, and sets direction for the AI/MCP platform.
- Expected to mentor and influence vendor and stakeholder decisions.
- Looking for a strong AWS & Kubernetes (EKS) expert to manage production-scale cloud infrastructure, platform reliability, performance tuning, and automation.
- Experience with AI/LLM platform operations, monitoring, troubleshooting, and cloud-native environments is highly preferred.
- This is a deep platform and performance-engineering role — owning production Kubernetes, AWS infrastructure, and the reliability of LLM-backed and agentic services in a regulated environment.
- It goes well beyond standard SRE: we are looking for an engineer who tunes clusters, diagnoses memory and runtime behaviour, and operates AI platform workloads at production scale
Mandatory Skills:
- Deep AWS expertise — End-to-end production deployment on AWS is a must — EC2, ECS, EKS, Lambda, S3, IAM, RDS, API Gateway, VPC, EFS, SNS, SQS, EventBridge, CodeBuild.
- Kubernetes (deep & mandatory) — Production EKS ownership — cluster configuration (CPU/ RAM sizing, node groups, autoscaling), workload fine-tuning (requests/limits, HPA/VPA, eviction policies),
and hands-on Helm chart management.
- Memory expertise — Container vs. runtime memory models, OOMKill diagnosis, cgroup behaviour, JVM/Python heap tuning, and database buffer-pool configuration.
- Serverless fine-tuning — Lambda memory/concurrency/cold-start optimization, ECS/Fargate task sizing, and serverless cost-performance tradeoffs.
- Database operations — Running and tuning RDS or equivalent in production; StatefulSet-based DB deployments in K8s a strong plus.
- AI platform & MCP environment — Hands-on deployment or operation of LLM-backed services, MCP servers, or agentic pipelines on cloud infrastructure.
Core Requirements:
- 5–10 years in software development, DevOps, SRE or production support roles.
- Working knowledge of GCP and/or Azure — multi-cloud integrations, platform differences, hybrid workloads.
- CI/CD with GitHub Actions, Jenkins or GitLab.
- Scripting proficiency in Python and/or Bash; YAML fluency.
- Terraform and/or CloudFormation for infrastructure provisioning.
- ITIL framework across Incident, Change, Problem and CAPA management.
- Monitoring with Splunk and/or Grafana — including infra-level resource and memory dashboards.
- ServiceNow and JIRA; robust ITSM discipline.
- Bachelor’s degree or equivalent practical experience
📌 Platform SRE (Bengaluru)
🏢 Yash Technologies
📍 Bengaluru