04 Aug
|
Yash Technologies
|
Bengaluru
04 Aug
Yash Technologies
Bengaluru
Role: Platform SRE + AI
Experience: 7–10 yrs
Location: Bangalore Senior Platform SRE: Owns reliability and performance of the platform end to end — leads incident response, drives cluster and runtime tuning standards, and sets direction for the AI/MCP platform.
Expected to mentor and influence vendor and stakeholder decisions.
Looking for a strong AWS & Kubernetes (EKS) expert to manage production-scale cloud infrastructure, platform reliability, performance tuning, and automation.
Experience with AI/LLM platform operations, monitoring, troubleshooting, and cloud-native environments is highly preferred.
This is a deep platform and performance-engineering role — owning production Kubernetes, AWS infrastructure, and the reliability of LLM-backed and agentic services in a regulated environment.
It goes well beyond standard SRE: we are looking for an engineer who tunes clusters, diagnoses memory and runtime behaviour, and operates AI platform workloads at production scale Mandatory Skills: Deep AWS expertise — End-to-end production deployment on AWS is a must — EC2, ECS, EKS, Lambda, S3, IAM, RDS, API Gateway, VPC, EFS, SNS, SQS, EventBridge, CodeBuild.
Kubernetes (deep & mandatory) — Production EKS ownership — cluster configuration (CPU/ RAM sizing, node groups, autoscaling), workload fine-tuning (requests/limits, HPA/VPA, eviction policies),
and hands-on Helm chart management.
Memory expertise — Container vs. runtime memory models, OOMKill diagnosis, cgroup behaviour, JVM/Python heap tuning, and database buffer-pool configuration.
Serverless fine-tuning — Lambda memory/concurrency/cold-start optimization, ECS/Fargate task sizing, and serverless cost-performance tradeoffs.
Database operations — Running and tuning RDS or equivalent in production; StatefulSet-based DB deployments in K8s a strong plus.
AI platform & MCP workplace — Hands-on deployment or operation of LLM-backed services, MCP servers, or agentic pipelines on cloud infrastructure.
Core Requirements: 5–10 years in software development, DevOps, SRE or production support roles.
Working knowledge of GCP and/or Azure — multi-cloud integrations, platform differences, hybrid workloads.
CI/CD with GitHub Actions, Jenkins or GitLab.
Scripting proficiency in Python and/or Bash; YAML fluency.
Terraform and/or CloudFormation for infrastructure provisioning.
ITIL framework across Incident, Change, Problem and CAPA management.
Monitoring with Splunk and/or Grafana — including infra-level resource and memory dashboards.
ServiceNow and JIRA; strong ITSM discipline.
Bachelor’s degree or equivalent practical experience
📌 Platform SRE (Bengaluru)
🏢 Yash Technologies
📍 Bengaluru