16 Sep
|
Acuity Analytics
|
Bengaluru
16 Sep
Acuity Analytics
Bengaluru
Acuity Analytics (the trading name of Acuity Knowledge Partners) is a global, tech-first organization helping financial institutions and corporates make better decisions through research, data, analytics and AI-enabled solutions. We combine deep financial services expertise with strong engineering, digital and AI capabilities to solve complex, real-world problems.
With a team of 7,200+ analysts, data specialists and technologists across 28 locations, we work with more than 800 organizations worldwide to drive efficiency, unlock insight and deliver measurable impact. Our success is built on the strength of our people—by investing in talent, encouraging collaboration and creating room to grow, we enable our teams to do their best work for clients.
Acuity became an independent business in 2019 following its acquisition from Moody’s Corporation by Equistone Partners Europe. In 2023, funds advised by global private equity firm Permira acquired a majority stake, with Equistone remaining a minority investor—supporting our continued growth and innovation.
For more information, visit www.acuityanalytics.com
Job Purpose
We are looking for a 12+ year experienced Senior Staff AI DevOps Lead / Engineering Manager to own the infrastructure, automation, and reliability of our AI/ML and GenAI platforms, while leading and growing a high-performing DevOps/MLOps team. The ideal candidate combines deep hands-on infrastructure expertise with strong engineering-management skills, and has experience operating LLM-based and agentic systems at production scale.
Key Responsibilities
- Define and own the DevOps/MLOps architecture and roadmap for enterprise AI and GenAI platforms.
- Design, build, and operate CI/CD pipelines for model training, fine-tuning, evaluation, and deployment.
- Lead infrastructure strategy for LLM inference, vector databases, RAG pipelines,
and agentic workflow services.
- Build and manage Kubernetes-based platforms for scalable, secure, cost-productive AI workloads (CPU and GPU).
- Establish observability, logging, tracing, and alerting standards across AI services (LLM latency, token usage, model drift, cost).
- Own infrastructure-as-code practices (Terraform, Helm, Pulumi) across cloud environments (AWS/Azure/GCP).
- Design and enforce security, secrets management, and access-control standards for AI/ML pipelines and data.
- Drive incident management, on-call practices, SLAs/SLOs, and disaster-recovery planning for production AI systems.
- Manage cost optimization and capacity planning for GPU/compute-intensive AI workloads.
- Hire, mentor, and grow a team of DevOps and MLOps engineers; conduct performance reviews and career development.
- Partner with AI/ML, Backend, Frontend, Data, Security, and Product teams to ship reliable, production-grade AI features.
- Own technical delivery from architecture and prototyping through production deployment, monitoring, and continuous improvement.
Key competencies
- 12+ years of software/infrastructure engineering experience, with 3+ years in a technical leadership or engineering-management capacity.
- Proven experience designing and operating CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, ArgoCD) for AI/ML workloads.
- Strong hands-on expertise with Kubernetes, Docker, Helm, and container orchestration at scale.
- Experience with MLOps/LLMOps tooling (MLflow, Kubeflow, SageMaker, Vertex AI, or equivalent) for model lifecycle management.
- Deep knowledge of infrastructure-as-code (Terraform, Pulumi) and cloud platforms (AWS, Azure, or GCP).
- Experience operating GPU infrastructure and optimizing cost/performance for LLM inference and training workloads.
- Strong understanding of observability stacks (Prometheus, Grafana, OpenTelemetry, Datadog) and incident response practices.
- Working knowledge of LLM APIs, RAG architectures, vector databases, and agentic workflow orchestration (e.g., Camunda, LangGraph, or similar).
- Experience with security best practices: secrets management, IAM, network policy, and compliance for regulated environments.
- Track record of building and leading engineering teams, including hiring, mentoring, and performance management.
- Excellent stakeholder management and communication skills, with the ability to work across Product, AI/ML, Data, Security, and Backend teams.
- Familiarity with financial-services or other regulated-industry data and security requirements is a plus.
Acuity Analytics has earned several prestigious industry recognitions, including Great Place to Work® certifications in India and Costa Rica, AVTAR Best Companies for Women in India, the AVTAR Most Inclusive Companies Index, and silver accreditation in the Workplace Equality Index. These accolades reflect our commitment to building an inclusive, supportive and high-performance workplace for our people.
Follow us on social media to stay updated with Acuity Analytics news
- LinkedIn: https://www.linkedin.com/company/acuityanalytics/
- Facebook: https://facebook.com/AcuityAnalytics/
- Twitter: https://x.com/acuityanalytic?s=21
- YouTube: https://www.youtube.com/@AcuityAnalytics
📌 AI Data Engineering Lead (Bengaluru)
🏢 Acuity Analytics
📍 Bengaluru