16 Sep
|
Eli Lilly
|
Hyderabad
16 Sep
Eli Lilly
Hyderabad
Role summary
You own the agentic resolution platform end to end - both the cloud-native substrate it runs on and the intelligent agents that run on top of it. From Kubernetes, CI/CD, and observability through agent runtime, LLM patterns, retrieval, evaluation, and human-in-the-loop boundaries, you write the production code, design the architecture, and set the engineering bar that lets automation deflect routine operational work before it ever reaches a human.
This is a senior individual-contributor role that is both strategic and deeply hands-on. You ship production platform services and agentic capabilities, make the build-vs-buy calls on tooling and model access, and partner with the senior architect on the patterns that scale. Success is measured by rising deflection, sustained toil reduction, platform availability and developer velocity, and the ability to scale agentic intelligence through systems and people - not individual effort.
What you'll be doing
1. Agentic resolution platform - intake to action
- Architect the agentic resolution platform end to end: intake, classification, action execution, verification, and human-in-the-loop fallback.
- Build the agent runtime and orchestration layer: agent state and memory, tool integration, multi-agent coordination patterns, and confidence-thresholded handoffs to humans.
- Define agent-decision observability, audit-ready posture, and the data contracts that let the platform consume durable fixes from upstream engineering teams.
- LLM application patterns, knowledge, and evaluation rigor
- Design production LLM patterns: prompt engineering, retrieval-augmented generation (RAG), structured outputs, multi-model routing, and hybrid retrieval over the knowledge corpus.
- Own the knowledge-base strategy as a compounding deflection lever - every resolved incident becomes training data and structured retrieval input for future automation.
- Establish evaluation and guardrail frameworks for non-deterministic systems: automated evals, quality scoring, drift detection,
and feedback loops that compound agent quality over time.
- Cloud-native platform - build and operate
- Architect and operate Kubernetes (EKS or equivalent) at scale for container and serverless workloads supporting agentic and LLM inference traffic patterns.
- Write production platform services and internal tooling (Python or Go) that automate provisioning, deployment, and operational workflows - not just infrastructure configuration.
- Define and maintain infrastructure as code (Terraform) integrated with a major cloud's AI stack (AWS Bedrock/SageMaker, Azure AI Foundry, or Vertex AI), with secrets management and audit-ready posture for regulated environments.
- CI/CD, observability, and developer experience
- Build and maintain CI/CD pipelines tuned for agentic and AI workloads: model and agent versioning, canary rollouts, evaluation gates, and rollback.
- Own the observability stack (Prometheus / Grafana / OpenTelemetry plus enterprise tooling) and instrument platform health, agent-decision telemetry, and model-inference metrics.
- Establish SLOs, SLIs, and reliability standards for both platform and agentic system health; design self-service patterns and golden-path templates that accelerate delivery.
- Cross-team partnership, security, and talent development
- Partner with the Reliability Engineering team on which production patterns become agent-assisted automations, and with senior architects on the patterns that scale.
- Own security posture: network policies, pod security, secrets rotation, vulnerability scanning, and access controls in a regulated pharmaceutical environment.
- Set the engineering bar through code quality standards, architectural reviews,
and role modeling; mentor senior engineers in agentic AI, LLM application patterns, and platform engineering. Influence engineering leaders to adopt automation-friendly patterns at the source, not just downstream.
How you will succeed
- Be recognized as the senior technical authority for agentic platform engineering in your area.
- Demonstrate measurable, sustained improvements: rising deflection rates, reduced toil, fewer recurring incidents, faster resolution, platform uptime, and deployment velocity.
- Ship production platform services and agentic capabilities that tangibly move the deflection and developer-experience numbers.
- Scale agentic intelligence through systems, standards, and people - not individual heroics.
What you should bring
Required
- 12+ years of progressive technology experience with substantial hands-on architecture and delivery of automation, AIOps, agentic systems, or cloud-native platforms at enterprise scale.
- Demonstrated ownership of measurable deflection or operational-toil-reduction outcomes in a production setting - with the numbers to show it.
- Deep technical fluency in agentic system design: agent runtimes, multi-agent orchestration, tool integration, memory and state management. Practical experience with at least one major framework (LangGraph, LangChain, LlamaIndex, or MCP).
- Production experience with LLM application patterns: prompt engineering, retrieval-augmented generation (RAG), structured outputs, and multi-model routing.
- Deep cloud-native background: Kubernetes (EKS or equivalent) at scale, container workloads, infrastructure as code (Terraform), and at least one major cloud's AI stack (AWS Bedrock/SageMaker, Azure AI Foundry, or Vertex AI).
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Senior Principal Agentic Engineer (Hyderabad)
🏢 Eli Lilly
📍 Hyderabad