Senior Software Engineer- AI/ML Infrastructure (India)

Senior Software Engineer- AI/ML Infrastructure (India)

17 Sep
|
Pilotcrew AI
|
India

17 Sep

Pilotcrew AI

India

Senior Software Engineer - AI/ML Infrastructure

Location: India

Employment Type:Full-time, Remote

Experience: 5-7 years

About PilotCrew AI

PilotCrew AI is building the operational layer for production AI agents.

We help teams build and connect AI agents, optimize their performance through closed-loop evaluations, monitor reliability and execution, and validate them through standardized benchmarks and AI Leagues.

Our platform brings together agent evaluation, optimization, observability, reliability, and validation, including automated and adversarial evaluations, failure analysis, trace-level monitoring, cost and reliability tracking, benchmarking, and human/expert assessment.

As AI agents move from demos into real-world production environments, our mission is to provide the infrastructure needed to measure, understand, improve, and validate their performance and reliability at scale.

About the Role

We're looking for a Senior Software Engineer - AI/ML Infrastructure to help build and scale the systems powering AI evaluation at PilotCrew AI

This role is for an engineer who enjoys solving complex engineering problems at scale. You should be comfortable taking ownership of an existing codebase, understanding it deeply, identifying architectural bottlenecks, and evolving it into reliable and scalable production infrastructure.

You'll work at the intersection of software engineering, distributed systems, cloud infrastructure, and AI/ML, building systems that enable us to evaluate increasingly capable AI agents reliably and at scale.

We're looking for someone who operates beyond implementing individual features. You should be comfortable designing systems, making architectural decisions, solving scalability challenges, and taking end-to-end technical ownership.

What You'll Do

- Take ownership of complex parts of our existing platform and improve their performance, reliability, scalability, and maintainability.
- Design and build distributed systems capable of handling growing evaluation workloads and data volumes.
- Architect backend services, APIs, data pipelines, and infrastructure for high-throughput AI evaluation workloads.
- Design systems that efficiently manage large numbers of concurrent and long-running evaluation jobs.
- Build and improve infrastructure for AI and ML evaluation, including automated evaluation pipelines, benchmarks, datasets, metrics, and experimentation workflows.
- Build systems supporting LLM and AI agent evaluation across different models, tools, environments, and workflows.
- Design asynchronous and event-driven systems for distributed AI workloads.
- Identify system bottlenecks and design solutions around latency, throughput, reliability, availability, and infrastructure cost.
- Build robust systems for retries, failure recovery, idempotency, rate limiting, timeouts, and graceful degradation.
- Improve data infrastructure for storing and querying agent traces, evaluation results, metrics, experiments, and model outputs.
- Build for fault tolerance, observability, monitoring,



and operational reliability.
- Debug complex production issues across services, queues, databases, cloud infrastructure, and external model APIs.
- Make thoughtful architectural trade-offs and contribute to the technical direction of the platform.
- Collaborate closely with ML, product, and engineering teams to turn ambiguous problems into robust technical solutions.
- Raise engineering standards through code reviews, technical design discussions, documentation, testing, and mentorship.

What You'll Own

- You'll have meaningful technical ownership across core parts of the PilotCrew AI platform, including:
- Distributed evaluation execution
- Backend and service architecture
- Workload and job orchestration
- AI agent and LLM evaluation infrastructure
- Trace and evaluation data infrastructure
- Reliability and fault tolerance
- Observability and production debugging
- Performance and scalability
- Cloud infrastructure and deployment systems

What We're Looking For

- 5+ years of skilled software engineering experience, with significant experience building backend, infrastructure, or distributed systems.
- Strong programming skills in Python, TypeScript, and/or JavaScript
- Strong fundamentals in data structures, algorithms, databases, APIs, concurrency, and backend engineering.
- Hands-on experience with system design and software architecture.
- Experience designing and scaling distributed systems.
- Experience taking an existing codebase or production system and significantly improving its architecture, performance, reliability, or scalability.
- Strong understanding of caching, queues, asynchronous processing, distributed execution, database scaling, fault tolerance, and high availability.
- Experience with cloud infrastructure, preferably AWS.
- Experience building and operating containerized production systems.
- Experience debugging distributed systems using logs, metrics, traces, and observability tooling.
- Strong understanding of API design, service boundaries, and production backend architecture.
- Experience working with AI/ML systems or infrastructure.
- Familiarity with LLM APIs, AI agents, tool calling, model inference, or AI application infrastructure.
- Ability to independently break down ambiguous technical problems and take ownership from design through implementation and production.
- Strong engineering judgment and the ability to make pragmatic architectural trade-offs.

Good to Have

- Experience building LLM evaluation or AI agent evaluation infrastructure
- Experience with automated evaluation pipelines, benchmarks, evaluation datasets, or model and agent metrics.
- Experience working with LLM APIs, Generative AI, AI agents,



tool calling, structured outputs, or multi-model systems.
- Experience with MCP or other agent/tool integration protocols.
- Experience building systems that capture and process AI agent traces, model outputs, evaluation results, and experiment data at scale.
- Experience designing distributed execution systems for long-running AI/ML workloads.
- Experience with model routing, inference infrastructure, or multi-provider AI systems.
- Experience with large-scale data processing or distributed compute.
- Experience with MongoDB or similar databases, including schema design, indexing, query optimization, and performance debugging.
- Experience with Redis, message queues, streaming systems, or event-driven architectures.
- Experience with AWS services such as ECS, EC2, S3, SQS, Batch, ECR, CloudFront, and ALB.
- Experience with Docker, Kubernetes, or similar container infrastructure.
- Experience with observability systems such as OpenTelemetry and distributed tracing.
- Experience working in an early-stage or fast-growing startup where engineers have significant ownership.

Our Tech Stack

- You do not need to have worked with every technology in our stack, but you should be comfortable working across backend systems, cloud infrastructure, distributed workloads, and AI/ML tooling.
- Frontend: React, TypeScript, Vite, Tailwind CSS
- Backend: Node.js, Express, Python, FastAPI
- Databases: MongoDB, Mongoose, PyMongo
- Caching & Job Queues: Redis, BullMQ
- Workload Orchestration: AWS SQS, AWS Batch
- Cloud Infrastructure:AWS EC2, ECS, S3, CloudFront, ALB, ECR
- Containerization: Docker
- CI/CD: AWS CodePipeline, CodeBuild, CodeDeploy
- Authentication: JWT, OAuth 2.0, Passport.js
- Secrets Management: AWS Systems Manager Parameter Store
- Observability: OpenTelemetry, OTLP, distributed tracing, structured logging
- AI & LLM Integrations:OpenAI, Anthropic, Google Gemini, AWS Bedrock, OpenRouter
- Agent Infrastructure: MCP, tool calling, agent workflows, evaluation pipelines
- ML & Data Tooling:Hugging Face Datasets, Transformers.js, NumPy, pandas, PyArrow
- Testing: Jest, Supertest, pytest
- Languages:TypeScript, JavaScript, Python

What We're Looking For in an Engineer

- Beyond technical experience, we're looking for someone who:
- Thinks in systems, not just features.
- Can enter an unfamiliar codebase and quickly understand how it works.
- Thinks about what happens when a system needs to operate at 10x or 100x scale.
- Can identify architectural bottlenecks before they become production problems.
- Is comfortable making architectural decisions and owning the consequences.
- Enjoys solving difficult and open-ended engineering problems.
- Knows when to optimize for simplicity, performance, reliability, or scale.
- Is comfortable working in a fast-moving environment with ambiguity.
- Takes end-to-end ownership and cares deeply about the quality of what goes into production.
- Is excited about the infrastructure challenges created by increasingly capable AI systems.

📌 Senior Software Engineer- AI/ML Infrastructure (India)
🏢 Pilotcrew AI
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior software engineer- ai/ml infrastructure (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: senior software engineer- ai/ml infrastructure (india) / india