01 Oct
|
Dautom
|
Bengaluru
About the Role
We are looking for a hands-on AI Solutions Engineer who can take generative and agentic AI solutions from design to production. You will own the full stack: LLM orchestration, retrieval, backend services, user interface, deployment, and evaluation. This is an engineering role first. You will write production code every day and be accountable for the reliability, security, and cost of what you ship.
Key Responsibilities
- Architect and build multi-agent systems with stateful orchestration, tool calling, routing, memory, and human-in-the-loop checkpoints
- Design retrieval pipelines covering document parsing, chunking, embeddings, hybrid search, and reranking, grounded in governed enterprise data
- Build high-performance Python backend services and streaming APIs that expose agents to web, chat, and voice channels
- Develop responsive React and TypeScript front ends with real-time streaming, rich visualizations, and clean UX for business users
- Deploy and operate AI workloads on the lakehouse platform and cloud, using containers, CI/CD, and infrastructure as code
- Implement LLMOps practices: offline and online evaluation, LLM-as-judge scoring, tracing, prompt versioning, and regression testing
- Engineer guardrails for prompt-injection defense, PII protection, output validation, and role-based data access
- Optimize latency, token cost, and throughput through model selection, caching, prompt design, and parallel execution
- Translate business requirements into technical designs, and document architecture decisions for review
- Mentor engineers and contribute reusable components, templates, and standards to the wider team
Technical Skills
Required
- Languages: Expert Python (async, typing, Pydantic, testing); working proficiency in TypeScript and SQL
- LLMs and GenAI: Hands-on with frontier models (Claude, GPT, or equivalent), structured outputs, function and tool calling, context engineering, and prompt design
- Agent frameworks: Production experience with LangGraph (strongly preferred) or comparable frameworks such as LangChain, LlamaIndex, or the OpenAI Agents SDK
- Retrieval: Vector databases and search services (Databricks Vector Search, Azure AI Search, pgvector, or similar), embedding models, hybrid retrieval, and reranking
- Backend: FastAPI, REST and streaming protocols (SSE, WebSockets), async I/O, and OAuth2 or Entra ID authentication
- Frontend: React, TypeScript, state management, and component libraries; able to build streaming chat and dashboard interfaces
- Data platform: Databricks (Model Serving, Unity Catalog, MLflow, Databricks Apps) or an equivalent enterprise data and AI platform
- Cloud and DevOps: Microsoft Azure, Docker, Git, and CI/CD using GitHub Actions or Azure DevOps
- LLMOps: Evaluation frameworks,
tracing and observability (MLflow Tracing, LangSmith, or similar), and cost and latency monitoring
Preferred
- Model Context Protocol (MCP) and agent-to-agent interoperability patterns
- Databricks Genie, Mosaic AI Agent Framework, or Agent Bricks
- Realtime voice and speech APIs (Azure AI Speech, GPT-Realtime, or similar)
- Multilingual or Arabic NLP
- Integration with enterprise systems such as CRM, ERP, Microsoft Teams, and Power Automate
- Kubernetes, Terraform, or Databricks Asset Bundles
Qualifications
- 4 to 10 years of software or ML engineering experience, including at least 2 years building LLM or generative AI applications
- A proven track record of shipping at least one GenAI or agentic application to production with real users
- Bachelor's or Master's degree in Computer Science, Engineering, or a related field, or equivalent practical experience
- Strong system design skills, with the ability to reason about reliability, security, and cost trade-offs
- Clear written and verbal communication with both technical and business stakeholders
- Relevant certifications (Databricks Generative AI Engineer, Azure AI Engineer) are a plus
Data Scientist / Machine Learning Engineer
About the Role
We are looking for a Data Scientist and ML Engineer who can build machine learning models and run them reliably in production. You will own the full model lifecycle: problem framing, feature engineering, model development, validation, deployment, and monitoring. You will also build the data pipelines that feed your models, so strong data engineering fundamentals are essential.
Key Responsibilities
- Frame business problems as ML problems, define success metrics, and agree on evaluation criteria with stakeholders
- Develop, tune, and validate models for forecasting, classification, regression, segmentation, anomaly detection, and recommendation
- Engineer robust features from large structured and time-series datasets, and publish them as reusable, governed feature tables
- Deploy models to production as real-time endpoints and scheduled batch inference jobs
- Build end-to-end MLOps workflows: experiment tracking, model registry, versioning, CI/CD, and champion-challenger promotion
- Monitor production models for data drift, prediction drift, and performance decay, and own retraining strategies
- Build and maintain scalable data pipelines using medallion architecture principles,
with data quality checks at every layer
- Explain model behavior using interpretability techniques, and communicate results clearly to non-technical audiences
- Apply statistical rigor through hypothesis testing, experiment design, and uncertainty quantification
- Contribute to team standards for ML code quality, reproducibility, documentation, and model governance
Technical Skills
Machine Learning Development (Core)
- Algorithms: Robust command of supervised and unsupervised learning, gradient boosting, time-series forecasting, clustering, and ensemble methods
- Frameworks: scikit-learn, XGBoost or LightGBM, statsmodels, and PyTorch or TensorFlow
- Model quality: Cross-validation strategies (including time-based splits), hyperparameter optimization (Optuna or similar), and explainability (SHAP)
- Statistics: Solid foundation in probability, inference, experimental design, and causal reasoning
Model Deployment and MLOps (Core)
- Lifecycle: MLflow for experiment tracking, model registry, and model packaging
- Serving: Real-time model serving endpoints and distributed batch inference at scale
- Operations: Model and data monitoring, drift detection, automated retraining, and alerting
- Engineering: Production-quality Python, unit and integration testing, Git, Docker, and CI/CD pipelines
Data Engineering
- Processing: PySpark and advanced SQL on large datasets
- Lakehouse: Delta Lake, Lakeflow Declarative Pipelines (DLT) or equivalent, and orchestration through Databricks Workflows or Azure Data Factory
- 1`1Governance: Unity Catalog or equivalent for access control, lineage, and data discovery
Preferred
- Databricks ML stack: Feature Engineering in Unity Catalog, Model Serving, and Lakehouse Monitoring
- Mathematical optimization with OR-Tools, PuLP, or Gurobi
- Marketing analytics: attribution, media mix modeling, customer lifetime value, or churn modeling
- Familiarity with GenAI and LLM integration in ML workflows
- Experience with ERP or CRM source data (for example, SAP)
Qualifications
- 4 to 10 years of experience in data science or ML engineering, with at least 2 years deploying and maintaining models in production
- A demonstrable track record of models that delivered measurable business impact, not only offline accuracy gains
- Bachelor's or Master's degree in Computer Science, Statistics, Mathematics, Engineering, or a related quantitative field
- Ability to own problems end to end, from raw data to a monitored production model
- Clear communication of technical findings to business stakeholders
- Relevant certifications (Databricks Machine Learning Professional, Azure Data Scientist Associate) are a plus
Location-Remote/India
Salary up to 4200 USD
📌 AI Solutions Engineer (Bengaluru)
🏢 Dautom
📍 Bengaluru