09 Aug
|
Albertsons Companies India
|
Bengaluru
09 Aug
Albertsons Companies India
Bengaluru
We are searching for someone with the following skills:
- 10+ years of industry experience in software engineering, machine learning engineering, AI systems development, or production data systems.
- Strong experience designing, building, deploying, and operating AI/ML systems for real-world production use cases at enterprise scale.
- Deep hands-on experience with Python and common ML frameworks and libraries such as scikit-learn, XGBoost, PyTorch, TensorFlow, Pandas, NumPy, and related production ML tooling.
- Proven experience building machine learning solutions for forecasting, anomaly detection, prediction, classification, clustering, ranking, recommendation, prioritization, or decision-support use cases.
- Strong understanding of time-series modeling techniques for forecasting, anomaly detection, capacity planning, operational prediction, and service health intelligence.
- Experience building or leading ML solutions for alert classification, incident prediction, event deduplication, signal correlation, prioritization, noise reduction, or operational intelligence.
- Strong knowledge of causal ML, causal inference, graph-based reasoning, dependency-aware analysis, and statistical modeling techniques for RCA and impact analysis.
- Hands-on experience building LLM-powered applications using frameworks such as LangChain, LangGraph, or similar agentic AI frameworks.
- Experience designing AI agents or multi-agent systems for reasoning, summarization, task orchestration, troubleshooting, knowledge retrieval, decision support, or assistant workflows.
- Strong understanding of prompt engineering, RAG architectures, embeddings, vector databases, tool calling, memory handling, context management, guardrails, and agent evaluation techniques.
- Experience integrating LLM systems with enterprise tools, APIs, knowledge repositories, observability platforms, ticketing systems, automation tools, and operational workflows.
- Strong experience building backend services, APIs, microservices, model-serving systems, inference services,
and agent orchestration platforms.
- Positive understanding of observability data, including logs, metrics, traces, topology, alerts, incidents, service maps, deployment events, and operational metadata.
- Solid data engineering knowledge, including feature engineering, data preprocessing, model pipelines, batch inference, streaming inference, data quality validation, and production data workflows.
- Familiarity with graph databases such as Neo4j and their use in dependency mapping, knowledge graphs, causal analysis, impact analysis, and knowledge-driven AI systems.
- Strong experience with REST APIs, microservices architecture, Docker, Kubernetes, cloud-native deployment patterns, distributed systems, and scalable architecture design.
- Experience with CI/CD, MLOps, model lifecycle management, experiment tracking, model versioning, feature stores, model monitoring, drift detection, and automated deployment practices.
- Knowledge of OpenTelemetry, monitoring systems, observability platforms, incident management systems, and SRE operating models is highly desirable.
- Strong understanding of software engineering fundamentals, system design, distributed systems, reliability engineering, testing strategies, and scalable architecture patterns.
- Proven ability to decompose ambiguous operational problems into measurable AI/ML solutions with clear success metrics and production impact. Excellent communication, collaboration, and technical leadership skills,
- with the ability to influence SREs, platform engineers, architects, product owners, and business stakeholders. Self-driven mindset with strong curiosity, innovation,
ownership, and the ability to evaluate and apply
- emerging AI techniques effectively in production environments.
We believe the successful candidate has these qualifications and experience:
- Bachelors degree in Computer Science, Engineering, Data Science, Artificial Intelligence, Information Systems, or a related field, or equivalent practical experience.
- 10+ years of overall industry experience in software engineering, machine learning engineering, AI system development, data platforms, or production-grade distributed systems.
- 6+ years of hands-on experience building, deploying, and operating machine learning systems in production environments.
- Experience designing or building AI/ML systems for observability, monitoring, AIOps, SRE, IT operations, incident management, or related operational domains would be a big plus.
- Strong hands-on experience in Python-based AI/ML development is required.
- Proven experience owning production ML systems end to end, including data pipelines, training workflows, evaluation, deployment, inference, monitoring, retraining, and operational support.
- Experience building LLM-powered applications, AI agents, or multi-agent workflows for enterprise use cases is highly preferred.
- Experience applying AI/ML to RCA, anomaly explanation, incident summarization, service health prediction, capacity risk prediction, alert intelligence, or remediation recommendations.
- Familiarity with knowledge graphs, graph databases, and graph-based ML techniques for dependency-aware intelligence and operational reasoning.
- Experience using vector databases and retrieval frameworks for enterprise search, RAG systems, and agentic AI applications. Experience integrating AI services with platforms such as ServiceNow, Grafana, Prometheus, Splunk, AppDynamics, Datadog, Dynatrace, PagerDuty, Jira, or similar enterprise tools.
- Familiarity with MCP-based client or agent integrations is a plus.
📌 Staff Engineer ML (Bengaluru)
🏢 Albertsons Companies India
📍 Bengaluru