Dear Candidate,
Please find the below job description and let us know if this opportunity interests you.
Role Description:
The Junior AI Observability Architect is an execution-focused engineer who designs, builds, and operates observability capabilities within a defined domain of the enterprise AI observability platform. Working under the strategic direction of the Senior AI Observability Architect , this role translates architecture blueprints into production-grade instrumentation, telemetry pipelines, dashboards, quality gates, and safety signals across agentic AI systems.
The junior architect is a hands-on engineer who codes, integrates, tests, and iterates owning feature-level delivery within one or more specialization tracks while developing a growing understanding of the full observability platform. They are a technical practitioner first, with an emerging architect mindset.
Description:
- Bachelor's or Master's degree in Computer Science, Software Engineering, AI/ML, Data Science, or a related technical field.
- 11+ years of experience in software engineering, platform engineering, or data engineering with at least 2 years of hands-on work in observability, monitoring, or distributed systems.
- Demonstrated ability to deliver production-grade software in a team environment; track record of completing complex technical features end-to-end.
- Python Proficiency: Strong Python engineering skills — writing clean, testable, maintainable production code; familiarity with async patterns, type hints, and contemporary Python tooling (Poetry, Ruff, pytest).
- Observability Fundamentals:
Solid working knowledge of the three pillars of observability (metrics, logs, traces); ability to instrument services with OpenTelemetry (OTEL) SDKs; understanding of trace context propagation and semantic conventions.
- Distributed Systems: Working knowledge of microservices, event streaming (Kafka or equivalent), REST/gRPC APIs, and containerized deployment (Docker, Kubernetes).
- Cloud Platforms: Hands-on experience with at least one major cloud provider (Azure, AWS, or GCP) — including managed services, IAM basics, and cost awareness.
- CI/CD & DevOps: Experience building or contributing to CI/CD pipelines; familiarity with GitOps, infrastructure-as-code concepts, and automated testing frameworks.
- Data Fundamentals: Ability to query, analyze, and visualize time-series and log data using tools such as Grafana, Datadog, Splunk, Prometheus, or equivalent.
- Hands-on experience with agentic AI frameworks (LangChain, LangGraph, AutoGen, Semantic Kernel, CrewAI, or equivalent).
- Contributions to open-source observability projects or OTEL community.
- Familiarity with reinforcement learning concepts, self-supervised learning, or model fine-tuning workflows.
- Experience with security tooling relevant to AI (adversarial robustness libraries, LLM safety frameworks, or red-team toolkits).
- Exposure to Responsible AI frameworks, fairness evaluation libraries (Arize, Fairlearn, AI Fairness 360), or explainability tools (SHAP, LIME).
- Experience in a fast-paced AI platform, MLOps, or LLMOps role with production deployment responsibilities.
Interested candidates can share their updated resumes at
[email protected]
📌 Deputy Director OR AI Architect (Telangana)
🏢 PepsiCo
📍 Telangana