Role Description :
We are seeking an experienced (8 to 10 years) AI Site Reliability Engineering to lead the reliability| scalability| and operational excellence of our AI/ML platforms and production systems.
This role bridges SRE| MLOps| and AI platform engineering| ensuring highly available| observable| and resilient AI-driven services at enterprise scale.
Lead enterprise MLOps frameworks covering experimentation| training| validation| deployment| monitoring| and retirement.
Ensure robust model versioning| reproducibility| governance| and auditability.
Partner with data science teams to productionize models with reliability| performance| and compliance in mind.
(Do share only relevant profiles that closely match the requirements. Kindly avoid submitting unsuitable or unrelated profiles.)
Technical Skills- Experience in SRE| Production Support of Application based on Java and SQL
Hands on experience of Linux/Windows operating system
Able to debug| troubleshoot| fix Java application and support during application failure
Defect analysis and Resolution
Developing scripts and automation must be able to do.
Prior experience on Investment Banking Domain
Oversee strategies for model monitoring| including drift| bias| accuracy decay| and data integrity.
Key Technologies PostGres| PL SQL| GCP| Python (Positive to have)Tools: Autosys| Cloud Observability Location - Pune/Chennai/Hyderabad/Bangalore Experience Required: 8+ years **If any one is interseted in this Postion Please share your Updated resume to this requried mail i'd :
Mail I'D :**
[email protected]
📌 SRE Site Reliability Engineering (Hyderabad)
🏢 VMC Soft Technologies
📍 Hyderabad