05 Sep
|
Tata Consultancy Services
|
Chennai
05 Sep
Tata Consultancy Services
Chennai
Job Description:
- Proficiency in Python/SQL, understanding of ML frameworks (PyTorch, TensorFlow), and data pre-processing logic.
- Core AI/ML Fundamentals knowledge
- Machine Learning Operations (MLOps)
- Should have good understanding about AI/ML Ecosystem Tools
- Strong understanding of GCP compute, storage, IAM, Vertex AI
- Exposure in managing GPU/TPU environments
- Debugging & Support Skills
- Ability to analyze logs, trace errors, and troubleshoot
- Knowledge of working with different APIs
- Identifying, resolving technical issues, and diagnosing the root cause of technical problems.
- Providing technical assistance to users, both internal and external, through various channels
- Experience working with Conversation Agents
- Communicating technical information clearly and concisely to users and stakeholders
- Support with 24x7 operations (Rotational Shifts)
- English language (verbal and written) Proficiency is must
1. Troubleshooting Model & Pipeline Issues
- The core of the role is diagnosing technical failures within the ML lifecycle.
- API & Integration Support: Debugging RESTful API integrations between the client's application and the AI model.
- Inference Failures: Investigating why a model is failing to provide predictions (e.g., timeout issues, memory overflows, or incorrect input data formatting).
- Workplace Configuration: Assisting clients with setup issues related to Docker, Kubernetes, or cloud-specific ML environments (AWS SageMaker, Azure ML, etc.).
2. Data & Performance Monitoring
- AI products are only as good as the data fed into them.
- Data Quality Checks: Helping customers identify if their input data is the cause of poor model performance (e.g.,
missing values, incorrect data types, or schema mismatches).
- Monitoring Drift: Assisting in identifying "Model Drift"where the AI’s performance degrades over time because real-world data has changed compared to the training data.
- Accuracy Inquiries: Explaining to customers why a model produced a specific result using interpretability tools (like SHAP or LIME) or logs.
3. Product Education & Technical Documentation
- Because AI is complex, the TSR serves as a technical teacher.
- Knowledge Base Authoring: Writing guides on "Best Practices for Prompt Engineering" or "How to Fine-tune Hyperparameters" for the specific platform.
- Customer Onboarding: Guiding new technical users through the initial setup of their ML experiments.
- Translating Documentation: Taking complex engineering release notes and making them understandable for the customer's IT team.
4. The "Feedback Bridge" to Engineering
- The TSR is the first to see patterns in how the product fails in the real world.
- Bug Reporting: Identifying and reproducing software bugs in the ML platform and escalating them to the ML Engineers or DevOps team.
- Feature Requests: Aggregating customer feedback regarding missing ML capabilities (e.g., "Customers are asking for support for PyTorch 2.0").
- Edge Case Discovery: Documenting unique edge cases where the AI model consistently fails, which helps the data science team improve future training sets.
- Responsibility for maintaining SLA (Service Level Agreements) and identifying product bugs for the engineering team.
- Translating complex AI concepts for non-technical users and managing high-pressure customer interactions.
📌 AI ML Engineer(Technical Support Representative) (Chennai)
🏢 Tata Consultancy Services
📍 Chennai