As a Senior Machine Learning Engineer within the Data & Analytics team, you will be responsible for managing and optimizing the training infrastructure for Large Language Models (LLMs). This role demands a deep understanding of GPU architecture, machine learning principles, and distributed computing. You will lead Gen AI initiatives in a cross-functional setup, ensuring efficient resource utilization and timely delivery of large-scale ML projects.
Key Responsibilities
Primary Responsibilities
- Lead Generative AI projects in a cross-functional team environment.
- Apply advanced machine learning principles and algorithms, particularly for LLMs such as GPT-4, BERT, and Transformers.
- Utilize deep learning frameworks like TensorFlow, PyTorch, and Keras for model training.
- Maximize GPU utilization and efficiency through deep knowledge of computer architecture.
- Manage and optimize cloud-based resources (AWS, Azure, GCP) for deep learning model training.
- Implement containerization and orchestration using Docker and Kubernetes.
- Apply parallel and distributed computing principles for scalable model training.
- Integrate big data technologies like Hadoop and Spark into ML workflows.
- Adopt MLOps practices and tools to manage the end-to-end ML lifecycle.
Secondary Responsibilities
- Manage infrastructure for multiple ML projects,
especially those involving deep learning models.
- Optimize performance and resource allocation for large-scale ML tasks.
- Handle GPU resource management both on-premises and in the cloud.
- Address challenges in training large models, including memory management, data loading optimization, and hardware troubleshooting.
- Collaborate closely with data scientists and ML engineers to understand infrastructure needs and deliver efficient solutions.
What We Are Looking For
Education
- Graduation in BSC or BCA or B.Tech.
Experience
- 4+ years of relevant experience in managing infrastructure for training large-scale ML models.
- Hands-on experience with LLMs and deep learning frameworks.
- Experience in cloud computing, containerization, and distributed systems.
- Prior involvement in Gen AI projects and cross-functional team collaboration.
Skills and Attributes
- Robust understanding of GPU architecture and optimization techniques.
- Proficiency in TensorFlow, PyTorch, Keras, Docker, Kubernetes, and cloud platforms.
- Knowledge of distributed computing frameworks like Hadoop and Spark.
- Familiarity with MLOps tools and practices.
- Excellent problem-solving and troubleshooting skills.
- Ability to lead technical aspects of projects and ensure error-free, timely deliverables.
📌 Senior ML Engineer - AI Labs (Bengaluru)
🏢 IDFC FIRST Bank
📍 Bengaluru
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.