18 Aug
|
IntraEdge
|
Hyderabad
18 Aug
IntraEdge
Hyderabad
Senior AI/ML Engineer
Location: Hyderabad
Experience: 6+ Years
Employment Type: Full-Time
Job Summary
We are looking for a highly skilled Senior AI/ML Engineer to design, build, deploy, and scale enterprise-grade Machine Learning and Artificial Intelligence solutions that deliver measurable business value.
The ideal candidate will have strong hands-on expertise in Machine Learning, MLOps, cloud platforms, ML model deployment, scalable ML pipelines, feature engineering, model monitoring, and software engineering practices . The candidate will work closely with Data Scientists, Data Engineers, Platform Engineers, SRE, DevOps, and Product teams to take AI/ML solutions from experimentation to reliable production systems.
This role requires a strong engineering mindset with the ability to build scalable, secure, observable, and maintainable ML solutions across the complete machine learning lifecycle.
Key Responsibilities
Machine Learning & AI Solution Development
- Design, develop, deploy, and scale machine learning and AI solutions for enterprise use cases.
- Translate business and product requirements into scalable AI/ML solutions.
- Work closely with Data Scientists to productionize machine learning models.
- Select appropriate ML algorithms, frameworks, and architectures based on business requirements.
- Develop reusable components and frameworks for AI/ML applications.
- Optimize models and ML applications for accuracy, performance, scalability, and cost.
ML Model Deployment & Productionization
- Deploy machine learning models into production environments across cloud and enterprise platforms.
- Build scalable real-time and batch inference solutions.
- Develop model-serving APIs and microservices.
- Implement model versioning, release management, rollback, and deployment strategies.
- Support production ML models and troubleshoot performance or reliability issues.
- Ensure models meet production SLAs for availability, latency, and throughput.
MLOps
- Design and implement end-to-end MLOps pipelines for model development, validation, deployment, and monitoring.
- Implement automated CI/CD pipelines for ML models and applications.
- Establish repeatable and reproducible ML workflows.
- Automate model training, testing, validation, deployment, and retraining processes.
- Implement model registry and experiment tracking solutions.
- Establish best practices for model lifecycle management and governance.
ML Pipeline Development
- Design scalable and reliable data and ML pipelines.
- Build automated workflows for data preparation, feature engineering, model training, and inference.
- Integrate data pipelines with ML workflows.
- Implement data validation and quality checks.
- Optimize pipelines for large-scale datasets and distributed processing.
- Support both batch and real-time ML processing requirements.
Feature Engineering & Feature Management
- Design and implement scalable feature engineering frameworks.
- Develop reusable feature transformation pipelines.
- Ensure consistency of features between training and production environments.
- Implement feature versioning and lifecycle management.
- Work with Data Engineering teams to build reliable feature pipelines.
- Exposure to Feature Store technologies is preferred.
Model Monitoring & Observability
- Design and implement comprehensive monitoring solutions for production ML models.
- Monitor:
- Model accuracy
- Data quality
- Data drift
- Concept drift
- Model performance
- Prediction latency
- Infrastructure health
- Develop dashboards, alerts, and operational metrics.
- Implement observability across ML pipelines and model-serving infrastructure.
- Identify model degradation and implement appropriate remediation strategies.
Cloud & Infrastructure
- Design and deploy AI/ML solutions on cloud platforms such as AWS, Azure, or GCP .
- Work with cloud-native services for ML, compute, storage, networking, and monitoring.
- Build scalable and highly available ML infrastructure.
- Implement secure cloud architectures for AI/ML workloads.
- Optimize cloud infrastructure and ML workloads for performance and cost.
- Exposure to Infrastructure as Code tools such as Terraform or CloudFormation is preferred.
Software Engineering
- Develop production-quality code primarily using Python .
- Build scalable APIs and microservices for AI/ML applications.
- Follow clean coding, design, testing, and documentation practices.
- Perform code reviews and contribute to engineering standards.
- Develop reusable libraries and components.
- Troubleshoot application, infrastructure, and ML pipeline issues.
CI/CD & DevOps
- Design and maintain CI/CD pipelines for ML applications and models.
- Automate build, test, validation, and deployment processes.
- Integrate ML workflows with enterprise DevOps practices.
- Implement automated quality and security checks.
- Work with Git-based source control and modern DevOps tools.
Model Governance & Security
- Implement appropriate model governance and lifecycle management processes.
- Maintain documentation for models, datasets, experiments, and deployments.
- Ensure ML solutions comply with enterprise security and privacy standards.
- Implement access controls and secure handling of sensitive data.
- Support auditability, traceability, and reproducibility of ML solutions.
- Contribute to responsible AI and model risk management practices.
Cross-Functional Collaboration
- Collaborate with:
- Data Scientists
- Data Engineers
- Platform Engineers
- SRE Teams
- DevOps Engineers
- Cloud Architects
- Product Managers
- Business Stakeholders
- Provide technical guidance on productionizing ML models.
- Participate in architecture and design discussions.
- Mentor junior AI/ML engineers and promote engineering best practices.
- Communicate complex technical concepts clearly to technical and business stakeholders.
Required Technical Skills
Machine Learning
- Robust understanding of machine learning concepts and algorithms.
- Experience with supervised and unsupervised learning.
- Knowledge of:
- Regression
- Classification
- Clustering
- Recommendation Systems
- Feature Engineering
- Model Evaluation
- Experience with ML frameworks such as:
- Scikit-learn
- XGBoost
- TensorFlow
- PyTorch
Programming
- Robust proficiency in Python .
- Good understanding of object-oriented programming and software engineering principles.
- Strong SQL knowledge.
- Experience developing REST APIs and microservices.
MLOps
- Hands-on experience with end-to-end ML lifecycle management.
- Experience with one or more:
- MLflow
- Kubeflow
- AWS SageMaker
- Azure ML
- Google Vertex AI
- Experience with experiment tracking and model registry.
- Model deployment and monitoring experience.
Cloud Strong experience with at least one major cloud platform:
- AWS
- Microsoft Azure
- Google Cloud Platform
AWS ML experience such as Amazon SageMaker, S3, Lambda, ECS/EKS, IAM, and CloudWatch is highly desirable.
Data Engineering
- Data ingestion and transformation.
- ETL/ELT pipelines.
- Feature engineering pipelines.
- Batch and real-time processing.
- Experience with Apache Spark/PySpark is preferred.
- Familiarity with data lakes and data warehouses.
DevOps & Containerization
- Docker
- Kubernetes
- Git
- CI/CD
- Jenkins / GitHub Actions / GitLab CI
- Terraform or CloudFormation
Monitoring & Observability
- Prometheus
- Grafana
- Datadog
- CloudWatch
- Logging and alerting frameworks
- Model performance monitoring
Preferred / Good-to-Have Skills
- Experience with Generative AI and LLM-based applications .
- Knowledge of RAG architectures and vector databases.
- Exposure to LangChain or LangGraph.
- Experience with AI agents and agentic workflows.
- Experience with Kafka or other event-streaming platforms.
- Knowledge of Feature Stores.
- Experience with model explainability and Responsible AI.
- Exposure to AI/ML solutions within Banking, Financial Services, Healthcare, or other regulated industries.
Qualifications
- Bachelor's or Master's degree in Computer Science, Engineering, Data Science, Artificial Intelligence, or a related technical field.
- 6+ years of relevant professional experience in AI/ML Engineering, Machine Learning Engineering, Data Science Engineering, or related software engineering roles.
- Demonstrated experience taking ML solutions from development/experimentation through production deployment.
- Strong understanding of software development lifecycle and Agile methodologies.
Soft Skills
- Strong analytical and problem-solving skills.
- Excellent communication and collaboration abilities.
- Strong ownership and accountability.
- Ability to work independently and across multiple technical teams.
- Ability to manage priorities in a fast-paced environment.
- Robust mentoring and technical leadership capabilities.
- Passion for learning and adopting emerging AI/ML technologies.
📌 Senior AI/ML Engineer (Hyderabad)
🏢 IntraEdge
📍 Hyderabad