Software Engineer - MLOps (Hyderabad)

Software Engineer - MLOps (Hyderabad)

18 Sep
|
Costco
|
Hyderabad

18 Sep

Costco

Hyderabad

Skills Required:

Python, GCP (Vertex AI, BigQuery, Dataproc, Dataflow, Cloud Composer, GKE, Pub/Sub), Vertex AI Pipelines / Kubeflow, MLflow, Feature Store, Docker, Kubernetes, Terraform, GitHub, TensorFlow / PyTorch Experience Range: 8+ years with 3 yrs of ML Ops experience Position Summary: The Data Science team is looking for a passionate MLOps Engineer eager to solve complex machine learning problems at a global scale. At Costco, we're building the next generation of retail technology, and we need talented individuals to drive our digital transformation. We're a company that not only delivers exceptional value to millions of members worldwide but also deeply values its employees, fosters a culture of innovation, and offers genuine opportunities for technical and career growth. The ideal candidate will perform complete cycle development work on ML pipelines, model serving on Vertex AI, associated data stores (BigQuery, alloydb), and integration with other Costco systems. As a Level 3 Engineer, the candidate will demonstrate the ability to independently own features and components within larger ML platform projects, and the knowledge and experience to design, build, debug, optimize and implement solutions. You'll join a high-caliber, Agile engineering cell where collective ownership and real-time collaboration are the norms. Together, we solve complex architectural challenges to deliver seamless, intelligent digital experiences for millions of members every single day.

This role will be responsible for:

- Delivering ML-driven capabilities that enhance the member experience across various digital touchpoints.
- Building ML pipeline and model serving components from the ground up.
- Ensuring the longevity, scalability and quality of our ML systems through continuous improvement, comprehensive documentation, meticulous profiling, and performance enhancements.
- Supporting and mentoring junior engineers, fostering a culture of continuous learning and improvement. Job Duties / Essential Functions :
- Contributes to the ML platform's architecture, applying principles that promote availability, reusability, interoperability, and security within the design framework.
- Builds and operates ML training, deployment, and inference pipelines on GCP. Implements and automates model deployment workflows, including A/B testing, rollouts, rollbacks, and auto-scaling.
- Implements monitoring and observability for ML systems, covering model performance, drift, latency, and business KPIs.




- Builds and maintains feature engineering pipelines and feature store integrations for batch and real-time serving. Builds streaming and event-driven data integrations for features, prediction logging, and asynchronous ML workflows.
- Follows and helps improve engineering best practices, coding standards, and CI/CD pipelines for ML.
- Mentors junior engineers by providing guidance, code review, and onboarding support.
- Understands the technology stack and underlying applications, services, and data stores in order to ensure optimal performance.
- Partners with data scientists to productionize models and reduce time from experiment to production.
- Applies governance, security, and compliance requirements to ML workloads, including access controls, lineage, and reproducibility standards.
- Performs development, debugging, optimization, and automation activities to support the implementation of the product/application. Uses test-driven development practices to detect defects early, including data validation and model quality gates.
- Conducts peer code reviews for the changes made by other engineers within the team.
- Improves reliability and cost efficiency of ML workloads through resource tuning and performance profiling.
- Contributes specifications and documentation across all phases of the product development cycle, from design to implementation. Works with the product and data science teams on refining requirements and delivery priorities.
- Estimates, plans, and manages assigned implementation tasks and reports on progress.
- Regular and reliable workplace attendance at your assigned location. Non-Essential Functions:
- Assists in other areas of the department as necessary.
- Assists in other areas of the company as necessary.
- Ability to operate vehicles, equipment or machinery
- Same as Essential Functions Experience, Skills, Education & Licenses/Certifications:
- 6+ years of overall software or data engineering experience, including 3+ years in MLOps, ML platform, or ML infrastructure engineering.
- 5+ years of hands-on Python development with solid software engineering fundamentals, including testing, version control, and modular design.
- 3+ years designing and deploying applications in a public cloud environment (GCP preferred),



with hands-on experience across Vertex AI, BigQuery, Dataflow, Dataproc, Cloud Composer, Pub/Sub, and GKE.3+ years building and orchestrating ML and data pipelines using Vertex AI Pipelines, Kubeflow, or Airflow / Cloud Composer.
- 3+ years with containerization and orchestration (Docker, Kubernetes) and Infrastructure as Code (Terraform).
- 3+ years with CI/CD tools and practices for ML: GitHub, Jenkins.
- 3+ years deploying and operating models in production, including model registry, versioning, rollout/rollback strategies, and auto-scaling inference. 2+ years implementing model observability: performance monitoring, drift and skew detection, dashboards, and alerting.
- 2+ years with event and streaming technologies such as GCP Pub/Sub, Dataflow, or Kafka.2+ years with experiment tracking and model management tooling (MLflow, Weights & Biases) and feature store usage Hands-on experience with ML frameworks such as TensorFlow, PyTorch, scikit-learn, or XGBoost, sufficient to productionize models built by data science teams. Solid knowledge of data application development in relational and no-SQL platforms, such as BigQuery or SpannerDB, including advanced SQL.
- Working knowledge of cloud security, compliance, and cost optimization practices for ML workloads. 3+ years of experience developing within an agile methodology. Strong verbal and written communication skills and be able to communicate to both technical and business audiences. Responsible, conscientious, organized, self-motivated and able to work with limited supervision, with strong problem-solving skills and the ability to work effectively under pressure.
- Able to support off-hours work as required, including weekends, holidays, and 24/7 on call responsibilities on a rotational basis. Bachelor's degree in Computer Science, Engineering, or a related field. Recommended: Experience in a retail ecommerce environment with recommendation, personalization, search ranking, forecasting, or pricing models.
- Experience with distributed training and GPU/TPU workload management. Practical experience operating ML systems in a high scale production environment, including performance analysis and optimization of inference services. Experience with real-time / streaming feature engineering and online serving architectures.
- Exposure to LLMOps, RAG pipelines, vector databases, or generative AI deployment patterns. Google Cloud Qualified Machine Learning Engineer certification. Master's degree in Computer Science, Machine Learning, or a related field..

📌 Software Engineer - MLOps (Hyderabad)
🏢 Costco
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: software engineer - mlops (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: software engineer - mlops (hyderabad) / hyderabad