Machine Learning Operation Engineer (MLOps Engineer)

KL Gateway, Kerinchi, Kuala Lumpur, Malaysia
Full Time
RTA
Experienced

About the Role

We are looking for a Machine Learning Operation Engineer to design, build and maintain the cloud, API and deployment infrastructure that powers ML solutions. You will work closely with data scientists, ML engineers, software engineers and DevOps teams to deploy reliable ML models across cloud, mobile and edge environments. The ideal candidate combines strong software engineering skills with hands-on experience in cloud infrastructure, Kubernetes, MLOps and ML deployment.

Key Responsibilities

  • Design, deploy and maintain scalable cloud and Kubernetes infrastructure for ML workloads, including container orchestration, autoscaling and production deployments.
  • Build and maintain RESTful/gRPC APIs and messaging services to support ML inference, data processing and real-time or asynchronous workloads.
  • Develop reliable ML deployment and MLOps pipelines, including Docker/Podman, CI/CD, model versioning, monitoring, logging and automated deployment/retraining workflows.
  • Support vector search, database and ML infrastructure, including PostgreSQL/SQL Server, embedding-based retrieval, FAISS/LanceDB/Qdrant and performance optimization.
  • Collaborate with data scientists, ML engineers and software teams to deploy and optimize models across cloud, mobile and edge devices, including model optimization and inference performance.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • Minimum 1 years of hands-on experience in cloud infrastructure (AWS, GCP or Azure) within a production environment.
  • Strong experience with Kubernetes, Helm, Docker/Podman and CI/CD.
  • Proficiency in Python and frameworks such as FastAPI, Flask or Litestar; experience with RESTful/gRPC APIs.
  • Strong knowledge of PostgreSQL/SQL Server, database design, query optimization and software architecture.
  • Experience with vector databases/search, embeddings and technologies such as FAISS, LanceDB or Qdrant.
  • Experience with MLOps tools such as MLflow, Airflow or Kubeflow is an advantage.
  • Experience with edge/mobile ML deployment, ONNX Runtime Mobile, NCNN or MNN is an advantage.
  • Familiarity with monitoring and logging tools such as Prometheus, Grafana, ELK or Loki.
  • Strong analytical, troubleshooting, communication and collaboration skills.
  • Experience with Go, C++ or Rust is a strong advantage, particularly for performance-critical applications.
  • Excellent communication and documentation habits.
Share

Apply for this position

Required*
We've received your resume. Click here to update it.
Attach resume as .pdf, .doc, .docx, .odt, .txt, or .rtf (limit 5MB) or Paste resume

Paste your resume here or Attach resume file

Human Check*