Machine Learning Platform Engineer

Posted 1hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

ML Platform Engineer building training, deployment, inference, and observability infrastructure. Supporting A1’s proactive smart assistant with scalable, reliable, cost-efficient AI systems.

Responsibilities:

  • Build and operate the ML infrastructure and platforms powering A1’s AI products
  • Design systems for model training, evaluation, deployment, inference, and experimentation
  • Build and optimise model serving and inference infrastructure for high-throughput and low-latency workloads
  • Improve reliability, scalability, latency, and cost efficiency of AI systems
  • Develop pipelines for data preparation, training, evaluation, model release, and continuous improvement
  • Build platforms and tooling that enable AI engineers and researchers to experiment, evaluate, and ship models faster
  • Develop evaluation and benchmarking infrastructure to measure model quality, performance, and regressions
  • Build production observability, monitoring, tracing, and alerting for AI/ML workloads
  • Identify bottlenecks across the ML stack and continuously improve system performance
  • Collaborate with AI engineers, researchers, and product teams to turn model requirements into production-ready infrastructure
  • Ensure AI infrastructure reliably supports production workloads at scale
  • Enable efficient model training, evaluation, deployment, and improvement
  • Deliver inference systems with strong latency, throughput, reliability, and cost efficiency
  • Ensure ML pipelines are reproducible, observable, maintainable, and robust
  • Detect and diagnose model and infrastructure regressions quickly
  • Create reusable ML infrastructure platform primitives
  • Enable the AI stack to evolve as new models, architectures, and inference techniques emerge

Requirements:

  • Strong software engineering fundamentals and experience building production systems
  • Experience building ML infrastructure, platforms, or production machine learning systems
  • Experience with model deployment, inference, evaluation, or data pipelines
  • Strong understanding of distributed systems and system reliability
  • Ability to write clean, maintainable, production-quality code
  • Experience with Python
  • Experience with PyTorch and/or JAX
  • Familiarity with LLM/ML serving infrastructure such as vLLM, SGLang, or TensorRT-LLM
  • Experience with cloud infrastructure
  • Experience with distributed systems
  • Experience with ML/data pipelines and workflow orchestration
  • Experience with GPU infrastructure and performance tooling
  • Experience with vector databases and retrieval infrastructure
  • Ability to work in ambiguous, fast-moving environments
  • Ownership, experimentation, and continuous improvement mindset