Machine Learning Platform Engineer
Posted 1hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
ML platform engineer building training, deployment, inference, and observability infrastructure. Powering A1’s proactive smart assistant with reliable, scalable, cost-efficient AI systems.
Responsibilities:
- Build and operate the ML infrastructure and platforms powering A1’s AI products
- Design systems for model training, evaluation, deployment, inference, and experimentation
- Build and optimise model serving and inference infrastructure for high-throughput and low-latency workloads
- Improve reliability, scalability, latency, and cost efficiency of AI systems
- Develop pipelines for data preparation, training, evaluation, model release, and continuous improvement
- Build platforms and tooling enabling AI engineers and researchers to experiment, evaluate, and ship models faster
- Develop evaluation and benchmarking infrastructure to measure model quality, performance, and regressions
- Build production observability, monitoring, tracing, and alerting for AI/ML workloads
- Identify bottlenecks across the ML stack and continuously improve system performance
- Collaborate with AI engineers, researchers, and product teams to turn model requirements into production-ready infrastructure
Requirements:
- Strong software engineering fundamentals and experience building production systems
- Experience building ML infrastructure, platforms, or production machine learning systems
- Experience with model deployment, inference, evaluation, or data pipelines
- Strong understanding of distributed systems and system reliability
- Ability to write clean, maintainable, production-quality code
- Comfortable working in ambiguous, fast-moving environments
- Python proficiency
- Experience with PyTorch or JAX
- Experience with ML serving infrastructure such as vLLM, SGLang, or TensorRT-LLM
- Knowledge of cloud infrastructure, distributed systems, ML/data pipelines, workflow orchestration, GPU infrastructure and performance tooling, vector databases, and retrieval infrastructure
- Professional/fluent English proficiency for technical communication
















