AI Platform Engineer
Posted 7hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
AI Platform Engineer operating and automating production AI services for Deluxe. Improving model deployment, observability, reliability, and delivery across global media workflows.
Responsibilities:
- Build and operate platforms for deploying, monitoring, and maintaining production AI models and services
- Support model-serving environments, inference pipelines, deployment workflows, and release automation
- Develop automation for model packaging, promotion, validation, rollout, rollback, and lifecycle management
- Implement monitoring, logging, alerting, and observability for AI services, including performance, latency, availability, error rates, and cost
- Partner with AI Engineers to move models, workflows, and AI services from prototype to production
- Partner with cloud platform, infrastructure, and application teams to ensure AI services run on secure, scalable, and reliable environments
- Help define operational standards for production AI systems, including runbooks, incident response, testing, and release readiness
- Support batch and real-time AI workloads where applicable
- Troubleshoot issues related to model serving, API behavior, environment configuration, infrastructure, and deployment pipelines
- Build reusable patterns for AI service deployment and integration across Deluxe applications
- Contribute to evaluation and adoption of MLOps tools, model-serving frameworks, observability platforms, and AI infrastructure technologies
Requirements:
- Experience in MLOps, AI platform engineering, DevOps, software engineering, or production ML operations
- Experience deploying or operating AI/ML models, inference services, data services, or API-based production systems
- Strong scripting or programming skills using Python or similar languages
- Experience with containers, CI/CD, cloud environments, monitoring, and operational automation
- Understanding of model deployment concepts, model lifecycle management, versioning, validation, and rollback
- Experience troubleshooting production systems across application, model, infrastructure, and deployment layers
- Preferred: Experience with model-serving frameworks, MLOps platforms, Kubernetes, Docker, ECS, EKS, MLflow, KServe, BentoML, Ray, Airflow, or similar technologies
- Preferred: Experience supporting LLMs, agentic workflows, RAG systems, speech models, translation models, or other applied AI services
- Preferred: Experience with GPU-backed inference, batch processing, distributed systems, or high-throughput workloads
- Preferred: Experience with AWS services used for compute, storage, networking, security, monitoring, and deployment
- Preferred: Experience with media, localization, dubbing, content workflows, ASR, MT, TTS, or language technologies
- Preferred: Experience evaluating or integrating commercial and open-source AI platforms



















