Senior ML Engineer
Posted 2hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior ML Engineer building CloudBolt’s recommendation engine for Kubernetes resource optimization. Reducing SaaS customers’ cloud spending through safe forecasting and production ML systems.
Responsibilities:
- Own the recommendation engine end to end, including model selection, algorithm design, preprocessing, and safety guardrails
- Design, evaluate, and productionize time-series forecasting and statistical models for right-sizing Kubernetes workloads across CPU, memory, GPU, and JVM heap
- Build and maintain the data-quality layer by detecting and filtering anomalies, load-test windows, startup spikes, and autoscaling artifacts from production telemetry
- Define and improve recommendation-quality measurement through regression testing, behavioral validation, and production accuracy and safety metrics
- Investigate and resolve recommendation-quality issues from customer environments by tracing data, preprocessing, and model behavior
- Serve as the team's machine learning authority and guide technical direction on ML questions and model-versus-heuristic tradeoffs
- Write production-grade Python for models and pipelines and share ownership of message consumption, metrics ingestion, and caching services
- Prototype and validate new optimization capabilities from research through gradual, feature-flagged rollout
- Stay current on time-series forecasting and resource optimization techniques and evaluate which are worth adopting
Requirements:
- Master's degree or higher in a quantitative field (Computer Science, Machine Learning, Statistics, Applied Mathematics)
- 5+ years of software engineering experience
- At least 3 years building and operating machine learning or statistical systems in production
- Expert-level Python, including typed, tested, production-grade code
- Fluency in numpy or similar array-based numerical computing
- Hands-on experience with time-series analysis and forecasting, including seasonality, trend decomposition, anomaly detection, and classical statistical methods
- Experience testing ML systems rigorously, including regression testing against known-good baselines, behavioral validation, and numerical reproducibility
- Working knowledge of Kubernetes, including resource requests and limits, autoscaling behavior, OOM kills, and CPU throttling
- Comfort owning a production service, including queues, caches, retries, observability, and debugging customer-environment issues from logs and metrics
- Clear written and verbal communication
- Experience with Prophet or similar forecasting libraries is beneficial
- Prometheus/PromQL and experience working with metrics at scale is beneficial
- Cloud cost optimization, capacity planning, or infrastructure efficiency background is beneficial
- AWS (S3, Managed Prometheus) experience is beneficial
- Experience being the ML domain expert on a team of generalists is beneficial
Benefits:
- Medical/Dental/Vision coverage
- 401k with Company Match
- Health & Dependent Care FSA
- Unlimited PTO
- 11 Company Holidays
- Volunteer/Community Engagement Day
- Tuition Reimbursement
- Paid Parental Leave
- Equity Grants
- Home internet Reimbursement


















