Senior ML Specialist
Posted 58mins ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior ML Specialist building and evaluating agentic AI for Sciene, empowering Quartile’s retail media optimization platform. Owning model strategy, evaluation, fine-tuning, and production operations.
Responsibilities:
- Own model strategy across the platform by evaluating and selecting models per task
- Use benchmarks and LLM-as-judge harnesses with statistically sound analysis to support model decisions
- Design and maintain evaluation methodology, including offline eval sets, judge calibration, regression benchmarks, and quality metrics
- Build classification and fine-tuned models for routing, categorization, and detection tasks
- Construct datasets, train, evaluate, and deploy models
- Design, develop, and ship agentic AI products, including agent identities, reusable skills, tool integrations, and structured outputs
- Perform prompt and context engineering, including system prompt assembly, live-data context injection, thread/memory management, and Pydantic structured outputs
- Own output quality through deterministic enforcers, validation quality gates, and LLM-as-judge evaluators
- Diagnose hallucination, inconsistency, and model-version drift and implement fixes through prompts, retrieval, guardrails, or model selection
- Build and extend agent tools querying Databricks, MongoDB, and external systems
- Integrate agents with the broader ecosystem through the Model Context Protocol (MCP)
- Operate shipped services with OpenTelemetry and monitor quality, latency, and cost in Grafana
- Collaborate with data engineers, software engineers, and product teams on end-to-end AI solutions
- Raise team ML fundamentals through reviews and knowledge sharing
- Critically evaluate research literature and apply validated techniques
- Document and maintain the codebase while ensuring code quality and best practices
Requirements:
- 5+ years of applied machine learning experience, including training, evaluating, and deploying models
- Strong foundation in statistics and probability, including hypothesis testing, confidence intervals, sampling, sample-size reasoning, bias/variance, and calibration
- Deep understanding of LLM and generative AI theory, including transformers, attention, tokenization, embeddings, training pipelines, decoding strategies, scaling behavior, and failure modes
- Hands-on classical ML and NLP experience, including classification, clustering, feature engineering, text classification, embeddings, and semantic similarity
- Experience fine-tuning models using LoRA/PEFT, supervised fine-tuning, or full fine-tuning; ability to build training datasets
- Experience designing AI evaluation methodology, including offline evaluation sets, LLM-as-judge, inter-rater agreement, regression benchmarks, and significance testing
- Production-grade Python experience with PyTorch, scikit-learn, Hugging Face, modern async Python, FastAPI, Pydantic, tests, and CI
- Hands-on experience with at least one major LLM provider API: OpenAI, Anthropic, or Google; structured outputs and tool/function calling
- Master's or PhD in Machine Learning, Statistics, Computer Science, Mathematics, or related quantitative field, or equivalent demonstrated depth through publications, competition results, released models, or research code
- Senior-level autonomy owning problems end to end
- Preferred: NLP/ML publications or open-source contributions; Kaggle or equivalent competition record
- Preferred: retrieval systems, agent evaluation, model cost/latency optimization, Databricks, Azure, Docker, observability stacks, MCP, and 0-to-1 product delivery experience
Benefits:
- Equal opportunity employer committed to diversity, inclusion, and belonging
- PJ contract arrangement

















