Applied ML Engineer

Posted 21hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Applied ML Engineer building model-evaluation and verification infrastructure for an AI technology startup. Translating research methods into reliable APIs, experiments, and product workflows.

Responsibilities:

  • Reproduce and evaluate research methods using open-weight and API-accessible models
  • Design evaluation datasets, probes, scoring methods, baselines, calibration tests, and experiment harnesses
  • Work with model weights, logits, hidden states, activations, model APIs, and inference infrastructure
  • Build and extend evaluation infrastructure, including runners, judges, persistence, experiment orchestration, and reporting
  • Turn research workflows into product experiences, including experiment configuration, runs, traces, comparisons, reports, and review workflows
  • Investigate verification methods under fine-tuning, merging, quantization, distillation, safety removal, and deliberate evasion
  • Design controlled experiments that separate meaningful signals from artifacts or confounders
  • Write technical reports distinguishing measured evidence, interpretation, and hypotheses
  • Ship production-quality systems with APIs, background jobs, observability, testing, and documentation
  • Reproduce a published model-provenance or verification method within six months
  • Build a repeatable model-verification runner with versioned inputs, artifacts, metrics, and reports
  • Add a verification workflow to Construct and make it accessible through the Eldros UI
  • Run controlled experiments across base, fine-tuned, merged, quantized, and distilled models
  • Improve understanding of when verification methods succeed, fail, and why

Requirements:

  • Strong Python engineering skills and hands-on experience with PyTorch and Hugging Face Transformers
  • Strong understanding of ML evaluation, including dataset design, baselines, metrics, calibration, false positives, false negatives, statistical uncertainty, and reproducibility
  • Ability to read ML research papers and implement methods from first principles
  • Experience building production software beyond notebooks, including APIs, asynchronous jobs, databases, logging, testing, and deployment
  • Comfort working with open-weight models and understanding modern LLM inference systems
  • Ability to work across backend and frontend boundaries; ability to work with React/TypeScript product surfaces
  • Strong technical judgment about experimental evidence
  • High agency and strong sense of ownership
  • Comfortable working in a fast-moving startup environment
  • Useful experience in model provenance, fingerprinting, watermarking, distillation detection, red-teaming, safety evaluations, interpretability, activation and representation analysis, DSPy, LiteLLM, Temporal, Ray, vLLM, PostgreSQL/pgvector, Next.js, React, TypeScript, data visualization, experiment dashboards, GPU model serving, and adversarial evaluations