Applied ML Engineer
Posted 21hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Applied ML Engineer building model-evaluation and verification infrastructure for an AI technology startup. Translating research methods into reliable APIs, experiments, and product workflows.
Responsibilities:
- Reproduce and evaluate research methods using open-weight and API-accessible models
- Design evaluation datasets, probes, scoring methods, baselines, calibration tests, and experiment harnesses
- Work with model weights, logits, hidden states, activations, model APIs, and inference infrastructure
- Build and extend evaluation infrastructure, including runners, judges, persistence, experiment orchestration, and reporting
- Turn research workflows into product experiences, including experiment configuration, runs, traces, comparisons, reports, and review workflows
- Investigate verification methods under fine-tuning, merging, quantization, distillation, safety removal, and deliberate evasion
- Design controlled experiments that separate meaningful signals from artifacts or confounders
- Write technical reports distinguishing measured evidence, interpretation, and hypotheses
- Ship production-quality systems with APIs, background jobs, observability, testing, and documentation
- Reproduce a published model-provenance or verification method within six months
- Build a repeatable model-verification runner with versioned inputs, artifacts, metrics, and reports
- Add a verification workflow to Construct and make it accessible through the Eldros UI
- Run controlled experiments across base, fine-tuned, merged, quantized, and distilled models
- Improve understanding of when verification methods succeed, fail, and why
Requirements:
- Strong Python engineering skills and hands-on experience with PyTorch and Hugging Face Transformers
- Strong understanding of ML evaluation, including dataset design, baselines, metrics, calibration, false positives, false negatives, statistical uncertainty, and reproducibility
- Ability to read ML research papers and implement methods from first principles
- Experience building production software beyond notebooks, including APIs, asynchronous jobs, databases, logging, testing, and deployment
- Comfort working with open-weight models and understanding modern LLM inference systems
- Ability to work across backend and frontend boundaries; ability to work with React/TypeScript product surfaces
- Strong technical judgment about experimental evidence
- High agency and strong sense of ownership
- Comfortable working in a fast-moving startup environment
- Useful experience in model provenance, fingerprinting, watermarking, distillation detection, red-teaming, safety evaluations, interpretability, activation and representation analysis, DSPy, LiteLLM, Temporal, Ray, vLLM, PostgreSQL/pgvector, Next.js, React, TypeScript, data visualization, experiment dashboards, GPU model serving, and adversarial evaluations

















