Member of Technical Staff, Inference
Posted 7hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Inference runtime engineer optimizing vLLM, the open-source AI inference engine. Improving LLM and diffusion-model serving across hardware and architectures.
Responsibilities:
- Push the boundaries of LLM and diffusion model serving
- Work at the core of vLLM to optimize model execution across diverse hardware and architectures
- Develop inference runtime innovations for mixture-of-experts, multimodal, and agentic architectures
- Implement inference techniques and model architectures from research papers
- Contribute performant and maintainable code
- Debug complex machine-learning codebases
- Directly improve how AI inference is run by making inference cheaper and faster
Requirements:
- Bachelor's degree or equivalent experience in computer science, engineering, or similar
- Deep understanding of transformer architectures and their variants
- Strong programming skills in Python with experience in PyTorch internals
- Experience with LLM inference systems such as vLLM, TensorRT-LLM, SGLang, or TGI
- Ability to read and implement model architectures and inference techniques from research papers
- Ability to contribute performant and maintainable code and debug in complex ML codebases
- Preferred: deep understanding of KV-cache memory management, prefix caching, and hybrid model serving
- Preferred: familiarity with RL frameworks and algorithms for LLMs
- Preferred: experience with multimodal inference across audio, image, video, and text
- Contributions to open-source ML or system infrastructure projects are preferred
- Bonus: core feature implementation in vLLM or other inference engine projects
- Bonus: contributions to vLLM integrations such as verl, OpenRLHF, Unsloth, or LlamaFactory
- Bonus: widely-shared technical blogs or side projects on vLLM or LLM inference
- Required application materials: resume, GitHub handle, and a link to a relevant personal project, open-source contribution, or technical blog post
Benefits:
- Competitive benefits appropriate to your location, including health coverage where applicable
- Equity
- Visa sponsorship on a case-by-case basis
- Fully remote work
- Timezone-flexible schedule with regular overlap with Pacific Time for critical syncs















