Research Engineer Intern – AI Systems
Posted 1ds ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Research Engineer Intern focused on AI Systems at Yotta Labs. Work on optimizing GPU kernels and LLM infrastructure in a remote setting.
Responsibilities:
- Implement and optimize compute kernels for Attention, GEMM, MoE, and quantization on NVIDIA, AMD, or AWS Trainium.
- Build custom operators using CUDA, Triton, ROCm/HIP, or the Neuron SDK with PyTorch/XLA.
- Profile and improve inference performance in vLLM, SGLang, and our custom runtimes — kernel fusion, scheduling, KV-cache and memory optimizations.
- Build benchmarks, chase down performance regressions, and turn profiler traces into concrete speedups.
- Ship code upstream to open-source AI infrastructure projects, with tests and documentation.
Requirements:
- Currently pursuing a BS, MS, or PhD in Computer Science, Computer Engineering, or a related field
- Solid programming skills in Python and familiarity with C++
- Understanding of GPU/accelerator architecture fundamentals (memory hierarchy, parallelism, occupancy) from coursework, research, or projects
- Experience writing CUDA, Triton, ROCm/HIP, or Neuron kernels — class projects and personal projects count
- Strong understanding of AI frameworks (e.g., PyTorch, Dynamo, LMCache), model architectures and profiling tools (e.g. Nsight, ROCm Profiler, or Neuron Profiler)
- Strong problem-solving skills and the ability to work independently in a collaborative, remote environment.
Benefits:
- Competitive internship compensation
- Flexible remote work environment
- Direct mentorship from engineers from leading institutions and tech companies
- Fast path to a full-time return offer for top performers




















