Senior Applied Research Scientist, Data Curation
Posted 23hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior Applied Research Scientist developing NVIDIA’s petabyte-scale multimodal data-curation and deduplication pipelines. Advancing document extraction for foundation-model training and NVIDIA Nemotron.
Responsibilities:
- Develop efficient and performant models and data pipelines that extract and curate multimodal data—documents, images, audio, and videos—for foundation-model training
- Build petabyte-scale extraction and content-deduplication pipelines, including document and HTML parsing, fuzzy and near-duplicate deduplication, semantic deduplication, and substring deduplication
- Expand and optimize curation methodologies for petabyte-scale multimodal data across hundred-node GPU clusters
- Explore and craft datasets, metrics, experiments, and validation scripts to develop standard research methodologies
- Help ML Engineers scale pipelines to production through NVIDIA Inference Microservices (NIMs) and blueprints
- Write papers, blog posts, documentation, and training materials for customers
- Keep up to date with data-curation developments in academia and industry
- Collaborate with Applied Research Scientists, Machine Learning Engineers, and MLOps Engineers
- Mentor junior engineers and interns
Requirements:
- Master's, Ph.D. or equivalent experience in data curation, document AI, information retrieval or multimodal research
- Track record of publication in leading conferences such as CVPR, ICCV, ECCV, KDD, etc.
- Hands-on experience developing computer vision and document-extraction models and pipelines, including layout analysis, OCR, and table, figure, or formula extraction
- Kaggle Grandmaster status or a strong record of top-tier results in machine learning competitions is a strong plus
- Understanding of the state of the art in data curation research, focused on multimodal content extraction and deduplication
- 10+ years of experience developing multimodal systems across a range of models and platforms
- Information retrieval experience is a big plus
- Expertise managing distributed data frameworks such as Ray, Spark, or Dask
- History of deploying massive, multi-node machine learning or data processing tasks within production environments
- Knowledge of batching, streaming, and scaling ingestion pipelines
- Excellent Python programming skills
- Strong hands-on experience with PyTorch or comparable modern deep learning frameworks
- Ability to communicate ideas through blog posts, papers, kernels, GitHub, etc.
- Excellent communication and interpersonal skills
- Ability to work in a dynamic, user-focused, distributed team
Benefits:
- Highly competitive salary
- Equity
- Comprehensive benefits package


















