Research Data Scientist

Posted 20hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Research Data Scientist advancing Innodata’s data engineering and AI services through GenAI, LLM, NLP, and multimodal research. Building evaluation frameworks, datasets, prototypes, and statistical insights.

Responsibilities:

  • Conduct independent and collaborative research in Generative AI, LLMs, NLP, multimodal AI, machine learning, model evaluation, and AI data.
  • Formulate research questions and translate complex AI/ML problems into structured research methodologies and experiments.
  • Design, execute, and analyze experiments to evaluate and improve AI/ML models and solutions.
  • Build analytical models, prototypes, and research pipelines using Python and relevant ML frameworks.
  • Develop and implement LLM evaluation frameworks, benchmarks, datasets, and evaluation criteria.
  • Evaluate models for accuracy, robustness, bias, hallucination, reasoning, relevance, response quality, and other performance dimensions.
  • Conduct model benchmarking, error analysis, comparative analysis, and performance evaluation.
  • Work on RAG, SFT, RLHF/DPO, prompt engineering, fine-tuning, embeddings, and LLM optimization, as applicable.
  • Collect, clean, analyze, and interpret large and complex structured and unstructured datasets.
  • Perform EDA, statistical analysis, hypothesis testing, significance testing, correlation analysis, sampling, and error analysis.
  • Develop and evaluate datasets, sampling methodologies, taxonomies, annotation frameworks, data quality frameworks, and evaluation criteria.
  • Analyze data quality and identify issues affecting model performance.
  • Collaborate with annotation, data engineering, AI/ML, research, domain expert, and delivery teams.
  • Contribute to research papers, technical reports, whitepapers, patents, benchmarks, internal publications, and other research outputs.
  • Present research findings, analytical insights, and technical recommendations to senior technical stakeholders.
  • Participate in client-facing technical discussions and presentations where required.
  • Translate business requirements into AI/ML solutions and complex research concepts into actionable recommendations.

Requirements:

  • Master’s or PhD in Computer Science, Artificial Intelligence, Machine Learning, Data Science, Statistics, Mathematics, Computational Science, or a related discipline.
  • 4–7 years of hands-on research experience in AI/ML, Data Science, NLP, Generative AI, LLMs, or related areas.
  • Strong demonstrated research experience with the ability to independently formulate research questions, design experiments, analyze results, and communicate findings.
  • Demonstrated research track record through research publications, patents, conference presentations, open-source contributions, or significant AI/ML research projects.
  • Strong proficiency in Python and SQL.
  • Strong hands-on experience with NumPy, Pandas, Scikit-learn, and preferably PyTorch/TensorFlow.
  • Strong understanding of machine learning algorithms, statistics and experimentation, data analysis and feature engineering, model evaluation and performance metrics, hypothesis testing, and statistical inference.
  • Hands-on exposure to LLMs, NLP, Generative AI, and multimodal AI.
  • Experience with one or more of RAG, LLM evaluation, prompt engineering, fine-tuning, SFT, RLHF/DPO, embeddings, or model benchmarking.
  • Experience working with large-scale structured and unstructured datasets.
  • Familiarity with Git and cloud platforms such as AWS, Azure, or GCP is desirable.
  • Bachelor’s/Master’s degree from IITs, NITs, or other premier engineering/research institutions is strongly preferred.
  • Candidates with publications in reputed conferences/journals and a strong academic/research profile will be preferred.