Data Scientist
Posted 3hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Data Scientist at Applaudo, an AI-native technology company, building entity-matching models with embeddings and LLMs. Evaluating scalable, cost-effective solutions on messy, multilingual datasets.
Responsibilities:
- Build and evaluate ML approaches for company/entity matching
- Develop embedding and LLM-based matching approaches
- Develop scoring and ranking methodologies to identify true matches and distinguish them from duplicates, lookalikes, and unrelated entities
- Work with messy data, including names, aliases, domains, websites, firmographic attributes, multilingual records, and data hierarchies
- Define benchmark datasets, metrics, baselines, and error-analysis processes
- Design and execute experiments to validate hypotheses
- Compare LLM-assisted approaches against lower-cost alternatives
- Analyze model behavior, edge cases, and trade-offs
- Consider inference economics and scalability from the beginning
- Communicate experimental findings and recommendations to engineering and business stakeholders
- Independently establish experimental pipelines and research approaches
- Clearly document both successful and unsuccessful experiments
Requirements:
- 5+ years of professional Data Science / Machine Learning experience
- Strong applied Machine Learning fundamentals
- Excellent Python and SQL skills
- Hands-on experience with embeddings and semantic similarity
- Practical experience applying LLMs to real-world problems
- Experience with supervised and unsupervised learning
- Strong experience with classification and NLP
- Working knowledge of neural networks and transformer architectures
- Hands-on experience with TensorFlow, PyTorch, PyCaret, or equivalent ML frameworks
- Experience retraining or maintaining classification models in production
- Strong experimental design and model evaluation skills
- Experience defining baselines, metrics, test sets, and error-analysis processes
- Ability to evaluate model quality and demonstrate measurable improvements
- Strong understanding of scalability and ML inference costs
- Strong English communication skills
- Nice-to-have: entity resolution, record linkage, or deduplication experience
- Nice-to-have: ranking and similarity scoring
- Nice-to-have: retrieval, clustering, or candidate-generation techniques
- Nice-to-have: LLM/embedding solutions designed for cost and scale constraints
- Nice-to-have: Spark, Snowflake, Databricks, or BigQuery
- Nice-to-have: experience with company, domain, website, or firmographic data
- Nice-to-have: experience working with multilingual datasets
- Strong analytical and experimental mindset
- Intellectual honesty and willingness to communicate negative results
- Strong autonomy and self-direction
- Excellent written and verbal communication
- Ability to defend technical recommendations with stakeholders
- Strong problem-solving skills
- Comfort working with ambiguity and large-scale datasets
- Ability to balance model quality, cost, and scalability
Benefits:
- Remote work option




















