All roles

Data Scientist - AI Evaluation

Remote · USA Full-time New today

About Wizard

Wizard is the top-performing AI Shopping Agent, delivering the best products from across the web with unmatched accuracy, quality, and trust.

The Role

We’re looking for a Data Scientist to own how we measure, understand and improve the accuracy of our AI agent. This role sits at the intersection of data science, machine learning and product and is focused on evaluation, experimentation and insight generation. You won’t be building models but you will make sure they work in real world scenarios. You will build the systems to measure what good looks like and partner closely with ML, AI Engineering and Product to continuously improve the agent’s performance.

What You’ll Do

  • Define and evolve accuracy metrics across the full shopping experience (retrieval, ranking, recommendations and outcomes)
  • Design and run experiments to measure improvements and regressions
  • Build and maintain evaluation datasets, benchmarks and scoring frameworks
  • Translate ambiguous product questions into clear, measurable hypotheses and analysis
  • Partner with ML Engineers to validate model changes and guide iteration
  • Identify failure modes and edge cases and drive improvements through data
  • Create dashboards and reporting that make agent performance visible, trusted and actionable

What Success Looks like

  • Clear, trusted accuracy metrics are consistently used across product and engineering
  • A robust automated evaluation framework exists for both offline and live experiments
  • Model and product changes are consistently measured before and after launch

Ideal Background

  • 4-6+ years in Data Science, ML Evaluation or Applied AI or similar roles
  • Deep experience evaluating AI/ML systems (ranking, recommendations, LLMs, etc)
  • Strong experience with experimentation (A/B testing, causal inference)
  • Experience working on consumer products or user facing systems and exposure to marketplace or e-commerce systems
  • Ability to translate messy problems into structured analysis and metrics
  • Strong product mindset, you care about real user outcomes
  • Clear communication with the ability to influence across engineering and product

Compensation & Benefits

The expected base salary range for this role is $225,000 - $280,000 USD, and will vary based on skills, experience, role level, and geographic location. Final compensation will be determined by considering these factors alongside overall role scope and responsibilities.

In addition to base salary, Wizard offers:

  • Equity in the form of stock options
  • Medical, dental, and vision coverage
  • 401(k) plan
  • Flexible PTO and company holidays
  • Fully remote work within the United States
  • Periodic company offsites and team gatherings

Wizard is committed to fair, transparent, and competitive compensation practices.

Apply To This Job

Related roles

Machine Learning Engineer - Relevance & Learning Systems

Remote · USA Full-time

Product Manager (GB)

Remote · USA Full-time

PR & Marketing Communications Manager (Waterloo, ON, CA, N2V 1C6)

Remote · USA Full-time

Account Executive (US)

Remote · USA Full-time

Customs Rater (Waterloo, ON, N2V 1C6)

Remote · USA Full-time

Senior Graphic Designer (Contract-to-Hire)

Remote · USA Full-time

Field Marketing Manager, Onsites

Remote · USA Full-time

Account Executive

Remote · USA Full-time

Lead Product Marketing Manager, Product

Remote · USA Full-time

Cyber Security (SME)

Remote · USA Full-time

[PART_TIME Remote] Part Time Data Entry Jobs/ Work-From Home/

Remote · USA Full-time

Sales Account Executive (Remote-San Diego, CA)

Remote · USA Full-time

Clinic Patient Access (Work from Home ) - Patient Service Agent

Remote · USA Full-time

Director, AI Sourcing

Remote · USA Full-time

Summer 2026 Legal Intern, LGBTQ & HIV Project

Remote · USA Full-time

Experienced Service Support Technician for Information Systems - Providing Proactive Technology Support to Summit County Government in Breckenridge, CO

Remote · USA Full-time

Remote Data Entry Specialist - Digital Information Management Professional (arenaflex Remote Opportunities)

Remote · USA Full-time

AP & Treasury Manager

Remote · USA Full-time

Experienced Overnight Remote Chat Support Representative - Flexible Part-Time Careers with blithequark, Enjoy Adaptability and Earn $25-$35/Hour

Remote · USA Full-time

Experienced Data Entry Specialist – Remote Opportunity with arenaflex

Remote · USA Full-time