San Francisco, California, 94102
Job description
**Job Title: AI Model Evaluation and Benchmarking Specialist**
In this role, you will be responsible for creating, executing, and validating innovative benchmarks or assessment methodologies for advanced AI models used in professional sectors. You will work closely with the APEX research team to evaluate model performance and incorporate your insights into the company’s assessment framework. Ideal candidates will possess a background in computer science, machine learning, statistics, or related disciplines, along with a clearly defined research proposal for a benchmark. You should be adept at thriving in a dynamic startup setting and prepared to dedicate a minimum of 20 hours each week.
STEMHUNTER is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, sexual orientation, gender identity, and gender expression), national origin, age, disability, genetic information, veteran status, or any other status protected by applicable federal, state, or local law. We comply with all applicable equal employment opportunity and affirmative action regulations.

