Speak AI Inc. logo

Assessment Design Lead

Job Overview

Location

Remote

Job Type

Full-time

Category

HR & Recruiting

Date Posted

August 11, 2026

Full Job Description

đź“‹ Description

  • • Own Assessment Design — Define what Speak measures, why, and how often across three distinct assessment types (Curriculum Mastery Assessment, Proficiency Test, Placement Test) — with the Proficiency Test as the immediate focus, expanding to the other two as the pod's priorities evolve.
  • • Define Constructs & Build the Rubric/Blueprint Layer — Translate fuzzy goals like "measure fluency" or "measure pronunciation" into concrete, scoreable constructs, item blueprints, and rubrics that an item writer can generate items against and an ML Engineer can build a grading model against.
  • • Own Validity & the Quality Bar — Sign off on content validity for every assessment that ships. Decide what "mastery" or a passing score operationally means, catch cases where an assessment is measuring the wrong thing before it ships, own the rubric/rater guidelines behind the human-labeled data our ML scoring models are evaluated against, and audit items/rubrics for bias across learner subgroups.
  • • Design and run the validity evidence plan — so validity is built into the process rather than checked only after launch. This includes concurrent/criterion studies benchmarking Speak’s assessments against external proficiency measures (CEFR-anchored exams, expert human ratings), so we can say what a Speak Score means in terms the outside world already trusts.
  • • Partner Tightly with Product and ML — Work closely with the Product Manager and ML Engineers on automated scoring, calibration, and feedback generation. You own the construct and quality bar, they own the model. Neither works without the other, and the loop between you is the product.

🎯 Requirements

  • • Assessment/Psychometric Design: 4+ years designing rubrics, blueprints, and item specs for a real, shipped language assessment product (or equivalent depth in closely related psychometric/measurement work) — not just academic theory.
  • • Language Proficiency Domain Expertise: Deep familiarity with frameworks like CEFR (or ACTFL, IELTS/TOEFL band descriptors) and what separates "did you learn what we taught you" from "how good is your speaking overall."
  • • Fairness Across Learner Populations: Can identify whether an item or rubric unfairly penalizes specific L1 backgrounds or accents (differential item functioning) — essential for a speech-based test serving learners across dozens of native languages.
  • • Translates Qualitative → Technical: Can turn a construct like "pronunciation quality" into something concrete enough for an ML engineer to build a scoring pipeline against, without either oversimplifying or getting lost in academic nuance.
  • • Quantitative Rigor: Comfortable running or interpreting the statistics behind a rubric or rater system — inter-rater reliability (e.g., Cohen's/Fleiss' kappa), classical test theory, and basic IRT concepts — enough to know whether a scoring system is actually reliable, not just plausible.
  • • Ownership of Quality Bar: Comfortable being the sign-off authority on content validity — makes the call clearly and follows through on it, rather than deferring to data alone or product pressure to ship.
  • • AI Fluency & Judgment: Uses AI tools directly in their own workflow (e.g., drafting item variants, testing rubric language, exploring construct definitions) and has real judgment about when AI-generated output is precise enough to ship vs. needs a human rewrite — distinct from spec'ing work for the ML Engineer to build.
  • • Comfort with Ambiguity: Comfortable operating in a 0-to-1 environment. Can wear multiple hats, take a fuzzy goal and turn it into a concrete plan, communicate tradeoffs clearly, and keep momentum without waiting for perfect clarity or team setup

🏖️ Benefits

  • • Competitive salary
  • • Opportunity to work on a high-impact project that can change people's lives
  • • Collaborative and dynamic work environment
  • • Flexible work arrangements
  • • Access to cutting-edge technology and tools
  • • Opportunities for professional growth and development
  • • Comprehensive benefits package

Skills & Technologies

Senior
Remote
Degree Required

Ready to Apply?

You will be redirected to an external site to apply.

AI Job Fit Analysis
Pro

See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.

Speak AI Inc. logo
Speak AI Inc.
Visit Website

About Speak AI Inc.

Speak AI Inc. provides an AI-powered platform that records, transcribes, and analyzes audio, video, and text data in over 70 languages. It converts unstructured conversations into searchable, shareable insights and automatically identifies keywords, topics, and sentiment. Teams use the service for qualitative research, sales calls, meetings, and academic interviews, gaining dashboards and reports without manual tagging. The company also offers an API and integrations with Zoom, Google Meet, and other tools, enabling embedded transcription and analytics workflows.

Get more remote jobs like this

Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.

Newsletter

Weekly remote jobs and featured talent.

No spam. Only curated remote roles and product updates. You can unsubscribe anytime.

Similar Opportunities

Remote - Ontario
Full-time
Expires Sep 8, 2026
Remote
$46k-63k

1 month ago

United States - Remote
Full-time
Expires Sep 8, 2026
GCP
Remote
Degree Required

1 month ago

Expired
Singapore
Full-time
Expired Jul 7, 2026
Hybrid

3 months ago

Expired
ICW Group Insurance Companies logo

ICW Group Insurance Companies

Remote
Full-time
Expired Jun 23, 2026
Remote

4 months ago