
Job Overview
Location
Remote
Job Type
Full-time
Category
HR & Recruiting
Date Posted
August 11, 2026
Full Job Description
đź“‹ Description
- • Own Assessment Design — Define what Speak measures, why, and how often across three distinct assessment types (Curriculum Mastery Assessment, Proficiency Test, Placement Test) — with the Proficiency Test as the immediate focus, expanding to the other two as the pod's priorities evolve.
- • Define Constructs & Build the Rubric/Blueprint Layer — Translate fuzzy goals like "measure fluency" or "measure pronunciation" into concrete, scoreable constructs, item blueprints, and rubrics that an item writer can generate items against and an ML Engineer can build a grading model against.
- • Own Validity & the Quality Bar — Sign off on content validity for every assessment that ships. Decide what "mastery" or a passing score operationally means, catch cases where an assessment is measuring the wrong thing before it ships, own the rubric/rater guidelines behind the human-labeled data our ML scoring models are evaluated against, and audit items/rubrics for bias across learner subgroups.
- • Design and run the validity evidence plan — so validity is built into the process rather than checked only after launch. This includes concurrent/criterion studies benchmarking Speak’s assessments against external proficiency measures (CEFR-anchored exams, expert human ratings), so we can say what a Speak Score means in terms the outside world already trusts.
- • Partner Tightly with Product and ML — Work closely with the Product Manager and ML Engineers on automated scoring, calibration, and feedback generation. You own the construct and quality bar, they own the model. Neither works without the other, and the loop between you is the product.
🎯 Requirements
- • Assessment/Psychometric Design: 4+ years designing rubrics, blueprints, and item specs for a real, shipped language assessment product (or equivalent depth in closely related psychometric/measurement work) — not just academic theory.
- • Language Proficiency Domain Expertise: Deep familiarity with frameworks like CEFR (or ACTFL, IELTS/TOEFL band descriptors) and what separates "did you learn what we taught you" from "how good is your speaking overall."
- • Fairness Across Learner Populations: Can identify whether an item or rubric unfairly penalizes specific L1 backgrounds or accents (differential item functioning) — essential for a speech-based test serving learners across dozens of native languages.
- • Translates Qualitative → Technical: Can turn a construct like "pronunciation quality" into something concrete enough for an ML engineer to build a scoring pipeline against, without either oversimplifying or getting lost in academic nuance.
- • Quantitative Rigor: Comfortable running or interpreting the statistics behind a rubric or rater system — inter-rater reliability (e.g., Cohen's/Fleiss' kappa), classical test theory, and basic IRT concepts — enough to know whether a scoring system is actually reliable, not just plausible.
- • Ownership of Quality Bar: Comfortable being the sign-off authority on content validity — makes the call clearly and follows through on it, rather than deferring to data alone or product pressure to ship.
- • AI Fluency & Judgment: Uses AI tools directly in their own workflow (e.g., drafting item variants, testing rubric language, exploring construct definitions) and has real judgment about when AI-generated output is precise enough to ship vs. needs a human rewrite — distinct from spec'ing work for the ML Engineer to build.
- • Comfort with Ambiguity: Comfortable operating in a 0-to-1 environment. Can wear multiple hats, take a fuzzy goal and turn it into a concrete plan, communicate tradeoffs clearly, and keep momentum without waiting for perfect clarity or team setup
🏖️ Benefits
- • Competitive salary
- • Opportunity to work on a high-impact project that can change people's lives
- • Collaborative and dynamic work environment
- • Flexible work arrangements
- • Access to cutting-edge technology and tools
- • Opportunities for professional growth and development
- • Comprehensive benefits package
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About Speak AI Inc.
Speak AI Inc. provides an AI-powered platform that records, transcribes, and analyzes audio, video, and text data in over 70 languages. It converts unstructured conversations into searchable, shareable insights and automatically identifies keywords, topics, and sentiment. Teams use the service for qualitative research, sales calls, meetings, and academic interviews, gaining dashboards and reports without manual tagging. The company also offers an API and integrations with Zoom, Google Meet, and other tools, enabling embedded transcription and analytics workflows.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

TransUnion LLC
1 month ago
1 month ago

OpenAI, Inc.
3 months ago

ICW Group Insurance Companies
4 months ago
