
Job Overview
Location
Palo Alto
Job Type
Full-time
Category
Data Science
Date Posted
August 8, 2026
Full Job Description
📋 Description
- • We're a small, fast-moving AI productivity startup (~25 people) building an autonomous AI executive assistant that operates across email, calendars, meetings, and business software. This role owns the measurement system that determines whether our AI agent is genuinely improving in ambiguous, real-world environments.
- • You'll partner closely with AI Agent Capabilities engineers to produce the evidence that drives product decisions, model choices, and release quality — turning hard questions about agent behavior into rigorous, actionable answers.
- • You'll architect and maintain automated evaluation pipelines that measure agent quality across product surfaces.
- • You'll translate product capabilities into explicit pass, partial-pass, and failure criteria for complex multi-step tasks.
- • You'll build representative gold datasets and regression suites covering real workflows, edge cases, and adversarial scenarios.
- • You'll define and track metrics including task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost, and reliability.
- • You'll design deterministic and model-based graders, calibrate LLM-as-a-judge systems, and monitor grader agreement.
- • You'll compare models, prompts, and implementations using rigorous offline experiments and production evidence.
- • You'll analyze traces and production outcomes to identify root causes and build a practical failure taxonomy.
- • You'll convert production failures into regression cases and continuously close gaps in evaluation coverage.
- • You'll build dashboards and release-quality signals that make results actionable for engineering, product, and leadership.
- • You'll partner with capability engineers to recommend improvements and verify that fixes raise quality without introducing unacceptable regressions.
🎯 Requirements
- • 4+ years in Applied Data Science or Machine Learning roles, with a focus on building and delivering evaluation systems, automated data pipelines, or production ML infrastructure.
- • Demonstrated experience designing and implementing automated evaluation frameworks, success criteria, and regression suites for complex AI/ML or agentic systems.
- • Production-grade proficiency in Python and SQL, with hands-on experience building and maintaining automated analytical pipelines on large datasets.
🏖️ Benefits
- • Competitive salary
- • Stock options
- • Flexible work arrangements
- • Professional development opportunities
- • Collaborative and dynamic work environment
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About Clera
Clera, Inc. operates as an AI Talent Agent, leveraging artificial intelligence to enhance and streamline various aspects of the talent acquisition and management process. The company's core offering appears to center on an AI-powered platform designed to connect skilled individuals with suitable opportunities or to assist organizations in efficiently sourcing and evaluating candidates. While comprehensive details regarding specific features, target industries, or the scale of its operations are not yet publicly available, the current online presence indicates that Clera's services are actively being prepared for launch. This suggests an upcoming introduction of innovative AI solutions aimed at optimizing the ecosystem where talent meets demand.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities
3 months ago

Hangar Aviation Technologies, Inc.
30 days ago

Singular Intelligence Ltd.
2 months ago

