
Job Overview
Location
San Francisco
Job Type
Full-time
Category
Software Engineering
Date Posted
June 21, 2026
Full Job Description
đź“‹ Description
- • Design and implement evaluation frameworks that enable Evaluation-Driven Development for AI systems deployed in customer environments
- • Define how system quality is measured in each domain, ensuring evaluation signals reflect real user needs, domain constraints, and business objectives
- • Build and maintain golden test cases and regression suites in Python using both human-authored and AI-assisted test generation to capture critical behaviors and edge cases, treating these suites as first-class system components that evolve alongside AI systems
- • Develop and maintain offline and online evaluation pipelines that integrate directly into system iteration loops, where evaluation results directly inform prompt design, agent logic, model selection, and release readiness
- • Define, calibrate, and operate LLM-based graders to align automated judgments with expert human assessments, investigating and refining grading approaches when evaluation signals diverge from real-world outcomes
- • Collaborate closely with Forward Deployed AI Engineers, Architects, Product Engineers, AI Strategists, and domain experts to ensure evaluation frameworks meaningfully guide system development and deployment in production
- • Write clean, maintainable Python code for production-grade evaluation and experimentation pipelines, treating evaluation code with the same rigor as application code
- • Translate human judgment from subject matter experts into scalable test cases, scoring functions, and graders that automate quality assessment across diverse domains
- • Operate within a systems-oriented mindset, understanding how evaluation interacts with prompts, agents, data, and deployment to support fast iteration while maintaining trust and safety in production
- • Use AI tools to generate tests, analyze failures, explore edge cases, and accelerate debugging and iteration as part of an AI-native working style
- • Travel between 10% and 50% of the time depending on project needs, role, and personal interest, supporting on-site collaboration with customers and internal teams
- • Work in a hybrid model requiring 3+ days per week (Tuesday–Thursday) in the San Francisco office
🎯 Requirements
- • 2+ years of software engineering experience
- • Strong Python Engineering Skills: Write clean, maintainable Python and are comfortable building evaluation and experimentation pipelines that run in production environments
- • Experience with Evaluation-Driven or Experiment-Driven Development: Use structured evaluation or experimentation frameworks to drive system iteration and understand pitfalls of overfitting to metrics that don’t reflect real outcomes
- • Ability to Translate Human Judgment into Code: Work with subject matter experts to elicit high-quality judgments and encode them into test cases, scoring functions, and graders that scale
- • Systems-Oriented Mindset: Understand how evaluation interacts with prompts, agents, data, and deployment to support fast iteration while maintaining trust and safety in production
- • AI-Native Working Style: Use AI tools to generate tests, analyze failures, explore edge cases, and accelerate debugging and iteration
🏖️ Benefits
- • Base salary range of $150K – $250K, depending on experience, location, and level, plus meaningful equity
- • 100% covered medical, dental, and vision for employees and dependents
- • 401(k) with additional perks including commuter benefits and in-office lunch
- • Access to state-of-the-art AI models and generous usage of modern AI tools
- • Ownership of high-impact projects across top enterprises
- • Mission-driven, fast-moving culture that prizes curiosity, pragmatism, and excellence
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About Distyl Inc.
Distyl is a cloud-native platform designed to simplify and accelerate the development and deployment of machine learning (ML) models. It provides a unified environment for data preparation, model training, versioning, and deployment, enabling data scientists and ML engineers to move from experimentation to production faster. The platform offers features such as automated data pipelines, managed training infrastructure, and scalable model serving. Distyl aims to reduce the complexity and operational overhead associated with MLOps, allowing organizations to focus on building and deploying impactful ML solutions. It supports various ML frameworks and integrates with existing cloud infrastructure.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

Harvey AI Inc.
1 month ago

Heidi Health Pty Ltd
3 months ago
1 month ago
1 month ago
