Prime Intellect, Inc. logo

Applied Research - Evals & Data

Job Overview

Location

San Francisco

Job Type

Full-time

Category

Engineering

Date Posted

July 10, 2026

Full Job Description

đź“‹ Description

  • • This is a customer-facing role at the intersection of cutting-edge RL/post-training methods, applied data, and agent systems.
  • • You’ll have a direct impact on shaping how advanced models are aligned, evaluated, deployed, and used in the real world by:
  • • Advancing Agent Capabilities: Designing and iterating on next-generation AI agents that tackle real workloads—workflow automation, reasoning-intensive tasks, and decision-making at scale.
  • • Building Robust Infrastructure: Developing the distributed systems, evaluation pipelines, and coordination frameworks that enable these agents to operate reliably, efficiently, and at massive scale.
  • • Bridge Between Customers & Research: Translating customer needs and insights from applied data into clear technical requirements that guide product and research priorities.
  • • Prototype in the Field: Rapidly designing and deploying agents, evals, and harnesses alongside customers to validate solutions.
  • • Work side-by-side with customers to deeply understand workflows, data sources, and bottlenecks.
  • • Prototype agents, data pipelines, and eval harnesses tailored to real use cases, then hand off hardened systems to core teams.
  • • Translate customer insights and evaluation results into roadmap and research direction.
  • • Design and implement novel RL and post-training methods (RLHF, RLVR, GRPO, etc.) to align large models with domain-specific tasks.
  • • Build evaluation harnesses and verifiers to measure reasoning, robustness, and agentic behavior in real-world workflows.
  • • Integrate applied data collection and analytics into the post-training process to surface regressions, emergent skills, and alignment opportunities.
  • • Prototype multi-agent and memory-augmented systems to expand capabilities for customer-facing solutions.
  • • Rapidly prototype and iterate on AI agents for automation, workflow orchestration, and decision-making.
  • • Extend and integrate with agent frameworks to support evolving feature requests and performance requirements.
  • • Architect and maintain distributed training and inference pipelines, ensuring scalability and cost efficiency.
  • • Develop observability and monitoring (Prometheus, Grafana, tracing) to ensure reliability and performance in production deployments.

🎯 Requirements

  • • Strong background in machine learning engineering, with experience in post-training, RL, or large-scale model alignment.
  • • Experience with applied data workflows and evaluation frameworks for large models or agents (e.g., SWE-Bench, HELM, EvalFlow, internal eval pipelines).
  • • Deep expertise in distributed training/inference frameworks (e.g., vLLM, sglang, Ray, Accelerate).
  • • Experience deploying containerized systems at scale (Docker, Kubernetes, Terraform).
  • • Track record of research contributions (publications, open-source contributions, benchmarks) in ML/RL.
  • • Passion for advancing the state-of-the-art in reasoning, measurement, and building practical, agentic AI systems.

🏖️ Benefits

  • • Cash Compensation Range of $150-300k + equity incentives
  • • Flexible Work (remote or San Francisco)
  • • Visa Sponsorship & relocation support
  • • Professional Development budget
  • • Team Off-sites & conference attendance

Skills & Technologies

Docker
Kubernetes
Terraform
Prometheus
Grafana
Remote

Ready to Apply?

You will be redirected to an external site to apply.

AI Job Fit Analysis
Pro

See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.

Prime Intellect, Inc. logo
Prime Intellect, Inc.
Visit Website

About Prime Intellect, Inc.

San Francisco–based startup building decentralized AI infrastructure that lets researchers pool compute and data to collaboratively train large models. Founded in 2023, the company offers open-source protocols and cloud orchestration tools that aggregate GPUs across providers, coordinate distributed training, and cryptographically verify contributions so participants share ownership and future rewards of the resulting models.

Get more remote jobs like this

Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.

Newsletter

Weekly remote jobs and featured talent.

No spam. Only curated remote roles and product updates. You can unsubscribe anytime.

Similar Opportunities

Dubai
Full-time
Expires Sep 14, 2026
Python
REST
Senior
+1 more

11 days ago

Remote - Munro, Argentina
Full-time
Expires Sep 2, 2026
Remote

23 days ago

Argentina
Full-time
Expires Sep 12, 2026
Python
TypeScript
AWS
+4 more

13 days ago

Expired
Argentina
Full-time
Expired Jul 27, 2026
Python
JavaScript
TypeScript
+4 more

2 months ago