
Job Overview
Location
San Francisco
Job Type
Full-time
Category
Product Management
Date Posted
June 21, 2026
Full Job Description
đź“‹ Description
- • Build and operate large-scale data ingestion systems for training frontier AI models, including web crawling, data extraction, normalization, versioning, and delivery to pre-training pipelines.
- • Design and implement specialized crawlers for high-priority data sources to acquire high-quality, diverse datasets from the open web and other large-scale sources.
- • Run experiments to evaluate crawling strategies, extraction methods, and ingestion tradeoffs, using measurable downstream impact on model performance to guide iterations.
- • Analyze ingested datasets to identify gaps, redundancy, and quality issues, and propose improvements to enhance training corpus effectiveness.
- • Develop scalable, observable, testable, and maintainable ingestion pipelines capable of handling multi-TB to PB-scale data campaigns.
- • Collaborate directly with AI researchers, pre-training teams, and data quality engineers to close the loop between data acquisition and model performance outcomes.
- • Debug production issues in ingestion infrastructure and continuously improve system reliability, efficiency, and performance.
- • Review code and contribute to engineering standards across the data ingestion stack, ensuring robustness and reproducibility in data pipelines.
- • Work with technologies such as Ray, Beam, Spark, or similar distributed systems to manage complex data acquisition workflows.
- • Communicate system behavior, design tradeoffs, and experimental results clearly to cross-functional teams including researchers, operations, and external partners.
- • Contribute to defining the foundation of open superintelligence by building the data infrastructure that enables training of next-generation foundational models.
- • Operate in a hybrid research–engineering environment where data decisions directly influence model capabilities and require rapid experimentation and iteration.
- • Participate in daily team lunches and dinners, regular off-sites, and team celebrations as part of a talent-dense, collaborative culture.
- • Receive relocation support and flexible paid time off to optimize personal and professional balance.
🎯 Requirements
- • Experience building web crawling, data ingestion, or large-scale data acquisition systems using Ray, Beam, Spark, or similar technologies.
- • Familiarity with how LLMs are trained and evaluated, and an intuition for what makes data useful for training.
- • Comfortable working with very large datasets (multi-TB to PB scale) and building systems that are observable, testable, and maintainable.
- • Comfortable designing experiments and using data to guide system improvements.
- • Excellent communication skills; able to explain system behavior and communicate tradeoffs clearly.
- • Ability to collaborate tightly across functions: researchers, infra, operations, and external partners.
🏖️ Benefits
- • Top-tier compensation: Salary and equity structured to recognize and retain the best talent globally.
- • Comprehensive medical, dental, vision, life, and disability insurance.
- • Fully paid parental leave for all new parents, including adoptive and surrogate journeys, with financial support for family planning.
- • Paid time off when needed, relocation support, and daily lunch and dinner provided.
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About ReflectionAI Inc.
ReflectionAI builds autonomous AI agents for enterprise process automation. The platform lets organizations create, deploy, and manage software agents that observe workflows, make decisions, and act across internal systems. Using reinforcement learning and large language models, agents learn from human guidance and adapt to changing environments. Customers use the technology for customer support triage, IT operations, compliance monitoring, and sales process automation, reducing repetitive manual tasks. The company offers cloud-hosted and on-premise deployments, role-based access controls, audit trails, and integrations with common business applications including Salesforce, ServiceNow, Jira, and Slack.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

Newfront Insurance Services Inc.
3 months ago

Centerstone
2 months ago

Abridge AI, Inc.
1 month ago

nCino, Inc.
1 month ago