
Job Overview
Location
San Francisco
Job Type
Full-time
Category
Product Management
Date Posted
June 21, 2026
Full Job Description
đź“‹ Description
- • Build and operate large-scale web crawling infrastructure capable of continuously collecting data across billions of URLs
- • Design and optimize URL discovery, prioritization, scheduling, and crawl orchestration systems
- • Develop distributed crawlers that efficiently acquire content while respecting site constraints and operational requirements
- • Build systems for content extraction, rendering, parsing, and normalization across diverse web formats including HTML, JavaScript-rendered pages, and dynamic content
- • Improve crawl coverage, freshness, efficiency, and quality through measurement, experimentation, and data-driven iteration
- • Design infrastructure for large-scale recrawling, change detection, and incremental updates to maintain dataset accuracy
- • Develop specialized crawlers for high-value domains, dynamic websites, and difficult-to-access content sources such as login-protected or JavaScript-heavy pages
- • Analyze crawl performance and web coverage to identify gaps, inefficiencies, and opportunities for improvement in data acquisition
- • Build observability, monitoring, and reliability systems for large-scale crawl operations to ensure uptime and data integrity
- • Debug production issues and continuously improve the performance, scalability, and resilience of crawling infrastructure
- • Collaborate closely with pre-training, infrastructure, and data quality teams to align crawling strategies with model training needs
- • Work directly with AI researchers to understand which parts of the web most impact model performance and prioritize crawling efforts accordingly
- • Balance tradeoffs between crawl quality, coverage, freshness, and operational efficiency to maximize value for foundational model training
- • Operate systems that process petabyte-scale datasets with high throughput and low latency
- • Design and execute experiments to improve crawl quality, coverage, and efficiency using empirical data and performance metrics
- • Communicate system tradeoffs and operational constraints clearly to cross-functional teams including researchers and engineers
🎯 Requirements
- • Experience building large-scale web crawling, search indexing, content acquisition, or internet-scale data collection systems
- • Strong understanding of crawling architectures, URL frontier management, scheduling, and distributed crawl coordination
- • Experience with large-scale distributed systems using technologies such as Ray, Spark, Beam, Flink, or similar frameworks
- • Familiarity with content extraction, HTML parsing, browser automation, rendering systems, and modern web technologies
- • Experience operating systems that process petabyte-scale datasets
- • Strong systems engineering skills, including reliability, observability, performance optimization, and debugging
🏖️ Benefits
- • Top-tier compensation: Salary and equity structured to recognize and retain the best talent globally
- • Comprehensive medical, dental, vision, life, and disability insurance
- • Fully paid parental leave for all new parents, including adoptive and surrogate journeys, with financial support for family planning
- • Paid time off when needed, relocation support, and daily lunch and dinner provided
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About ReflectionAI Inc.
ReflectionAI builds autonomous AI agents for enterprise process automation. The platform lets organizations create, deploy, and manage software agents that observe workflows, make decisions, and act across internal systems. Using reinforcement learning and large language models, agents learn from human guidance and adapt to changing environments. Customers use the technology for customer support triage, IT operations, compliance monitoring, and sales process automation, reducing repetitive manual tasks. The company offers cloud-hosted and on-premise deployments, role-based access controls, audit trails, and integrations with common business applications including Salesforce, ServiceNow, Jira, and Slack.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

Hightop Health Inc.
1 month ago

Syneos Health, Inc.
3 months ago

Accordion Partners, LLC
1 month ago
