ReflectionAI Inc. logo

Member of Technical Staff - Web Crawl Engineer

Job Overview

Location

San Francisco

Job Type

Full-time

Category

Product Management

Date Posted

June 21, 2026

Full Job Description

đź“‹ Description

  • • Build and operate large-scale web crawling infrastructure capable of continuously collecting data across billions of URLs
  • • Design and optimize URL discovery, prioritization, scheduling, and crawl orchestration systems
  • • Develop distributed crawlers that efficiently acquire content while respecting site constraints and operational requirements
  • • Build systems for content extraction, rendering, parsing, and normalization across diverse web formats including HTML, JavaScript-rendered pages, and dynamic content
  • • Improve crawl coverage, freshness, efficiency, and quality through measurement, experimentation, and data-driven iteration
  • • Design infrastructure for large-scale recrawling, change detection, and incremental updates to maintain dataset accuracy
  • • Develop specialized crawlers for high-value domains, dynamic websites, and difficult-to-access content sources such as login-protected or JavaScript-heavy pages
  • • Analyze crawl performance and web coverage to identify gaps, inefficiencies, and opportunities for improvement in data acquisition
  • • Build observability, monitoring, and reliability systems for large-scale crawl operations to ensure uptime and data integrity
  • • Debug production issues and continuously improve the performance, scalability, and resilience of crawling infrastructure
  • • Collaborate closely with pre-training, infrastructure, and data quality teams to align crawling strategies with model training needs
  • • Work directly with AI researchers to understand which parts of the web most impact model performance and prioritize crawling efforts accordingly
  • • Balance tradeoffs between crawl quality, coverage, freshness, and operational efficiency to maximize value for foundational model training
  • • Operate systems that process petabyte-scale datasets with high throughput and low latency
  • • Design and execute experiments to improve crawl quality, coverage, and efficiency using empirical data and performance metrics
  • • Communicate system tradeoffs and operational constraints clearly to cross-functional teams including researchers and engineers

🎯 Requirements

  • • Experience building large-scale web crawling, search indexing, content acquisition, or internet-scale data collection systems
  • • Strong understanding of crawling architectures, URL frontier management, scheduling, and distributed crawl coordination
  • • Experience with large-scale distributed systems using technologies such as Ray, Spark, Beam, Flink, or similar frameworks
  • • Familiarity with content extraction, HTML parsing, browser automation, rendering systems, and modern web technologies
  • • Experience operating systems that process petabyte-scale datasets
  • • Strong systems engineering skills, including reliability, observability, performance optimization, and debugging

🏖️ Benefits

  • • Top-tier compensation: Salary and equity structured to recognize and retain the best talent globally
  • • Comprehensive medical, dental, vision, life, and disability insurance
  • • Fully paid parental leave for all new parents, including adoptive and surrogate journeys, with financial support for family planning
  • • Paid time off when needed, relocation support, and daily lunch and dinner provided

Skills & Technologies

Apache Spark
Senior
Onsite

Ready to Apply?

You will be redirected to an external site to apply.

AI Job Fit Analysis
Pro

See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.

ReflectionAI Inc. logo
ReflectionAI Inc.
Visit Website

About ReflectionAI Inc.

ReflectionAI builds autonomous AI agents for enterprise process automation. The platform lets organizations create, deploy, and manage software agents that observe workflows, make decisions, and act across internal systems. Using reinforcement learning and large language models, agents learn from human guidance and adapt to changing environments. Customers use the technology for customer support triage, IT operations, compliance monitoring, and sales process automation, reducing repetitive manual tasks. The company offers cloud-hosted and on-premise deployments, role-based access controls, audit trails, and integrations with common business applications including Salesforce, ServiceNow, Jira, and Slack.

Get more remote jobs like this

Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.

Newsletter

Weekly remote jobs and featured talent.

No spam. Only curated remote roles and product updates. You can unsubscribe anytime.

Similar Opportunities

Hightop Health Inc. logo

Hightop Health Inc.

Roswell, GA
Full-time
Expires Aug 20, 2026
Backend
Hybrid
Degree Required

1 month ago

Expired
ARG-Remote
Full-time
Expired Jul 6, 2026
Senior
Remote

3 months ago

Ruby Labs Ltd. logo

Ruby Labs Ltd.

Poland
Full-time
Expires Aug 23, 2026
Go
Ruby
Remote

1 month ago

Accordion Partners, LLC logo

Accordion Partners, LLC

Dallas
Full-time
Expires Aug 19, 2026
Remote
Degree Required

1 month ago