Research Engineer - Web Crawlers
Build and scale distributed web crawlers to source high-quality, multilingual data from the open web for training frontier AI models. Design targeted pipelines for audio, video, and text, and develop tooling for researchers to monitor and access crawled data efficiently. Focus on reliability, content extraction, deduplication, and compliance at web scale.