Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
MO

Data Scraping Genius

Moovsoon
Posted 6 hours ago
🇵🇭Philippines🏠Remote📁Data & Analytics
Is this job info correct?

Full-Time | Remote | Growing Startup


We are looking for someone who is exceptionally good at large-scale web data collection.


Not someone who has built a few Scrapy spiders. Not someone whose answer is “we can use Bright Data.” Not someone who plans to point an AI agent at websites and hope it works.


We need an engineer who has already collected property listing data across the United States at meaningful scale and understands what happens when a scraper has to operate every day, across thousands of markets, against constantly changing websites and anti bot systems.


Previous experience scraping U.S. real estate or property listing sites is required. Please do not apply without it.


You Must Submit A Quick Note With Your Resume That Higlights This Previous Experience.


What We Are Building


We need near real time coverage of active U.S. property listings.

The objective is straightforward:


Capture the newest property listing activity across the entire United States, every day, reliably and cost-effectively.


This includes discovering new listings quickly, detecting changes to existing listings, maintaining broad geographic coverage, normalizing records across sources, and operating the collection infrastructure continuously.


We already understand this problem and have existing scraping infrastructure. You will be speaking with technical people who have done this before.


We are hiring someone who can take the system significantly further.



What You Will Own


You will eventually own the company's data acquisition and scraping infrastructure.


That includes:

  • Nationwide property listing collection
  • New listing discovery
  • Incremental and high-frequency crawling
  • Listing update detection
  • Property detail extraction
  • Image and media collection where appropriate
  • Source prioritization
  • Geographic coverage monitoring
  • Data normalization
  • Duplicate detection and entity resolution
  • Scraper health monitoring
  • Failure detection and automatic recovery
  • Proxy infrastructure
  • Browser automation infrastructure
  • Anti-bot resilience
  • Cost optimization
  • Queueing and distributed crawling
  • Data freshness measurement
  • Source-specific maintenance
  • Infrastructure for adding new data sources quickly


The role will expand beyond property listings as our company adds additional proprietary datasets.



You Should Already Understand


We expect candidates to be comfortable discussing, in detail:



Proxy Infrastructure

  • Residential vs ISP vs datacenter proxies
  • Sticky sessions
  • IP rotation strategies
  • Geo-targeted proxies
  • ASN diversity
  • Proxy reputation
  • Success rate by provider and source
  • Bandwidth economics
  • Cost per successful request
  • When a premium provider such as Bright Data, Oxylabs, Decodo, NetNut, etc. makes sense
  • When building additional infrastructure around those services is necessary


Simply saying “use Bright Data Web Unlocker” is not an architecture.

Managed unblockers can be part of the stack, but you should understand when they become prohibitively expensive, where they fail, and how to build a more customized collection system around them.



Fingerprinting and Anti-Bot Systems

You should understand modern detection systems including concepts such as:

  • TLS fingerprints
  • HTTP fingerprints
  • Browser fingerprints
  • Header consistency
  • User agent consistency
  • Cookies and session state
  • JavaScript challenges
  • Browser behavior
  • Request velocity
  • IP reputation
  • ASN detection
  • Geographic inconsistencies
  • Headless browser detection
  • CAPTCHA systems
  • Rate limiting
  • Behavioral detection
  • Session management


You should know why simply rotating IP addresses is usually not enough.



Browser and Request Architecture

You should know when to use:

  • Direct HTTP requests
  • Scrapy
  • Playwright
  • Puppeteer
  • Chromium
  • Browser pools
  • Remote browsers
  • Managed unblockers
  • Source APIs
  • Mobile or alternate endpoints where legitimately available
  • Hybrid request/browser architectures


The best solution is not necessarily the most complicated one. We care about reliability, freshness, and cost per usable record.



Scale Matters


The biggest requirement for this position is understanding the difference between:


“I scraped a real estate website.” and “I operated nationwide real estate collection infrastructure every day.”


They are completely different problems.


At scale you need to think about:

  • Crawl scheduling
  • Incremental crawling
  • Prioritizing high velocity markets
  • Avoiding unnecessary recrawls
  • Identifying listings without repeatedly crawling entire inventories
  • Detecting stale records
  • Distributed workers
  • Queue architecture
  • Backpressure
  • Retries
  • Circuit breakers
  • Request budgets
  • Proxy budgets
  • Source specific success rates
  • Schema changes
  • DOM changes
  • Silent extraction failures
  • Data validation
  • Monitoring
  • Alerting
  • Historical state
  • Deduplication
  • Address normalization
  • Listing identity across multiple sources
  • Freshness SLAs


We care just as much about how intelligently you crawl as how successfully you get a page to load.



Cost Is Part of the Engineering Problem

A solution that technically works but costs an enormous amount per million pages is not a successful solution.


You should be able to reason about:


Cost per request → success rate → records extracted → unique usable listings → cost per usable listing.


The Interview


This will be a technical interview.


We will ask you to walk us through how you would design a system to collect the latest property listing activity across the entire United States every day.


Be prepared to get specific.


We are going to push beyond surface-level answers.



Ideal Background


You are likely a fit if you have:

  • 4+ years of backend, data acquisition, crawling, scraping, or infrastructure experience
  • Direct experience collecting U.S. property listing data
  • Experience running production crawlers continuously
  • Python expertise
  • Strong knowledge of HTTP and browser behavior
  • Experience with distributed systems
  • Experience with proxy networks
  • Experience with Playwright, Puppeteer, Scrapy, or comparable tooling
  • Experience with queues such as Kafka, RabbitMQ, SQS, Redis, or similar systems
  • Experience with large-scale data pipelines
  • Strong SQL skills
  • Experience monitoring scraping reliability and data quality
  • A demonstrated ability to reduce infrastructure and proxy costs


Experience owning hundreds of millions of requests, millions of records, or similarly large collection workloads is a major plus.



What This Role Can Become


This is not a maintenance scraping position.


You would start with one of the most important datasets in our business and have the opportunity to become the technical owner of our broader data infrastructure and acquisition platform.


As we expand, the mandate expands with it.


We are a growing startup with a strong engineering culture. We move quickly; we already know enough about this space to recognize hand-waving, and we are looking for someone who knows considerably more about large-scale data collection than we do.

Similar jobs

Similar jobs

SA

Learning Experience Designer – AI-Enabled Learning & Development

Satelliteoffice

🇵🇭Philippines4 hours ago
Accessoffshoring logo

Service Admin (AO-14140)

Accessoffshoring

🇵🇭Philippines1 hour ago
Hedgeserv logo

Fund Accounting Manager

Hedgeserv

🌍Philippines, United States3 hours ago
Virtualcolleague logo

General Virtual Assistant

Virtualcolleague

🇵🇭Philippines3 hours ago
Virtualcolleague logo

Executive Virtual Assistant (EA) – Executive Support, Coordination & LinkedIn Outreach

Virtualcolleague

🇵🇭Philippines3 hours ago
Livenation logo

Senior Recruiter (Remote, Philippines)

Livenation

🇵🇭Philippines3 hours ago