Senior Data Infrastructure Engineer (Scraping & Scale)
- Hiring from
- Probably Worldwide
- Work type
- Remote
- Posted
Is this job info correct?
532,070 remote jobs, straight from company career pages
100% free · New jobs every hour
Show job descriptionHide job description
About Us
Founded in 2018, AWISEE is a global digital marketing agency specializing in SEO, Digital PR, and KOL & Influencer Marketing. We help brands enter new markets or expand within existing ones, combining short-term visibility strategies and long-term growth solutions to drive impactful, measurable results.
Global Reach & Industry Expertise Operating across Europe, North America, Asia, and the Middle East, AWISEE tailors strategies to each market for cultural fit and compliance. We specialize in Tech & SaaS, Fintech, E-commerce, Crypto, iGaming, Travel - driving impactful growth worldwide.
Job Description
This is a remote position.
We are looking for a Senior Data Infrastructure Engineer specializing in web scraping, anti-bot evasion, and large-scale data ingestion. In this role, you will build and maintain the core ingestion engine powering our social media data lab. You will overcome complex platform defenses to deliver millions of profile, post, and video records daily with near-zero downtime.
Responsibilities
- Anti-Bot Evasion Architecture: Build and manage stealth scraping clusters using residential proxy networks, TLS fingerprinting, headful/headless browser farms (Playwright, Puppeteer), and session rotation.
- High-Throughput Pipelines: Build fault-tolerant, scalable web-scraping pipelines that extract data from Instagram, TikTok, YouTube, X, and web sources.
- Pipeline Orchestration: Design distributed queues and workflow engines (Temporal, Ray, Apache Kafka, Celery) to manage millions of asynchronous scraping tasks daily.
- Storage & Data Lake Management: Architect structured and unstructured storage environments (Parquet, Apache Iceberg, Snowflake, S3) for downstream AI modeling.
- Monitoring & Evasion Recovery: Implement automated alerts for platform UI/API changes, blocking patterns, and proxy failures.
Requirements
- 4+ years of experience in high-volume web scraping, data engineering, or reverse engineering.
- Deep experience defeating advanced anti-bot providers (Cloudflare, Akamai, PerimeterX) via TLS impersonation, browser automation, and proxy management.
- Mastery of Python or Go, with deep knowledge of Playwright, Puppeteer, Scrapy, or Selenium.
- Experience with distributed systems and task queues (Temporal, Ray, Kafka, Redis).
- Strong SQL skills and experience with modern analytical data lakes / data warehouses.
Preferred Qualifications
- Direct experience extracting short-form video content and user metadata from TikTok, Instagram, and YouTube.
- Experience integrating scraping outputs directly into vector databases and real-time AI processing queues.