Case Study · Personal / R&D

Multi-Platform Scraping & Analytics Platform

Production data pipeline that scrapes TikTok Shop & EchoTik behind anti-bot defenses and lands audit-grade analytics.

Multi-Platform Scraping & Analytics Platform — Win Naing Soe project case study
Role
Full-Stack / Systems Engineer
Period
2025 — Present
Client
Personal / R&D
Domain
Data Engineering · E-Commerce Analytics · Web Scraping
/ Overview

Designed and built a production-oriented data pipeline that scrapes e-commerce and creator data from TikTok Shop and EchoTik, and lands it in a canonical, audit-grade analytics database for downstream reporting. The system spans three independently deployable services — a React admin console, a NestJS/BullMQ scraping engine, and a Spring Boot/Oracle analytics API — chosen deliberately because the three layers have genuinely different concerns: user-facing job control, browser-automation orchestration under anti-bot pressure, and long-term relational data integrity.

/ Problem & Challenge

Scraping competitor intelligence across TikTok Shop and EchoTik required extracting deeply authenticated, rate-limited, anti-bot-protected data without triggering account suspensions or creating noisy duplicate rows in analytical records.

/ Architecture & System Decisions

Separated concerns across 3 independently scalable services: a React admin UI, a NestJS/BullMQ scraper with Redis Lua-scripted locks and circuit breaker, and a Spring Boot/Oracle analytics API with isolated REQUIRES_NEW audit transactions.

/ Engineering Trade-offs
⚖ TRADE-OFF DECISIONBullMQ/Redis vs Database Job Polling: Kept fast, atomic, short-lived execution state in Redis and permanent business data in Oracle, matching each datastore strictly to its operational strengths.
⚖ TRADE-OFF DECISIONIndependent Audit Bookkeeping: Engineered import pipeline so job/audit bookkeeping commits in its own transaction (REQUIRES_NEW) independent of master-upsert steps, guaranteeing audit trail survival even on write failures.
/ Impact

Architected a three-service system along real operational boundaries, surviving hostile anti-bot defenses while producing data clean enough to trust for analytics — with an audit trail that survives failure, not just success.

/ Highlights
  • ◆Built a session-coordination layer (Redis-backed) that serializes all EchoTik scraping to a single active session across 5 parallel workers, using atomic Lua-scripted locks, jittered delays, per-session API-call budgets, and a circuit breaker with exponential backoff.
  • ◆Chose BullMQ/Redis over a polling database table for job orchestration — keeping fast, atomic, short-lived execution state in Redis and permanent business data in Oracle, each in the datastore best suited to its purpose.
  • ◆Every scraped record is content-hashed and checked against a freshness window before syncing, so re-running a scrape against unchanged data doesn't create redundant writes or noisy history.
  • ◆Designed the import pipeline so job/audit bookkeeping commits in its own transaction (REQUIRES_NEW), independent of the master-upsert step — so the audit trail survives even when the write itself fails and rolls back.
  • ◆Separated master business identity (CREATOR → CREATOR_PLATFORM) from insert-only daily snapshots (CREATOR_METRIC, CREATOR_COMMERCE) in a 3NF-oriented schema, so historical trend data is never overwritten.
  • ◆Bridged three identifier spaces (BullMQ job id, analytics SCRAPE_JOB audit row, platform natural business key) explicitly at each hand-off, keeping each layer's identifier scheme simple and fit for its own purpose.