/ Overview
Designed and built a production-oriented data pipeline that scrapes e-commerce and creator data from TikTok Shop and EchoTik, and lands it in a canonical, audit-grade analytics database for downstream reporting. The system spans three independently deployable services — a React admin console, a NestJS/BullMQ scraping engine, and a Spring Boot/Oracle analytics API — chosen deliberately because the three layers have genuinely different concerns: user-facing job control, browser-automation orchestration under anti-bot pressure, and long-term relational data integrity.
/ Problem & Challenge
Scraping competitor intelligence across TikTok Shop and EchoTik required extracting deeply authenticated, rate-limited, anti-bot-protected data without triggering account suspensions or creating noisy duplicate rows in analytical records.
/ Architecture & System Decisions
Separated concerns across 3 independently scalable services: a React admin UI, a NestJS/BullMQ scraper with Redis Lua-scripted locks and circuit breaker, and a Spring Boot/Oracle analytics API with isolated REQUIRES_NEW audit transactions.
/ Engineering Trade-offs
⚖ TRADE-OFF DECISIONBullMQ/Redis vs Database Job Polling: Kept fast, atomic, short-lived execution state in Redis and permanent business data in Oracle, matching each datastore strictly to its operational strengths.
⚖ TRADE-OFF DECISIONIndependent Audit Bookkeeping: Engineered import pipeline so job/audit bookkeeping commits in its own transaction (REQUIRES_NEW) independent of master-upsert steps, guaranteeing audit trail survival even on write failures.
/ Impact
Architected a three-service system along real operational boundaries, surviving hostile anti-bot defenses while producing data clean enough to trust for analytics — with an audit trail that survives failure, not just success.
/ Highlights
- ◆Built a session-coordination layer (Redis-backed) that serializes all EchoTik scraping to a single active session across 5 parallel workers, using atomic Lua-scripted locks, jittered delays, per-session API-call budgets, and a circuit breaker with exponential backoff.
- ◆Chose BullMQ/Redis over a polling database table for job orchestration — keeping fast, atomic, short-lived execution state in Redis and permanent business data in Oracle, each in the datastore best suited to its purpose.
- ◆Every scraped record is content-hashed and checked against a freshness window before syncing, so re-running a scrape against unchanged data doesn't create redundant writes or noisy history.
- ◆Designed the import pipeline so job/audit bookkeeping commits in its own transaction (REQUIRES_NEW), independent of the master-upsert step — so the audit trail survives even when the write itself fails and rolls back.
- ◆Separated master business identity (CREATOR → CREATOR_PLATFORM) from insert-only daily snapshots (CREATOR_METRIC, CREATOR_COMMERCE) in a 3NF-oriented schema, so historical trend data is never overwritten.
- ◆Bridged three identifier spaces (BullMQ job id, analytics SCRAPE_JOB audit row, platform natural business key) explicitly at each hand-off, keeping each layer's identifier scheme simple and fit for its own purpose.