Files
yt-channel-scraper/docs/superpowers/specs/2026-07-27-discovery-only-scrape-design.md

3.5 KiB

Discovery-Only Scrape Design

Goal

Add manual discovery controls for one channel or all tracked channels. Discovery checks the complete recent video list returned by YouTube, registers unknown videos as pending, and does not download transcripts, Markdown, audio, thumbnails, or avatars.

Existing Context

The webapp already has a single-worker JobManager, SSE job events, and a POST /api/scrape endpoint. The current channel job combines discovery with pipeline.process_video(), so every pending video is immediately extracted and rendered. The channels table has per-channel actions, while the videos view can be filtered to one channel.

Architecture

  • Extend the existing job dispatch with opts.mode == "discover".
  • channel_id selects one tracked channel. A missing channel_id means all tracked channels for discovery jobs only.
  • Discovery uses the existing discover_channel() call without a playlist cap. It compares the returned IDs with the SQLite catalog and upserts only catalog metadata. It never calls process_video().
  • The existing scrape mode remains unchanged for users who want discovery plus transcript/Markdown processing.
  • All discovery jobs stay in the existing sequential queue and use the existing SSE stream, cancellation, progress, and history.

Persistence and Results

  • New VideoRef rows are inserted with the existing default status pending.
  • Existing video rows retain their status and processed data; rediscovery only refreshes title, upload date, and duration through the existing upsert behavior.
  • The channel record is refreshed with its name, handle, catalog count, and last_scraped; no avatar fetch/cache is performed by discovery-only jobs.
  • A per-channel progress event reports new_videos, known_videos, and the channel ID.
  • The terminal event reports total channels, discovered entries, new videos, known videos, and errors.
  • A one-channel discovery failure marks the job error. An all-channel job continues after individual failures and finishes with a summary so one broken channel does not prevent other channels from being scanned.

UI

  • Channels view: add Investigar todos in the header and Investigar in each channel row.
  • Videos view: add an investigation button in the header. It investigates the selected channel when filters.channel is set, otherwise all channels.
  • Buttons use the existing job widget/SSE stream, are disabled while an active job is being followed, and show the number of new videos on completion.
  • Completion refreshes channels, videos, and dashboard data. Existing .md, Process, and Audio controls remain separate.

Error Handling

  • Empty channel catalogs complete successfully with zero channels scanned.
  • Discovery errors are emitted in the job log and do not invoke video processing.
  • A cancellation checks the existing cancellation set between channels and leaves already inserted catalog rows intact.
  • Duplicate IDs from a discovery response are counted once by the catalog comparison.

Testing

  • Store tests verify that upserting an existing and a new reference reports only the new reference.
  • Job tests fake discover_channel(), verify new rows are pending, existing rows retain their status, and process_video() is never called.
  • Job tests verify all-channel discovery continues after one channel fails and reports the successful channel.
  • The full existing Python test suite remains the regression check. The static UI is verified by code inspection and the existing webapp smoke path; no frontend build step exists.