Files
yt-channel-scraper/docs/superpowers/specs/2026-07-27-discovery-only-scrape-design.md
T

48 lines
3.5 KiB
Markdown

# Discovery-Only Scrape Design
## Goal
Add manual discovery controls for one channel or all tracked channels. Discovery checks the complete recent video list returned by YouTube, registers unknown videos as `pending`, and does not download transcripts, Markdown, audio, thumbnails, or avatars.
## Existing Context
The webapp already has a single-worker `JobManager`, SSE job events, and a `POST /api/scrape` endpoint. The current channel job combines discovery with `pipeline.process_video()`, so every pending video is immediately extracted and rendered. The channels table has per-channel actions, while the videos view can be filtered to one channel.
## Architecture
- Extend the existing job dispatch with `opts.mode == "discover"`.
- `channel_id` selects one tracked channel. A missing `channel_id` means all tracked channels for discovery jobs only.
- Discovery uses the existing `discover_channel()` call without a playlist cap. It compares the returned IDs with the SQLite catalog and upserts only catalog metadata. It never calls `process_video()`.
- The existing scrape mode remains unchanged for users who want discovery plus transcript/Markdown processing.
- All discovery jobs stay in the existing sequential queue and use the existing SSE stream, cancellation, progress, and history.
## Persistence and Results
- New `VideoRef` rows are inserted with the existing default status `pending`.
- Existing video rows retain their status and processed data; rediscovery only refreshes title, upload date, and duration through the existing upsert behavior.
- The channel record is refreshed with its name, handle, catalog count, and `last_scraped`; no avatar fetch/cache is performed by discovery-only jobs.
- A per-channel progress event reports `new_videos`, `known_videos`, and the channel ID.
- The terminal event reports total channels, discovered entries, new videos, known videos, and errors.
- A one-channel discovery failure marks the job `error`. An all-channel job continues after individual failures and finishes with a summary so one broken channel does not prevent other channels from being scanned.
## UI
- Channels view: add `Investigar todos` in the header and `Investigar` in each channel row.
- Videos view: add an investigation button in the header. It investigates the selected channel when `filters.channel` is set, otherwise all channels.
- Buttons use the existing job widget/SSE stream, are disabled while an active job is being followed, and show the number of new videos on completion.
- Completion refreshes channels, videos, and dashboard data. Existing `.md`, `Process`, and `Audio` controls remain separate.
## Error Handling
- Empty channel catalogs complete successfully with zero channels scanned.
- Discovery errors are emitted in the job log and do not invoke video processing.
- A cancellation checks the existing cancellation set between channels and leaves already inserted catalog rows intact.
- Duplicate IDs from a discovery response are counted once by the catalog comparison.
## Testing
- Store tests verify that upserting an existing and a new reference reports only the new reference.
- Job tests fake `discover_channel()`, verify new rows are `pending`, existing rows retain their status, and `process_video()` is never called.
- Job tests verify all-channel discovery continues after one channel fails and reports the successful channel.
- The full existing Python test suite remains the regression check. The static UI is verified by code inspection and the existing webapp smoke path; no frontend build step exists.