Scraper de canales de YouTube hacia notas Markdown para base de conocimiento (Obsidian-ready), con plataforma web local. Engine + CLI (Workstream A): - Modular pipeline: discover/extract/parse/chapters/render/store + ratelimit - SQLite store con migración idempotente: FTS5 (transcript search), columnas de metadata enriquecida, tablas cookies_meta y scrape_jobs - Módulos: segments, cookies (Netscape vault), export (json/csv/srt/html), analysis (word freq/timeline/wordcloud), monitor (watch loop), pipeline - CLI Click group: search, export, audio, channels, watch, analyze, re-render - Fix del bug de scoping de cookies en cli.py Webapp local (Workstream B): - FastAPI backend: dashboard, channels, videos facetado, transcript, search FTS, analysis, scrape jobs con SSE, cookies drag-and-drop, exports, folders (abrir en OS), tools (re-render, formato) - SPA no-build (Alpine.js + Tailwind + Chart.js por CDN): 9 vistas, tema dark command-center con acento rojo→rosa, cookie vault drag-drop, consola de scrapeo con progreso live vía SSE Launcher + subagentes (Workstream C): - start-server.bat / stop-server.bat con auto port-scan + browser open - .opencode/agent/webapp-builder.md + .opencode/goals/webapp-build.md Tests: 38 pytest verdes. Sin funcionalidad de IA (enfoque data-mining).
17 KiB
Design: yt-channel-scraper → Local Content-Mining Platform
Date: 2026-07-26 Status: Approved (design review) Stack decision: FastAPI + Alpine.js + Tailwind (CDN), no Node/npm build step Scope rule: Data-mining + content focus. No AI (excluded: LLM summaries, RAG, NER, speaker diarization).
1. Goal
Turn the existing yt-dlp-based CLI scraper into a local "intelligent platform" for scraping and mining content-creator data. Add the high-value CLI features from OPPORTUNITIES.md (minus AI), a visually attractive local webapp with drag-and-drop cookie session login, .bat launchers with automatic port handling, and an opencode subagent + sub-goal that owns the webapp implementation.
Everything runs 100% locally — no remote deploy, no auth, no rate limits beyond YouTube's own.
2. Context (current state)
- Modular pipeline:
cli → discover → extract → parse → chapters → render → store, plusratelimitandconfig. - SQLite (
data/state.db): tableschannels(2 rows),videos(883 rows). Statuses:pending/done/no_subtitles/error. - Rich metadata (views, likes, tags, thumbnail, description) currently lives only in Markdown frontmatter, not in the DB.
- 33 transcripts already rendered as
.mdunderdata/markdown/<channel>/. - A valid Netscape cookie file exists at
cookies/253506e4-387c-4cda-9f11-e95548f8ca0b.txt. - Known bug:
cli.py:228-229—_process_one()referencescookiesandcookies_from_browserthat are local tomain()and therefore out of scope. Cookies are currently non-functional. This design fixes it.
3. Architecture decision
Approach A — Monorepo, in-package webapp (chosen).
The FastAPI app lives inside the existing package at src/yt_scraper/webapp/ and reuses Store + pipeline modules directly. The same FastAPI process serves the SPA static files, the JSON API, and SSE. One repo, one process, zero query duplication.
Rejected alternatives:
- B (separate top-level package importing
yt_scraper): doubles plumbing for no local benefit. - C (separate processes sharing only DB): duplicates query logic, breaks the "one platform" feel.
4. New repo layout
src/yt_scraper/
cli.py +subcommands: search, export, audio, channels, watch, analyze, re-render
store.py +schema migration, +segments/metadata/cookies/jobs queries
segments.py NEW populate transcript_segments + FTS5; backfill from .md
export.py NEW json/csv/srt/html exporters
analysis.py NEW word freq, top-N terms, term-over-time timeline (matplotlib)
monitor.py NEW watch-mode loop (interval discovery + process)
cookies.py NEW Netscape parse/validate/vault; resolve active cookie path
webapp/
__init__.py
app.py FastAPI factory: static mount, CORS off (local), /healthz, SSE mount
api.py REST routers (channels, videos, search, transcript, analysis, scrape, cookies, exports)
jobs.py in-process scrape job runner + SSE event bus (queue/stream/cancel)
static/
index.html SPA shell (Tailwind CDN, Alpine.js, Chart.js CDN)
app.js Alpine app: views, state, API client, SSE consumer
styles.css custom dark "command-center" theme
pipeline glue: extract.py / cli.py now accept the active cookie path
templates/
video.md.j2 unchanged (re-render reuses it)
gallery.html.j2 NEW template for `export --format html`
docs/superpowers/specs/2026-07-26-platform-design.md this file
.opencode/
agent/webapp-builder.md NEW subagent owning Workstream B
goals/webapp-build.md NEW sub-goal the subagent executes
start-server.bat NEW auto port-scan + launch + open browser (repo root)
stop-server.bat NEW graceful shutdown + port cleanup (repo root)
.run/ gitignored runtime: server.info (port/PID)
data/analysis/ gitignored output of `analyze` (PNG/CSV)
pyproject.toml gains optional deps: fastapi, uvicorn[standard], matplotlib, wordcloud, and ffmpeg is a system requirement for audio.
5. Data model (non-destructive, idempotent migration)
Run automatically on every Store instantiation via _init_schema + ALTER TABLE ADD COLUMN guards.
5.1 videos — added columns
view_count INTEGER
like_count INTEGER
tags TEXT -- JSON array
thumbnail TEXT
description TEXT
chapters_json TEXT -- JSON of chapters (for re-render)
segments_json TEXT -- JSON of segments (for re-render, avoids re-download)
5.2 New transcript_segments + FTS5
CREATE TABLE transcript_segments (
video_id TEXT NOT NULL,
idx INTEGER NOT NULL,
start_sec REAL NOT NULL,
end_sec REAL NOT NULL,
text TEXT NOT NULL,
PRIMARY KEY (video_id, idx),
FOREIGN KEY (video_id) REFERENCES videos(video_id)
);
CREATE INDEX idx_segments_video ON transcript_segments(video_id);
-- FTS5 external-content table over transcript_segments(text)
CREATE VIRTUAL TABLE transcript_fts USING fts5(
text, content='transcript_segments', content_rowid='rowid'
);
Populated during scrape (after parse) and by a one-time backfill that parses the 33 existing .md files. Drives search (CLI) and webapp search (snippet + timestamp).
5.3 New cookies_meta
CREATE TABLE cookies_meta (
id TEXT PRIMARY KEY, -- uuid (matches filename stem)
filename TEXT NOT NULL,
label TEXT,
added_at TEXT NOT NULL,
expires_at TEXT, -- earliest cookie expiry, ISO
is_active INTEGER DEFAULT 0, -- exactly one row may be 1
has_session INTEGER DEFAULT 0, -- 1 if SID/SAPISID/LOGIN_INFO present
cookie_count INTEGER DEFAULT 0
);
Cookie files are stored in cookies/. Exactly one cookie may be active globally. cookies.py resolves the active path for yt-dlp's cookiefile.
5.4 New scrape_jobs
CREATE TABLE scrape_jobs (
id TEXT PRIMARY KEY,
channel_id TEXT,
opts_json TEXT,
status TEXT NOT NULL, -- queued/running/done/cancelled/error
started_at TEXT,
finished_at TEXT,
total INTEGER DEFAULT 0,
completed INTEGER DEFAULT 0,
last_error TEXT
);
The scrape console + SSE consumer read/write this table.
5.5 Backfill
One-time segments.backfill_from_markdown() invoked on first run (or via re-render --backfill): parse each done .md, extract segments + chapters + frontmatter metadata, write into transcript_segments/transcript_fts and the new videos columns. Idempotent (skip videos already having segments).
6. Workstream A — Core engine + CLI
6.1 CLI subcommands (new)
| Command | Behavior |
|---|---|
search "query" [-c CHANNEL] [-l N] |
FTS5 MATCH → prints video_id, start timestamp, snippet. |
export --format {json,csv,srt,html} [-c CHANNEL] [-o DIR] |
JSON = full dump; CSV = metadata table; SRT = subtitles from segments; HTML = Jinja2 gallery (gallery.html.j2). |
audio [-c CHANNEL] [--limit N] [--from DATE] |
yt-dlp bestaudio → mp3 via FFmpegExtractAudio. Skips videos already having audio unless --force. Records nothing new in DB beyond existing done. |
re-render [-c CHANNEL] |
Rebuild all .md from segments_json+chapters_json+metadata using the current video.md.j2. Errors if segments missing (hint: run backfill). |
channels add <URL> / channels list / `channels remove <id |
handle>` |
--all-channels (global flag) |
When set, search/export/analyze/re-render operate across every channel instead of one. Mutually exclusive with -c. |
watch --interval 6h [-c CHANNEL] [--once] |
Loop: discover_channel → upsert → process pending → sleep. --once runs a single iteration (useful for Task Scheduler). |
analyze --wordcloud --top-words N --timeline TERM [-c CHANNEL] [-o DIR] |
Saves wordcloud.png, top_words.csv, timeline.png to data/analysis/. Pure frequency mining (regex tokenization, Spanish stopwords). |
6.2 Bug fix
_process_one() and extract_video() already accept cookies_file/cookies_from_browser — the bug is that main() doesn't pass its local cookies/cookies_from_browser into _process_one(). Fix: thread the two values through _process_one()'s signature. Additionally, when neither flag is given, fall back to the active cookie from the vault (cookies.resolve_active_path()), so the webapp's drag-and-drop selection works for CLI too.
6.3 Module responsibilities
segments.py:store_segments(video_id, segments),populate_fts(video_id),search(query, channel_id, limit) -> list[Hit],backfill_from_markdown(store, md_dir).export.py:export_json/csv/srt/html(store, segments_store, channel_id, out_dir).analysis.py:word_frequency(channel_id),top_words(n),term_timeline(term)— returns data structures; matplotlib rendering optional behind arender=Trueflag so the webapp can call the same functions for Chart.js data without PNG.monitor.py:watch_loop(config, interval, once).cookies.py:parse_netscape(path) -> CookieFile,validate(meta) -> {valid, has_session, expires_at},import_upload(text_or_path) -> id,set_active(id),resolve_active_path() -> str|None,list_vault() -> list[CookieMeta].
7. Workstream B — Webapp (owned by the webapp-builder subagent)
7.1 Backend (FastAPI)
app.py: factory mounts/static(SPA), includes API routers under/api, exposes/healthz, and an SSE endpoint/api/scrape/{job_id}/stream.api.pyrouters:GET /api/dashboard→ aggregate stats (per-channel counts, total duration, date range, uploads-over-time buckets, top tags, duration histogram, status breakdown).GET /api/channels,POST /api/channels {url},DELETE /api/channels/{id}.GET /api/videos?channel=&status=&from=&to=&min_dur=&q=&sort=&page=&size=.GET /api/videos/{id}→ full metadata + chapters.GET /api/videos/{id}/transcript→ segments list (for the reader).GET /api/search?q=&channel=&limit=→ FTS hits with snippet + timestamp.GET /api/analysis?channel=&type=wordcloud|topwords|timeline&term=→ JSON for Chart.js.POST /api/scrape {channel_id, opts}→ creates ascrape_jobsrow, enqueues injobs.py, returns{job_id}.GET /api/scrape/{job_id}/stream(SSE) → progress + log lines.DELETE /api/scrape/{job_id}→ cancel.POST /api/cookies/upload(multipart/text) → import;GET /api/cookies;POST /api/cookies/{id}/activate;DELETE /api/cookies/{id};POST /api/cookies/{id}/test.GET /api/export?format=&channel=→ streams the file.
jobs.py: a single-threadedqueue.Queueworker runs scrape jobs sequentially (avoids concurrent yt-dlp hammering). Emits SSE events (progress,log,done,error). Persisted inscrape_jobs.- The active cookie from the vault is passed to every scrape job automatically.
7.2 Frontend (Alpine.js SPA)
Design language: dark "data command center." Near-black canvas (zinc-950), glassy cards (backdrop-blur, subtle borders), red→rose accent gradient (creator-content nod), monospace data accents, Chart.js viz, smooth Alpine x-transitions. Responsive (works on the laptop you launch it on). No auth.
Views (hash routes):
- Dashboard — channel cards + global charts (uploads-over-time line, top-tags bar, duration histogram, status donut).
- Channels — list / add (paste URL) / remove / drill-in to a channel detail.
- Videos — faceted table/grid: filter by channel, status, date range, min duration, free-text; thumbnail grid or list; sort; click → detail.
- Video detail — metadata header, chapters sidebar, full transcript reader with per-segment timestamps, in-page term highlight, "open on YouTube at t", per-video export (srt/json), audio-download trigger.
- Search — FTS across all transcripts → result cards with highlighted snippet + timestamp deep-link into the video detail at that segment.
- Analysis — word cloud (canvas), top-N terms bar, term-over-time timeline.
- Scrape console — launch a job (pick channel + options: limit, since, languages, shorts, which cookie), live progress via SSE, streaming log, cancel button, watch-mode toggle.
- Cookie vault — drag-and-drop zone (see §8), vault list, test-session.
- Exports — pick format + scope → download.
7.3 "Intelligent" without AI
Smart faceted filters, cross-channel term mining, term-over-time timelines, saved searches (localStorage), live monitoring + auto-processing, auto-suggestions from top terms. Feels like a command center for content-creator data mining.
8. Cookie drag-and-drop UX
- Drop zone accepts multiple
.txtfiles or pasted Netscape cookie text. - On upload (
POST /api/cookies/upload):cookies.parse_netscape()validates the format, extracts the earliest expiry, detects session cookies (SID,SAPISID,__Secure-3PSID,LOGIN_INFO), and returns a card:{label, expires_at, has_session, cookie_count, size}. - Vault list with radio = active (exactly one active globally), rename, delete.
- The active cookie path is passed to yt-dlp as
cookiefilefor every scrape (members-only, age-gate, etc.). - "Test session" → lightweight yt-dlp call (fetch the logged-in user's identity / a known members-only-free check) → reports
OK / expired / weak. - The existing
cookies/253506e4-...txtis auto-imported into the vault on first server start.
9. Launcher .bat (automatic port system)
PowerShell-assisted .bat files at the repo root.
start-server.bat- Resolve project root (script dir).
- Find a free TCP port: scan from
8000upward untilbindsucceeds (PowerShellTcpListenerprobe). - Write
port,PID, and ISO timestamp to.run/server.info. - Launch
python -m uvicorn yt_scraper.webapp.app:app --port <port>viastart(minimized). - Poll
http://localhost:<port>/healthzuntil 200 (or timeout ~30s). - Open
http://localhost:<port>in the default browser. - Echo a clear success banner (port, URL) or an error.
- Idempotent: if
.run/server.infoalready describes a live server, just re-open the browser.
stop-server.bat- Read
.run/server.info. - Graceful terminate the uvicorn process (taskkill by PID, then
/Tfor children). - Confirm the port is freed (re-probe).
- Remove
.run/server.info. - Echo a clear banner.
- Read
.run/ is added to .gitignore.
10. Subagent + sub-goal
.opencode/agent/webapp-builder.md— an opencode agent specialized for this webapp. It knows: the FastAPI + Alpine + Tailwind (CDN, no-build) stack; theStoreAPI surface; the design language (§7.2); and its verification steps (start server,curl /healthz, smoke-test each SPA view, runpytest). It receives Workstream B scope and is expected to reuseStore,segments,analysis,cookiesrather than duplicating logic..opencode/goals/webapp-build.md— the sub-goal: "Implement the local webapp (Workstream B) perdocs/superpowers/specs/2026-07-26-platform-design.md. Build backend routers, SSE job runner, and the 9-view Alpine SPA. Verify with/healthz, smoke tests, andpytest."
11. Dependencies
Add to pyproject.toml [project.optional-dependencies]:
web = ["fastapi>=0.110", "uvicorn[standard]>=0.27"]
analysis = ["matplotlib>=3.8", "wordcloud>=1.9"]
System requirement for audio: ffmpeg (winget install ffmpeg or choco install ffmpeg). The CLI will detect ffmpeg and print a clear hint if missing.
cookies/ is added to .gitignore (real session secrets must never be committed).
12. Build order
A → B → C:
- A1 Schema migration + backfill + fix cookies bug.
- A2
segments/export/analysis/monitor/cookiesmodules. - A3 CLI subcommands (
search,export,audio,channels,watch,analyze,re-render). - A4 Tests;
pytestgreen. - B Webapp (subagent) — backend routers, job runner, SPA, drag-drop cookies. Verify
/healthz+ smoke. - C
.batlauncher +.opencodeagent/goal files.
13. Testing & verification
pytest tests/ -vcovering: schema migration idempotency, FTS search, segment backfill from.md, cookie Netscape parse/validate, exporters (json/csv/srt/html), analysis frequency/timeline, multi-channel store ops.- Webapp:
/healthz200; representative endpoint per router; SPA loads; SSE streams a fake job. - Manual: run
start-server.bat, confirm browser opens, drag-drop the real cookie, trigger a small scrape (--limit 1) against the existing Nostal Vlad channel, verify a new.md+ FTS row appear and search returns the new content. python -m yt_scraper.cli search "..."returns hits after backfill.
14. Out of scope (explicitly excluded)
- AI features: LLM summaries, Q&A/RAG, NER (spaCy), speaker diarization (pyannote).
- CLI
stats,clip,thumbnails,translation(dropped per user scope decision — aggregate queries still exist internally to feed the dashboard). - Remote deployment, auth, multi-user, cloud sync.