Extraccion autenticada:
- import_from_browser (Brave) con fallback CDP headless para cookies app-bound v20
- extract_via_watch_page: GET plano + ytInitialPlayerResponse cuando yt-dlp falla
con sesion logueada (members-only); regex y opener cacheados
- js_runtimes (node/deno/bun/quickjs) propagado a todos los ydl_opts
- rutas de Brave multiplataforma (Windows/macOS/Linux)
Webapp UX: chips de filtros removibles, skeleton loaders, estado de vista en URL,
memoria de scroll, copyMd/openMd, import de cookies desde navegador, no-cache de statics
Rendimiento:
- entorno Jinja2 cacheado por directorio de plantilla (antes 1 por nota)
- _rank_unranked con guarda (antes full-scan en cada arranque/import)
- upsert_videos con executemany; dashboard sin N+1 (GROUP BY + conteo de tags en SQL)
- thumbnails en paralelo (6 hilos, CDN ytimg); handlers bloqueantes -> def (threadpool)
- reconcile de arranque en hilo daemon: uvicorn arriba al instante (0.95s con 1503 md),
healthz expone reconcile_done
- Store.transaction(): escrituras por video agrupadas (~6 commits -> 3)
Refactor: helpers unicos (extract_handle->discover, safe_dirname/filename->render,
order_pending->store, keep_ref->config, seconds_to_ts solo en segments);
re-render del CLI delega en pipeline.re_render_videos (retira huerfanos y marca done);
fuera wrappers muertos de segments.py
- Video detail (done videos): added 'Download audio' button; removed the
manual 'Thumbnail' button (thumbnails auto-download now, so it was redundant).
- New 'GET /api/videos/{id}/audio' endpoint (also serves HEAD for probing),
serving data/audio/<video_id>.mp3. Audio job outtmpl switched to %(id)s so
files are addressable per video.
- Integrated <audio> player appears in the detail view once the MP3 is present
(probed via HEAD on openVideo; polled after a download job until ready).
- Synced transcript follower: as audio plays, the matching segment is
highlighted (seg-active) and auto-scrolled into view like a karaoke/lyrics
tracker. Clicking any segment timestamp or chapter seeks the audio to that
point (falls back to scroll when no audio). Playback indicator pulses while
playing.
- Bulk 'Download .md' now also caches thumbnails (one action does both);
removed the redundant standalone Thumbnails button from the bulk action bar.
- Scrape UX optimization: batch job skips videos whose .md already exists
(just caches their thumbnail) instead of re-extracting.
- Auto-thumbnails: on addChannel, thumbnails cache in the background so a
newly-attached channel's catalog looks populated immediately, even before
any .md is generated.
- Thumbnail fallback: pending videos without a stored thumbnail URL use the
canonical https://i.ytimg.com/vi/<id>/hqdefault.jpg, so every video shows
a thumbnail (and /api/thumbnails/{id} never 404s — redirects if uncached).
- New helpers in pipeline.py: thumbnail_url_for() + cache_thumbnail().
Scraper de canales de YouTube hacia notas Markdown para base de conocimiento
(Obsidian-ready), con plataforma web local.
Engine + CLI (Workstream A):
- Modular pipeline: discover/extract/parse/chapters/render/store + ratelimit
- SQLite store con migración idempotente: FTS5 (transcript search), columnas
de metadata enriquecida, tablas cookies_meta y scrape_jobs
- Módulos: segments, cookies (Netscape vault), export (json/csv/srt/html),
analysis (word freq/timeline/wordcloud), monitor (watch loop), pipeline
- CLI Click group: search, export, audio, channels, watch, analyze, re-render
- Fix del bug de scoping de cookies en cli.py
Webapp local (Workstream B):
- FastAPI backend: dashboard, channels, videos facetado, transcript, search
FTS, analysis, scrape jobs con SSE, cookies drag-and-drop, exports,
folders (abrir en OS), tools (re-render, formato)
- SPA no-build (Alpine.js + Tailwind + Chart.js por CDN): 9 vistas, tema
dark command-center con acento rojo→rosa, cookie vault drag-drop,
consola de scrapeo con progreso live vía SSE
Launcher + subagentes (Workstream C):
- start-server.bat / stop-server.bat con auto port-scan + browser open
- .opencode/agent/webapp-builder.md + .opencode/goals/webapp-build.md
Tests: 38 pytest verdes. Sin funcionalidad de IA (enfoque data-mining).