wip: estado de trabajo pendiente antes de la vista grid (suite 230 verde)
This commit is contained in:
@@ -0,0 +1,7 @@
|
|||||||
|
{
|
||||||
|
"permissions": {
|
||||||
|
"allow": [
|
||||||
|
"PowerShell(python \"C:\\\\Users\\\\URIELJ~1\\\\AppData\\\\Local\\\\Temp\\\\claude\\\\d--yt-channel-scraper\\\\22447059-a4b2-482a-92ce-78bd6f3f9010\\\\scratchpad\\\\probe_failing.py\")"
|
||||||
|
]
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,276 @@
|
|||||||
|
# CLAUDE.md
|
||||||
|
|
||||||
|
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
|
||||||
|
|
||||||
|
## Qué es esto
|
||||||
|
|
||||||
|
Plataforma **100% local** de minería de contenido de creadores de YouTube: `yt-dlp` → transcripciones + metadatos → SQLite (con FTS5) → notas Markdown estilo Obsidian, más una webapp FastAPI/Alpine que expone todo eso. Sin auth, sin deploy remoto, sin features de IA (regla de alcance explícita en `docs/superpowers/specs/2026-07-26-platform-design.md` §14).
|
||||||
|
|
||||||
|
## Comandos
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pip install -e ".[dev,web,analysis]" # Python >= 3.10; ffmpeg es requisito del sistema para audio
|
||||||
|
|
||||||
|
python -m pytest tests/ -q # suite completa (61 tests, sin red)
|
||||||
|
python -m pytest tests/test_store_platform.py -v # un archivo
|
||||||
|
python -m pytest tests/test_store_platform.py::test_dashboard_aggregates -v # un test
|
||||||
|
python -m pytest -m "not integration" # marker declarado en pyproject (aún sin uso)
|
||||||
|
|
||||||
|
yt-scraper --help # CLI (equivale a: python -m yt_scraper.cli)
|
||||||
|
yt-scraper scrape --dry-run --limit 10 # discovery sin descargar
|
||||||
|
yt-scraper search "vaporwave" # FTS5 sobre transcripciones
|
||||||
|
yt-scraper re-render --backfill # reimportar .md a la DB y regenerar Markdown
|
||||||
|
|
||||||
|
start-server.bat # webapp: busca puerto libre 8000-8100, arranca uvicorn, abre browser
|
||||||
|
stop-server.bat # mata el proceso registrado en .run/server.info y libera el puerto
|
||||||
|
python -m uvicorn yt_scraper.webapp.app:app --port 8765 # arranque manual (debug)
|
||||||
|
```
|
||||||
|
|
||||||
|
No hay linter ni formatter configurado. Los `.bat` de la raíz solo envuelven `scripts/start-server.ps1` / `scripts/stop-server.ps1`.
|
||||||
|
|
||||||
|
**Ojo:** `yt-scraper` sin subcomando ejecuta un scrape completo del canal de `config.yaml` (`invoke_without_command=True` en [cli.py:39](src/yt_scraper/cli.py#L39)).
|
||||||
|
|
||||||
|
## Orientación antes de leer código
|
||||||
|
|
||||||
|
Existe un grafo de conocimiento en [graphify-out/](graphify-out/) y un hook que exige usarlo antes de leer/grepear fuentes: `graphify query "<pregunta>"`, `graphify explain "<concepto>"`, `graphify path "<A>" "<B>"`. Si el grafo está desactualizado respecto a los archivos, `graphify update`. [graphify-out/GRAPH_REPORT.md](graphify-out/GRAPH_REPORT.md) resume comunidades y god nodes (`Store` es el más conectado con diferencia).
|
||||||
|
|
||||||
|
## Arquitectura
|
||||||
|
|
||||||
|
### Tres superficies, un solo pipeline
|
||||||
|
|
||||||
|
`cli.py` (Click+Rich), `webapp/` (FastAPI+SSE) y `monitor.py` (watch loop) son **fachadas**; las tres convergen en `pipeline.process_video(row, cfg, store, ...)`, la unidad de trabajo por vídeo: extract → parse → align chapters → persistir en DB → renderizar `.md` → `mark_done`. Devuelve `"done" | "no_subtitles" | "error"` y nunca lanza por fallo de un vídeo.
|
||||||
|
|
||||||
|
Consecuencia práctica: una feature nueva se implementa en `pipeline`/`store` y **se expone dos veces** (subcomando en `cli.py` + endpoint en `webapp/api.py`). No dupliques lógica en la capa web.
|
||||||
|
|
||||||
|
```
|
||||||
|
discover.py yt-dlp --flat-playlist → VideoRef[] (+ avatar; deep_channel_avatar como fallback)
|
||||||
|
extract.py yt-dlp extract_info → info dict + pick_subtitle (prioridad json3 > srv1 > srv3 > vtt > ttml)
|
||||||
|
parse.py JSON3/VTT → Segment{start,end,text}, con merge_adjacent
|
||||||
|
chapters.py align_chapters: capítulos ↔ segmentos → Section[]
|
||||||
|
render.py Jinja2 → .md (filtros format_timestamp / quote_yaml / to_json)
|
||||||
|
store.py SQLite: todo el SQL a mano, filas como dataclasses
|
||||||
|
segments.py camino inverso: .md → DB (backfill)
|
||||||
|
```
|
||||||
|
|
||||||
|
### `Store` es la única capa de datos
|
||||||
|
|
||||||
|
SQLite sin ORM, `sqlite3.Row`, WAL, `foreign_keys=ON`. Cada método abre y cierra su conexión vía el contextmanager `_cursor()` (commit al salir), así que es seguro desde el worker thread de jobs.
|
||||||
|
|
||||||
|
**Migraciones:** `_init_schema()` corre en **cada** instanciación de `Store`. Es `SCHEMA` base + `ALTER TABLE ADD COLUMN` idempotente guiado por los dicts `_VIDEO_COLUMNS` / `_CHANNEL_COLUMNS` + `_EXTRA_SCHEMA`. Para añadir una columna: agrégala al dict, no escribas script de migración.
|
||||||
|
|
||||||
|
**FTS5:** `transcript_fts` es una tabla **standalone** (columnas duplicadas como `UNINDEXED`), no external-content — pese a lo que dice el spec. Por eso `store_segments()` borra e inserta en `transcript_segments` **y** en `transcript_fts`; si escribes segmentos por otra vía, replica ambas. Las queries de usuario pasan por `_sanitize_fts()` (tokens entre comillas unidos con `AND`).
|
||||||
|
|
||||||
|
**Doble persistencia de transcripciones**, intencional: filas en `transcript_segments` (búsqueda, análisis, reader) *y* `segments_json`/`chapters_json` en `videos` (permite `re-render` sin volver a descargar). `pipeline.process_video()` escribe las dos.
|
||||||
|
|
||||||
|
### El Markdown es fuente de datos, no solo salida
|
||||||
|
|
||||||
|
`segments.backfill_from_markdown()` parsea `data/markdown/**/*.md` de vuelta a la DB (idempotente, salta vídeos que ya tienen segmentos). Se invoca al arrancar la webapp ([webapp/app.py:31](src/yt_scraper/webapp/app.py#L31)) y con `re-render --backfill`.
|
||||||
|
|
||||||
|
Eso convierte el formato del `.md` en un **contrato bidireccional**: [templates/video.md.j2](templates/video.md.j2) escribe `**MM:SS** · texto` y `### Título (MM:SS)`; los regex `_SEG_LINE` / `_CHAPTER` / `_FRONTMATTER_KEY` de [segments.py:17](src/yt_scraper/segments.py#L17) los leen. Si tocas la plantilla, actualiza los regex en el mismo cambio o el backfill se rompe en silencio.
|
||||||
|
|
||||||
|
### Sincronización incremental (no recorras el canal entero)
|
||||||
|
|
||||||
|
Re-escanear un canal ya trackeado **no** pagina todo el canal. `discover.discover_incremental()` lee la pestaña `/videos` (cronológica inversa) con `playlistend` y corta en cuanto ve `sync.overlap` vídeos consecutivos que ya están en la DB. Si toda la ventana resulta nueva, la duplica (`window` → `max_window`) en vez de perderse subidas. Medido sobre los 4 canales reales: **21.3 s / 1661 entradas → 3.3 s / 120 entradas**, mismos vídeos nuevos detectados.
|
||||||
|
|
||||||
|
Detalle que condiciona el diseño: en modo flat, yt-dlp **no** devuelve `upload_date` ni `timestamp` para los entries de YouTube (verificado). Por eso el corte real se hace por **solapamiento de IDs**, no por fecha; `since=` existe como condición secundaria por si el extractor sí las trae. La fecha del último vídeo se guarda como marca de agua en `channels.last_video_date` (poblada desde `videos.upload_date`, que sí se conoce tras la extracción) y es lo que se muestra en la UI y en el CLI.
|
||||||
|
|
||||||
|
- El `keep=` que se pasa a `discover_incremental` **debe** replicar el filtro shorts/live bajo el que se pobló la DB (`jobs._keep_ref`). Si no, la cola de la ventana se llena de entradas que nunca podrán ser "conocidas" y la ventana crece sin motivo.
|
||||||
|
- `upsert_videos` usa `COALESCE` en `upload_date`/`duration`: el discovery flat manda `None` y sin eso cada sync borraría las fechas aprendidas en la extracción — justo la marca de agua.
|
||||||
|
- `video_count` se recalcula con `store.mark_channel_synced()`, nunca con `len(refs)`: la ventana son 30 y el canal puede tener 861.
|
||||||
|
- Sin historial local (canal nuevo) se recorre el canal completo, que es lo correcto en la primera pasada.
|
||||||
|
- Escotilla de escape: `--full` en el CLI, `opts["full"]` en los jobs, botones **Full rescan** / **Full** y el checkbox *Full channel rescan* en la webapp. Config en `sync:` de `config.yaml`.
|
||||||
|
- `/tools/sync-channels` y `/tools/avatars` solo quieren metadatos+avatar del objeto canal, así que llaman a `discover_channel(..., limit=1)`: una petición en vez de paginar el canal.
|
||||||
|
|
||||||
|
### El orden de la lista es cronológico, y la fecha casi siempre es inferida
|
||||||
|
|
||||||
|
La sección de vídeos ordena por defecto como YouTube: subida más reciente primero, **haya o no `.md`**. El problema es que la mayoría de las filas no tienen fecha — medido sobre la biblioteca real, **4541 de 4959** — porque el discovery flat no la trae y solo aparece con la extracción.
|
||||||
|
|
||||||
|
`videos.channel_seq` es la posición del vídeo en la pestaña `/videos` (mayor = más nuevo) y es lo único que sitúa a un vídeo sin extraer. `store.SORT_DATE_SQL` deriva de ahí la fecha de orden, en cascada: (1) su `upload_date`; (2) la del vídeo con fecha **inmediatamente superior** en su canal — no es anterior a esa, y el rank desempata; (3) la fecha más nueva conocida en el canal, para la racha que está por encima de todo lo datado; (4) `NO_DATE_SENTINEL` (`"00000000"`) cuando el canal entero está sin extraer. El resultado se expone como `sort_date` + `date_estimated`, y la UI lo pinta con `~`. El sentinel no sale nunca al cliente: es un rank, no una fecha.
|
||||||
|
|
||||||
|
- **El fallback anterior era `discovered_at`**, que fechaba "hoy" todo lo no scrapeado y clavaba el backlog entero encima de los vídeos realmente recientes. Es la razón de ser de todo esto.
|
||||||
|
- El paso (3) también es deliberado: con `'99999999'` un canal sin un solo vídeo extraído (Víctor Pérez, 212 vídeos) se adueñaba de la página 1 sobre canales con fechas reales. Un vídeo sin evidencia no adelanta a uno con evidencia; dentro de su canal el rank lo sigue ordenando bien.
|
||||||
|
- `upsert_videos` asigna `channel_seq` **por encima del máximo actual del canal**, no desde cero: un sync solo ve la ventana más nueva, y todo lo que no trajo es más viejo por construcción. Así ambas mitades quedan ordenadas sin repaginar el canal, y sincronizar repetidamente no reordena nada.
|
||||||
|
- Los dos lookups de `SORT_DATE_SQL` corren una vez por fila sin fecha, así que van sobre **índices parciales** (`idx_videos_dated_seq`, `idx_videos_dated`) que solo indexan las filas datadas. Sin el `WHERE` parcial, buscar "el vídeo datado que tengo encima" recorre todas las filas intermedias mirando la tabla: 2500 filas dentro de un canal cuyos 8 vídeos datados están arriba del todo es cuadrático. Medido: **576 ms → 22 ms por página**. El `WHERE` de la query tiene que estar escrito igual que el del índice o SQLite lo descarta en silencio; hay un test sobre `EXPLAIN QUERY PLAN` que lo fija.
|
||||||
|
- El orden lleva desempate total (`channel_seq`, `video_id`). Sin él, `LIMIT/OFFSET` repite o se salta filas entre páginas.
|
||||||
|
- `_order_clause` acepta las dos ortografías de cada clave. La webapp mandaba `view_count`/`like_count`/`duration` y el store solo conocía `views_desc`/`duration_desc`: **todo** sort que no fuera fecha caía en un `upload_date DESC` crudo que ignoraba la inferencia.
|
||||||
|
- Bases anteriores a la columna se siembran una sola vez en `_seed_channel_seq()`, rankeando por `discovered_at DESC, rowid ASC` — el orden en que se aprendieron: un batch de discovery comparte timestamp y se inserta de nuevo a viejo, y un sync incremental solo añade ids más nuevos que todo lo guardado.
|
||||||
|
|
||||||
|
**Verificado contra la lista real de YouTube** (`discover_channel(limit=N)` sobre los 7 canales, 1 petición por ventana), no contra la intuición:
|
||||||
|
|
||||||
|
| Canal | Filas comparadas | Pares invertidos |
|
||||||
|
|---|---|---|
|
||||||
|
| Alex Hormozi | 28 | 0 / 378 |
|
||||||
|
| Benjamín Cordero | 29 | 0 / 406 |
|
||||||
|
| Fazt | 26 | 0 / 325 |
|
||||||
|
| Nostal Vlad | 30 | 0 / 435 |
|
||||||
|
| Platzi | 29 | 0 / 406 |
|
||||||
|
| Víctor Pérez | **212 (canal completo, 0 fechas reales)** | **0 / 22 366** |
|
||||||
|
| Código Espinoza | 30 | 9 / 435 → **0 tras un sync** |
|
||||||
|
| Código Espinoza (fondo) | 200 | 0 / 19 900 |
|
||||||
|
|
||||||
|
- Víctor Pérez es la prueba fuerte del seed: 212 vídeos **sin una sola fecha real**, orden derivado enteramente del rank reconstruido, y **cada fila en el índice exacto** de YouTube.
|
||||||
|
- El único desvío (Espinoza, 9 pares) venía de un seed erróneo en la ventana más nueva, no del mecanismo: los rangos ahí no se habían observado. **Un pase normal de discovery (1 petición, 30 entradas) lo dejó en coincidencia exacta**, porque `upsert_videos` re-rankea toda la ventana, no solo lo nuevo. Para el fondo de un canal hace falta un **Full rescan**.
|
||||||
|
- Ese caso también decide el diseño: con rank puro (`channel_seq DESC` solo) Espinoza salía **peor**. La fecha real va primero justo para que lo que costó una extracción corrija un rank reconstruido, y el rank solo desempata cuando la fecha —que es solo día, sin hora— no puede.
|
||||||
|
- Scripts de la medición: comparan `query_videos` contra `discover_channel` y cuentan pares invertidos. Repetible con ~1 petición por canal.
|
||||||
|
|
||||||
|
**Lo que la coincidencia exacta NO demuestra.** Una auditoría posterior sobre la base viva encontró que **`channel_seq` no contiene ni una sola posición observada de YouTube**: el 100 % es la reconstrucción de `_rank_unranked`, y los rangos son perfectamente contiguos por canal (`MIN=1`, `MAX=n`, sin huecos) — la firma del seed, no de un sync, que deja huecos al solapar ventanas. Que acierte es una propiedad del historial de discovery, no del diseño. Y **Nostal Vlad la rompe**: sus 3 vídeos más nuevos tienen `channel_seq` **1, 2, 3** — el fondo del canal — porque un lote posterior trajo vídeos más viejos, justo la premisa que el seed asume falsa. Se muestra bien **solo porque esos 3 tienen fecha real** y la fecha manda. Sin fecha se hundirían. Un discovery lo re-observa.
|
||||||
|
|
||||||
|
#### Tres bugs del orden que la prueba de campo no cubría
|
||||||
|
|
||||||
|
Los tres se encontraron auditando la base viva, no razonando:
|
||||||
|
|
||||||
|
- **`sort=oldest` abría con lo que no tiene fecha.** `NO_DATE_SENTINEL` (`"00000000"`) es el valor más bajo, así que lo que lo mantiene fuera de la portada bajo el orden por defecto es exactamente lo que lo ponía **primero** en ascendente: 212 vídeos de un canal sin extraer por delante de una subida real de 2017. `_OLDEST_FIRST` lleva ahora `(sort_date = '00000000')` como primer término. "No lo sabemos" no es "el principio de los tiempos".
|
||||||
|
- **El desempate global rankeaba por tamaño de catálogo.** `channel_seq` cuenta hasta el número de vídeos del canal, así que compararlo **entre** canales ordena por quién tiene más. Con el **92 %** de los pares adyacentes empatados en `sort_date`, eso decidía casi toda la lista: un canal entero delante de otro solo porque 575 > 476. El desempate es ahora `channel_id` **y luego** `channel_seq`, de modo que cada bloque empatado queda contiguo por canal y el orden interno —lo que tiene que cuadrar con YouTube— no se toca. Verificado: 373 bloques de empate, **0** con un canal partido.
|
||||||
|
- **Una fila sin rank no queda desordenada, queda mal colocada.** `COALESCE(channel_seq, -1)` hace que la regla 2 herede el vídeo datado **más viejo** del canal, así que un vídeo que discovery acaba de encontrar —de los más nuevos— se muestra el último. Medido: dos subidas nuevas de Hormozi con `sort_date=20180720`, penúltima y última de 513. Pasa cuando un server de larga vida sigue con el código previo a la columna después de migrar. `_rank_unranked()` corre **siempre**, no solo al añadir la columna, y es no-op si no hay NULLs.
|
||||||
|
|
||||||
|
**Ojo al importar `yt_scraper.webapp.app`:** tiene `app = create_app()` a nivel de módulo ([app.py:102](src/yt_scraper/webapp/app.py#L102)), así que **el simple `import` abre la base real del proyecto** vía `config.yaml`, corre la migración, `auto_import_dir` y `reconcile_markdown` — aunque después le pases un `Config` distinto a `create_app()`. Para tocar solo una copia, importa `yt_scraper.store` / `yt_scraper.discover` directamente y nunca `webapp.app`.
|
||||||
|
|
||||||
|
### El botón de `.md` descarga `.md`
|
||||||
|
|
||||||
|
`_run_batch` ya no cachea miniaturas. Las traen los caminos de canal (alta de canal, herramienta *Download thumbnails*) y `/api/thumbnails/{id}` redirige al CDN lo que no esté en disco, así que colgarlas del batch gastaba una petición por vídeo por una imagen que la UI ya podía mostrar. Un vídeo cuyo `.md` ya existe en disco ahora se salta entero.
|
||||||
|
|
||||||
|
### Estados terminales y frescura (la UI tiene que reflejar DB + disco)
|
||||||
|
|
||||||
|
`no_subtitles` **no** significa "este vídeo no tiene subtítulos". Se escribe siempre que `data.segments` viene vacío ([pipeline.py:52](src/yt_scraper/pipeline.py#L52)), lo que mezcla tres causas muy distintas: el vídeo no tiene pistas, la política de idiomas rechazó las que sí tiene, o la descarga vino vacía por throttling. `extract.describe_missing_subtitle()` distingue los casos y el motivo se guarda en `videos.error_msg` vía `mark_status(vid, status, reason)`.
|
||||||
|
|
||||||
|
Incidente que motivó esto: 511 vídeos de un canal quedaron en `no_subtitles` porque `config.example.yaml` ponía los idiomas en modo `manual` y el canal solo publica subtítulos automáticos. `_sources_for("manual")` no hace fallback. El default es ahora `any` (manual primero, auto después) — `manual` es opt-in explícito.
|
||||||
|
|
||||||
|
- **Ambos estados son reintentables.** `Store.RETRYABLE_STATUSES` = `("error", "no_subtitles")`. `reset_videos()` los devuelve a `pending`; `reset_errors()` conserva su significado estrecho. `done` nunca se toca. Expuesto en `POST /api/videos/reset`, `GET /api/videos-retryable` y `yt-scraper reset`.
|
||||||
|
- **Contenido bloqueado se detecta en el discovery, no al fallar.** Los entries planos de yt-dlp traen `availability`; `subscriber_only` = vídeo de membresía. Medido: 120 entradas en una petición, 4 marcadas, exactamente las mismas que la DB había aprendido a base de extracciones fallidas. Se guarda en `videos.availability` y `VideoRow.block_reason` lo traduce, con fallback al `error_msg` para las filas antiguas. `get_pending()` las excluye (los `_run_batch` por ID explícito sí las intentan, que es lo que hace recuperable el caso "compré la membresía"). Filtro `blocked` en `query_videos` / `GET /api/videos?blocked=true`, badge `.st-locked` en la UI.
|
||||||
|
- El enum completo es `private | premium_only | subscriber_only | needs_auth | unlisted | public` ([yt_dlp/extractor/common.py:414](https://github.com/yt-dlp/yt-dlp)). Solo los cuatro primeros bloquean: **`unlisted` se descarga sin problema** y marcarlo escondería vídeos que sí se pueden tener. En el SQL del filtro, `COALESCE` es imprescindible: con `availability` NULL, `NOT (NULL OR ...)` es NULL y la rama negada devolvería cero filas.
|
||||||
|
- **Salvo los permanentes.** `Store.PERMANENT_ERROR_PATTERNS` (members-only, private, removed) se excluyen del reset: reintentarlos no los va a desbloquear y gasta peticiones que necesitan los que sí pueden salir. `retryable_counts()` los devuelve aparte en la clave `permanent`. Mantén la lista estrecha: el mensaje de throttling de YouTube ("rate-limited … try again later") **sí** es reintentable y no debe caer ahí.
|
||||||
|
- **`backfill_from_markdown` no arregla estados**: puebla segmentos y metadatos, pero nunca `status` ni `markdown_path`. Para eso está `segments.reconcile_markdown()`, que corre al arrancar la webapp, en `POST /api/tools/reconcile` y en `yt-scraper reconcile`.
|
||||||
|
- `reconcile` es aditivo por defecto. `prune=True` es opt-in porque es destructivo: borra duplicados obsoletos y degrada `done` cuyo `.md` desapareció. Se niega a degradar nada si el árbol de markdown está vacío (root mal configurado).
|
||||||
|
- **La identidad de una nota es el `video_id`, nunca el nombre del fichero.** El nombre sale de `{upload_date}_{slug}` y las dos partes son inestables: YouTube sirve los títulos localizados (el mismo vídeo volvió como "La controversia de Claude Fable 5" en una pasada y "The Claude Fable controversy 5" en la siguiente) y los creadores renombran. Cada cambio de stem escribía un fichero nuevo y dejaba el anterior huérfano.
|
||||||
|
- Las dos rutas que renderizan pasan ahora por `pipeline._render_and_retire()`, que borra el fichero al que apuntaba `markdown_path` si el stem cambió. Existe **porque ya divergieron una vez**: `re_render_videos` usaba la fecha cruda (`20240519_`) y `process_video` la normalizada (`2024-05-19_`) → 94 `.md` para 61 filas. Medido tras arreglarlo: 33 duplicados reales en disco, todos con canónico existente, mezclando las dos causas (`20240519_why-did-the-2000s-look-like-this` contra `2024-05-19_por-que-los-2000-se-veian-asi`). Cero notas únicas perdidas.
|
||||||
|
- `build_filename_stem` acepta `{video_id}` en `filename_template` para quien quiera que el fichero se identifique solo, sin la DB. Es opt-in: el default no cambia, para no renombrar bibliotecas existentes.
|
||||||
|
- `mark_done` sigue siendo el **único** escritor de `markdown_path`.
|
||||||
|
- Al leer `.md`, captura `UnicodeDecodeError` además de `OSError`: es un `ValueError`, y dejarlo escapar abortaba el escaneo entero saltándose todos los ficheros posteriores.
|
||||||
|
|
||||||
|
En el frontend, `refreshLiveState()` es el único punto de invalidación: lo llaman el handler `done` del SSE, un poll de 5 s mientras hay job activo, y `visibilitychange`/`focus`. `setView` recarga **siempre**, no solo cuando la lista está vacía.
|
||||||
|
|
||||||
|
### Rate limiting: `ratelimit.py` es la política, y no es opcional
|
||||||
|
|
||||||
|
Tres piezas distintas, no intercambiables:
|
||||||
|
|
||||||
|
- **`Pacer` (`GLOBAL_PACER`)** — separación mínima entre peticiones, **global al proceso**. Existe porque el `JobManager` serializa *jobs* pero los `/api/tools/*` corren fuera de él, en el threadpool: sin un pacer compartido, dos consumidores machacan YouTube creyendo cada uno que va educado. `wait(cost=N)` reserva N ranuras — `extract_info` de un vídeo son **2** peticiones (watch + player) y cobrarlo como 1 hacía que el pacer contase la mitad.
|
||||||
|
- **`backoff_delay`** — `min(base * 2**n + jitter, cap)`, el algoritmo que Google documenta para sus propias APIs, con el jitter re-sorteado en cada intento.
|
||||||
|
- **`ThrottleGuard`** — el circuit breaker. Cuenta rate-limits **consecutivos**; al llegar a `throttle_threshold` (3) el job para. Lo que no tocó sigue en `pending`, que es el estado recuperable.
|
||||||
|
|
||||||
|
**El incidente que lo motiva:** 343 de los 350 `error` de la DB decían literalmente *"The current session has been rate-limited by YouTube for up to an hour"*. No es que los vídeos fallaran: el scraper se ganaba un ban de una hora y luego **quemaba el resto de la cola en cascada** marcándolos como fallidos. Sin breaker, un throttling se convierte en cientos de filas envenenadas.
|
||||||
|
|
||||||
|
- Un fallo **no** de throttling (vídeo privado, sin subtítulos) **resetea** el contador. Si no, tres vídeos de membresía seguidos abortarían un scrape sano.
|
||||||
|
- `is_rate_limited` e `is_quota_exhausted` son cosas distintas: Google documenta la cuota como diaria (AIP-194), así que reintentar no la arregla y el guard corta en seco sin backoff.
|
||||||
|
- Los mensajes reales de la DB están fijados verbatim en [tests/test_ratelimit.py](tests/test_ratelimit.py). Si el detector deja de reconocerlos, el breaker es decorativo.
|
||||||
|
- **No amplíes `Store.PERMANENT_ERROR_PATTERNS`** con el mensaje de throttling: tiene que seguir siendo reintentable.
|
||||||
|
|
||||||
|
#### Los nombres de las opciones de yt-dlp se validan, no se revisan
|
||||||
|
|
||||||
|
Durante toda la historia del proyecto, los cuatro puntos de red pasaron `sleep_subrequests` a yt-dlp. **Esa opción no existe.** yt-dlp ignora en silencio las claves que no conoce, así que nunca hubo ni un milisegundo de pausa entre las sub-peticiones de una extracción, mientras `config.yaml` aparentaba tenerlo configurado. Lo mismo con `extract_flat_args`.
|
||||||
|
|
||||||
|
El nombre real es **`sleep_interval_requests`**, y es la **única** palanca que afecta a la extracción: `sleep_interval` / `max_sleep_interval` se disparan en el downloader de ficheros y **nunca** actúan bajo `skip_download`, que es todo el camino de metadatos.
|
||||||
|
|
||||||
|
Todo `ydl_opts` de politeness sale ahora de `ratelimit.ydl_throttle_opts()`, y [tests/test_request_economy.py](tests/test_request_economy.py) valida cada clave contra `yt_dlp.parse_options([]).ydl_opts` (172 nombres válidos). Si inventas una opción, el test falla.
|
||||||
|
|
||||||
|
**yt-dlp no reintenta 403/429 en YouTube**: su extractor los excluye explícitamente del `RetryManager`. `retries=10` no te protege de nada; el backoff ante throttling es responsabilidad nuestra.
|
||||||
|
|
||||||
|
#### Economía de peticiones: lo medido, para no re-litigarlo
|
||||||
|
|
||||||
|
| Operación | Peticiones | Nota |
|
||||||
|
|---|---|---|
|
||||||
|
| Extracción de un vídeo | **2** | watch + `youtubei/v1/player`. Es el suelo. |
|
||||||
|
| Transcripción (`timedtext`) | 1 | vía `yt_get` |
|
||||||
|
| Miniatura | 1 | `i.ytimg.com`, CDN estático, no cuenta contra el throttle de la API |
|
||||||
|
| `discover_channel(limit=30)` | **1** | |
|
||||||
|
| Discovery completo de canal de 2564 vídeos | ~86 | ~30 entradas por petición |
|
||||||
|
| `deep_channel_avatar` | **3-4** | antes **735 y subiendo** cuando la medición lo abortó |
|
||||||
|
|
||||||
|
- **`deep_channel_avatar` extraía el canal entero para leer una URL de imagen.** yt-dlp redirige una URL de canal pelada a `/videos`, recorre además `/streams` y `/shorts`, y `download=False` solo evita bajar el media, no la extracción. `extract_flat` + `playlistend: 1` son **carga estructural, no tuning**.
|
||||||
|
- **Restringir `player_client` no ahorra nada y rompe cosas.** Medido: `web_safari` solo → 1 petición pero **0 pistas de subtítulos**. `player_skip=webpage` → pierde subtítulos *y* `duration`/`channel_id`/`tags`. El default de yt-dlp (`android_vr` + `web_safari`, 2 peticiones) ya es el óptimo.
|
||||||
|
- **`playlistend` se ignora en silencio con `process=False`.** Medido sobre 2564 vídeos: procesado + `playlistend=60` = **2** peticiones; sin procesar, el generador perezoso ignora el límite y recorrerlo cuesta **86**. Hay un test que fija `process is not False`.
|
||||||
|
- `check_formats` cuesta **una petición HTTP por formato**. Nunca lo actives.
|
||||||
|
- No desactives `cachedir`: yt-dlp cachea ahí el player JS resuelto y quitarlo añade peticiones.
|
||||||
|
- `POST /api/tools/thumbnails` con solo `channel_id` acota a `THUMBNAIL_AUTO_LIMIT` (60). Antes, añadir un canal disparaba una petición al CDN **por cada vídeo del catálogo** — 2564 de golpe en Platzi.
|
||||||
|
|
||||||
|
#### La transcripción tiene que ser el idioma que se habla, no una traducción
|
||||||
|
|
||||||
|
`languages` es una preferencia **entre idiomas que sabes leer**, no una orden de aceptar una traducción automática cuando el transcript real está ahí. `pick_subtitle` pone delante el idioma hablado si está entre los configurados; solo si no lo está manda el orden del dict.
|
||||||
|
|
||||||
|
Cómo se detecta el original: YouTube publica el ASR como **`<lang>-orig`** y luego una cola larga de traducciones con el código pelado — incluida **una traducción al propio idioma del vídeo**. En un vídeo español existen `es-orig` *y* `es`, y solo el primero es el transcript real. `_is_original_track()` mira ese sufijo, y `original_language()` lo prefiere sobre `info["language"]` porque el sufijo es evidencia de la lista de pistas mientras que `language` es metadato que YouTube localiza.
|
||||||
|
|
||||||
|
**El fallo:** `_normalize_lang("es-orig") == "es"`, así que el sufijo — lo único que distingue el ASR real de una traducción — se tiraba antes de comparar. Con `{es, es-419, en}` sobre el canal de Alex Hormozi (inglés), el picker enganchaba `es` y guardaba una traducción máquina del inglés hablado. Prueba: las URLs de `timedtext` llevaban `lang=en&kind=asr&variant=gemini&tlang=es` — **`tlang=` es la marca de traducción**. Los canales en español acertaban **por casualidad**: yt-dlp lista `es-orig` antes que `es`. 3 filas de 119 afectadas; las otras 116 son `es-orig` legítimo.
|
||||||
|
|
||||||
|
- `extractor_args: {youtube: {skip: ["translated_subs"]}}` **no arregla esto**. El gate está anidado dentro de `if is_manual_subs` (`_video.py:4284-4285`) y `is_manual_subs` es False para pistas ASR. Medido: 158 claves con el flag y sin él, idénticas.
|
||||||
|
- `subtitleslangs` **no influye** en la selección: `pick_subtitle` lee `info["automatic_captions"]` crudo, mientras yt-dlp confina su propia selección a `requested_subtitles`, clave que este repo nunca lee.
|
||||||
|
- Recuperar una fila así **no se puede por las vías normales**: `reset_videos` nunca toca `done`, y `re-render` regeneraría el idioma equivocado desde `segments_json`. Hay que limpiar `transcript_segments`, `transcript_fts`, `segments_json` **y borrar el `.md`** — si lo dejas, `reconcile_markdown()` lo re-marca `done` al arrancar.
|
||||||
|
|
||||||
|
**Pendiente, sin resolver:** los títulos también vienen localizados. El canal Platzi tiene títulos en inglés en la DB ("The Claude Fable controversy 5" por "La controversia de Claude Fable 5") mientras sus `.md` se llamaron en español. No está determinado cuál de los dos lados —el listado flat del tab o la extracción— es el localizado; hace falta una medición contra la red para saberlo.
|
||||||
|
|
||||||
|
#### El detector de throttling no puede depender de la prosa
|
||||||
|
|
||||||
|
Hay dos ortografías y **son cadenas distintas**: yt-dlp lanza `HTTP Error 429: ...` y `requests` lanza `429 Client Error: ... for url: ...`. `_HTTP_STATUS` solo cubría la primera, así que la segunda se detectaba únicamente por el substring `"too many requests"` — y un 429 servido **sin reason phrase**, que es lo normal en HTTP/2, no lleva esa prosa y atravesaba el breaker sin contarse. Igual con 408, que Google documenta como reintentable.
|
||||||
|
|
||||||
|
Ahora `_download_subtitle` normaliza el mensaje anteponiendo `HTTP Error <status>:` desde `exc.response.status_code`, y el detector reconoce ambas formas. Está cubierto por tests parametrizados con las dos ortografías y con los casos que **no** deben disparar.
|
||||||
|
|
||||||
|
#### Los eventos de error usan `message`; los de log usan `msg`
|
||||||
|
|
||||||
|
`jobs.py` emite `{"message": ...}` en los cinco sitios donde manda un evento `error`. El frontend leía `d.msg`, que no existe en esos eventos, así que **todo** fallo de job salía como `"Job failed: connection error"` — incluida la explicación detallada del breaker. Corregido en `app.js` leyendo `d.message || d.msg`, y el flag `throttled` cambia el texto: parar por rate limiting no es un crash, es una parada deliberada con todo lo no alcanzado aún en `pending`.
|
||||||
|
|
||||||
|
#### Los metadatos se persisten antes de la salida por "sin transcripción"
|
||||||
|
|
||||||
|
`process_video` guarda `view_count`/`description`/`thumbnail`/`tags`/`upload_date` **encima** del `return "no_subtitles"`. La extracción ya pagó sus dos peticiones y el `info` está en memoria; tirarlo porque falló la descarga *separada* de subtítulos hace que el reintento las vuelva a gastar para obtener datos que ya teníamos. Medido tras el incidente: cinco filas quedaron con `upload_date`, `view_count`, `description` y `thumbnail` todos NULL.
|
||||||
|
|
||||||
|
Relacionado: `mark_status` trunca el motivo a **2000** caracteres, no 500. A 500 el corte caía veinte caracteres antes del `tlang=` que probaba el diagnóstico.
|
||||||
|
|
||||||
|
#### Los handlers que hacen I/O de red van en `def`, no en `async def`
|
||||||
|
|
||||||
|
FastAPI ejecuta los `async def` **en el event loop**; los `def` van al threadpool. `add_channel`, `sync_channels`, `download_thumbnails`, `download_avatars` y `reconcile` bloquean con yt-dlp o `requests`, así que siendo `async` congelaban el servidor entero — incluido el SSE del job en curso — durante toda su duración. Verificado en producción: añadir Platzi bloquea 295 s, y con el cambio `/api/dashboard` respondió 146/146 sondas con mediana de 0,000 s.
|
||||||
|
|
||||||
|
`stream_job` **sí** debe seguir siendo `async`: devuelve el `EventSourceResponse`.
|
||||||
|
|
||||||
|
#### Calibración
|
||||||
|
|
||||||
|
El único techo publicado es el de la wiki de yt-dlp: **~300 vídeos/hora (~1000 peticiones/hora)** en sesión sin cuenta. Google no documenta límites para acceso sin API, y los umbrales del bot-check no están documentados en ninguna parte: cualquier cifra concreta es prudencia, no norma.
|
||||||
|
|
||||||
|
`config.yaml` apunta por debajo de eso. Medido en producción sobre Platzi: **13,25 s/vídeo → ~272 vídeos/hora, ~815 peticiones/hora**. Si cambias los tiempos, vuelve a medir: 3 peticiones a `youtube.com` por vídeo es la constante de la que sale todo lo demás.
|
||||||
|
|
||||||
|
### Cookies
|
||||||
|
|
||||||
|
Vault híbrido: metadatos en `cookies_meta`, archivos Netscape en `cookies/<uuid>.txt` (gitignored). **Exactamente una** cookie activa a la vez. Cuando no se pasa `--cookies`, CLI, webapp y monitor caen en `cookies.resolve_active_path(store)`. `auto_import_dir()` adopta `.txt` sueltos al arrancar CLI y webapp.
|
||||||
|
|
||||||
|
### Job runner de la webapp
|
||||||
|
|
||||||
|
`JobManager` = un único thread daemon + `deque`, deliberadamente secuencial para no martillear a YouTube. Los eventos van a una lista en memoria por job (`_events`) que el endpoint SSE **poletea** con un cursor cada 250 ms; no hay pub/sub real, y los eventos se pierden al reiniciar (el estado durable está en `scrape_jobs`).
|
||||||
|
|
||||||
|
Despacho por `opts` en `_run_job()`: `mode == "audio"` → `_run_audio`; hay `video_ids` → `_run_batch` (por vídeo: `.md` + thumbnail, saltando los que ya tienen `.md` en disco); si no → `_run_channel` (discovery + pendientes + `polite_sleep` entre vídeos). La cancelación es cooperativa: `_cancel` es un set que los loops consultan.
|
||||||
|
|
||||||
|
### Frontend
|
||||||
|
|
||||||
|
Sin build step, por decisión fija ([.opencode/agent/webapp-builder.md](.opencode/agent/webapp-builder.md)): Tailwind, Alpine 3 y Chart.js por CDN. Todo el estado vive en **un solo** componente Alpine, `window.platform()` en [static/app.js](src/yt_scraper/webapp/static/app.js), y **debe** quedar definido antes de que Alpine inicialice — el orden de los `<script defer>` en `index.html` (app.js antes de alpinejs) es carga funcional, no estilo. Las vistas se sincronizan con la URL (`hydrateURL`/`syncURL`), no con hash routes. Identidad visual: dark "command center", `#0a0a0f` + acento rose `#f43f5e`.
|
||||||
|
|
||||||
|
### Rutas de datos y convenciones
|
||||||
|
|
||||||
|
El data root se deriva siempre como `Path(cfg.output_dir_resolved).parent` → `data/{markdown,audio,thumbnails,avatars,exports,analysis}` y `data/state.db`. `videos.markdown_path` se guarda **relativo a ese root**.
|
||||||
|
|
||||||
|
Gotcha real: el job de audio de la webapp guarda `data/audio/<video_id>.mp3` (`outtmpl` con `%(id)s`) y `GET /api/videos/{id}/audio` solo encuentra ese nombre, mientras que `yt-scraper audio` (CLI) escribe por título. Los MP3 bajados por CLI no los sirve la API.
|
||||||
|
|
||||||
|
### Cualquier petición HTTP a CDNs de YouTube va por `_yt_http.yt_get()`
|
||||||
|
|
||||||
|
Headers compartidos (UA de Chrome + `Referer: https://www.youtube.com/`). Sin eso, `yt3.ggpht.com` (avatares) y `timedtext` (subtítulos) devuelven 403. `yt-dlp` es la **única** interfaz con YouTube: no añadas llamadas directas a InnerTube.
|
||||||
|
|
||||||
|
### Idiomas: dict por-idioma con compatibilidad legacy
|
||||||
|
|
||||||
|
`languages` puede ser lista (`["es","en"]`, todos usan `prefer_manual`) o dict `{lang: "manual"|"auto"|"any"}`. La normalización está duplicada a propósito: `config.parse_languages()` y `extract._coerce_languages()` (para que `extract` no dependa de `config`). Si cambias las reglas, cambia las dos.
|
||||||
|
|
||||||
|
Estados de vídeo: `pending` / `done` / `no_subtitles` / `error`. Estados terminales de job: `Store.TERMINAL_STATUSES`.
|
||||||
|
|
||||||
|
## Tests
|
||||||
|
|
||||||
|
Solo `tmp_path` + `monkeypatch`, cero red: `yt_dlp.YoutubeDL` se sustituye por un fake (ver [tests/test_webapp_jobs.py](tests/test_webapp_jobs.py)) y `pick_subtitle` se testea con dicts `info` sintéticos. Los tests del store construyen un `Store` sobre `tmp_path`, así que también cubren la migración idempotente.
|
||||||
|
|
||||||
|
## Documentos de referencia
|
||||||
|
|
||||||
|
- [docs/superpowers/specs/2026-07-26-platform-design.md](docs/superpowers/specs/2026-07-26-platform-design.md) — diseño de referencia (esquema, endpoints, alcance, exclusiones). Es la fuente de verdad de las decisiones; donde el código difiere, gana el código.
|
||||||
|
- [OPPORTUNITIES.md](OPPORTUNITIES.md) — backlog de features con esfuerzo estimado.
|
||||||
|
- [YOUTUBE-TRANSCRIPT-AUDIT.md](YOUTUBE-TRANSCRIPT-AUDIT.md) — ingeniería inversa de Obsidian Web Clipper (mecanismo InnerTube). Contexto de *por qué* `yt-dlp` es la ruta elegida.
|
||||||
|
- `data/`, `cookies/`, `.run/` están gitignored y contienen datos reales (883 vídeos, sesión de YouTube). No los commitees ni pegues su contenido en respuestas.
|
||||||
@@ -0,0 +1,505 @@
|
|||||||
|
# Auditoría · Cómo la extensión obtiene la transcripción / información de vídeos de YouTube
|
||||||
|
|
||||||
|
> **Producto auditado:** *Obsidian Web Clipper* (Chrome / Chromium / Firefox / Safari) — versión `1.7.1` del paquete `D:\Obsidian Web Clipper - Chrome Web Store 1.7.1.0`.
|
||||||
|
> **Tipo de extensión:** MV3 (manifest v3) con service worker (`background.js`).
|
||||||
|
> **Alcance de la auditoría:** mecanismo end‑to‑end por el que la extensión extrae metadatos y la transcripción de un vídeo de YouTube (incluye short `youtu.be`, `youtube.com/watch?v=…` y `youtube.com/shorts/…`).
|
||||||
|
> **Fecha:** 2026‑07‑26.
|
||||||
|
> **Audiencia del documento:** LLMs / agentes de mantenimiento. Estructura deliberadamente declarativa, sin prosa narrativa.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 0 · TL;DR (resumen ejecutable)
|
||||||
|
|
||||||
|
1. La extensión **no usa `timedtext`, `youtube-transcript` web, ni scraping de `ytd-transcript-segment-renderer` como ruta principal** cuando la URL es un watch normal: usa la **API privada `youtubei/v1/player`** (InnerTube) con cabeceras que imitan clientes oficiales de YouTube (ANDROID, IOS, WEB).
|
||||||
|
2. Para llegar a esa API sin ser bloqueada por CORS / firma, el `service worker` declara una regla `declarativeNetRequest` (id `9002`, nombre interno `enableYouTubeInnertubeRule`) que **fuerza `Origin: https://www.youtube.com` y `Referer: https://www.youtube.com/`** en toda petición XHR iniciada por la propia extensión hacia `||youtube.com/youtubei/`.
|
||||||
|
3. La capa de extracción es una clase `YoutubeExtractor` (en `popup.js` y replicada en `reader-page.js`, ambos `webpack` bundles de Defuddle) que:
|
||||||
|
- 1️⃣ parsea el JSON embebido `ytInitialPlayerResponse` del DOM para sacar `captionTracks` y `baseUrl` sin red.
|
||||||
|
- 2️⃣ si falla, abre el panel "Mostrar transcripción" del propio YouTube haciendo `click()` y espera con `MutationObserver`‑style polling (DOM scraping fallback).
|
||||||
|
- 3️⃣ si la transcripción automática no está disponible, llama a `youtubei/v1/player` con 3 identidades de cliente en cascada (ANDROID → IOS → WEB) hasta que una devuelve `captions.playerCaptionsTracklistRenderer.captionTracks`.
|
||||||
|
- 4️⃣ descarga la pista (`timedtext`-like `baseUrl` con sufijo `&fmt=…`) y la parsea como XML.
|
||||||
|
4. Los capítulos se extraen con una segunda ruta: `youtubei/v1/next` (también con cabeceras de cliente), o desde `ytInitialData` embebido (`playerOverlays.playerOverlayRenderer.decoratedPlayerBarRenderer.multiMarkersPlayerBarRenderer.markersMap`).
|
||||||
|
5. Toda la red de la popup se hace con `globalThis.fetch` directo (no hay proxy interno), aprovechando la regla DNR 9002. El background solo ofrece un *fallback* `sendNativeMessage` para hosts que devuelven CORS (por ejemplo Bilibili).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1 · Vista general de componentes (mapa de archivos)
|
||||||
|
|
||||||
|
| Archivo | Rol respecto a YouTube | Tamaño aprox. | Notas |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `manifest.json` | Declara `host_permissions: ["<all_urls>","http://*/*","https://*/*"]` y `declarativeNetRequest`. | 89 líneas | Sin URL allow‑list específica de YouTube. |
|
||||||
|
| `background.js` (service worker) | Define la **regla DNR 9002** `enableYouTubeInnertubeRule` (set Origin/Referer para `||youtube.com/youtubei/`). Contiene además `enableYouTubeEmbedRule` (9001) que pone `Referer: https://obsidian.md/` en iframes `||youtube.com/embed/`. | ~1 archivo compilado | Comentario interno: `initiatorDomains: [chrome.runtime.id]` ⇒ sólo afecta peticiones de la propia extensión. |
|
||||||
|
| `content.js` (content script) | Sólo contiene el glue de highlights (`getClosestTextBlock` ignora elementos con clase `transcript-segment` para no romper la selección). **No extrae la transcripción.** | 1 bundle webpack | No realiza llamadas a YouTube. |
|
||||||
|
| `popup.js` | Contiene la clase `YoutubeExtractor` real (minificada) y todo el código de extracción. Se carga como `popup.html` y como `side-panel.html` (ver `<script type="module" src="popup.js">`). | ~2.5 MB minificado | Aquí vive toda la lógica de transcripción. |
|
||||||
|
| `reader-page.js` | Réplica exacta de los extractores (incluye otra copia de `YoutubeExtractor`). Se usa cuando se abre la URL en modo *Reader* (`reader.html?url=…`). | ~2.5 MB minificado | Mismo binario que popup. |
|
||||||
|
| `highlighter.js` | Sólo lógica de resaltado (no relevante para transcripción). | — | — |
|
||||||
|
| `reader-script.js` | Inyectado por background con `scripting.executeScript` para modo Reader. | — | — |
|
||||||
|
| `_locales/*/messages.json` | i18n; incluye claves `readerTranscripts`, `readerPinPlayer`, `readerHighlightActiveLine` (configuración visual de la transcripción, no de extracción). | — | — |
|
||||||
|
| `web_accessible_resources` | Lista `reader.css`, `reader-script.js`, `browser-polyfill.min.js`, `style.css`, `side-panel.html`, `flatten-shadow-dom.js`, `highlighter.css`. | — | `popup.js` **no** está en `web_accessible_resources`; por tanto la extracción no se hace desde un script inyectado en la página, sino desde la propia página de extensión. |
|
||||||
|
|
||||||
|
### 1.1 Flujo de control (quién llama a quién)
|
||||||
|
|
||||||
|
```
|
||||||
|
[user clicks action / shortcut / context menu]
|
||||||
|
└─ background.js (service worker)
|
||||||
|
└─ browser.action.openPopup() OR tabs.sendMessage("openPopup")
|
||||||
|
└─ popup.html (extension page, chrome-extension://<id>/popup.html)
|
||||||
|
└─ popup.js (module)
|
||||||
|
├─ Defuddle.parse(doc) ← extractor genérico
|
||||||
|
│ └─ para URL que matchea "youtube.com" o "youtu.be"
|
||||||
|
│ └─ new YoutubeExtractor(document, url, schemaOrg, options)
|
||||||
|
│ └─ extractAsync() ⇒ runExtractor()
|
||||||
|
│ ├─ extractTranscriptFromExistingDom() (1ª opción)
|
||||||
|
│ ├─ fetchTranscript() (2ª opción: red)
|
||||||
|
│ └─ extractTranscriptFromOpenedDom() (3ª opción: click en panel)
|
||||||
|
└─ resultado ⇒ variables { transcript, language } ⇒ se inyecta en la nota Markdown
|
||||||
|
|
||||||
|
(En paralelo, la regla DNR 9002 reescribe Origin/Referer de las XHR
|
||||||
|
lanzadas por la propia extensión hacia youtube.com/youtubei/v1/…)
|
||||||
|
```
|
||||||
|
|
||||||
|
> **Punto importante para LLMs:** la extracción de YouTube **no se ejecuta dentro de la página de YouTube** ni desde el content script. Se ejecuta en el contexto privilegiado de la extensión (`chrome-extension://`). La página de YouTube solo aporta el `document` con el HTML actual y, opcionalmente, el JSON embebido en los `<script>`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2 · Punto de entrada: ¿cuándo se invoca la extracción?
|
||||||
|
|
||||||
|
`background.js` ofrece cuatro formas de abrir la popup, todas convergen al mismo punto:
|
||||||
|
|
||||||
|
| Acción del usuario | Mensaje / llamada en `background.js` | Destino final |
|
||||||
|
|---|---|---|
|
||||||
|
| Click en icono de la extensión | `action.onClicked` → `openPopup()` | `popup.html` |
|
||||||
|
| Atajo `Ctrl+Shift+O` (`_execute_action`) | `commands.onCommand` → `openPopup()` | `popup.html` |
|
||||||
|
| Atajo `Alt+Shift+O` (`quick_clip`) | `commands.onCommand` → `openPopup()` + 500 ms `triggerQuickClip` | `popup.html` |
|
||||||
|
| Menú contextual "Save this page" / "Add to highlights" | `contextMenus.onClicked` → `openPopup()` | `popup.html` |
|
||||||
|
| Behavior `embedded` | `tabs.sendMessage("toggle-iframe")` → `side-panel.html?context=iframe` | `side-panel.html` |
|
||||||
|
|
||||||
|
> `side-panel.html` y `popup.html` cargan **el mismo `popup.js`** (`<script type="module" src="popup.js">`). Por tanto, la lógica de YouTube es única y se invoca desde dos contenedores distintos.
|
||||||
|
|
||||||
|
Cuando la popup se carga, en `popup.js` se hace algo equivalente a:
|
||||||
|
|
||||||
|
```js
|
||||||
|
const defuddle = new Defuddle(document, { url: location.href });
|
||||||
|
const result = defuddle.parse(); // extracción síncrona
|
||||||
|
const asyncVars = await defuddle.fetchAsyncVariables({ language, fetch }); // asíncrono
|
||||||
|
```
|
||||||
|
|
||||||
|
`Defuddle` consulta su `ExtractorRegistry` (poblado en `ExtractorRegistry.initialize()`); para YouTube registra:
|
||||||
|
|
||||||
|
```js
|
||||||
|
this.register({ patterns: ["youtube.com","youtu.be"], extractor: YoutubeExtractor });
|
||||||
|
this.register({ patterns: ["m.youtube.com"], extractor: YoutubeExtractor }); // implícito
|
||||||
|
this.register({ patterns: [/youtube\.com\/shorts\//], extractor: YoutubeExtractor });
|
||||||
|
```
|
||||||
|
|
||||||
|
El extractor se instancia con `new YoutubeExtractor(document, url, schemaOrgData, options)`. Las `options` que recibe la popup le inyectan `language` (preferida por el usuario) y un `fetch` opcional; si no se inyecta, usa `globalThis.fetch`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3 · `YoutubeExtractor` (la clase clave) — Anatomía
|
||||||
|
|
||||||
|
> **Ubicación física del código (minificado):**
|
||||||
|
> - En `popup.js`, la clase aparece aproximadamente entre los offsets `382 000`–`405 000` del bundle (texto buscado: `class YoutubeExtractor` o el alias `class A extends o.BaseExtractor`).
|
||||||
|
> - En `reader-page.js` es la **misma clase** con nombre `A` (mismo fingerprint de strings `transcript-segment-view-model`, `ytwTranscriptSegmentViewModelTimestamp`, `ytInitialPlayerResponse`).
|
||||||
|
> - En origen viene del paquete npm `defuddle` (≥ v0.x) — la extensión lo reempaqueta con webpack.
|
||||||
|
|
||||||
|
### 3.1 Identidad de cliente (constantes globales)
|
||||||
|
|
||||||
|
```js
|
||||||
|
// constantes a nivel de módulo, dentro del bundle de popup.js / reader-page.js
|
||||||
|
const TIMEOUT_MS = 4000; // f = 4e3
|
||||||
|
const PLAYER_URL = "https://www.youtube.com/youtubei/v1/player?prettyPrint=false"; // g
|
||||||
|
const NEXT_URL = "https://www.youtube.com/youtubei/v1/next?prettyPrint=false"; // usado en fetchChapters
|
||||||
|
const ANDROID_UA = "com.google.android.youtube/20.10.38 (Linux; U; Android 14)"; // v, usada en b/x
|
||||||
|
|
||||||
|
// Contextos de cliente (probados en cascada)
|
||||||
|
const ANDROID_CLIENT = { client: { clientName: "ANDROID", clientVersion: "20.10.38" } }; // b / x
|
||||||
|
const IOS_CLIENT = { client: { clientName: "IOS", clientVersion: "20.10.3" } }; // y
|
||||||
|
const WEB_CLIENT = { client: { clientName: "WEB", clientVersion: "2.20240101.00.00" } }; // w
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3.2 Selectores DOM (definidos como objetos)
|
||||||
|
|
||||||
|
```js
|
||||||
|
const DESKTOP_SELECTORS = {
|
||||||
|
segments: "ytd-transcript-segment-renderer",
|
||||||
|
timestamp: ".segment-timestamp",
|
||||||
|
text: ".segment-text",
|
||||||
|
};
|
||||||
|
const MOBILE_SELECTORS = {
|
||||||
|
segments: "transcript-segment-view-model",
|
||||||
|
timestamp: ".ytwTranscriptSegmentViewModelTimestamp",
|
||||||
|
text: "span.yt-core-attributed-string",
|
||||||
|
chapters: "timeline-chapter-view-model h3",
|
||||||
|
};
|
||||||
|
```
|
||||||
|
|
||||||
|
`getTranscriptSelectors(container)` elige uno u otro mirando qué nodos existen. Si no existe ninguno devuelve `undefined` (⇒ no hay transcripción en el DOM todavía).
|
||||||
|
|
||||||
|
### 3.3 Métodos principales (firmas y propósito)
|
||||||
|
|
||||||
|
| Método | Tipo | Propósito |
|
||||||
|
|---|---|---|
|
||||||
|
| `getVideoId()` | síncrono | Devuelve el id de 11 chars. Soporta `youtube.com/watch?v=…`, `youtu.be/…`, `youtube.com/shorts/…`. Cachea en `this._videoId`. |
|
||||||
|
| `canExtractAsync()` | síncrono | Devuelve `true` si la URL es de YouTube. |
|
||||||
|
| `extractAsync()` | async | Punto de entrada. Cadena: `extractTranscriptFromExistingDom()` → si vacío `fetchTranscript()` → si vacío `extractTranscriptFromOpenedDom()`. Devuelve `{ html, text, languageCode, … }` o `null`. |
|
||||||
|
| `extractTranscriptFromExistingDom()` | try/catch | Lee los segmentos si YouTube ya renderizó el panel de transcripción (panel abierto por el usuario o cargado por interacción previa). |
|
||||||
|
| `getTranscriptContainer()` | síncrono | Selector: `'ytd-engagement-panel-section-list-renderer[target-id="engagement-panel-searchable-transcript"] #segments-container'`; en `m.youtube.com`: `ytm-macro-markers-list-renderer .ytm-macro-markers-list-container`. |
|
||||||
|
| `buildTranscriptFromContainer(container, chapters)` | síncrono | Itera cada segmento, parsea timestamp (`parseTimestamp` acepta `h:mm:ss` o `mm:ss`), agrupa por hablante si detecta patrón (`groupTranscriptSegments` → `groupBySpeaker` o `groupBySentence`), y emite HTML `<p class="transcript-segment">` y texto plano `**HH:MM:SS** · texto`. Llama a `buildTranscript("youtube", groups, chapters)`. |
|
||||||
|
| `extractTranscriptFromOpenedDom()` | async | Si el panel no estaba abierto pero el DOM lo permite (`canOpenTranscriptPanel()`: `typeof MutationObserver === "function"`): hace **click programático** en `ytd-video-description-transcript-section-renderer button` (o equivalente mobile), espera el contenedor y re-usa `buildTranscriptFromContainer`. |
|
||||||
|
| `openMobileTranscriptPanel()` | async | Variante `m.youtube.com`: clicks en `button[aria-label="Show more"]` → espera `button[aria-label="View all"]` → click → espera segmentos. |
|
||||||
|
| `fetchTranscript()` | async | **Ruta de red principal**. Ver §3.4. |
|
||||||
|
| `fetchPlayerData(videoId)` | async | POST a `youtubei/v1/player` probando 3 clientes en cascada. Ver §3.4. |
|
||||||
|
| `fetchChapters(videoId)` | async | POST a `youtubei/v1/next` (cliente WEB) → extrae capítulos de `playerOverlays…multiMarkersPlayerBarRenderer.markersMap[*].value.chapters[*].chapterRenderer`; fallback a `engagementPanels[*].engagementPanelSectionListRenderer.content.macroMarkersListRenderer.contents[*].macroMarkersListItemRenderer`. Si ya están embebidos en `ytInitialData` no se hace red. |
|
||||||
|
| `getValidatedPlayerResponse()` | síncrono | Devuelve el JSON parseado de `ytInitialPlayerResponse` (parseado inline desde el `<script>`) **solo si** su `videoDetails.videoId` o `microformat.playerMicroformatRenderer.externalVideoId` coincide con `getVideoId()`. |
|
||||||
|
| `parseInlineJson(varName)` | síncrono | Itera todos los `<script>` del documento, encuentra el primero cuyo `textContent` contiene la variable global, **balancea llaves** manualmente y hace `JSON.parse`. Cachea el resultado en `this.inlineJsonCache` (Map). |
|
||||||
|
| `getCaptionTracks(playerResponse)` | síncrono | `playerResponse?.captions?.playerCaptionsTracklistRenderer?.captionTracks` (devuelve `[]` si no es array). |
|
||||||
|
| `pickCaptionTrack(tracks)` | síncrono | Si hay `options.language`: prefiere la pista exacta (`code === lang`) ⇒ mismo idioma base (`code.split("-")[0] === lang.split("-")[0]`) ⇒ mismo prefijo de idioma. Filtra `kind === "asr"` (auto‑generadas) si hay manuales. Si no, devuelve la primera no‑`asr`, o una con `languageCode === "en"`, o la primera. |
|
||||||
|
| `findPreferredCaptionTrack(tracks, lang)` | síncrono | Igual que el anterior pero con scoring explícito. |
|
||||||
|
| `getInlineCaptionTrack()` | síncrono | `getValidatedPlayerResponse()` → `getCaptionTracks()` → `pickCaptionTrack()`. Si hay `baseUrl`, devuelve la pista. |
|
||||||
|
| `fetchCaptionXml(track, chaptersPromise)` | async | `fetch(track.baseUrl, { headers: { "User-Agent":"Mozilla/5.0", "Accept-Language": lang } })` con `AbortSignal.timeout(4000)`. Devuelve solo si la URL acaba en `.youtube.com`. Pasa el texto a `parseTranscriptXml`. |
|
||||||
|
| `parseTranscriptXml(xml, lang, chapters)` | síncrono | Dos regex: `<p t="N">…<s>…</s>…</p>` (formato nuevo) y `<text start="N">…</text>` (formato legacy). Decodifica entidades (`decodeEntities`). Llama a `groupTranscriptSegments` y `buildTranscript`. |
|
||||||
|
| `decodeEntities(str)` | síncrono | Reemplazos para `& < > " ' ' &#xHH; &#NN;`. |
|
||||||
|
| `groupTranscriptSegments(segs)` | síncrono | Decide por regex CJK: si hay mezcla CJK/Latín → `groupBySpeaker` (split por `:` al inicio de línea); si no → `groupBySentence`. |
|
||||||
|
| `getVideoData()` | síncrono | Lee `<script type="application/ld+json">` buscando un `VideoObject` cuyo `embedUrl`/`url`/`@id` contenga el videoId. Fallback a `meta[property="og:title|og:description|og:image|og:url"]`. |
|
||||||
|
| `getChannelNameFromDom()` / `getChannelNameFromPlayerResponse()` | síncrono | `[itemprop="name"]` o `videoDetails.author`/`ownerChannelName`/`microformat.playerMicroformatRenderer.ownerChannelName`. |
|
||||||
|
| `getTranscriptLanguageCodeFromDom()` | síncrono | Lee el botón de `yt-sort-filter-sub-menu-renderer` en el footer del panel y compara con `name.simpleText`/`name.runs[].text` de cada caption track. |
|
||||||
|
| `getInlineChapters()` | síncrono | Desde `ytInitialData`; valida que el videoId en el JSON coincida con el de la URL; cae a `extractChaptersFromEngagementPanels`. |
|
||||||
|
| `buildResult(transcript)` | síncrono | Empaqueta `{ title, author, site:"YouTube", image, published, description }`, añade `transcript` y `language` al `variables`, prepende un `<iframe src="https://www.youtube.com/embed/{id}" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen>`. |
|
||||||
|
| `formatDescription(text)` | síncrono | `<p>…</p>` con `escapeHtml` y `<br>` por `\n`. |
|
||||||
|
|
||||||
|
### 3.4 `fetchTranscript()` — núcleo de la ruta de red
|
||||||
|
|
||||||
|
```js
|
||||||
|
async fetchTranscript() {
|
||||||
|
const videoId = this.getVideoId();
|
||||||
|
const chapters = this.fetchChapters(videoId); // (a)
|
||||||
|
const inlineTrack = this.getInlineCaptionTrack(); // (b) caption del JSON embebido
|
||||||
|
const inlineFetch = inlineTrack ? this.fetchCaptionXml(inlineTrack, chapters) : undefined; // (c)
|
||||||
|
const playerData = await this.fetchPlayerData(videoId); // (d) red: youtubei/v1/player
|
||||||
|
const picked = playerData ? this.pickCaptionTrack(this.getCaptionTracks(playerData)) : undefined;
|
||||||
|
const remoteFetch = (picked?.baseUrl && picked.baseUrl !== inlineTrack?.baseUrl)
|
||||||
|
? this.fetchCaptionXml(picked, chapters) // (e)
|
||||||
|
: undefined;
|
||||||
|
return (await remoteFetch) || (await inlineFetch);
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Y `fetchPlayerData()`:
|
||||||
|
|
||||||
|
```js
|
||||||
|
async fetchPlayerData(videoId) {
|
||||||
|
// 1º intento: cliente IOS
|
||||||
|
try {
|
||||||
|
const r = await this.fetch(PLAYER_URL, {
|
||||||
|
method: "POST",
|
||||||
|
headers: { "Content-Type":"application/json", ...(lang && {"Accept-Language": lang}) },
|
||||||
|
signal: AbortSignal.timeout(TIMEOUT_MS),
|
||||||
|
body: JSON.stringify({ context: IOS_CLIENT, videoId })
|
||||||
|
});
|
||||||
|
if (r.ok) { const t = await r.json(); if (this.getCaptionTracks(t).length) return t; }
|
||||||
|
} catch {}
|
||||||
|
|
||||||
|
// 2º intento: cliente ANDROID con UA
|
||||||
|
try {
|
||||||
|
const r = await this.fetch(PLAYER_URL, {
|
||||||
|
method: "POST",
|
||||||
|
headers: { "Content-Type":"application/json", "User-Agent": ANDROID_UA, ...(lang && {"Accept-Language": lang}) },
|
||||||
|
signal: AbortSignal.timeout(TIMEOUT_MS),
|
||||||
|
body: JSON.stringify({ context: ANDROID_CLIENT, videoId })
|
||||||
|
});
|
||||||
|
if (r.ok) { const t = await r.json(); if (this.getCaptionTracks(t).length) return t; }
|
||||||
|
} catch {}
|
||||||
|
|
||||||
|
// 3º intento: cliente WEB clásico
|
||||||
|
try {
|
||||||
|
const r = await this.fetch(PLAYER_URL, {
|
||||||
|
method: "POST",
|
||||||
|
headers: { "Content-Type":"application/json" },
|
||||||
|
signal: AbortSignal.timeout(TIMEOUT_MS),
|
||||||
|
body: JSON.stringify({ context: WEB_CLIENT, videoId })
|
||||||
|
});
|
||||||
|
if (r.ok) { const t = await r.json(); if (this.getCaptionTracks(t).length) return t; }
|
||||||
|
} catch {}
|
||||||
|
|
||||||
|
// 4º fallback: parsear el JSON embebido en el HTML (sin red)
|
||||||
|
const inline = this.parseInlineJson("ytInitialPlayerResponse");
|
||||||
|
if (this.getCaptionTracks(inline).length) return inline;
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
> **Para LLMs:** los 3 contextos de cliente y el `User-Agent` ANDROID son los mismos que usa `youtube-dl`, `yt-dlp` y la mayoría de librerías no oficiales. Si YouTube empieza a rechazar la firma, el orden de los 3 intentos es lo que hay que tocar; también se puede añadir `TVHTML5_SIMPLY_EMBEDDED_PLAYER` u otros.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4 · Mecanismo de bypass de CORS / firma de YouTube
|
||||||
|
|
||||||
|
YouTube no expone `youtubei/v1/player` con CORS abierto, así que la extensión necesita que las peticiones se *vean* como originadas desde la propia web de YouTube. Hay **dos piezas** que lo permiten:
|
||||||
|
|
||||||
|
### 4.1 `enableYouTubeInnertubeRule` (DNR id `9002`)
|
||||||
|
|
||||||
|
En `background.js` (en la función `initialize()`):
|
||||||
|
|
||||||
|
```js
|
||||||
|
await dnr.updateSessionRules({
|
||||||
|
removeRuleIds: [9002],
|
||||||
|
addRules: [{
|
||||||
|
id: 9002,
|
||||||
|
priority: 1,
|
||||||
|
action: {
|
||||||
|
type: "modifyHeaders",
|
||||||
|
requestHeaders: [
|
||||||
|
{ header: "Origin", operation: "set", value: "https://www.youtube.com" },
|
||||||
|
{ header: "Referer", operation: "set", value: "https://www.youtube.com/" }
|
||||||
|
]
|
||||||
|
},
|
||||||
|
condition: {
|
||||||
|
urlFilter: "||youtube.com/youtubei/",
|
||||||
|
resourceTypes: ["xmlhttprequest"],
|
||||||
|
initiatorDomains: [ chrome.runtime.id ].filter(Boolean)
|
||||||
|
}
|
||||||
|
}]
|
||||||
|
});
|
||||||
|
```
|
||||||
|
|
||||||
|
> Como el popup (`chrome-extension://<id>/popup.html`) lanza `fetch` desde el contexto de la extensión, el `initiator` de la petición es el `extension_id` ⇒ la regla **solo** se aplica a las peticiones que dispara la propia extensión. Las peticiones que haga el content script de la página de YouTube no se tocan aquí.
|
||||||
|
|
||||||
|
### 4.2 `webRequest.onBeforeSendHeaders` (sólo Firefox / WebExtensions)
|
||||||
|
|
||||||
|
En `background.js` también hay (entre `try { … } catch {}` para tolerancia a Safari):
|
||||||
|
|
||||||
|
```js
|
||||||
|
browser_polyfill.webRequest.onBeforeSendHeaders.addListener(details => {
|
||||||
|
if (details.tabId > 0) {
|
||||||
|
const refHeader = details.requestHeaders.find(h => h.name.toLowerCase() === "referer");
|
||||||
|
const refValue = refHeader?.value || "";
|
||||||
|
const originHdr = details.requestHeaders.find(h => h.name.toLowerCase() === "origin");
|
||||||
|
const originValue = originHdr?.value || "";
|
||||||
|
if (!(refValue.startsWith("moz-extension://") || refValue.startsWith("safari-web-extension://"))) {
|
||||||
|
return { requestHeaders: details.requestHeaders }; // no tocar: es la propia web
|
||||||
|
}
|
||||||
|
}
|
||||||
|
const headers = details.requestHeaders || [];
|
||||||
|
const setHeader = (name, value) => {
|
||||||
|
const existing = headers.find(h => h.name.toLowerCase() === name.toLowerCase());
|
||||||
|
existing ? existing.value = value : headers.push({ name, value });
|
||||||
|
};
|
||||||
|
setHeader("Origin", "https://www.youtube.com");
|
||||||
|
setHeader("Referer", "https://www.youtube.com/");
|
||||||
|
return { requestHeaders: headers };
|
||||||
|
}, { urls: ["*://www.youtube.com/*"] }, ["blocking","requestHeaders"]);
|
||||||
|
```
|
||||||
|
|
||||||
|
> **Para LLMs:** este listener sólo aplica a Firefox MV2 / WebExtensions, donde `declarativeNetRequest` no soporta `modifyHeaders` de la misma forma. **No se ejecuta en Chrome** (donde ya tenemos DNR 9002). En Safari se ignora silenciosamente (`catch` lo traga).
|
||||||
|
|
||||||
|
### 4.3 `enableYouTubeEmbedRule` (DNR id `9001`) — caso especial, no transcripción
|
||||||
|
|
||||||
|
No participa en la extracción de transcripción, pero la documentamos para que el lector no se confunda: cuando la popup muestra el `<iframe src="https://www.youtube.com/embed/…">` en modo *Reader*, background fuerza `Referer: https://obsidian.md/` para que el embed no se rompa en vídeos con restricción por referer.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5 · `fetchProxy` y `nativeFetch` — por qué existen (y por qué YouTube NO los usa)
|
||||||
|
|
||||||
|
En `background.js`:
|
||||||
|
|
||||||
|
```js
|
||||||
|
browser_polyfill.runtime.onMessage.addListener(request => {
|
||||||
|
if (request.action !== "fetchProxy") return;
|
||||||
|
return fetch(request.url, request.options)
|
||||||
|
.then(async resp => {
|
||||||
|
const text = await resp.text();
|
||||||
|
const looksLikeHTML = !resp.ok && (text.includes("Sorry") || text.includes("<html"));
|
||||||
|
if (!looksLikeHTML) return { ok: resp.ok, status: resp.status, text, finalUrl: resp.url };
|
||||||
|
return browser_polyfill.runtime.sendNativeMessage ? nativeFetch(request.url, request.options) : { ok:false, status:0, error:"CORS_PERMISSION_NEEDED" };
|
||||||
|
})
|
||||||
|
.catch(() => browser_polyfill.runtime.sendNativeMessage ? nativeFetch(request.url, request.options) : { ok:false, error:"CORS_PERMISSION_NEEDED" });
|
||||||
|
});
|
||||||
|
```
|
||||||
|
|
||||||
|
`nativeFetch` usa `runtime.sendNativeMessage("application.id", { type:"fetchRequest", url, method, headers, body })` que solo está disponible en Safari (App‑bound messaging) — sirve para que la app nativa de Mac de Safari haga la petición y devuelva el cuerpo.
|
||||||
|
|
||||||
|
> **Implicación para YouTube:** la popup **nunca** enruta sus llamadas a YouTube por `fetchProxy` ni por `nativeFetch`. Las llamadas van por `globalThis.fetch` directo, protegidas por la regla DNR 9002.
|
||||||
|
|
||||||
|
`fetchProxy` se usa en el extractor de **Bilibili** (`BilibiliExtractor`):
|
||||||
|
- Llama a `https://api.bilibili.com/x/player/wbi/v2?bvid=…&cid=…` desde la popup.
|
||||||
|
- Bilibili suele devolver CORS abierto, así que normalmente no hace falta el proxy. El proxy queda como fallback de seguridad.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6 · Estructura del output (qué se mete en la nota Markdown)
|
||||||
|
|
||||||
|
`YoutubeExtractor.buildResult(transcript)` produce:
|
||||||
|
|
||||||
|
```js
|
||||||
|
{
|
||||||
|
content: ` <iframe … src="https://www.youtube.com/embed/{id}" …></iframe><p>{descripción}</p>{transcript.html}`,
|
||||||
|
contentHtml: idem,
|
||||||
|
extractedContent: { videoId, author },
|
||||||
|
variables: {
|
||||||
|
title: e.name || "", // videoDetails.title
|
||||||
|
author: r, // canal
|
||||||
|
site: "YouTube",
|
||||||
|
image: Array.isArray(e.thumbnailUrl) ? e.thumbnailUrl[0] : "",
|
||||||
|
published: e.uploadDate, // microformat.publishDate o uploadDate
|
||||||
|
description: n.slice(0, 200).trim(),
|
||||||
|
transcript: transcript.text, // sólo si hay
|
||||||
|
language: transcript.languageCode // "en", "es", "es-419", ...
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
`transcript.text` es texto plano con formato Markdown:
|
||||||
|
|
||||||
|
```
|
||||||
|
**00:00** · Primera frase hablada
|
||||||
|
|
||||||
|
**00:04** · Segunda frase hablada
|
||||||
|
```
|
||||||
|
|
||||||
|
`transcript.html` es:
|
||||||
|
|
||||||
|
```html
|
||||||
|
<div class="youtube transcript">
|
||||||
|
<h2>Transcript</h2>
|
||||||
|
<p class="transcript-segment"><strong><span class="timestamp" data-timestamp="0">00:00</span></strong> · Primera frase hablada</p>
|
||||||
|
…
|
||||||
|
</div>
|
||||||
|
```
|
||||||
|
|
||||||
|
Los `transcript-segment` con `data-timestamp` permiten que el modo *Reader* haga **highlight de la línea activa mientras reproduce** (`readerHighlightActiveLine`, `readerPinPlayer`, `readerAutoScroll` — ver `_locales/*/messages.json`).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7 · Manejo de errores y degradación
|
||||||
|
|
||||||
|
| Escenario | Comportamiento observado |
|
||||||
|
|---|---|
|
||||||
|
| Vídeo sin subtítulos manuales ni auto‑generados | `getCaptionTracks([])` ⇒ `pickCaptionTrack(undefined)` ⇒ `null` ⇒ `buildResult` se llama con `transcript = undefined` ⇒ `variables.transcript` queda ausente. El iframe y los metadatos sí se guardan. |
|
||||||
|
| Vídeo con subtítulos sólo auto (`kind:"asr"`) | `pickCaptionTrack` la prefiere si no hay otra o si la opción `language` coincide; en otro caso la salta. |
|
||||||
|
| `fetchPlayerData` falla en los 3 clientes y no hay JSON embebido | `transcript` queda `undefined`. La popup sigue mostrando la nota con metadatos. |
|
||||||
|
| Red lenta / timeout 4 s | `AbortSignal.timeout(4000)` corta la petición. Se prueba el siguiente cliente. |
|
||||||
|
| CORS aún bloqueado pese a DNR 9002 | En Chrome MV3 no hay fallback automático desde background (no usa `fetchProxy` para YouTube). El error queda silenciado dentro del `try { … } catch {}` y la nota se guarda sin transcripción. |
|
||||||
|
| Página no es `youtube.com` (p.ej. embed de otro dominio) | `canExtractAsync()` devuelve `false`; Defuddle cae a su extractor genérico (`BbcodeDataExtractor` según `ExtractorRegistry.mappings[last]`). |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8 · Datos estáticos y constantes que un LLM debe conocer
|
||||||
|
|
||||||
|
### 8.1 Versiones de cliente (claves de la firma)
|
||||||
|
|
||||||
|
| Variable | Valor | Notas |
|
||||||
|
|---|---|---|
|
||||||
|
| `ANDROID_CLIENT_VERSION` | `20.10.38` | Cliente ANDROID. |
|
||||||
|
| `IOS_CLIENT_VERSION` | `20.10.3` | Cliente IOS. |
|
||||||
|
| `WEB_CLIENT_VERSION` | `2.20240101.00.00` | Cliente WEB clásico. |
|
||||||
|
| `ANDROID_USER_AGENT` | `com.google.android.youtube/20.10.38 (Linux; U; Android 14)` | Solo se envía en el 2º intento. |
|
||||||
|
| `PLAYER_ENDPOINT` | `https://www.youtube.com/youtubei/v1/player?prettyPrint=false` | — |
|
||||||
|
| `NEXT_ENDPOINT` | `https://www.youtube.com/youtubei/v1/next?prettyPrint=false` | Solo `fetchChapters`. |
|
||||||
|
| `TIMEOUT_MS` | `4000` | `AbortSignal.timeout`. |
|
||||||
|
| `LANG_HEADER` | `Accept-Language` (opcional) | Se añade sólo si `options.language` está definido. |
|
||||||
|
|
||||||
|
### 8.2 Reglas DNR declaradas por background
|
||||||
|
|
||||||
|
| id | Nombre interno | Trigger | Acción | Uso |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| `9001` | `enableYouTubeEmbedRule` | `urlFilter:"||youtube.com/embed/"`, `resourceTypes:["sub_frame"]`, `tabIds:[<sender tab>]` | `set Referer: https://obsidian.md/` | Solo embeds (no transcripción). |
|
||||||
|
| `9002` | `enableYouTubeInnertubeRule` | `urlFilter:"||youtube.com/youtubei/"`, `resourceTypes:["xmlhttprequest"]`, `initiatorDomains:[chrome.runtime.id]` | `set Origin: https://www.youtube.com`, `set Referer: https://www.youtube.com/` | **Clave para que funcione la extracción.** |
|
||||||
|
|
||||||
|
### 8.3 Selectores DOM relevantes
|
||||||
|
|
||||||
|
| Uso | Selector |
|
||||||
|
|---|---|
|
||||||
|
| Panel de transcripción (desktop) | `ytd-engagement-panel-section-list-renderer[target-id="engagement-panel-searchable-transcript"] #segments-container` |
|
||||||
|
| Botón "Mostrar transcripción" | `ytd-video-description-transcript-section-renderer button` |
|
||||||
|
| Idioma seleccionado en panel | `ytd-engagement-panel-section-list-renderer[target-id="engagement-panel-searchable-transcript"] #footer yt-sort-filter-sub-menu-renderer yt-dropdown-menu button` |
|
||||||
|
| Panel mobile | `ytm-macro-markers-list-renderer .ytm-macro-markers-list-container` |
|
||||||
|
| Botón "Show more" mobile | `button[aria-label="Show more"]` |
|
||||||
|
| Botón "View all" mobile | `button[aria-label="View all"]` |
|
||||||
|
| Segmento desktop | `ytd-transcript-segment-renderer` (timestamp `.segment-timestamp`, texto `.segment-text`) |
|
||||||
|
| Segmento mobile | `transcript-segment-view-model` (timestamp `.ytwTranscriptSegmentViewModelTimestamp`, texto `span.yt-core-attributed-string`) |
|
||||||
|
| Metadatos (LD+JSON) | `script[type="application/ld+json"]` buscando `VideoObject` |
|
||||||
|
| OpenGraph | `meta[property="og:title|og:description|og:image|og:url"]` |
|
||||||
|
| Nombre de canal (DOM) | `[itemprop="name"]` / `link[itemprop="name"]` / `a, span` |
|
||||||
|
| JSON embebido | `<script>` con `ytInitialPlayerResponse` o `ytInitialData` |
|
||||||
|
|
||||||
|
### 8.4 Mensajes i18n ligados a la transcripción
|
||||||
|
|
||||||
|
`_locales/*/messages.json` contiene (en cada idioma) entradas como:
|
||||||
|
|
||||||
|
- `readerTranscripts` → "Transcripts" / "Transcripciones" / "Transcriptions" / …
|
||||||
|
- `readerHighlightActiveLine` / `…Description` → toggle de la línea activa
|
||||||
|
- `readerPinPlayer` / `…Description` → fija el reproductor
|
||||||
|
- `readerAutoScroll` / `…Description` → auto‑scroll durante reproducción
|
||||||
|
- `readerThemeSection` → tema de la transcripción en el Reader
|
||||||
|
|
||||||
|
(Solo configuran el render, no el proceso de extracción.)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9 · Diagrama textual del flujo de red
|
||||||
|
|
||||||
|
```
|
||||||
|
┌──────────────────────────────────────────────────────┐
|
||||||
|
│ popup.html / side-panel.html (chrome-extension://) │
|
||||||
|
│ ┌────────────────┐ │
|
||||||
|
│ │ popup.js │ │
|
||||||
|
│ │ Defuddle + │ │
|
||||||
|
│ │ YoutubeExtractor │
|
||||||
|
│ └───┬────────────┘ │
|
||||||
|
│ │ globalThis.fetch (XHR) │
|
||||||
|
│ ▼ │
|
||||||
|
│ https://www.youtube.com/youtubei/v1/player?… │
|
||||||
|
│ Body: { context: <IOS|ANDROID|WEB>, videoId } │
|
||||||
|
└──────────────────────────────────────────────────────┘
|
||||||
|
│ (DNR rule 9002)
|
||||||
|
│ set Origin: https://www.youtube.com
|
||||||
|
│ set Referer: https://www.youtube.com/
|
||||||
|
▼
|
||||||
|
YouTube InnerTube API
|
||||||
|
│ JSON con captionTracks[*].baseUrl
|
||||||
|
▼
|
||||||
|
https://www.youtube.com/api/timedtext?…&fmt=…&v=…&lang=… (track.baseUrl)
|
||||||
|
│ Headers: User-Agent: Mozilla/5.0
|
||||||
|
│ Accept-Language: <opción del usuario>
|
||||||
|
▼
|
||||||
|
XML de subtítulos
|
||||||
|
│ parseTranscriptXml()
|
||||||
|
▼
|
||||||
|
{ html, text, languageCode }
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
variables.transcript + variables.language
|
||||||
|
⇒ se inyecta en la nota Markdown
|
||||||
|
```
|
||||||
|
|
||||||
|
Paralelo: `fetchChapters` → `https://www.youtube.com/youtubei/v1/next?…` con `client: WEB` ⇒ `playerOverlays…markersMap` ⇒ capítulos embebidos como `## H2` dentro del HTML de la transcripción.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 10 · Riesgos y consideraciones para mantenimiento
|
||||||
|
|
||||||
|
1. **Versiones hard‑coded de cliente (`20.10.38`, `20.10.3`, `2.20240101.00.00`)**: YouTube rota las firmas mensualmente. Si la transcripción deja de funcionar, lo primero a actualizar son estas tres constantes en `popup.js` y `reader-page.js` (buscar `clientVersion` y `20.10.38`).
|
||||||
|
2. **DNR `9002` solo en `xmlhttprequest`**: si YouTube migra a `fetch` puro o cambia el `resourceType`, la regla deja de aplicar. Verificar `chrome.declarativeNetRequest.getEnabledRulesets()` en la consola de la extensión.
|
||||||
|
3. **`fetchProxy` NO se usa para YouTube**: no intentes enrutar por ahí; el flujo correcto es `globalThis.fetch` + DNR.
|
||||||
|
4. **El content script no extrae transcripción**: si la popup falla, no hay fallback desde `content.js`. Mejorar la extracción implica tocar `popup.js` y `reader-page.js` (que son el mismo bundle).
|
||||||
|
5. **Cache `Map<key, transcript>`**: existe en `BilibiliExtractor.transcriptCache` (LRU con cap 300). `YoutubeExtractor` **no** cachea transcripciones; cada `extract` rehace la red.
|
||||||
|
6. **Permisos**: el manifest pide `<all_urls>` y `declarativeNetRequest`. Si se reduce a un `optional_host_permissions` específico de YouTube, la popup seguirá funcionando porque `youtubei` está cubierto por `host_permissions`, pero el `webRequest` listener de Firefox puede dejar de aplicar.
|
||||||
|
7. **Time limit de service worker (MV3)**: el SW se duerme tras 30 s. `enableYouTubeInnertubeRule` se registra en `initialize()` al arrancar y se elimina solo si se pide `disableYouTubeInnertubeRule` (que no existe en el código actual). En la práctica, la regla es *session‑scoped* y dura lo que dure la sesión de Chrome.
|
||||||
|
8. **`ytInitialPlayerResponse` puede no estar presente** en páginas con cookie consent previo o si el usuario está en `consent.youtube.com`. La cascada ANDROID/IOS/WEB lo cubre.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 11 · Resumen para indexar/embeddings
|
||||||
|
|
||||||
|
> Texto generado para ser embeddings‑friendly. 7 frases autocontenidas.
|
||||||
|
|
||||||
|
1. La extensión **Obsidian Web Clipper 1.7.1** extrae la transcripción de YouTube desde la **popup** (contexto de la extensión, `chrome-extension://`), nunca desde el content script.
|
||||||
|
2. La clase **`YoutubeExtractor`** (minificada en `popup.js` y duplicada en `reader-page.js`, originalmente de Defuddle) ofrece una cascada de tres rutas: (a) parseo del JSON embebido `ytInitialPlayerResponse`, (b) scraping del panel de transcripción DOM con selectores `ytd-transcript-segment-renderer` / `transcript-segment-view-model`, (c) peticiones a la API privada **`youtubei/v1/player`** con contextos de cliente ANDROID / IOS / WEB.
|
||||||
|
3. El bypass de CORS / firma se hace con una regla **`declarativeNetRequest` id 9002** que fuerza `Origin: https://www.youtube.com` y `Referer: https://www.youtube.com/` para todas las XHR a `||youtube.com/youtubei/` iniciadas por la propia extensión; Firefox usa `webRequest.onBeforeSendHeaders` en su lugar.
|
||||||
|
4. La pista de subtítulos descargada (`track.baseUrl` con sufijo `fmt=`) es XML y se parsea con dos regex: `<p t="N">…<s>…</s>…</p>` (formato moderno) y `<text start="N">…</text>` (formato legacy), produciendo HTML con clase `transcript-segment` y texto plano con timestamps `HH:MM:SS`.
|
||||||
|
5. Los **capítulos** se extraen de `ytInitialData` o, en su defecto, de `youtubei/v1/next` con cliente WEB, parseando `playerOverlays.playerOverlayRenderer.decoratedPlayerBarRenderer.multiMarkersPlayerBarRenderer.markersMap`.
|
||||||
|
6. La popup usa `globalThis.fetch` directo (sin proxy) para YouTube; el `fetchProxy`/`nativeFetch` de `background.js` solo se utiliza como fallback CORS para Bilibili y otros hosts, no para YouTube.
|
||||||
|
7. El resultado (`{transcript.text, language}`) se inyecta en la nota Markdown final bajo la variable `transcript` / `language`; la nota incluye también un `<iframe src="https://www.youtube.com/embed/{id}">` y metadatos (`title`, `author`, `image`, `published`, `description`).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
*Fin del informe.*
|
||||||
+51
-5
@@ -1,20 +1,66 @@
|
|||||||
channel_url: "https://www.youtube.com/@Nostal-Vlad/videos"
|
channel_url: "https://www.youtube.com/@Nostal-Vlad/videos"
|
||||||
|
|
||||||
languages: ["es", "es-419", "en"]
|
# Per-language subtitle preference. Either a list (legacy) or a dict.
|
||||||
prefer_manual: true
|
# Dict form (preferred): per-lang mode. Modes are
|
||||||
|
# "manual" - ONLY human-uploaded captions; a video with just auto captions is
|
||||||
|
# skipped and stored as `no_subtitles`
|
||||||
|
# "auto" - ONLY YouTube-auto-generated captions
|
||||||
|
# "any" - try manual first (per `prefer_manual`), fall back to auto
|
||||||
|
# List form (legacy): `languages: ["es", "en"]` maps EVERY entry to the hard
|
||||||
|
# "manual"/"auto" mode picked by `prefer_manual` - note this means manual-ONLY
|
||||||
|
# when prefer_manual is true, which yields nothing on channels that publish
|
||||||
|
# only auto-generated captions. Prefer the dict form below.
|
||||||
|
languages:
|
||||||
|
es: any
|
||||||
|
es-419: any
|
||||||
|
en: any
|
||||||
|
prefer_manual: true # with mode "any": try manual first, then auto
|
||||||
include_shorts: false
|
include_shorts: false
|
||||||
include_live: true
|
include_live: true
|
||||||
min_duration_sec: 30
|
min_duration_sec: 30
|
||||||
|
|
||||||
|
# Pacing and what to do when YouTube starts refusing.
|
||||||
|
#
|
||||||
|
# YouTube publishes no rate limits, so none of these numbers are official. The
|
||||||
|
# backoff shape is: yt-dlp's own docs put a guest session at roughly 300 videos
|
||||||
|
# an hour (~1000 requests) before the bot wall, and Google documents truncated
|
||||||
|
# exponential backoff with jitter for its own APIs, which is what backoff_base
|
||||||
|
# and backoff_cap feed.
|
||||||
delay:
|
delay:
|
||||||
min_seconds: 1.5
|
min_seconds: 1.5 # randomised gap between videos / channels
|
||||||
max_seconds: 3.5
|
max_seconds: 3.5
|
||||||
backoff_base: 2.0
|
backoff_base: 2.0 # after a rate-limit: min(base * 2**n + jitter, cap)
|
||||||
backoff_cap: 60.0
|
backoff_cap: 60.0
|
||||||
|
throttle_threshold: 3 # consecutive rate-limit responses that abort the run
|
||||||
|
# Floor on the gap between ANY two requests to YouTube from this process,
|
||||||
|
# including the /api/tools/* endpoints that run outside the job runner.
|
||||||
|
# 0 = off. Raise it if you still hit the bot wall; it is the one setting that
|
||||||
|
# applies everywhere at once.
|
||||||
|
min_request_interval: 0.0
|
||||||
|
audio_rate_limit: 0 # bytes/sec ceiling for audio downloads, 0 = unlimited
|
||||||
|
|
||||||
yt_dlp:
|
yt_dlp:
|
||||||
retries: 10
|
retries: 10 # download retries (only bites on the audio path)
|
||||||
|
# Seconds between the individual HTTP calls inside one extraction — the watch
|
||||||
|
# page, the InnerTube player call, each continuation page of a channel tab.
|
||||||
|
# Forwarded to yt-dlp as `sleep_interval_requests`.
|
||||||
sleep_subrequests: 2
|
sleep_subrequests: 2
|
||||||
|
extractor_retries: 3 # retries during extraction (5xx and network only:
|
||||||
|
# yt-dlp's YouTube extractor never retries 403/429)
|
||||||
|
socket_timeout: 30.0 # without this a hung connection blocks the worker
|
||||||
|
|
||||||
|
# Incremental channel sync. Re-scanning a tracked channel reads the /videos tab
|
||||||
|
# newest-first and stops once it sees `overlap` videos already in the DB, so a
|
||||||
|
# routine sync costs one page instead of paginating the whole channel. If every
|
||||||
|
# video in the window turns out to be new, the window doubles (up to
|
||||||
|
# max_window) rather than silently missing uploads.
|
||||||
|
# Set incremental: false — or pass --full / tick "Full channel rescan" — to walk
|
||||||
|
# every page, which is only needed when the local catalog is incomplete.
|
||||||
|
sync:
|
||||||
|
incremental: true
|
||||||
|
window: 30 # entries read on the first pass
|
||||||
|
max_window: 300 # ceiling before reporting a truncated scan
|
||||||
|
overlap: 3 # consecutive known videos that prove we caught up
|
||||||
|
|
||||||
database_path: "data/state.db"
|
database_path: "data/state.db"
|
||||||
output_dir: "data/markdown"
|
output_dir: "data/markdown"
|
||||||
|
|||||||
+67
@@ -0,0 +1,67 @@
|
|||||||
|
channel_url: "https://www.youtube.com/@Nostal-Vlad/videos"
|
||||||
|
|
||||||
|
languages:
|
||||||
|
es: any
|
||||||
|
es-419: any
|
||||||
|
en: any
|
||||||
|
prefer_manual: true
|
||||||
|
include_shorts: false
|
||||||
|
include_live: true
|
||||||
|
min_duration_sec: 30
|
||||||
|
|
||||||
|
# --------------------------------------------------------------------------
|
||||||
|
# Ritmo contra YouTube.
|
||||||
|
#
|
||||||
|
# El unico techo publicado que existe es el de la wiki de yt-dlp: una sesion
|
||||||
|
# sin cuenta aguanta ~300 videos/hora (~1000 peticiones/hora) antes del muro
|
||||||
|
# de bot. Google no documenta limites para acceso sin API, asi que ese numero
|
||||||
|
# es la referencia menos mala. Los valores de abajo apuntan deliberadamente
|
||||||
|
# por debajo.
|
||||||
|
#
|
||||||
|
# Cuentas: un video cuesta 3 peticiones a youtube.com
|
||||||
|
# 1. la pagina /watch
|
||||||
|
# 2. la llamada InnerTube /youtubei/v1/player (separada de la anterior por
|
||||||
|
# sleep_subrequests)
|
||||||
|
# 3. /api/timedtext, la transcripcion
|
||||||
|
# (la miniatura va a i.ytimg.com, que es CDN estatico y no cuenta contra esto)
|
||||||
|
#
|
||||||
|
# Esas 3 son irreducibles: se midio si se podia bajar de ahi restringiendo
|
||||||
|
# player_client o saltando la pagina /watch, y no se puede. Con solo
|
||||||
|
# web_safari, o con player_skip=webpage, yt-dlp devuelve CERO pistas de
|
||||||
|
# subtitulos y pierde duration/channel_id/tags. El default de yt-dlp
|
||||||
|
# (android_vr + web_safari, 2 peticiones) ya es el optimo. No re-litigar.
|
||||||
|
#
|
||||||
|
# Medido con estos valores: 12.4 s/video => ~870 peticiones/hora, ~290
|
||||||
|
# videos/hora. Por debajo del techo, con margen.
|
||||||
|
#
|
||||||
|
# Si vas con prisa y aceptas el riesgo, baja delay.min/max_seconds y
|
||||||
|
# min_request_interval. Si aun asi te topas con el muro, subelos:
|
||||||
|
# min_request_interval es el unico ajuste que aplica a la vez a los jobs y
|
||||||
|
# a los /api/tools/*, que corren fuera del JobManager.
|
||||||
|
# --------------------------------------------------------------------------
|
||||||
|
delay:
|
||||||
|
min_seconds: 5.0
|
||||||
|
max_seconds: 10.0
|
||||||
|
backoff_base: 2.0 # tras un rate-limit: min(2 * 2**n + jitter, cap)
|
||||||
|
backoff_cap: 60.0
|
||||||
|
throttle_threshold: 3 # 3 rate-limits seguidos y el job para
|
||||||
|
min_request_interval: 2.5
|
||||||
|
audio_rate_limit: 0
|
||||||
|
|
||||||
|
yt_dlp:
|
||||||
|
retries: 10
|
||||||
|
sleep_subrequests: 3.0 # -> sleep_interval_requests de yt-dlp
|
||||||
|
extractor_retries: 3
|
||||||
|
socket_timeout: 30.0
|
||||||
|
|
||||||
|
sync:
|
||||||
|
incremental: true
|
||||||
|
window: 30
|
||||||
|
max_window: 300
|
||||||
|
overlap: 3
|
||||||
|
|
||||||
|
database_path: "data/state.db"
|
||||||
|
output_dir: "data/markdown"
|
||||||
|
template_path: "templates/video.md.j2"
|
||||||
|
|
||||||
|
filename_template: "{upload_date}_{slug}"
|
||||||
@@ -0,0 +1,308 @@
|
|||||||
|
# Discovery-Only Scrape Implementation Plan
|
||||||
|
|
||||||
|
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
||||||
|
|
||||||
|
**Goal:** Add one-channel and all-channel discovery jobs that register new YouTube videos as `pending` without downloading or processing their content.
|
||||||
|
|
||||||
|
**Architecture:** Extend the existing sequential `JobManager` with `opts.mode == "discover"`. Reuse `POST /api/scrape`, its SSE stream, and its job widget; add UI actions in Channels and Videos that pass either a channel ID or no channel ID. Keep the existing full scrape and processing endpoints unchanged.
|
||||||
|
|
||||||
|
**Tech Stack:** Python 3.10+, FastAPI, SQLite via the existing `Store`, yt-dlp flat discovery, Alpine.js CDN SPA, pytest.
|
||||||
|
|
||||||
|
## Global Constraints
|
||||||
|
|
||||||
|
- Discovery-only jobs must not call `process_video()`.
|
||||||
|
- Discovery-only jobs must not download transcripts, Markdown, audio, thumbnails, or avatars.
|
||||||
|
- All HTTP requests must use the existing job queue and SSE event stream.
|
||||||
|
- All database access must remain inside `Store` methods.
|
||||||
|
- Preserve existing user changes in the dirty worktree and edit only the files listed below.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Task 1: Make catalog insertion counts accurate
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `src/yt_scraper/store.py:266-283`
|
||||||
|
- Test: `tests/test_store_platform.py`
|
||||||
|
|
||||||
|
**Interfaces:**
|
||||||
|
- Consumes: existing `Store.upsert_videos(refs: list[VideoRef])` callers.
|
||||||
|
- Produces: the same integer return type, now equal to the number of distinct references that were not already present.
|
||||||
|
|
||||||
|
- [ ] **Step 1: Write the failing test**
|
||||||
|
|
||||||
|
Add this test after the existing store filtering tests:
|
||||||
|
|
||||||
|
```python
|
||||||
|
def test_upsert_videos_reports_only_new_rows(seeded_store):
|
||||||
|
inserted = seeded_store.upsert_videos([
|
||||||
|
VideoRef("v1", "UC1", "Updated title", "https://y/watch?v=v1", "20240101", 120),
|
||||||
|
VideoRef("v4", "UC1", "New video", "https://y/watch?v=v4", "20240501", 180),
|
||||||
|
])
|
||||||
|
|
||||||
|
assert inserted == 1
|
||||||
|
assert seeded_store.get_video("v1").title == "Updated title"
|
||||||
|
assert seeded_store.get_video("v4").status == "pending"
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 2: Run the focused test to verify it fails**
|
||||||
|
|
||||||
|
Run: `python -m pytest tests/test_store_platform.py::test_upsert_videos_reports_only_new_rows -q`
|
||||||
|
|
||||||
|
Expected: FAIL because SQLite's `ON CONFLICT DO UPDATE` currently makes `rowcount` nonzero for an existing row.
|
||||||
|
|
||||||
|
- [ ] **Step 3: Implement the minimal count fix**
|
||||||
|
|
||||||
|
Inside `Store.upsert_videos`, load the existing IDs for the incoming references before the loop, deduplicate the incoming IDs with a set, and increment `inserted` only when an incoming distinct ID was absent. Keep the current upsert SQL so existing rows still refresh title, upload date, and duration.
|
||||||
|
|
||||||
|
The essential implementation shape is:
|
||||||
|
|
||||||
|
```python
|
||||||
|
incoming = {r.video_id for r in refs}
|
||||||
|
with self._cursor() as cur:
|
||||||
|
existing = set()
|
||||||
|
if incoming:
|
||||||
|
placeholders = ",".join("?" for _ in incoming)
|
||||||
|
cur.execute(
|
||||||
|
f"SELECT video_id FROM videos WHERE video_id IN ({placeholders})",
|
||||||
|
list(incoming),
|
||||||
|
)
|
||||||
|
existing = {row["video_id"] for row in cur.fetchall()}
|
||||||
|
for r in refs:
|
||||||
|
# existing SQL remains here
|
||||||
|
if r.video_id not in existing:
|
||||||
|
inserted += 1
|
||||||
|
existing.add(r.video_id)
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 4: Run the focused and regression tests**
|
||||||
|
|
||||||
|
Run: `python -m pytest tests/test_store_platform.py -q`
|
||||||
|
|
||||||
|
Expected: PASS for all store tests.
|
||||||
|
|
||||||
|
- [ ] **Step 5: Commit**
|
||||||
|
|
||||||
|
Do not commit automatically in this workspace unless the user explicitly requests a commit. Leave the focused diff ready for review.
|
||||||
|
|
||||||
|
### Task 2: Add discovery-only job execution
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `src/yt_scraper/webapp/jobs.py:88-269`
|
||||||
|
- Test: `tests/test_webapp_jobs.py`
|
||||||
|
|
||||||
|
**Interfaces:**
|
||||||
|
- Consumes: `JobManager.enqueue(channel_id, opts)` with `opts={"mode": "discover"}` and optional `channel_id`.
|
||||||
|
- Produces: queued jobs whose SSE events contain per-channel discovery counts and whose terminal event contains a summary.
|
||||||
|
|
||||||
|
- [ ] **Step 1: Write failing job tests**
|
||||||
|
|
||||||
|
Append tests using a fake `discover_channel` and a fake `process_video`:
|
||||||
|
|
||||||
|
```python
|
||||||
|
def test_discovery_job_registers_new_videos_without_processing(tmp_path, monkeypatch):
|
||||||
|
store = Store(tmp_path / "state.db")
|
||||||
|
store.upsert_channel("UC1", "@alpha", "Alpha", 1)
|
||||||
|
store.upsert_videos([
|
||||||
|
VideoRef("old", "UC1", "Old", "https://y/watch?v=old", "20240101", 60)
|
||||||
|
])
|
||||||
|
store.mark_status("old", "done")
|
||||||
|
store.create_job("discover-job", "UC1", {"mode": "discover"})
|
||||||
|
calls = []
|
||||||
|
|
||||||
|
def fake_discover(url, sleep_subrequests=2.0):
|
||||||
|
calls.append(url)
|
||||||
|
return ("UC1", "Alpha", None, [
|
||||||
|
VideoRef("new", "UC1", "New", "https://y/watch?v=new", "20240501", 90),
|
||||||
|
VideoRef("old", "UC1", "Old", "https://y/watch?v=old", "20240101", 60),
|
||||||
|
])
|
||||||
|
|
||||||
|
processed = []
|
||||||
|
monkeypatch.setattr("yt_scraper.webapp.jobs.discover_channel", fake_discover)
|
||||||
|
monkeypatch.setattr("yt_scraper.webapp.jobs.process_video", lambda *a, **k: processed.append(a))
|
||||||
|
manager = JobManager(store, Config(database_path=str(tmp_path / "state.db"), output_dir=str(tmp_path / "markdown")))
|
||||||
|
|
||||||
|
manager._run_job("discover-job")
|
||||||
|
|
||||||
|
assert calls == ["https://www.youtube.com/@alpha/videos"]
|
||||||
|
assert store.get_video("new").status == "pending"
|
||||||
|
assert store.get_video("old").status == "done"
|
||||||
|
assert processed == []
|
||||||
|
assert store.get_job("discover-job").status == "done"
|
||||||
|
assert any(event["event"] == "done" and event["data"]["new_videos"] == 1
|
||||||
|
for event in manager.events_since("discover-job", 0))
|
||||||
|
|
||||||
|
|
||||||
|
def test_all_channel_discovery_continues_after_one_error(tmp_path, monkeypatch):
|
||||||
|
store = Store(tmp_path / "state.db")
|
||||||
|
store.upsert_channel("UC1", "@one", "One", 0)
|
||||||
|
store.upsert_channel("UC2", "@two", "Two", 0)
|
||||||
|
store.create_job("discover-all", None, {"mode": "discover"})
|
||||||
|
|
||||||
|
def fake_discover(url, sleep_subrequests=2.0):
|
||||||
|
if "@one" in url:
|
||||||
|
raise RuntimeError("temporary failure")
|
||||||
|
return ("UC2", "Two", None, [VideoRef("new2", "UC2", "New", "https://y/watch?v=new2")])
|
||||||
|
|
||||||
|
monkeypatch.setattr("yt_scraper.webapp.jobs.discover_channel", fake_discover)
|
||||||
|
manager = JobManager(store, Config(database_path=str(tmp_path / "state.db"), output_dir=str(tmp_path / "markdown")))
|
||||||
|
|
||||||
|
manager._run_job("discover-all")
|
||||||
|
|
||||||
|
assert store.get_video("new2").status == "pending"
|
||||||
|
assert store.get_job("discover-all").status == "done"
|
||||||
|
done = [e for e in manager.events_since("discover-all", 0) if e["event"] == "done"][-1]
|
||||||
|
assert done["data"]["errors"] == 1
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 2: Run the focused tests to verify they fail**
|
||||||
|
|
||||||
|
Run: `python -m pytest tests/test_webapp_jobs.py::test_discovery_job_registers_new_videos_without_processing tests/test_webapp_jobs.py::test_all_channel_discovery_continues_after_one_error -q`
|
||||||
|
|
||||||
|
Expected: FAIL because `_run_job` currently routes both cases to `_run_channel`, which tries to process pending videos and cannot interpret an all-channel discovery job.
|
||||||
|
|
||||||
|
- [ ] **Step 3: Implement discovery dispatch and execution**
|
||||||
|
|
||||||
|
In `JobManager._run_job`, route `opts.get("mode") == "discover"` to a new `_run_discovery(job_id)` before the existing audio/video/channel branches.
|
||||||
|
|
||||||
|
Implement `_run_discovery` with this behavior:
|
||||||
|
|
||||||
|
```python
|
||||||
|
def _run_discovery(self, job_id: str) -> None:
|
||||||
|
job = self.store.get_job(job_id)
|
||||||
|
if not job:
|
||||||
|
return
|
||||||
|
channels = ([self.store.get_channel(job.channel_id)] if job.channel_id
|
||||||
|
else self.store.list_channels())
|
||||||
|
channels = [c for c in channels if c]
|
||||||
|
self.store.update_job(job_id, status="running", total=len(channels), completed=0)
|
||||||
|
completed = 0
|
||||||
|
totals = {"new_videos": 0, "known_videos": 0, "errors": 0}
|
||||||
|
for channel in channels:
|
||||||
|
if job_id in self._cancel:
|
||||||
|
self.store.update_job(job_id, status="cancelled", completed=completed, finished=True)
|
||||||
|
self._emit(job_id, "cancelled", {"completed": completed, "total": len(channels)})
|
||||||
|
return
|
||||||
|
try:
|
||||||
|
# Build the stored handle/channel URL, call discover_channel, filter
|
||||||
|
# using cfg.include_shorts/cfg.include_live, upsert the channel and refs.
|
||||||
|
# Do not call deep_channel_avatar, cache helpers, or process_video.
|
||||||
|
new_count = self.store.upsert_videos(refs)
|
||||||
|
known_count = len({r.video_id for r in refs}) - new_count
|
||||||
|
totals["new_videos"] += new_count
|
||||||
|
totals["known_videos"] += known_count
|
||||||
|
self._emit(job_id, "progress", {"channel_id": channel["channel_id"], "new_videos": new_count, "known_videos": known_count, "completed": completed + 1, "total": len(channels)})
|
||||||
|
except Exception as exc:
|
||||||
|
totals["errors"] += 1
|
||||||
|
self._emit(job_id, "log", {"msg": f"discovery failed for {channel.get('name') or channel['channel_id']}: {exc}"})
|
||||||
|
completed += 1
|
||||||
|
self.store.update_job(job_id, completed=completed)
|
||||||
|
self.store.update_job(job_id, status="done", completed=completed, finished=True)
|
||||||
|
self._emit(job_id, "done", {"completed": completed, "total": len(channels), **totals})
|
||||||
|
```
|
||||||
|
|
||||||
|
Use the existing `_resolve_channel_url` for each stored channel. Preserve the existing processing path untouched. For a single-channel exception, emit an error terminal event/status; for all-channel jobs, continue and report the error count as above.
|
||||||
|
|
||||||
|
- [ ] **Step 4: Run focused job tests**
|
||||||
|
|
||||||
|
Run: `python -m pytest tests/test_webapp_jobs.py -q`
|
||||||
|
|
||||||
|
Expected: PASS for both existing audio tests and the new discovery tests.
|
||||||
|
|
||||||
|
- [ ] **Step 5: Run the full Python suite**
|
||||||
|
|
||||||
|
Run: `python -m pytest tests/ -q`
|
||||||
|
|
||||||
|
Expected: all tests pass with no network access.
|
||||||
|
|
||||||
|
### Task 3: Expose the discovery mode through the API
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `src/yt_scraper/webapp/api.py:111-117`
|
||||||
|
- Test: `tests/test_webapp_jobs.py` (job boundary coverage; no API test fixture exists in the repository).
|
||||||
|
|
||||||
|
**Interfaces:**
|
||||||
|
- Consumes: JSON `{"channel_id": "UC...", "opts": {"mode": "discover"}}` or the same body without `channel_id` for all channels.
|
||||||
|
- Produces: the existing `{"job_id": "..."}` response and a queued `JobManager` job.
|
||||||
|
|
||||||
|
- [ ] **Step 1: Add request-shape coverage at the job boundary**
|
||||||
|
|
||||||
|
Extend the discovery job tests to enqueue `manager.enqueue(None, {"mode": "discover"})` and verify that it completes as an all-channel job. This covers the exact payload shape the API forwards; the repository has no existing FastAPI router test fixture.
|
||||||
|
|
||||||
|
- [ ] **Step 2: Implement the minimal API change**
|
||||||
|
|
||||||
|
Keep the endpoint response unchanged. Normalize discovery options and permit a missing channel only for discovery:
|
||||||
|
|
||||||
|
```python
|
||||||
|
opts = (payload or {}).get("opts", {}) or {}
|
||||||
|
if opts.get("mode") == "discover":
|
||||||
|
opts = {"mode": "discover"}
|
||||||
|
job_id = jobs.enqueue(channel_id, opts)
|
||||||
|
return {"job_id": job_id}
|
||||||
|
```
|
||||||
|
|
||||||
|
Do not add a second synchronous discovery endpoint.
|
||||||
|
|
||||||
|
- [ ] **Step 3: Run the API/job regression tests**
|
||||||
|
|
||||||
|
Run: `python -m pytest tests/test_webapp_jobs.py tests/test_store_platform.py -q`
|
||||||
|
|
||||||
|
Expected: PASS.
|
||||||
|
|
||||||
|
### Task 4: Add one-channel and all-channel UI controls
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `src/yt_scraper/webapp/static/app.js:45-73, 758-815`
|
||||||
|
- Modify: `src/yt_scraper/webapp/static/index.html:167-229, 231-319`
|
||||||
|
|
||||||
|
**Interfaces:**
|
||||||
|
- Consumes: `POST /api/scrape` discovery jobs and SSE `progress`/`done` payloads.
|
||||||
|
- Produces: buttons for channel-specific and all-channel discovery, with refresh and result feedback.
|
||||||
|
|
||||||
|
- [ ] **Step 1: Add the failing static contract checks**
|
||||||
|
|
||||||
|
Before editing, use a lightweight repository check to confirm the new labels and method are absent:
|
||||||
|
|
||||||
|
Run: `rg "startDiscovery|Investigar nuevos|Investigar todos" src/yt_scraper/webapp/static`
|
||||||
|
|
||||||
|
Expected: no matches.
|
||||||
|
|
||||||
|
- [ ] **Step 2: Add the Alpine discovery action**
|
||||||
|
|
||||||
|
Add `startDiscovery(channelId)` near `startScrape()`. It posts `{channel_id: channelId || null, opts: {mode: "discover"}}`, subscribes with kind `discovery`, and prevents duplicate active jobs. Extend the SSE `done` handler to parse the payload and show `N video(s) nuevo(s) encontrado(s)` for discovery jobs. Refresh channels, videos, and dashboard after completion.
|
||||||
|
|
||||||
|
- [ ] **Step 3: Add controls to Channels and Videos**
|
||||||
|
|
||||||
|
In Channels, add the global header button and one row button calling `startDiscovery(c.channel_id)`. In Videos, add a header action calling `startDiscovery(filters.channel || null)`. Disable each while `jobActive()` is true and while the channel list is empty. Preserve `.md`, `Process`, and `Audio` actions.
|
||||||
|
|
||||||
|
- [ ] **Step 4: Verify the static contract and inspect the rendered paths**
|
||||||
|
|
||||||
|
Run: `rg "startDiscovery|Investigar nuevos|Investigar todos|mode: \"discover\"" src/yt_scraper/webapp/static`
|
||||||
|
|
||||||
|
Expected: matches in `app.js` and `index.html` for the method, payload, and both scopes. Inspect the changed Alpine expressions to ensure the existing `app.js` script remains before Alpine in `index.html`.
|
||||||
|
|
||||||
|
### Task 5: Final verification and review
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Review: `src/yt_scraper/store.py`
|
||||||
|
- Review: `src/yt_scraper/webapp/jobs.py`
|
||||||
|
- Review: `src/yt_scraper/webapp/api.py`
|
||||||
|
- Review: `src/yt_scraper/webapp/static/app.js`
|
||||||
|
- Review: `src/yt_scraper/webapp/static/index.html`
|
||||||
|
|
||||||
|
- [ ] **Step 1: Run the complete test suite**
|
||||||
|
|
||||||
|
Run: `python -m pytest tests/ -q`
|
||||||
|
|
||||||
|
Expected: all tests pass.
|
||||||
|
|
||||||
|
- [ ] **Step 2: Review the diff for scope and unintended downloads**
|
||||||
|
|
||||||
|
Run: `git diff -- src/yt_scraper/store.py src/yt_scraper/webapp/jobs.py src/yt_scraper/webapp/api.py src/yt_scraper/webapp/static/app.js src/yt_scraper/webapp/static/index.html tests/test_store_platform.py tests/test_webapp_jobs.py`
|
||||||
|
|
||||||
|
Confirm discovery code contains no calls to `process_video`, `cache_thumbnail`, `cache_channel_avatar`, `extract_video`, or audio routes.
|
||||||
|
|
||||||
|
- [ ] **Step 3: Check worktree status**
|
||||||
|
|
||||||
|
Run: `git status --short`
|
||||||
|
|
||||||
|
Confirm only intended feature files and the two planning documents are changed or untracked; do not modify or revert unrelated existing user changes.
|
||||||
@@ -0,0 +1,250 @@
|
|||||||
|
# Vista Grid estilo YouTube (sección Vídeos) — Implementation Plan
|
||||||
|
|
||||||
|
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
||||||
|
|
||||||
|
**Goal:** Añadir a la sección Vídeos un conmutador tabla ⇄ grid donde el grid replica el look de YouTube (tarjetas 16:9 con miniatura, duración, título y metadatos) manteniendo filtros, orden, selección masiva y paginación compartidos.
|
||||||
|
|
||||||
|
**Architecture:** Solo frontend. El estado `videos.view` (`'table'|'grid'`) vive en el componente Alpine único `window.platform()`; la elección se persiste en `localStorage["videos-view"]`. En `index.html` los dos layouts son bloques hermanos alternados con `<template x-if>`; el toggle es una fila fina sobre ellos.
|
||||||
|
|
||||||
|
**Tech Stack:** Alpine 3 + Tailwind CDN (sin build step), estáticos servidos por FastAPI.
|
||||||
|
|
||||||
|
**Spec:** [docs/superpowers/specs/2026-08-22-videos-grid-view-design.md](../specs/2026-08-22-videos-grid-view-design.md)
|
||||||
|
|
||||||
|
## Global Constraints
|
||||||
|
|
||||||
|
- Sin build step: todo cambio es HTML/JS/CSS estático editado a mano ([.opencode/agent/webapp-builder.md](../../.opencode/agent/webapp-builder.md)).
|
||||||
|
- Todo el estado vive en **un solo** componente Alpine, `window.platform()` en `src/yt_scraper/webapp/static/app.js`; no crear componentes nuevos.
|
||||||
|
- El orden de `<script defer>` en `index.html` (app.js antes de alpinejs) es carga funcional: no reordenar scripts.
|
||||||
|
- No tocar backend, endpoints ni `store.py`: `/api/videos` ya devuelve todos los campos que necesita la tarjeta.
|
||||||
|
- Sin tests automatizados de frontend (el repo no tiene infra JS): verificación = `node --check` + smoke manual con `start-server.bat`.
|
||||||
|
- Estilo visual existente: dark "command center", `glass`, `accent-grad btn-primary` como tratamiento activo (mismo patrón que la paginación).
|
||||||
|
- Sin acciones por tarjeta (`.md`/`Clip` quedan en tabla y detalle).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Task 1: Estado y persistencia en app.js
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `src/yt_scraper/webapp/static/app.js` (línea 39 store `videos`; `init()` líneas 92-93; método nuevo junto a los helpers de selección ~línea 463)
|
||||||
|
|
||||||
|
**Interfaces:**
|
||||||
|
- Consumes: nada nuevo.
|
||||||
|
- Produces: `this.videos.view` (`'table' | 'grid'`, default `'table'`) y `setVideoView(v)` — el HTML del Task 2/3 los usa vía `videos.view` / `@click="setVideoView('grid')"`.
|
||||||
|
|
||||||
|
- [ ] **Step 1: Añadir `view` al store**
|
||||||
|
|
||||||
|
En `app.js:39`, cambiar:
|
||||||
|
|
||||||
|
```js
|
||||||
|
videos: { items: [], total: 0, page: 1, size: 25, selected: [] },
|
||||||
|
```
|
||||||
|
|
||||||
|
por:
|
||||||
|
|
||||||
|
```js
|
||||||
|
videos: { items: [], total: 0, page: 1, size: 25, selected: [], view: "table" },
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 2: Restaurar la elección en init()**
|
||||||
|
|
||||||
|
En `app.js` dentro de `init()`, justo después del bloque:
|
||||||
|
|
||||||
|
```js
|
||||||
|
const size = Number(localStorage.getItem("videos-size"));
|
||||||
|
if ([10, 25, 50, 100].includes(size)) this.videos.size = size;
|
||||||
|
```
|
||||||
|
|
||||||
|
añadir:
|
||||||
|
|
||||||
|
```js
|
||||||
|
const savedView = localStorage.getItem("videos-view");
|
||||||
|
if (savedView === "table" || savedView === "grid") this.videos.view = savedView;
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 3: Añadir setVideoView()**
|
||||||
|
|
||||||
|
En `app.js`, inmediatamente antes de `toggleSelect(id) {` (la sección de helpers de selección, ~línea 463), añadir:
|
||||||
|
|
||||||
|
```js
|
||||||
|
// View mode of the Videos section ('table' | 'grid'). UI preference like
|
||||||
|
// videos-size: persisted in localStorage, never in the URL.
|
||||||
|
setVideoView(v) {
|
||||||
|
if (v !== "table" && v !== "grid") return;
|
||||||
|
this.videos.view = v;
|
||||||
|
try { localStorage.setItem("videos-view", v); } catch (_) {}
|
||||||
|
},
|
||||||
|
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 4: Verificar sintaxis JS**
|
||||||
|
|
||||||
|
Run: `node --check src/yt_scraper/webapp/static/app.js`
|
||||||
|
Expected: sin salida (exit 0). Si `node` no está disponible, verificar abriendo la webapp en el Task 4.
|
||||||
|
|
||||||
|
- [ ] **Step 5: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add src/yt_scraper/webapp/static/app.js
|
||||||
|
git commit -m "feat(webapp): estado y persistencia de vista tabla/grid en Videos"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Task 2: Conmutador de vista + alternancia x-if en index.html
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `src/yt_scraper/webapp/static/index.html` (contenedor de resultados, líneas 301-352)
|
||||||
|
|
||||||
|
**Interfaces:**
|
||||||
|
- Consumes: `videos.view` y `setVideoView(v)` del Task 1.
|
||||||
|
- Produces: la estructura de dos bloques `<template x-if>` que el Task 3 completa con el grid (este task deja la tabla funcional tal cual).
|
||||||
|
|
||||||
|
- [ ] **Step 1: Envolver la tabla en template x-if**
|
||||||
|
|
||||||
|
En `index.html`, la tabla vive en:
|
||||||
|
|
||||||
|
```html
|
||||||
|
<div class="glass overflow-hidden">
|
||||||
|
<table class="tbl">
|
||||||
|
```
|
||||||
|
|
||||||
|
(que cierra con `</table>` + `</div>` en las líneas ~351-352, justo antes del comentario `<!-- pagination -->`). Envolver ese `<div class="glass overflow-hidden">…</div>` completo en:
|
||||||
|
|
||||||
|
```html
|
||||||
|
<template x-if="videos.view==='table'">
|
||||||
|
<div class="glass overflow-hidden">
|
||||||
|
… contenido actual de la tabla SIN CAMBIOS …
|
||||||
|
</div>
|
||||||
|
</template>
|
||||||
|
```
|
||||||
|
|
||||||
|
Reindentar una nivel el interior. Ningún atributo ni clase de la tabla cambia.
|
||||||
|
|
||||||
|
- [ ] **Step 2: Insertar la fila del toggle**
|
||||||
|
|
||||||
|
Justo encima del `<template x-if="videos.view==='table'">` recién creado, añadir:
|
||||||
|
|
||||||
|
```html
|
||||||
|
<!-- table ⇄ grid switch -->
|
||||||
|
<div class="flex justify-end">
|
||||||
|
<div class="inline-flex rounded-lg border border-zinc-800 bg-zinc-900/60 p-1 gap-1" role="group" aria-label="View mode">
|
||||||
|
<button type="button" class="btn !py-1 !px-2" :class="videos.view==='table' ? 'accent-grad btn-primary' : 'btn-ghost'" :aria-pressed="(videos.view==='table').toString()" @click="setVideoView('table')" title="Table view">
|
||||||
|
<svg class="w-4 h-4" viewBox="0 0 24 24" fill="currentColor"><path d="M3 5h18v2H3zm0 6h18v2H3zm0 6h18v2H3z"/></svg>
|
||||||
|
</button>
|
||||||
|
<button type="button" class="btn !py-1 !px-2" :class="videos.view==='grid' ? 'accent-grad btn-primary' : 'btn-ghost'" :aria-pressed="(videos.view==='grid').toString()" @click="setVideoView('grid')" title="Grid view">
|
||||||
|
<svg class="w-4 h-4" viewBox="0 0 24 24" fill="currentColor"><path d="M3 3h8v8H3zm10 0h8v8h-8zM3 13h8v8H3zm10 0h8v8h-8z"/></svg>
|
||||||
|
</button>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 3: Smoke manual**
|
||||||
|
|
||||||
|
Run: `start-server.bat`, abrir Vídeos.
|
||||||
|
Expected: la tabla se ve idéntica a antes; el toggle aparece arriba a la derecha; pulsar el icono de grid NO cambia nada visible todavía pero el botón grid queda activo (accent), se refresca la página y sigue activo (localStorage); volver a tabla también persiste.
|
||||||
|
|
||||||
|
- [ ] **Step 4: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add src/yt_scraper/webapp/static/index.html
|
||||||
|
git commit -m "feat(webapp): conmutador de vista tabla/grid en Videos"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Task 3: Bloque grid de tarjetas estilo YouTube
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `src/yt_scraper/webapp/static/index.html` (insertar tras el cierre `</template>` del bloque tabla, antes de `<!-- pagination -->`)
|
||||||
|
|
||||||
|
**Interfaces:**
|
||||||
|
- Consumes: `videos.view`, `setVideoView` (Task 1); helpers existentes de `app.js`: `ts(sec)`, `fmtNum(n)`, `chanName(id)`, `videoDate(v)`, `videoDateTitle(v)`, `statusClass(s)`, `isBlocked(v)`, `blockLabel(v)`, `blockTitle(v)`, `isSelected(id)`, `toggleSelect(id)`, `openVideo(id)`, `loading.videos`, `videos.items`.
|
||||||
|
- Produces: la vista grid completa (fin de feature).
|
||||||
|
|
||||||
|
- [ ] **Step 1: Insertar el bloque grid**
|
||||||
|
|
||||||
|
Entre el `</template>` que cierra el bloque tabla y `<!-- pagination -->`, añadir:
|
||||||
|
|
||||||
|
```html
|
||||||
|
<template x-if="videos.view==='grid'">
|
||||||
|
<div>
|
||||||
|
<template x-if="loading.videos"><div class="text-center text-zinc-500 py-10">loading…</div></template>
|
||||||
|
<template x-if="!loading.videos && videos.items.length===0"><div class="text-center text-zinc-500 py-10"><div>No videos match these filters.</div><button class="btn btn-ghost mt-3" @click.stop="filters={ channel:'', status:'', from:'', to:'', min_dur:'', q:'', sort:'upload_date' }; loadVideos(1)">Clear filters</button></div></template>
|
||||||
|
<template x-if="!loading.videos && videos.items.length>0">
|
||||||
|
<div class="grid grid-cols-2 sm:grid-cols-3 lg:grid-cols-4 2xl:grid-cols-5 gap-4">
|
||||||
|
<template x-for="v in videos.items" :key="v.video_id">
|
||||||
|
<div class="glass group relative cursor-pointer hover:border-rose-500/40 transition-colors overflow-hidden"
|
||||||
|
:class="isSelected(v.video_id) ? 'row-selected' : ''"
|
||||||
|
role="button" tabindex="0"
|
||||||
|
@click="openVideo(v.video_id)"
|
||||||
|
@keydown.enter.prevent="openVideo(v.video_id)">
|
||||||
|
<div class="relative aspect-video bg-zinc-900">
|
||||||
|
<img class="absolute inset-0 w-full h-full object-cover" :src="'/api/thumbnails/'+v.video_id" :alt="v.title" onerror="this.style.visibility='hidden'" />
|
||||||
|
<span x-show="v.duration" class="absolute bottom-1.5 right-1.5 font-mono text-[0.7rem] leading-none px-1.5 py-1 rounded bg-black/80 text-white" x-text="ts(v.duration)"></span>
|
||||||
|
<button type="button" @click.stop="toggleSelect(v.video_id)"
|
||||||
|
class="absolute top-1.5 left-1.5 w-6 h-6 rounded-md border flex items-center justify-center transition-opacity"
|
||||||
|
:class="isSelected(v.video_id) ? 'bg-rose-500 border-rose-400 opacity-100' : 'bg-black/70 border-zinc-300/70 opacity-0 group-hover:opacity-100'"
|
||||||
|
:aria-pressed="isSelected(v.video_id).toString()" title="Select video">
|
||||||
|
<svg x-show="isSelected(v.video_id)" class="w-4 h-4 text-white" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="3"><path stroke-linecap="round" stroke-linejoin="round" d="M5 13l4 4L19 7"/></svg>
|
||||||
|
</button>
|
||||||
|
</div>
|
||||||
|
<div class="p-3 space-y-1.5">
|
||||||
|
<div class="text-sm font-medium text-zinc-100 line-clamp-2 leading-snug" x-text="v.title" :title="v.title"></div>
|
||||||
|
<div class="text-xs text-zinc-400 truncate" x-text="chanName(v.channel_id)" :title="chanName(v.channel_id)"></div>
|
||||||
|
<div class="flex items-center gap-2 font-mono text-xs text-zinc-400 flex-wrap">
|
||||||
|
<span x-text="fmtNum(v.view_count)"></span>
|
||||||
|
<span class="text-zinc-600">·</span>
|
||||||
|
<span :class="v.upload_date ? '' : 'italic text-zinc-600'" x-text="videoDate(v)" :title="videoDateTitle(v)"></span>
|
||||||
|
</div>
|
||||||
|
<div class="flex items-center gap-1 flex-wrap pt-0.5">
|
||||||
|
<span class="pill" :class="statusClass(v.status)" x-text="v.status" :title="v.error_msg || v.status"></span>
|
||||||
|
<span x-show="isBlocked(v)" class="pill st-locked" :title="blockTitle(v)">
|
||||||
|
<svg class="w-3 h-3" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" aria-hidden="true"><rect x="4" y="10" width="16" height="10" rx="2"/><path d="M8 10V7a4 4 0 1 1 8 0v3"/></svg>
|
||||||
|
<span x-text="blockLabel(v)"></span>
|
||||||
|
</span>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</template>
|
||||||
|
</div>
|
||||||
|
</template>
|
||||||
|
</div>
|
||||||
|
</template>
|
||||||
|
```
|
||||||
|
|
||||||
|
Notas de diseño ya decididas:
|
||||||
|
- `.row-selected` existe en styles.css:179 con `!important` → aplica igual sobre la tarjeta div que sobre el `<tr>`.
|
||||||
|
- `line-clamp-2` ya se usa en index.html:553 → Tailwind CDN lo resuelve.
|
||||||
|
- Checkbox estilo YouTube: aparece al hover (`group-hover`) o permanece si está seleccionada; click con `.stop` para no abrir el detalle.
|
||||||
|
- La selección masiva, barra bulk, paginación y filtros NO se tocan: viven fuera de estos bloques.
|
||||||
|
|
||||||
|
- [ ] **Step 2: Verificación manual completa (checklist del spec)**
|
||||||
|
|
||||||
|
Run: `start-server.bat`
|
||||||
|
|
||||||
|
1. Alternar tabla ⇄ grid: ambas muestran los mismos vídeos con el filtro activo.
|
||||||
|
2. En grid: pasar el ratón sobre una tarjeta → checkbox visible; marcar 3 → barra bulk aparece con "3 selected" → "Download .md" encola el job.
|
||||||
|
3. Click en tarjeta (fuera del checkbox) → abre el detalle; Back → vuelve a grid conservando página.
|
||||||
|
4. Paginar en grid: Prev/Next/números funcionan, scroll arriba, sin repeticiones.
|
||||||
|
5. Recargar página: la vista elegida se conserva; borrar `localStorage["videos-view"]` → cae a tabla.
|
||||||
|
6. Tarjeta de vídeo blocked: pill candado con label; badge de duración visible en tarjetas con duración.
|
||||||
|
7. Miniatura rota: `onerror` la oculta sin romper el layout (contenedor aspect-video mantiene proporción).
|
||||||
|
8. La tabla sigue funcionando exactamente igual que antes.
|
||||||
|
|
||||||
|
Run: `python -m pytest tests/ -q`
|
||||||
|
Expected: suite verde (backend intacto; sanity check barato).
|
||||||
|
|
||||||
|
- [ ] **Step 3: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add src/yt_scraper/webapp/static/index.html
|
||||||
|
git commit -m "feat(webapp): grid de tarjetas estilo YouTube en Videos"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Self-review hecho
|
||||||
|
|
||||||
|
- **Cobertura del spec:** estado+persistencia (Task 1), toggle (Task 2), tarjetas/clamps/badges/pills/loading/empty (Task 3), verificación manual ítem a ítem (Task 3 Step 2). Fuera de alcance respetado: cero cambios backend, sin acciones por tarjeta.
|
||||||
|
- **Placeholders:** ninguno; todo step lleva código literal o comando exacto.
|
||||||
|
- **Consistencia de nombres:** `videos.view`, `setVideoView`, `"videos-view"` usados igual en Tasks 1-3; helpers referenciados existen en app.js (verificado: ts:1265, fmtNum:1293, videoDate:1317, videoDateTitle:1324, chanName:1345, blockLabel:1355, blockTitle:1365, isBlocked:1374, statusClass:1376, toggleSelect:464, isSelected:469).
|
||||||
@@ -0,0 +1,47 @@
|
|||||||
|
# Discovery-Only Scrape Design
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
|
||||||
|
Add manual discovery controls for one channel or all tracked channels. Discovery checks the complete recent video list returned by YouTube, registers unknown videos as `pending`, and does not download transcripts, Markdown, audio, thumbnails, or avatars.
|
||||||
|
|
||||||
|
## Existing Context
|
||||||
|
|
||||||
|
The webapp already has a single-worker `JobManager`, SSE job events, and a `POST /api/scrape` endpoint. The current channel job combines discovery with `pipeline.process_video()`, so every pending video is immediately extracted and rendered. The channels table has per-channel actions, while the videos view can be filtered to one channel.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
- Extend the existing job dispatch with `opts.mode == "discover"`.
|
||||||
|
- `channel_id` selects one tracked channel. A missing `channel_id` means all tracked channels for discovery jobs only.
|
||||||
|
- Discovery uses the existing `discover_channel()` call without a playlist cap. It compares the returned IDs with the SQLite catalog and upserts only catalog metadata. It never calls `process_video()`.
|
||||||
|
- The existing scrape mode remains unchanged for users who want discovery plus transcript/Markdown processing.
|
||||||
|
- All discovery jobs stay in the existing sequential queue and use the existing SSE stream, cancellation, progress, and history.
|
||||||
|
|
||||||
|
## Persistence and Results
|
||||||
|
|
||||||
|
- New `VideoRef` rows are inserted with the existing default status `pending`.
|
||||||
|
- Existing video rows retain their status and processed data; rediscovery only refreshes title, upload date, and duration through the existing upsert behavior.
|
||||||
|
- The channel record is refreshed with its name, handle, catalog count, and `last_scraped`; no avatar fetch/cache is performed by discovery-only jobs.
|
||||||
|
- A per-channel progress event reports `new_videos`, `known_videos`, and the channel ID.
|
||||||
|
- The terminal event reports total channels, discovered entries, new videos, known videos, and errors.
|
||||||
|
- A one-channel discovery failure marks the job `error`. An all-channel job continues after individual failures and finishes with a summary so one broken channel does not prevent other channels from being scanned.
|
||||||
|
|
||||||
|
## UI
|
||||||
|
|
||||||
|
- Channels view: add `Investigar todos` in the header and `Investigar` in each channel row.
|
||||||
|
- Videos view: add an investigation button in the header. It investigates the selected channel when `filters.channel` is set, otherwise all channels.
|
||||||
|
- Buttons use the existing job widget/SSE stream, are disabled while an active job is being followed, and show the number of new videos on completion.
|
||||||
|
- Completion refreshes channels, videos, and dashboard data. Existing `.md`, `Process`, and `Audio` controls remain separate.
|
||||||
|
|
||||||
|
## Error Handling
|
||||||
|
|
||||||
|
- Empty channel catalogs complete successfully with zero channels scanned.
|
||||||
|
- Discovery errors are emitted in the job log and do not invoke video processing.
|
||||||
|
- A cancellation checks the existing cancellation set between channels and leaves already inserted catalog rows intact.
|
||||||
|
- Duplicate IDs from a discovery response are counted once by the catalog comparison.
|
||||||
|
|
||||||
|
## Testing
|
||||||
|
|
||||||
|
- Store tests verify that upserting an existing and a new reference reports only the new reference.
|
||||||
|
- Job tests fake `discover_channel()`, verify new rows are `pending`, existing rows retain their status, and `process_video()` is never called.
|
||||||
|
- Job tests verify all-channel discovery continues after one channel fails and reports the successful channel.
|
||||||
|
- The full existing Python test suite remains the regression check. The static UI is verified by code inspection and the existing webapp smoke path; no frontend build step exists.
|
||||||
@@ -1 +1,44 @@
|
|||||||
{"0": "Webapp Frontend (JS)", "1": "Transcript Extraction & Parsing", "2": "Config & Discovery Pipeline", "3": "Store & Export Layer", "4": "CLI Command Layer", "5": "Platform Design & Proposals", "6": "Segments Search & Backfill", "7": "Store Platform Tests", "8": "Content Analysis", "9": "README & Config Docs", "10": "Server Start Script", "11": "Package Manifest", "12": "Server Stop Script", "13": "Webapp Package Init"}
|
{
|
||||||
|
"0": "Webapp Frontend (JS)",
|
||||||
|
"1": "Transcript Extraction & Parsing",
|
||||||
|
"2": "Config & Discovery Pipeline",
|
||||||
|
"3": "Store & Export Layer",
|
||||||
|
"4": "CLI Command Layer",
|
||||||
|
"5": "Platform Design & Proposals",
|
||||||
|
"6": "Segments Search & Backfill",
|
||||||
|
"7": "Store Platform Tests",
|
||||||
|
"8": "Content Analysis",
|
||||||
|
"9": "README & Config Docs",
|
||||||
|
"10": "Server Start Script",
|
||||||
|
"11": "Package Manifest",
|
||||||
|
"12": "Server Stop Script",
|
||||||
|
"13": "Webapp Package Init",
|
||||||
|
"14": "test_extract.py",
|
||||||
|
"15": "Auditoría · Cómo la extensión obtiene la transcripción / información de vídeos de YouTube",
|
||||||
|
"16": "api",
|
||||||
|
"17": "pipeline.py",
|
||||||
|
"18": "test_blocked_videos.py",
|
||||||
|
"19": "config.py",
|
||||||
|
"20": "setView",
|
||||||
|
"21": "Discovery-Only Scrape Design",
|
||||||
|
"22": "renderDashboardCharts",
|
||||||
|
"23": "Global Constraints",
|
||||||
|
"24": "Config",
|
||||||
|
"25": "init",
|
||||||
|
"26": "render.py",
|
||||||
|
"27": "loadChannels",
|
||||||
|
"28": "jobs.py",
|
||||||
|
"29": "Segment",
|
||||||
|
"30": "cookies.py",
|
||||||
|
"31": "discover.py",
|
||||||
|
"32": "export.py",
|
||||||
|
"33": "pipeline.py",
|
||||||
|
"34": "_channel_targets",
|
||||||
|
"35": "test_features.py",
|
||||||
|
"36": "VideoRow",
|
||||||
|
"37": "test_since_filter.py",
|
||||||
|
"38": ".reset_videos",
|
||||||
|
"39": "._connect",
|
||||||
|
"40": "re_render_cmd",
|
||||||
|
"41": "handleFiles"
|
||||||
|
}
|
||||||
|
|||||||
@@ -1 +1 @@
|
|||||||
D:\yt-channel-scraper
|
.
|
||||||
@@ -0,0 +1,27 @@
|
|||||||
|
{
|
||||||
|
"0": "Webapp Frontend (JS)",
|
||||||
|
"1": "Transcript Extraction & Parsing",
|
||||||
|
"2": "Config & Discovery Pipeline",
|
||||||
|
"3": "Store & Export Layer",
|
||||||
|
"4": "CLI Command Layer",
|
||||||
|
"5": "Platform Design & Proposals",
|
||||||
|
"6": "Segments Search & Backfill",
|
||||||
|
"7": "Store Platform Tests",
|
||||||
|
"8": "Content Analysis",
|
||||||
|
"9": "README & Config Docs",
|
||||||
|
"10": "Server Start Script",
|
||||||
|
"11": "Package Manifest",
|
||||||
|
"12": "Server Stop Script",
|
||||||
|
"13": "Webapp Package Init",
|
||||||
|
"14": "test_extract.py",
|
||||||
|
"15": "Auditoría · Cómo la extensión obtiene la transcripción / información de vídeos de YouTube",
|
||||||
|
"16": "api",
|
||||||
|
"17": "pipeline.py",
|
||||||
|
"18": "render.py",
|
||||||
|
"19": "config.py",
|
||||||
|
"20": "setView",
|
||||||
|
"21": "Discovery-Only Scrape Design",
|
||||||
|
"22": "renderDashboardCharts",
|
||||||
|
"23": "Global Constraints",
|
||||||
|
"24": "loadJobs"
|
||||||
|
}
|
||||||
@@ -0,0 +1,182 @@
|
|||||||
|
# Graph Report - yt-channel-scraper (2026-07-31)
|
||||||
|
|
||||||
|
## Corpus Check
|
||||||
|
- 45 files · ~46,603 words
|
||||||
|
- Verdict: corpus is large enough that graph structure adds value.
|
||||||
|
|
||||||
|
## Summary
|
||||||
|
- 612 nodes · 1392 edges · 25 communities (22 shown, 3 thin omitted)
|
||||||
|
- Extraction: 91% EXTRACTED · 9% INFERRED · 0% AMBIGUOUS · INFERRED: 129 edges (avg confidence: 0.78)
|
||||||
|
- Token cost: 0 input · 0 output
|
||||||
|
|
||||||
|
## Graph Freshness
|
||||||
|
- Built from commit: `06497299`
|
||||||
|
- Run `git rev-parse HEAD` and compare to check if the graph is stale.
|
||||||
|
- Run `graphify update .` after code changes (no API cost).
|
||||||
|
|
||||||
|
## Community Hubs (Navigation)
|
||||||
|
- Webapp Frontend (JS)
|
||||||
|
- Transcript Extraction & Parsing
|
||||||
|
- Config & Discovery Pipeline
|
||||||
|
- Store & Export Layer
|
||||||
|
- CLI Command Layer
|
||||||
|
- Platform Design & Proposals
|
||||||
|
- Segments Search & Backfill
|
||||||
|
- Store Platform Tests
|
||||||
|
- Content Analysis
|
||||||
|
- README & Config Docs
|
||||||
|
- Server Start Script
|
||||||
|
- Package Manifest
|
||||||
|
- Server Stop Script
|
||||||
|
- test_extract.py
|
||||||
|
- Auditoría · Cómo la extensión obtiene la transcripción / información de vídeos de YouTube
|
||||||
|
- api
|
||||||
|
- pipeline.py
|
||||||
|
- render.py
|
||||||
|
- config.py
|
||||||
|
- setView
|
||||||
|
- Discovery-Only Scrape Design
|
||||||
|
- renderDashboardCharts
|
||||||
|
- Global Constraints
|
||||||
|
- loadJobs
|
||||||
|
|
||||||
|
## God Nodes (most connected - your core abstractions)
|
||||||
|
1. `Store` - 96 edges
|
||||||
|
2. `api()` - 36 edges
|
||||||
|
3. `toast()` - 36 edges
|
||||||
|
4. `Config` - 34 edges
|
||||||
|
5. `Segment` - 32 edges
|
||||||
|
6. `JobManager` - 26 edges
|
||||||
|
7. `VideoRef` - 25 edges
|
||||||
|
8. `discover_incremental()` - 21 edges
|
||||||
|
9. `process_video()` - 19 edges
|
||||||
|
10. `subscribeJob()` - 18 edges
|
||||||
|
|
||||||
|
## Surprising Connections (you probably didn't know these)
|
||||||
|
- `test_config_languages_defaults_to_dict()` --calls--> `Config` [INFERRED]
|
||||||
|
tests/test_extract.py → src/yt_scraper/config.py
|
||||||
|
- `test_config_prefer_manual_defaults_true()` --calls--> `Config` [INFERRED]
|
||||||
|
tests/test_extract.py → src/yt_scraper/config.py
|
||||||
|
- `test_dashboard_aggregates()` --calls--> `Segment` [INFERRED]
|
||||||
|
tests/test_store_platform.py → src/yt_scraper/parse.py
|
||||||
|
- `test_store_and_search_segments()` --calls--> `Segment` [INFERRED]
|
||||||
|
tests/test_store_platform.py → src/yt_scraper/parse.py
|
||||||
|
- `test_store_segments_overwrites()` --calls--> `Segment` [INFERRED]
|
||||||
|
tests/test_store_platform.py → src/yt_scraper/parse.py
|
||||||
|
|
||||||
|
## Import Cycles
|
||||||
|
- None detected.
|
||||||
|
|
||||||
|
## Hyperedges (group relationships)
|
||||||
|
- **Workstream B webapp stack (FastAPI + Alpine SPA + subagent)** — opencode_agent_webapp_builder, opencode_goals_webapp_build, src_yt_scraper_webapp_static_index, docs_superpowers_specs_2026_07_26_platform_design_dark_command_center, docs_superpowers_specs_2026_07_26_platform_design_sse_job_runner [EXTRACTED 0.95]
|
||||||
|
- **FTS5 transcript search flow (proposal -> spec -> SPA view)** — opportunities_search_fts5, docs_superpowers_specs_2026_07_26_platform_design_transcript_fts5, src_yt_scraper_webapp_static_index [INFERRED 0.85]
|
||||||
|
- **Cookie auth chain (vault -> resolve_active_path -> scrape job)** — docs_superpowers_specs_2026_07_26_platform_design_cookie_vault_ux, docs_superpowers_specs_2026_07_26_platform_design_sse_job_runner, docs_superpowers_specs_2026_07_26_platform_design_cookies_bug_fix, src_yt_scraper_webapp_static_index [INFERRED 0.85]
|
||||||
|
|
||||||
|
## Communities (25 total, 3 thin omitted)
|
||||||
|
|
||||||
|
### Community 0 - "Webapp Frontend (JS)"
|
||||||
|
Cohesion: 0.07
|
||||||
|
Nodes (20): _armAutoHide(), closeClip(), closeJobWidget(), closeSidebar(), closeStats(), _disarmAutoHide(), downloadClipTxt(), downloadFile() (+12 more)
|
||||||
|
|
||||||
|
### Community 1 - "Transcript Extraction & Parsing"
|
||||||
|
Cohesion: 0.16
|
||||||
|
Nodes (17): merge_adjacent(), parse_auto_dump(), parse_json3(), parse_vtt(), Segment, _vtt_ts_to_seconds(), store(), test_export_json_csv() (+9 more)
|
||||||
|
|
||||||
|
### Community 2 - "Config & Discovery Pipeline"
|
||||||
|
Cohesion: 0.08
|
||||||
|
Nodes (34): Config, Path, deep_channel_avatar(), Robust avatar recovery via non-flat yt-dlp call on the channel root URL., _extract_handle(), _keep_ref(), Same shorts/live rules the store was populated under — see jobs._keep_ref., Discover + process pending videos for a channel, optionally on an interval. (+26 more)
|
||||||
|
|
||||||
|
### Community 3 - "Store & Export Layer"
|
||||||
|
Cohesion: 0.06
|
||||||
|
Nodes (43): Connection, Cursor, Environment, Row, auto_import_dir(), cookies_dir(), delete(), import_file() (+35 more)
|
||||||
|
|
||||||
|
### Community 4 - "CLI Command Layer"
|
||||||
|
Cohesion: 0.08
|
||||||
|
Nodes (37): analyze_cmd(), _apply_filters(), audio_cmd(), _channel_targets(), channels_add(), cli(), export_cmd(), _extract_handle() (+29 more)
|
||||||
|
|
||||||
|
### Community 5 - "Platform Design & Proposals"
|
||||||
|
Cohesion: 0.12
|
||||||
|
Nodes (20): Platform Implementation Plan, Platform Design Spec, Cookie drag-and-drop vault UX, cli.py cookies scope bug fix, Design language: dark data command center, Architecture: Monorepo in-package webapp (Approach A), Scope rule: No AI features, Single-threaded scrape job runner + SSE event bus (+12 more)
|
||||||
|
|
||||||
|
### Community 6 - "Segments Search & Backfill"
|
||||||
|
Cohesion: 0.10
|
||||||
|
Nodes (26): APIRouter, FastAPI, cache_thumbnail(), Thumbnail URL for a video: stored URL, else the canonical YouTube one derived fr, Download + cache a video's thumbnail to <thumbnails_dir>/<video_id>.jpg. Returns, thumbnail_url_for(), backfill_from_markdown(), parse_markdown() (+18 more)
|
||||||
|
|
||||||
|
### Community 7 - "Store Platform Tests"
|
||||||
|
Cohesion: 0.06
|
||||||
|
Nodes (44): Collection, _channel_root_url(), discover_channel(), discover_incremental(), _flatten_entries(), IncrementalDiscovery, _pick_channel_avatar(), Any (+36 more)
|
||||||
|
|
||||||
|
### Community 8 - "Content Analysis"
|
||||||
|
Cohesion: 0.27
|
||||||
|
Nodes (10): _iter_texts(), Path, Return [(YYYY-MM, count)] of months where `term` appears in transcripts., render_timeline_chart(), render_top_words_chart(), render_wordcloud(), term_timeline(), _tokenize() (+2 more)
|
||||||
|
|
||||||
|
### Community 10 - "Server Start Script"
|
||||||
|
Cohesion: 0.38
|
||||||
|
Nodes (3): Open-Browser(), Out-Line(), Show-State()
|
||||||
|
|
||||||
|
### Community 14 - "test_extract.py"
|
||||||
|
Cohesion: 0.08
|
||||||
|
Nodes (46): Response, parse_languages(), Any, Normalise legacy list / new dict / None into ``{lang: mode}``., _coerce_languages(), _download_subtitle(), extract_video(), _normalize_lang() (+38 more)
|
||||||
|
|
||||||
|
### Community 15 - "Auditoría · Cómo la extensión obtiene la transcripción / información de vídeos de YouTube"
|
||||||
|
Cohesion: 0.05
|
||||||
|
Nodes (41): Arquitectura, Comandos, Cookies, Cualquier petición HTTP a CDNs de YouTube va por `_yt_http.yt_get()`, Documentos de referencia, El Markdown es fuente de datos, no solo salida, Frontend, Idiomas: dict por-idioma con compatibilidad legacy (+33 more)
|
||||||
|
|
||||||
|
### Community 16 - "api"
|
||||||
|
Cohesion: 0.14
|
||||||
|
Nodes (34): activateCookie(), addChannel(), api(), checkAudio(), copyClip(), copyText(), downloadAudio(), downloadAudioOne() (+26 more)
|
||||||
|
|
||||||
|
### Community 17 - "pipeline.py"
|
||||||
|
Cohesion: 0.24
|
||||||
|
Nodes (15): align_chapters(), Chapter, chapters_from_info(), Section, _normalize_date(), process_video(), Regenerate .md for done videos that have stored segments. Returns count re-rende, Extract + parse + store + render a single video. Returns final status string. (+7 more)
|
||||||
|
|
||||||
|
### Community 18 - "render.py"
|
||||||
|
Cohesion: 0.20
|
||||||
|
Nodes (16): build_filename_stem(), format_timestamp(), _make_env(), Path, quote_yaml(), render_markdown(), to_json(), test_build_filename_stem() (+8 more)
|
||||||
|
|
||||||
|
### Community 19 - "config.py"
|
||||||
|
Cohesion: 0.27
|
||||||
|
Nodes (13): _build_config(), DelayConfig, load_config(), Incremental channel sync — how far back a routine re-scan looks. The /video, SyncConfig, YtDlpConfig, Tests for the YAML config loader. Focuses on the ``languages`` and ``prefer_man, test_load_legacy_languages_list() (+5 more)
|
||||||
|
|
||||||
|
### Community 20 - "setView"
|
||||||
|
Cohesion: 0.18
|
||||||
|
Nodes (14): checkHealth(), cycleSort(), destroyCharts(), hydrateURL(), init(), loadFolders(), loadFormat(), loadSearch() (+6 more)
|
||||||
|
|
||||||
|
### Community 21 - "Discovery-Only Scrape Design"
|
||||||
|
Cohesion: 0.22
|
||||||
|
Nodes (8): Architecture, Discovery-Only Scrape Design, Error Handling, Existing Context, Goal, Persistence and Results, Testing, UI
|
||||||
|
|
||||||
|
### Community 22 - "renderDashboardCharts"
|
||||||
|
Cohesion: 0.36
|
||||||
|
Nodes (9): axisOpts(), barOpts(), fmtMonth(), lineOpts(), loadAnalysis(), makeChart(), renderDashboardCharts(), renderTimelineChart() (+1 more)
|
||||||
|
|
||||||
|
### Community 23 - "Global Constraints"
|
||||||
|
Cohesion: 0.25
|
||||||
|
Nodes (7): Discovery-Only Scrape Implementation Plan, Global Constraints, Task 1: Make catalog insertion counts accurate, Task 2: Add discovery-only job execution, Task 3: Expose the discovery mode through the API, Task 4: Add one-channel and all-channel UI controls, Task 5: Final verification and review
|
||||||
|
|
||||||
|
### Community 24 - "loadJobs"
|
||||||
|
Cohesion: 0.29
|
||||||
|
Nodes (7): cancelScrape(), clearJobHistory(), closeStream(), deleteJob(), jobActionLabel(), jobIsTerminal(), loadJobs()
|
||||||
|
|
||||||
|
## Knowledge Gaps
|
||||||
|
- **57 isolated node(s):** `yt-channel-scraper`, `Qué es esto`, `Comandos`, `Orientación antes de leer código`, `Tres superficies, un solo pipeline` (+52 more)
|
||||||
|
These have ≤1 connection - possible missing edges or undocumented components.
|
||||||
|
- **3 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
|
||||||
|
|
||||||
|
## Suggested Questions
|
||||||
|
_Questions this graph is uniquely positioned to answer:_
|
||||||
|
|
||||||
|
- **Why does `Store` connect `Store & Export Layer` to `Transcript Extraction & Parsing`, `Config & Discovery Pipeline`, `CLI Command Layer`, `Segments Search & Backfill`, `Store Platform Tests`, `Content Analysis`, `pipeline.py`?**
|
||||||
|
_High betweenness centrality (0.146) - this node is a cross-community bridge._
|
||||||
|
- **Why does `Segment` connect `Transcript Extraction & Parsing` to `CLI Command Layer`, `Segments Search & Backfill`, `Store Platform Tests`, `test_extract.py`, `pipeline.py`, `render.py`?**
|
||||||
|
_High betweenness centrality (0.035) - this node is a cross-community bridge._
|
||||||
|
- **Why does `VideoRef` connect `Store Platform Tests` to `Transcript Extraction & Parsing`, `Config & Discovery Pipeline`, `Store & Export Layer`, `CLI Command Layer`?**
|
||||||
|
_High betweenness centrality (0.027) - this node is a cross-community bridge._
|
||||||
|
- **Are the 12 inferred relationships involving `Store` (e.g. with `ParsedMarkdown` and `store()`) actually correct?**
|
||||||
|
_`Store` has 12 INFERRED edges - model-reasoned connections that need verification._
|
||||||
|
- **Are the 9 inferred relationships involving `Config` (e.g. with `test_config_languages_defaults_to_dict()` and `test_config_prefer_manual_defaults_true()`) actually correct?**
|
||||||
|
_`Config` has 9 INFERRED edges - model-reasoned connections that need verification._
|
||||||
|
- **Are the 17 inferred relationships involving `Segment` (e.g. with `Chapter` and `Section`) actually correct?**
|
||||||
|
_`Segment` has 17 INFERRED edges - model-reasoned connections that need verification._
|
||||||
|
- **What connects `yt-channel-scraper`, `Qué es esto`, `Comandos` to the rest of the system?**
|
||||||
|
_57 weakly-connected nodes found - possible documentation gaps or missing edges._
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
{
|
||||||
|
"runs": [
|
||||||
|
{
|
||||||
|
"date": "2026-07-27T05:54:56.822038+00:00",
|
||||||
|
"input_tokens": 11800,
|
||||||
|
"output_tokens": 4100,
|
||||||
|
"files": 37
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"total_input_tokens": 11800,
|
||||||
|
"total_output_tokens": 4100
|
||||||
|
}
|
||||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,237 @@
|
|||||||
|
{
|
||||||
|
"pyproject.toml": {
|
||||||
|
"mtime": 1785126752.9759955,
|
||||||
|
"ast_hash": "f80dc007b3b67c7f7a78d9df01a81ead",
|
||||||
|
"semantic_hash": "f80dc007b3b67c7f7a78d9df01a81ead"
|
||||||
|
},
|
||||||
|
"scripts/start-server.ps1": {
|
||||||
|
"mtime": 1785132936.4055703,
|
||||||
|
"ast_hash": "7390a228e676bff301ae85c5a4e7449e",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"scripts/stop-server.ps1": {
|
||||||
|
"mtime": 1785132762.2959402,
|
||||||
|
"ast_hash": "4a62db3fa3dca6b7109fcaadc1219226",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/__init__.py": {
|
||||||
|
"mtime": 1785120608.3336751,
|
||||||
|
"ast_hash": "4867131295172353fe6c2295851defc7",
|
||||||
|
"semantic_hash": "4867131295172353fe6c2295851defc7"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/analysis.py": {
|
||||||
|
"mtime": 1785126463.1545029,
|
||||||
|
"ast_hash": "c28d08131e6594e1e7c6111ff01b2d37",
|
||||||
|
"semantic_hash": "c28d08131e6594e1e7c6111ff01b2d37"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/chapters.py": {
|
||||||
|
"mtime": 1785120675.3308973,
|
||||||
|
"ast_hash": "1bf51b7c13486b9db9bcc3420fba2593",
|
||||||
|
"semantic_hash": "1bf51b7c13486b9db9bcc3420fba2593"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/cli.py": {
|
||||||
|
"mtime": 1785555646.793181,
|
||||||
|
"ast_hash": "36369eb5843535426c93cb5e0e14343a",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/config.py": {
|
||||||
|
"mtime": 1785555391.6008003,
|
||||||
|
"ast_hash": "b7a36354f41dbbaa3e4bbc7b635be4a2",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/cookies.py": {
|
||||||
|
"mtime": 1785126417.4218228,
|
||||||
|
"ast_hash": "dd5ec44bb4c6e6375be80f21307fc73d",
|
||||||
|
"semantic_hash": "dd5ec44bb4c6e6375be80f21307fc73d"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/discover.py": {
|
||||||
|
"mtime": 1785555427.3430922,
|
||||||
|
"ast_hash": "fce9799bd4076f35d4e5b4b1359b9782",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/export.py": {
|
||||||
|
"mtime": 1785126439.6829402,
|
||||||
|
"ast_hash": "f1cff31fb1004f08e997655f838e3a58",
|
||||||
|
"semantic_hash": "f1cff31fb1004f08e997655f838e3a58"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/extract.py": {
|
||||||
|
"mtime": 1785134973.9585488,
|
||||||
|
"ast_hash": "a9cff81702b263b80f72087a14864d7a",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/monitor.py": {
|
||||||
|
"mtime": 1785555591.7611978,
|
||||||
|
"ast_hash": "e7948e61b74b84b246708518f02dbcae",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/parse.py": {
|
||||||
|
"mtime": 1785120665.3384771,
|
||||||
|
"ast_hash": "25354365ac1ab52f1561c97256573806",
|
||||||
|
"semantic_hash": "25354365ac1ab52f1561c97256573806"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/pipeline.py": {
|
||||||
|
"mtime": 1785134990.2165341,
|
||||||
|
"ast_hash": "22d9c0d4630559f97ab68b652d9e1ef3",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/ratelimit.py": {
|
||||||
|
"mtime": 1785120650.3795395,
|
||||||
|
"ast_hash": "6082befe56d342fc43abd5299dd6c6ea",
|
||||||
|
"semantic_hash": "6082befe56d342fc43abd5299dd6c6ea"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/render.py": {
|
||||||
|
"mtime": 1785120688.3247268,
|
||||||
|
"ast_hash": "ff83764f3a8c8874557993e11eab1f89",
|
||||||
|
"semantic_hash": "ff83764f3a8c8874557993e11eab1f89"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/segments.py": {
|
||||||
|
"mtime": 1785128385.6652262,
|
||||||
|
"ast_hash": "09c7da35bc09a3f75fd0a627c0534746",
|
||||||
|
"semantic_hash": "09c7da35bc09a3f75fd0a627c0534746"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/store.py": {
|
||||||
|
"mtime": 1785555451.541978,
|
||||||
|
"ast_hash": "1821f8382f14edc0fe16e27c947caa3b",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/webapp/__init__.py": {
|
||||||
|
"mtime": 1785126827.4957752,
|
||||||
|
"ast_hash": "d41d8cd98f00b204e9800998ecf8427e",
|
||||||
|
"semantic_hash": "d41d8cd98f00b204e9800998ecf8427e"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/webapp/api.py": {
|
||||||
|
"mtime": 1785555568.718664,
|
||||||
|
"ast_hash": "be2bce8aacf98625a43edd840439a5a5",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/webapp/app.py": {
|
||||||
|
"mtime": 1785126888.8032887,
|
||||||
|
"ast_hash": "e0b1925b4318305dcf746d360373ff50",
|
||||||
|
"semantic_hash": "e0b1925b4318305dcf746d360373ff50"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/webapp/jobs.py": {
|
||||||
|
"mtime": 1785555552.558772,
|
||||||
|
"ast_hash": "8ca8f21136d8a4564f49927821244344",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/webapp/static/app.js": {
|
||||||
|
"mtime": 1785555832.5020685,
|
||||||
|
"ast_hash": "6932efd1a6044e66d227a2fc93042187",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_chapters.py": {
|
||||||
|
"mtime": 1785120789.3300896,
|
||||||
|
"ast_hash": "6d5303fa86de8dc99af6ae2ee17da033",
|
||||||
|
"semantic_hash": "6d5303fa86de8dc99af6ae2ee17da033"
|
||||||
|
},
|
||||||
|
"tests/test_features.py": {
|
||||||
|
"mtime": 1785126709.65728,
|
||||||
|
"ast_hash": "21f9c44ce7cc1d618ece402ab7c2b057",
|
||||||
|
"semantic_hash": "21f9c44ce7cc1d618ece402ab7c2b057"
|
||||||
|
},
|
||||||
|
"tests/test_parse.py": {
|
||||||
|
"mtime": 1785120974.3283188,
|
||||||
|
"ast_hash": "57ce74ee10fe6f9a90126ec7d9fd96b0",
|
||||||
|
"semantic_hash": "57ce74ee10fe6f9a90126ec7d9fd96b0"
|
||||||
|
},
|
||||||
|
"tests/test_render.py": {
|
||||||
|
"mtime": 1785120806.3260238,
|
||||||
|
"ast_hash": "27125960e235813724fd669001faab59",
|
||||||
|
"semantic_hash": "27125960e235813724fd669001faab59"
|
||||||
|
},
|
||||||
|
"tests/test_store_platform.py": {
|
||||||
|
"mtime": 1785194588.4186954,
|
||||||
|
"ast_hash": "44069db75a444906fb5001880dedf180",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
".opencode/agent/webapp-builder.md": {
|
||||||
|
"mtime": 1785126783.7152567,
|
||||||
|
"ast_hash": "58e92641724400370d7823dce1e02ca3",
|
||||||
|
"semantic_hash": "58e92641724400370d7823dce1e02ca3"
|
||||||
|
},
|
||||||
|
".opencode/goals/webapp-build.md": {
|
||||||
|
"mtime": 1785126796.8402326,
|
||||||
|
"ast_hash": "478e7a2350ded2fd273cfdb0714d228d",
|
||||||
|
"semantic_hash": "478e7a2350ded2fd273cfdb0714d228d"
|
||||||
|
},
|
||||||
|
"OPPORTUNITIES.md": {
|
||||||
|
"mtime": 1785122706.3466253,
|
||||||
|
"ast_hash": "4af60e3106b394664fbf5b5c5f645804",
|
||||||
|
"semantic_hash": "4af60e3106b394664fbf5b5c5f645804"
|
||||||
|
},
|
||||||
|
"README.md": {
|
||||||
|
"mtime": 1785121499.3282604,
|
||||||
|
"ast_hash": "a0a8e6b62a991fce5d6c58171f80075e",
|
||||||
|
"semantic_hash": "a0a8e6b62a991fce5d6c58171f80075e"
|
||||||
|
},
|
||||||
|
"config.example.yaml": {
|
||||||
|
"mtime": 1785555848.8664799,
|
||||||
|
"ast_hash": "f65a5089bf286e47326e916ad19ca572",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"docs/superpowers/plans/2026-07-26-platform-build.md": {
|
||||||
|
"mtime": 1785126240.4211702,
|
||||||
|
"ast_hash": "8097056ef23068f22c48151746de305b",
|
||||||
|
"semantic_hash": "8097056ef23068f22c48151746de305b"
|
||||||
|
},
|
||||||
|
"docs/superpowers/specs/2026-07-26-platform-design.md": {
|
||||||
|
"mtime": 1785126024.4879057,
|
||||||
|
"ast_hash": "9b85f02fd45c93cdabfd2e1c6d1e0c9e",
|
||||||
|
"semantic_hash": "9b85f02fd45c93cdabfd2e1c6d1e0c9e"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/webapp/static/index.html": {
|
||||||
|
"mtime": 1785555838.550363,
|
||||||
|
"ast_hash": "372b8d77aae448ab1c478253782579ce",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/_yt_http.py": {
|
||||||
|
"mtime": 1785134945.1512043,
|
||||||
|
"ast_hash": "8af224580322df5aeeba10062592d961",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_config.py": {
|
||||||
|
"mtime": 1785135145.3334274,
|
||||||
|
"ast_hash": "5588a7ad8e1c866b1571a818d816f7a7",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_discover_incremental.py": {
|
||||||
|
"mtime": 1785555704.1061945,
|
||||||
|
"ast_hash": "e1c382a9b8f5c927a85d67814d3dbb09",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_extract.py": {
|
||||||
|
"mtime": 1785135157.3141267,
|
||||||
|
"ast_hash": "0bd7c2f3fb9f6bdfe855758b17e3bcf0",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_store_sync_watermark.py": {
|
||||||
|
"mtime": 1785555725.6495814,
|
||||||
|
"ast_hash": "9f9943f5b914078cee49bf4aaa060564",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_webapp_jobs.py": {
|
||||||
|
"mtime": 1785555746.4912646,
|
||||||
|
"ast_hash": "7688c701256ff0cdd2a7d78d7090c8b3",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"CLAUDE.md": {
|
||||||
|
"mtime": 1785556046.8310063,
|
||||||
|
"ast_hash": "631eba15392019298d5fbfcce89c6611",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"YOUTUBE-TRANSCRIPT-AUDIT.md": {
|
||||||
|
"mtime": 1785116715.5118847,
|
||||||
|
"ast_hash": "67c5b7e3b4661ca8c8a7280e63b13ba9",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"docs/superpowers/plans/2026-07-27-discovery-only-scrape.md": {
|
||||||
|
"mtime": 1785192916.5353482,
|
||||||
|
"ast_hash": "3cf41d6331bd67e39ccd96b0e6138cf2",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"docs/superpowers/specs/2026-07-27-discovery-only-scrape-design.md": {
|
||||||
|
"mtime": 1785192898.2889626,
|
||||||
|
"ast_hash": "529b39e9bd2adcafd17a2dc30263281c",
|
||||||
|
"semantic_hash": ""
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,26 @@
|
|||||||
|
{
|
||||||
|
"0": "Webapp Frontend (JS)",
|
||||||
|
"1": "Transcript Extraction & Parsing",
|
||||||
|
"2": "Config & Discovery Pipeline",
|
||||||
|
"3": "Store & Export Layer",
|
||||||
|
"4": "CLI Command Layer",
|
||||||
|
"5": "Platform Design & Proposals",
|
||||||
|
"6": "Segments Search & Backfill",
|
||||||
|
"7": "Store Platform Tests",
|
||||||
|
"8": "Content Analysis",
|
||||||
|
"9": "README & Config Docs",
|
||||||
|
"10": "Server Start Script",
|
||||||
|
"11": "Package Manifest",
|
||||||
|
"12": "Server Stop Script",
|
||||||
|
"13": "Webapp Package Init",
|
||||||
|
"14": "test_extract.py",
|
||||||
|
"15": "Auditoría · Cómo la extensión obtiene la transcripción / información de vídeos de YouTube",
|
||||||
|
"16": "api",
|
||||||
|
"17": "pipeline.py",
|
||||||
|
"20": "setView",
|
||||||
|
"21": "Discovery-Only Scrape Design",
|
||||||
|
"22": "renderDashboardCharts",
|
||||||
|
"23": "Global Constraints",
|
||||||
|
"25": "init",
|
||||||
|
"27": "loadChannels"
|
||||||
|
}
|
||||||
@@ -0,0 +1,177 @@
|
|||||||
|
# Graph Report - yt-channel-scraper (2026-08-01)
|
||||||
|
|
||||||
|
## Corpus Check
|
||||||
|
- 48 files · ~51,184 words
|
||||||
|
- Verdict: corpus is large enough that graph structure adds value.
|
||||||
|
|
||||||
|
## Summary
|
||||||
|
- 669 nodes · 1538 edges · 24 communities (21 shown, 3 thin omitted)
|
||||||
|
- Extraction: 89% EXTRACTED · 11% INFERRED · 0% AMBIGUOUS · INFERRED: 162 edges (avg confidence: 0.78)
|
||||||
|
- Token cost: 0 input · 0 output
|
||||||
|
|
||||||
|
## Graph Freshness
|
||||||
|
- Built from commit: `06497299`
|
||||||
|
- Run `git rev-parse HEAD` and compare to check if the graph is stale.
|
||||||
|
- Run `graphify update .` after code changes (no API cost).
|
||||||
|
|
||||||
|
## Community Hubs (Navigation)
|
||||||
|
- Webapp Frontend (JS)
|
||||||
|
- Transcript Extraction & Parsing
|
||||||
|
- Config & Discovery Pipeline
|
||||||
|
- Store & Export Layer
|
||||||
|
- CLI Command Layer
|
||||||
|
- Platform Design & Proposals
|
||||||
|
- Segments Search & Backfill
|
||||||
|
- Store Platform Tests
|
||||||
|
- Content Analysis
|
||||||
|
- README & Config Docs
|
||||||
|
- Server Start Script
|
||||||
|
- Package Manifest
|
||||||
|
- Server Stop Script
|
||||||
|
- test_extract.py
|
||||||
|
- Auditoría · Cómo la extensión obtiene la transcripción / información de vídeos de YouTube
|
||||||
|
- api
|
||||||
|
- pipeline.py
|
||||||
|
- setView
|
||||||
|
- Discovery-Only Scrape Design
|
||||||
|
- renderDashboardCharts
|
||||||
|
- Global Constraints
|
||||||
|
- init
|
||||||
|
- loadChannels
|
||||||
|
|
||||||
|
## God Nodes (most connected - your core abstractions)
|
||||||
|
1. `Store` - 101 edges
|
||||||
|
2. `VideoRef` - 42 edges
|
||||||
|
3. `api()` - 39 edges
|
||||||
|
4. `toast()` - 38 edges
|
||||||
|
5. `Config` - 36 edges
|
||||||
|
6. `Segment` - 32 edges
|
||||||
|
7. `JobManager` - 26 edges
|
||||||
|
8. `discover_incremental()` - 21 edges
|
||||||
|
9. `process_video()` - 19 edges
|
||||||
|
10. `reconcile_markdown()` - 19 edges
|
||||||
|
|
||||||
|
## Surprising Connections (you probably didn't know these)
|
||||||
|
- `seeded_store()` --calls--> `VideoRef` [INFERRED]
|
||||||
|
tests/test_store_platform.py → src/yt_scraper/store.py
|
||||||
|
- `test_upload_sort_uses_discovered_time_when_date_is_missing()` --calls--> `VideoRef` [INFERRED]
|
||||||
|
tests/test_store_platform.py → src/yt_scraper/store.py
|
||||||
|
- `test_upsert_videos_reports_only_new_rows()` --calls--> `VideoRef` [INFERRED]
|
||||||
|
tests/test_store_platform.py → src/yt_scraper/store.py
|
||||||
|
- `store()` --calls--> `Store` [INFERRED]
|
||||||
|
tests/test_store_platform.py → src/yt_scraper/store.py
|
||||||
|
- `test_migration_idempotent()` --calls--> `Store` [INFERRED]
|
||||||
|
tests/test_store_platform.py → src/yt_scraper/store.py
|
||||||
|
|
||||||
|
## Import Cycles
|
||||||
|
- None detected.
|
||||||
|
|
||||||
|
## Hyperedges (group relationships)
|
||||||
|
- **Workstream B webapp stack (FastAPI + Alpine SPA + subagent)** — opencode_agent_webapp_builder, opencode_goals_webapp_build, src_yt_scraper_webapp_static_index, docs_superpowers_specs_2026_07_26_platform_design_dark_command_center, docs_superpowers_specs_2026_07_26_platform_design_sse_job_runner [EXTRACTED 0.95]
|
||||||
|
- **FTS5 transcript search flow (proposal -> spec -> SPA view)** — opportunities_search_fts5, docs_superpowers_specs_2026_07_26_platform_design_transcript_fts5, src_yt_scraper_webapp_static_index [INFERRED 0.85]
|
||||||
|
- **Cookie auth chain (vault -> resolve_active_path -> scrape job)** — docs_superpowers_specs_2026_07_26_platform_design_cookie_vault_ux, docs_superpowers_specs_2026_07_26_platform_design_sse_job_runner, docs_superpowers_specs_2026_07_26_platform_design_cookies_bug_fix, src_yt_scraper_webapp_static_index [INFERRED 0.85]
|
||||||
|
|
||||||
|
## Communities (24 total, 3 thin omitted)
|
||||||
|
|
||||||
|
### Community 0 - "Webapp Frontend (JS)"
|
||||||
|
Cohesion: 0.06
|
||||||
|
Nodes (22): _armAutoHide(), cancelScrape(), closeClip(), closeJobWidget(), closeSidebar(), closeStats(), closeStream(), deleteJob() (+14 more)
|
||||||
|
|
||||||
|
### Community 1 - "Transcript Extraction & Parsing"
|
||||||
|
Cohesion: 0.27
|
||||||
|
Nodes (10): _iter_texts(), Path, Return [(YYYY-MM, count)] of months where `term` appears in transcripts., render_timeline_chart(), render_top_words_chart(), render_wordcloud(), term_timeline(), _tokenize() (+2 more)
|
||||||
|
|
||||||
|
### Community 2 - "Config & Discovery Pipeline"
|
||||||
|
Cohesion: 0.06
|
||||||
|
Nodes (46): _build_config(), Config, DelayConfig, load_config(), Path, Incremental channel sync — how far back a routine re-scan looks. The /video, SyncConfig, YtDlpConfig (+38 more)
|
||||||
|
|
||||||
|
### Community 3 - "Store & Export Layer"
|
||||||
|
Cohesion: 0.05
|
||||||
|
Nodes (46): Connection, Cursor, Environment, Row, auto_import_dir(), cookies_dir(), delete(), import_file() (+38 more)
|
||||||
|
|
||||||
|
### Community 4 - "CLI Command Layer"
|
||||||
|
Cohesion: 0.07
|
||||||
|
Nodes (44): analyze_cmd(), _apply_filters(), audio_cmd(), _channel_targets(), channels_add(), cli(), export_cmd(), _extract_handle() (+36 more)
|
||||||
|
|
||||||
|
### Community 5 - "Platform Design & Proposals"
|
||||||
|
Cohesion: 0.12
|
||||||
|
Nodes (20): Platform Implementation Plan, Platform Design Spec, Cookie drag-and-drop vault UX, cli.py cookies scope bug fix, Design language: dark data command center, Architecture: Monorepo in-package webapp (Approach A), Scope rule: No AI features, Single-threaded scrape job runner + SSE event bus (+12 more)
|
||||||
|
|
||||||
|
### Community 6 - "Segments Search & Backfill"
|
||||||
|
Cohesion: 0.09
|
||||||
|
Nodes (28): APIRouter, FastAPI, cache_channel_avatar(), cache_thumbnail(), Thumbnail URL for a video: stored URL, else the canonical YouTube one derived fr, Download + cache a video's thumbnail to <thumbnails_dir>/<video_id>.jpg. Returns, Download + cache a channel avatar to <avatars_dir>/<channel_id>.jpg. Returns Tru, thumbnail_url_for() (+20 more)
|
||||||
|
|
||||||
|
### Community 7 - "Store Platform Tests"
|
||||||
|
Cohesion: 0.13
|
||||||
|
Nodes (28): Collection, _channel_root_url(), deep_channel_avatar(), discover_channel(), discover_incremental(), _flatten_entries(), IncrementalDiscovery, _pick_channel_avatar() (+20 more)
|
||||||
|
|
||||||
|
### Community 8 - "Content Analysis"
|
||||||
|
Cohesion: 0.10
|
||||||
|
Nodes (41): Re-escanear data/markdown y hacer que la DB coincida con el disco., reconcile_cmd(), Make the DB agree with what is actually on disk. `backfill_from_markdown` f, reconcile_markdown(), VideoRef, Recovery paths: retryable statuses, recorded skip reasons, disk<->DB reconcile., The throttling message is the one that must stay retryable., The exact symptom the user reported: a .md exists but the row still shows a (+33 more)
|
||||||
|
|
||||||
|
### Community 10 - "Server Start Script"
|
||||||
|
Cohesion: 0.38
|
||||||
|
Nodes (3): Open-Browser(), Out-Line(), Show-State()
|
||||||
|
|
||||||
|
### Community 14 - "test_extract.py"
|
||||||
|
Cohesion: 0.07
|
||||||
|
Nodes (48): Response, parse_languages(), Any, Normalise legacy list / new dict / None into ``{lang: mode}``., _coerce_languages(), describe_missing_subtitle(), _download_subtitle(), extract_video() (+40 more)
|
||||||
|
|
||||||
|
### Community 15 - "Auditoría · Cómo la extensión obtiene la transcripción / información de vídeos de YouTube"
|
||||||
|
Cohesion: 0.04
|
||||||
|
Nodes (42): Arquitectura, Comandos, Cookies, Cualquier petición HTTP a CDNs de YouTube va por `_yt_http.yt_get()`, Documentos de referencia, El Markdown es fuente de datos, no solo salida, Estados terminales y frescura (la UI tiene que reflejar DB + disco), Frontend (+34 more)
|
||||||
|
|
||||||
|
### Community 16 - "api"
|
||||||
|
Cohesion: 0.15
|
||||||
|
Nodes (29): activateCookie(), api(), checkAudio(), clearJobHistory(), copyClip(), copyText(), downloadAudio(), downloadAudioOne() (+21 more)
|
||||||
|
|
||||||
|
### Community 17 - "pipeline.py"
|
||||||
|
Cohesion: 0.05
|
||||||
|
Nodes (56): align_chapters(), Chapter, chapters_from_info(), Section, merge_adjacent(), parse_auto_dump(), parse_json3(), parse_vtt() (+48 more)
|
||||||
|
|
||||||
|
### Community 20 - "setView"
|
||||||
|
Cohesion: 0.24
|
||||||
|
Nodes (10): cycleSort(), destroyCharts(), loadFolders(), loadFormat(), loadSearch(), loadVideos(), openChannel(), resetSearch() (+2 more)
|
||||||
|
|
||||||
|
### Community 21 - "Discovery-Only Scrape Design"
|
||||||
|
Cohesion: 0.22
|
||||||
|
Nodes (8): Architecture, Discovery-Only Scrape Design, Error Handling, Existing Context, Goal, Persistence and Results, Testing, UI
|
||||||
|
|
||||||
|
### Community 22 - "renderDashboardCharts"
|
||||||
|
Cohesion: 0.36
|
||||||
|
Nodes (9): axisOpts(), barOpts(), fmtMonth(), lineOpts(), loadAnalysis(), makeChart(), renderDashboardCharts(), renderTimelineChart() (+1 more)
|
||||||
|
|
||||||
|
### Community 23 - "Global Constraints"
|
||||||
|
Cohesion: 0.25
|
||||||
|
Nodes (7): Discovery-Only Scrape Implementation Plan, Global Constraints, Task 1: Make catalog insertion counts accurate, Task 2: Add discovery-only job execution, Task 3: Expose the discovery mode through the API, Task 4: Add one-channel and all-channel UI controls, Task 5: Final verification and review
|
||||||
|
|
||||||
|
### Community 25 - "init"
|
||||||
|
Cohesion: 0.20
|
||||||
|
Nodes (11): checkHealth(), hydrateURL(), init(), jobActive(), loadRetryable(), reconcile(), refreshLiveState(), resetVideos() (+3 more)
|
||||||
|
|
||||||
|
### Community 27 - "loadChannels"
|
||||||
|
Cohesion: 0.32
|
||||||
|
Nodes (8): addChannel(), downloadAvatarsScope(), loadChannelPending(), loadChannels(), loadDashboard(), removeChannel(), _setSyncResult(), syncChannels()
|
||||||
|
|
||||||
|
## Knowledge Gaps
|
||||||
|
- **58 isolated node(s):** `yt-channel-scraper`, `Qué es esto`, `Comandos`, `Orientación antes de leer código`, `Tres superficies, un solo pipeline` (+53 more)
|
||||||
|
These have ≤1 connection - possible missing edges or undocumented components.
|
||||||
|
- **3 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
|
||||||
|
|
||||||
|
## Suggested Questions
|
||||||
|
_Questions this graph is uniquely positioned to answer:_
|
||||||
|
|
||||||
|
- **Why does `Store` connect `Store & Export Layer` to `Transcript Extraction & Parsing`, `Config & Discovery Pipeline`, `CLI Command Layer`, `Segments Search & Backfill`, `Content Analysis`, `pipeline.py`?**
|
||||||
|
_High betweenness centrality (0.151) - this node is a cross-community bridge._
|
||||||
|
- **Why does `VideoRef` connect `Content Analysis` to `Config & Discovery Pipeline`, `Store & Export Layer`, `CLI Command Layer`, `Store Platform Tests`, `pipeline.py`?**
|
||||||
|
_High betweenness centrality (0.048) - this node is a cross-community bridge._
|
||||||
|
- **Why does `Segment` connect `pipeline.py` to `CLI Command Layer`, `test_extract.py`, `Segments Search & Backfill`?**
|
||||||
|
_High betweenness centrality (0.032) - this node is a cross-community bridge._
|
||||||
|
- **Are the 13 inferred relationships involving `Store` (e.g. with `ParsedMarkdown` and `store()`) actually correct?**
|
||||||
|
_`Store` has 13 INFERRED edges - model-reasoned connections that need verification._
|
||||||
|
- **Are the 31 inferred relationships involving `VideoRef` (e.g. with `IncrementalDiscovery` and `store()`) actually correct?**
|
||||||
|
_`VideoRef` has 31 INFERRED edges - model-reasoned connections that need verification._
|
||||||
|
- **Are the 11 inferred relationships involving `Config` (e.g. with `test_config_languages_defaults_to_dict()` and `test_config_prefer_manual_defaults_true()`) actually correct?**
|
||||||
|
_`Config` has 11 INFERRED edges - model-reasoned connections that need verification._
|
||||||
|
- **What connects `yt-channel-scraper`, `Qué es esto`, `Comandos` to the rest of the system?**
|
||||||
|
_58 weakly-connected nodes found - possible documentation gaps or missing edges._
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
{
|
||||||
|
"runs": [
|
||||||
|
{
|
||||||
|
"date": "2026-07-27T05:54:56.822038+00:00",
|
||||||
|
"input_tokens": 11800,
|
||||||
|
"output_tokens": 4100,
|
||||||
|
"files": 37
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"total_input_tokens": 11800,
|
||||||
|
"total_output_tokens": 4100
|
||||||
|
}
|
||||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,252 @@
|
|||||||
|
{
|
||||||
|
"pyproject.toml": {
|
||||||
|
"mtime": 1785126752.9759955,
|
||||||
|
"ast_hash": "f80dc007b3b67c7f7a78d9df01a81ead",
|
||||||
|
"semantic_hash": "f80dc007b3b67c7f7a78d9df01a81ead"
|
||||||
|
},
|
||||||
|
"scripts/start-server.ps1": {
|
||||||
|
"mtime": 1785132936.4055703,
|
||||||
|
"ast_hash": "7390a228e676bff301ae85c5a4e7449e",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"scripts/stop-server.ps1": {
|
||||||
|
"mtime": 1785132762.2959402,
|
||||||
|
"ast_hash": "4a62db3fa3dca6b7109fcaadc1219226",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/__init__.py": {
|
||||||
|
"mtime": 1785120608.3336751,
|
||||||
|
"ast_hash": "4867131295172353fe6c2295851defc7",
|
||||||
|
"semantic_hash": "4867131295172353fe6c2295851defc7"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/analysis.py": {
|
||||||
|
"mtime": 1785126463.1545029,
|
||||||
|
"ast_hash": "c28d08131e6594e1e7c6111ff01b2d37",
|
||||||
|
"semantic_hash": "c28d08131e6594e1e7c6111ff01b2d37"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/chapters.py": {
|
||||||
|
"mtime": 1785120675.3308973,
|
||||||
|
"ast_hash": "1bf51b7c13486b9db9bcc3420fba2593",
|
||||||
|
"semantic_hash": "1bf51b7c13486b9db9bcc3420fba2593"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/cli.py": {
|
||||||
|
"mtime": 1785561823.2458293,
|
||||||
|
"ast_hash": "100604834f369047b55641bd327333d0",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/config.py": {
|
||||||
|
"mtime": 1785561416.2574291,
|
||||||
|
"ast_hash": "712dddce27149cb3a5ceb34f78fac2fd",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/cookies.py": {
|
||||||
|
"mtime": 1785126417.4218228,
|
||||||
|
"ast_hash": "dd5ec44bb4c6e6375be80f21307fc73d",
|
||||||
|
"semantic_hash": "dd5ec44bb4c6e6375be80f21307fc73d"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/discover.py": {
|
||||||
|
"mtime": 1785555427.3430922,
|
||||||
|
"ast_hash": "fce9799bd4076f35d4e5b4b1359b9782",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/export.py": {
|
||||||
|
"mtime": 1785126439.6829402,
|
||||||
|
"ast_hash": "f1cff31fb1004f08e997655f838e3a58",
|
||||||
|
"semantic_hash": "f1cff31fb1004f08e997655f838e3a58"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/extract.py": {
|
||||||
|
"mtime": 1785561382.7161057,
|
||||||
|
"ast_hash": "a46b96123cfef40a10436e6e65b45d4b",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/monitor.py": {
|
||||||
|
"mtime": 1785555591.7611978,
|
||||||
|
"ast_hash": "e7948e61b74b84b246708518f02dbcae",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/parse.py": {
|
||||||
|
"mtime": 1785120665.3384771,
|
||||||
|
"ast_hash": "25354365ac1ab52f1561c97256573806",
|
||||||
|
"semantic_hash": "25354365ac1ab52f1561c97256573806"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/pipeline.py": {
|
||||||
|
"mtime": 1785562085.460622,
|
||||||
|
"ast_hash": "7dae8872284b33015654fb356c008e67",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/ratelimit.py": {
|
||||||
|
"mtime": 1785120650.3795395,
|
||||||
|
"ast_hash": "6082befe56d342fc43abd5299dd6c6ea",
|
||||||
|
"semantic_hash": "6082befe56d342fc43abd5299dd6c6ea"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/render.py": {
|
||||||
|
"mtime": 1785120688.3247268,
|
||||||
|
"ast_hash": "ff83764f3a8c8874557993e11eab1f89",
|
||||||
|
"semantic_hash": "ff83764f3a8c8874557993e11eab1f89"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/segments.py": {
|
||||||
|
"mtime": 1785562183.2376134,
|
||||||
|
"ast_hash": "902d851e321b83c99d4cae30c1ae39fc",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/store.py": {
|
||||||
|
"mtime": 1785569257.9151454,
|
||||||
|
"ast_hash": "fffc4c7aa5265bd53cd7ca2856b453ef",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/webapp/__init__.py": {
|
||||||
|
"mtime": 1785126827.4957752,
|
||||||
|
"ast_hash": "d41d8cd98f00b204e9800998ecf8427e",
|
||||||
|
"semantic_hash": "d41d8cd98f00b204e9800998ecf8427e"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/webapp/api.py": {
|
||||||
|
"mtime": 1785569267.6258545,
|
||||||
|
"ast_hash": "cadbcafc755f0e1daa574709bc54be84",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/webapp/app.py": {
|
||||||
|
"mtime": 1785562107.3107448,
|
||||||
|
"ast_hash": "80e375becb41e202f78239138252c5c8",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/webapp/jobs.py": {
|
||||||
|
"mtime": 1785556090.2023566,
|
||||||
|
"ast_hash": "45ba7f10b13b847c213f04441a5e741d",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/webapp/static/app.js": {
|
||||||
|
"mtime": 1785561992.7035935,
|
||||||
|
"ast_hash": "94250d573a31bf9bfb29cb54dc4c01e0",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_chapters.py": {
|
||||||
|
"mtime": 1785120789.3300896,
|
||||||
|
"ast_hash": "6d5303fa86de8dc99af6ae2ee17da033",
|
||||||
|
"semantic_hash": "6d5303fa86de8dc99af6ae2ee17da033"
|
||||||
|
},
|
||||||
|
"tests/test_features.py": {
|
||||||
|
"mtime": 1785126709.65728,
|
||||||
|
"ast_hash": "21f9c44ce7cc1d618ece402ab7c2b057",
|
||||||
|
"semantic_hash": "21f9c44ce7cc1d618ece402ab7c2b057"
|
||||||
|
},
|
||||||
|
"tests/test_parse.py": {
|
||||||
|
"mtime": 1785120974.3283188,
|
||||||
|
"ast_hash": "57ce74ee10fe6f9a90126ec7d9fd96b0",
|
||||||
|
"semantic_hash": "57ce74ee10fe6f9a90126ec7d9fd96b0"
|
||||||
|
},
|
||||||
|
"tests/test_render.py": {
|
||||||
|
"mtime": 1785120806.3260238,
|
||||||
|
"ast_hash": "27125960e235813724fd669001faab59",
|
||||||
|
"semantic_hash": "27125960e235813724fd669001faab59"
|
||||||
|
},
|
||||||
|
"tests/test_store_platform.py": {
|
||||||
|
"mtime": 1785194588.4186954,
|
||||||
|
"ast_hash": "44069db75a444906fb5001880dedf180",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
".opencode/agent/webapp-builder.md": {
|
||||||
|
"mtime": 1785126783.7152567,
|
||||||
|
"ast_hash": "58e92641724400370d7823dce1e02ca3",
|
||||||
|
"semantic_hash": "58e92641724400370d7823dce1e02ca3"
|
||||||
|
},
|
||||||
|
".opencode/goals/webapp-build.md": {
|
||||||
|
"mtime": 1785126796.8402326,
|
||||||
|
"ast_hash": "478e7a2350ded2fd273cfdb0714d228d",
|
||||||
|
"semantic_hash": "478e7a2350ded2fd273cfdb0714d228d"
|
||||||
|
},
|
||||||
|
"OPPORTUNITIES.md": {
|
||||||
|
"mtime": 1785122706.3466253,
|
||||||
|
"ast_hash": "4af60e3106b394664fbf5b5c5f645804",
|
||||||
|
"semantic_hash": "4af60e3106b394664fbf5b5c5f645804"
|
||||||
|
},
|
||||||
|
"README.md": {
|
||||||
|
"mtime": 1785121499.3282604,
|
||||||
|
"ast_hash": "a0a8e6b62a991fce5d6c58171f80075e",
|
||||||
|
"semantic_hash": "a0a8e6b62a991fce5d6c58171f80075e"
|
||||||
|
},
|
||||||
|
"config.example.yaml": {
|
||||||
|
"mtime": 1785561451.814817,
|
||||||
|
"ast_hash": "43661d5c00d05c922bcf884e8ac707c3",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"docs/superpowers/plans/2026-07-26-platform-build.md": {
|
||||||
|
"mtime": 1785126240.4211702,
|
||||||
|
"ast_hash": "8097056ef23068f22c48151746de305b",
|
||||||
|
"semantic_hash": "8097056ef23068f22c48151746de305b"
|
||||||
|
},
|
||||||
|
"docs/superpowers/specs/2026-07-26-platform-design.md": {
|
||||||
|
"mtime": 1785126024.4879057,
|
||||||
|
"ast_hash": "9b85f02fd45c93cdabfd2e1c6d1e0c9e",
|
||||||
|
"semantic_hash": "9b85f02fd45c93cdabfd2e1c6d1e0c9e"
|
||||||
|
},
|
||||||
|
"src/yt_scraper/webapp/static/index.html": {
|
||||||
|
"mtime": 1785569275.8425553,
|
||||||
|
"ast_hash": "4f3d8f27f4f663c6f741be3db2312df3",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/_yt_http.py": {
|
||||||
|
"mtime": 1785134945.1512043,
|
||||||
|
"ast_hash": "8af224580322df5aeeba10062592d961",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_config.py": {
|
||||||
|
"mtime": 1785561477.3907576,
|
||||||
|
"ast_hash": "46f3eefb24bcf44535c17977ac175509",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_discover_incremental.py": {
|
||||||
|
"mtime": 1785555704.1061945,
|
||||||
|
"ast_hash": "e1c382a9b8f5c927a85d67814d3dbb09",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_extract.py": {
|
||||||
|
"mtime": 1785561483.422817,
|
||||||
|
"ast_hash": "9949ef4f6e7fe8cbdf12bc8d1dc71e3f",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_store_sync_watermark.py": {
|
||||||
|
"mtime": 1785555725.6495814,
|
||||||
|
"ast_hash": "9f9943f5b914078cee49bf4aaa060564",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_webapp_jobs.py": {
|
||||||
|
"mtime": 1785555746.4912646,
|
||||||
|
"ast_hash": "7688c701256ff0cdd2a7d78d7090c8b3",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"CLAUDE.md": {
|
||||||
|
"mtime": 1785569349.167527,
|
||||||
|
"ast_hash": "a161232f286c0a4f674715f54b66fa76",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"YOUTUBE-TRANSCRIPT-AUDIT.md": {
|
||||||
|
"mtime": 1785116715.5118847,
|
||||||
|
"ast_hash": "67c5b7e3b4661ca8c8a7280e63b13ba9",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"docs/superpowers/plans/2026-07-27-discovery-only-scrape.md": {
|
||||||
|
"mtime": 1785192916.5353482,
|
||||||
|
"ast_hash": "3cf41d6331bd67e39ccd96b0e6138cf2",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"docs/superpowers/specs/2026-07-27-discovery-only-scrape-design.md": {
|
||||||
|
"mtime": 1785192898.2889626,
|
||||||
|
"ast_hash": "529b39e9bd2adcafd17a2dc30263281c",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_recovery.py": {
|
||||||
|
"mtime": 1785569307.786051,
|
||||||
|
"ast_hash": "53f25aeda88cbf3755f45f2377970b62",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_since_filter.py": {
|
||||||
|
"mtime": 1785556104.7475204,
|
||||||
|
"ast_hash": "1d827c61156ab488e838dac91261624f",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
".claude/settings.json": {
|
||||||
|
"mtime": 1785569155.4443069,
|
||||||
|
"ast_hash": "dc072413449915d630bc773c9fb57dc7",
|
||||||
|
"semantic_hash": ""
|
||||||
|
}
|
||||||
|
}
|
||||||
+200
-56
@@ -1,12 +1,18 @@
|
|||||||
# Graph Report - . (2026-07-26)
|
# Graph Report - yt-channel-scraper (2026-08-01)
|
||||||
|
|
||||||
## Corpus Check
|
## Corpus Check
|
||||||
- Corpus is ~28,854 words - fits in a single context window. You may not need a graph.
|
- 49 files · ~52,574 words
|
||||||
|
- Verdict: corpus is large enough that graph structure adds value.
|
||||||
|
|
||||||
## Summary
|
## Summary
|
||||||
- 392 nodes · 951 edges · 14 communities (12 shown, 2 thin omitted)
|
- 696 nodes · 1591 edges · 42 communities (37 shown, 5 thin omitted)
|
||||||
- Extraction: 93% EXTRACTED · 7% INFERRED · 0% AMBIGUOUS · INFERRED: 65 edges (avg confidence: 0.76)
|
- Extraction: 90% EXTRACTED · 10% INFERRED · 0% AMBIGUOUS · INFERRED: 163 edges (avg confidence: 0.78)
|
||||||
- Token cost: 11,800 input · 4,100 output
|
- Token cost: 0 input · 0 output
|
||||||
|
|
||||||
|
## Graph Freshness
|
||||||
|
- Built from commit: `06497299`
|
||||||
|
- Run `git rev-parse HEAD` and compare to check if the graph is stale.
|
||||||
|
- Run `graphify update .` after code changes (no API cost).
|
||||||
|
|
||||||
## Community Hubs (Navigation)
|
## Community Hubs (Navigation)
|
||||||
- Webapp Frontend (JS)
|
- Webapp Frontend (JS)
|
||||||
@@ -19,31 +25,61 @@
|
|||||||
- Store Platform Tests
|
- Store Platform Tests
|
||||||
- Content Analysis
|
- Content Analysis
|
||||||
- README & Config Docs
|
- README & Config Docs
|
||||||
|
- Server Start Script
|
||||||
- Package Manifest
|
- Package Manifest
|
||||||
|
- Server Stop Script
|
||||||
|
- test_extract.py
|
||||||
|
- Auditoría · Cómo la extensión obtiene la transcripción / información de vídeos de YouTube
|
||||||
|
- api
|
||||||
|
- pipeline.py
|
||||||
|
- test_blocked_videos.py
|
||||||
|
- config.py
|
||||||
|
- setView
|
||||||
|
- Discovery-Only Scrape Design
|
||||||
|
- renderDashboardCharts
|
||||||
|
- Global Constraints
|
||||||
|
- Config
|
||||||
|
- init
|
||||||
|
- render.py
|
||||||
|
- loadChannels
|
||||||
|
- jobs.py
|
||||||
|
- Segment
|
||||||
|
- cookies.py
|
||||||
|
- discover.py
|
||||||
|
- export.py
|
||||||
|
- pipeline.py
|
||||||
|
- _channel_targets
|
||||||
|
- test_features.py
|
||||||
|
- VideoRow
|
||||||
|
- test_since_filter.py
|
||||||
|
- .reset_videos
|
||||||
|
- ._connect
|
||||||
|
- re_render_cmd
|
||||||
|
- handleFiles
|
||||||
|
|
||||||
## God Nodes (most connected - your core abstractions)
|
## God Nodes (most connected - your core abstractions)
|
||||||
1. `Store` - 81 edges
|
1. `Store` - 103 edges
|
||||||
2. `Segment` - 32 edges
|
2. `VideoRef` - 43 edges
|
||||||
3. `api()` - 31 edges
|
3. `api()` - 39 edges
|
||||||
4. `toast()` - 29 edges
|
4. `toast()` - 38 edges
|
||||||
5. `Config` - 23 edges
|
5. `Config` - 36 edges
|
||||||
6. `process_video()` - 19 edges
|
6. `Segment` - 32 edges
|
||||||
7. `JobManager` - 17 edges
|
7. `JobManager` - 26 edges
|
||||||
8. `resolve_active_path()` - 13 edges
|
8. `discover_incremental()` - 21 edges
|
||||||
9. `discover_channel()` - 13 edges
|
9. `process_video()` - 19 edges
|
||||||
10. `backfill_from_markdown()` - 13 edges
|
10. `reconcile_markdown()` - 19 edges
|
||||||
|
|
||||||
## Surprising Connections (you probably didn't know these)
|
## Surprising Connections (you probably didn't know these)
|
||||||
- `test_dashboard_aggregates()` --calls--> `Segment` [INFERRED]
|
- `test_config_languages_defaults_to_dict()` --calls--> `Config` [INFERRED]
|
||||||
tests/test_store_platform.py → src/yt_scraper/parse.py
|
tests/test_extract.py → src/yt_scraper/config.py
|
||||||
- `test_store_and_search_segments()` --calls--> `Segment` [INFERRED]
|
- `test_config_prefer_manual_defaults_true()` --calls--> `Config` [INFERRED]
|
||||||
tests/test_store_platform.py → src/yt_scraper/parse.py
|
tests/test_extract.py → src/yt_scraper/config.py
|
||||||
- `test_store_segments_overwrites()` --calls--> `Segment` [INFERRED]
|
- `test_export_json_csv()` --calls--> `Segment` [INFERRED]
|
||||||
tests/test_store_platform.py → src/yt_scraper/parse.py
|
tests/test_features.py → src/yt_scraper/parse.py
|
||||||
- `store()` --calls--> `Store` [INFERRED]
|
- `test_export_srt()` --calls--> `Segment` [INFERRED]
|
||||||
tests/test_store_platform.py → src/yt_scraper/store.py
|
tests/test_features.py → src/yt_scraper/parse.py
|
||||||
- `test_migration_idempotent()` --calls--> `Store` [INFERRED]
|
- `test_term_timeline()` --calls--> `Segment` [INFERRED]
|
||||||
tests/test_store_platform.py → src/yt_scraper/store.py
|
tests/test_features.py → src/yt_scraper/parse.py
|
||||||
|
|
||||||
## Import Cycles
|
## Import Cycles
|
||||||
- None detected.
|
- None detected.
|
||||||
@@ -53,63 +89,171 @@
|
|||||||
- **FTS5 transcript search flow (proposal -> spec -> SPA view)** — opportunities_search_fts5, docs_superpowers_specs_2026_07_26_platform_design_transcript_fts5, src_yt_scraper_webapp_static_index [INFERRED 0.85]
|
- **FTS5 transcript search flow (proposal -> spec -> SPA view)** — opportunities_search_fts5, docs_superpowers_specs_2026_07_26_platform_design_transcript_fts5, src_yt_scraper_webapp_static_index [INFERRED 0.85]
|
||||||
- **Cookie auth chain (vault -> resolve_active_path -> scrape job)** — docs_superpowers_specs_2026_07_26_platform_design_cookie_vault_ux, docs_superpowers_specs_2026_07_26_platform_design_sse_job_runner, docs_superpowers_specs_2026_07_26_platform_design_cookies_bug_fix, src_yt_scraper_webapp_static_index [INFERRED 0.85]
|
- **Cookie auth chain (vault -> resolve_active_path -> scrape job)** — docs_superpowers_specs_2026_07_26_platform_design_cookie_vault_ux, docs_superpowers_specs_2026_07_26_platform_design_sse_job_runner, docs_superpowers_specs_2026_07_26_platform_design_cookies_bug_fix, src_yt_scraper_webapp_static_index [INFERRED 0.85]
|
||||||
|
|
||||||
## Communities (14 total, 2 thin omitted)
|
## Communities (42 total, 5 thin omitted)
|
||||||
|
|
||||||
### Community 0 - "Webapp Frontend (JS)"
|
### Community 0 - "Webapp Frontend (JS)"
|
||||||
Cohesion: 0.06
|
Cohesion: 0.06
|
||||||
Nodes (59): activateCookie(), addChannel(), api(), axisOpts(), barOpts(), cancelScrape(), checkHealth(), closeStream() (+51 more)
|
Nodes (22): _armAutoHide(), blockLabel(), blockTitle(), cancelScrape(), closeClip(), closeJobWidget(), closeSidebar(), closeStats() (+14 more)
|
||||||
|
|
||||||
### Community 1 - "Transcript Extraction & Parsing"
|
### Community 1 - "Transcript Extraction & Parsing"
|
||||||
Cohesion: 0.07
|
Cohesion: 0.27
|
||||||
Nodes (57): Environment, align_chapters(), Chapter, chapters_from_info(), Section, _parse_tags(), Regenerar Markdown desde segmentos almacenados., re_render_cmd() (+49 more)
|
Nodes (10): _iter_texts(), Path, Return [(YYYY-MM, count)] of months where `term` appears in transcripts., render_timeline_chart(), render_top_words_chart(), render_wordcloud(), term_timeline(), _tokenize() (+2 more)
|
||||||
|
|
||||||
### Community 2 - "Config & Discovery Pipeline"
|
### Community 2 - "Config & Discovery Pipeline"
|
||||||
Cohesion: 0.06
|
Cohesion: 0.17
|
||||||
Nodes (40): APIRouter, FastAPI, _build_config(), Config, DelayConfig, load_config(), Any, Path (+32 more)
|
Nodes (10): _clone_config(), _handle(), JobManager, _keep_ref(), Any, Single-worker scrape job runner with an in-memory event log per job (SSE-polled), Refresh the catalog without extracting or downloading video content., Refresh one channel's catalog. Incremental by default: only the newest (+2 more)
|
||||||
|
|
||||||
### Community 3 - "Store & Export Layer"
|
### Community 3 - "Store & Export Layer"
|
||||||
Cohesion: 0.07
|
Cohesion: 0.10
|
||||||
Nodes (28): Connection, Cursor, Row, is_expired(), list_vault(), export_csv(), export_html(), export_json() (+20 more)
|
Nodes (10): Cursor, _now_iso(), Any, Every video id already recorded for a channel — the boundary an incremen, Newest known upload date (YYYYMMDD) for a channel, or None., Refresh a channel's counters from what is actually stored. Incremental, Update only the fields explicitly passed. Preserves last_scraped and video_count, Set a video's status, optionally recording why. `no_subtitles` used to (+2 more)
|
||||||
|
|
||||||
### Community 4 - "CLI Command Layer"
|
### Community 4 - "CLI Command Layer"
|
||||||
Cohesion: 0.08
|
Cohesion: 0.12
|
||||||
Nodes (41): analyze_cmd(), _apply_filters(), audio_cmd(), _channel_targets(), channels_add(), cli(), export_cmd(), _extract_handle() (+33 more)
|
Nodes (24): channels_add(), cli(), _extract_handle(), _fmt_date(), _fmt_duration(), _known_channel_for(), main(), _order_pending() (+16 more)
|
||||||
|
|
||||||
### Community 5 - "Platform Design & Proposals"
|
### Community 5 - "Platform Design & Proposals"
|
||||||
Cohesion: 0.12
|
Cohesion: 0.12
|
||||||
Nodes (20): Platform Implementation Plan, Platform Design Spec, Cookie drag-and-drop vault UX, cli.py cookies scope bug fix, Design language: dark data command center, Architecture: Monorepo in-package webapp (Approach A), Scope rule: No AI features, Single-threaded scrape job runner + SSE event bus (+12 more)
|
Nodes (20): Platform Implementation Plan, Platform Design Spec, Cookie drag-and-drop vault UX, cli.py cookies scope bug fix, Design language: dark data command center, Architecture: Monorepo in-package webapp (Approach A), Scope rule: No AI features, Single-threaded scrape job runner + SSE event bus (+12 more)
|
||||||
|
|
||||||
### Community 6 - "Segments Search & Backfill"
|
### Community 6 - "Segments Search & Backfill"
|
||||||
Cohesion: 0.26
|
Cohesion: 0.13
|
||||||
Nodes (11): backfill_from_markdown(), parse_markdown(), ParsedMarkdown, Path, Parse all done .md files under md_root, populate segments/FTS/metadata. Idempote, search(), store_segments(), _strip_quotes() (+3 more)
|
Nodes (17): APIRouter, FastAPI, cache_channel_avatar(), cache_thumbnail(), Thumbnail URL for a video: stored URL, else the canonical YouTube one derived fr, Download + cache a video's thumbnail to <thumbnails_dir>/<video_id>.jpg. Returns, Download + cache a channel avatar to <avatars_dir>/<channel_id>.jpg. Returns Tru, thumbnail_url_for() (+9 more)
|
||||||
|
|
||||||
### Community 7 - "Store Platform Tests"
|
### Community 7 - "Store Platform Tests"
|
||||||
Cohesion: 0.17
|
Cohesion: 0.22
|
||||||
Nodes (5): store(), test_dashboard_aggregates(), test_migration_idempotent(), test_store_and_search_segments(), test_store_segments_overwrites()
|
Nodes (18): Collection, discover_incremental(), IncrementalDiscovery, Fetch only the newest slice of a channel instead of paginating all of it. A, Outcome of a windowed channel sync., _fake_channel(), Windowed channel sync: only fetch what is newer than what we already have. The, Shorts the store never recorded must not keep the window widening. (+10 more)
|
||||||
|
|
||||||
### Community 8 - "Content Analysis"
|
### Community 8 - "Content Analysis"
|
||||||
|
Cohesion: 0.07
|
||||||
|
Nodes (52): Re-escanear data/markdown y hacer que la DB coincida con el disco., reconcile_cmd(), backfill_from_markdown(), parse_markdown(), ParsedMarkdown, Path, Parse all done .md files under md_root, populate segments/FTS/metadata. Idempote, Make the DB agree with what is actually on disk. `backfill_from_markdown` f (+44 more)
|
||||||
|
|
||||||
|
### Community 10 - "Server Start Script"
|
||||||
|
Cohesion: 0.38
|
||||||
|
Nodes (3): Open-Browser(), Out-Line(), Show-State()
|
||||||
|
|
||||||
|
### Community 14 - "test_extract.py"
|
||||||
|
Cohesion: 0.07
|
||||||
|
Nodes (50): Response, parse_languages(), Any, Normalise legacy list / new dict / None into ``{lang: mode}``., _coerce_languages(), describe_missing_subtitle(), _download_subtitle(), extract_video() (+42 more)
|
||||||
|
|
||||||
|
### Community 15 - "Auditoría · Cómo la extensión obtiene la transcripción / información de vídeos de YouTube"
|
||||||
|
Cohesion: 0.04
|
||||||
|
Nodes (42): Arquitectura, Comandos, Cookies, Cualquier petición HTTP a CDNs de YouTube va por `_yt_http.yt_get()`, Documentos de referencia, El Markdown es fuente de datos, no solo salida, Estados terminales y frescura (la UI tiene que reflejar DB + disco), Frontend (+34 more)
|
||||||
|
|
||||||
|
### Community 16 - "api"
|
||||||
|
Cohesion: 0.16
|
||||||
|
Nodes (29): activateCookie(), api(), checkAudio(), clearJobHistory(), copyClip(), copyText(), deleteJob(), downloadAudio() (+21 more)
|
||||||
|
|
||||||
|
### Community 17 - "pipeline.py"
|
||||||
|
Cohesion: 0.10
|
||||||
|
Nodes (19): merge_adjacent(), parse_auto_dump(), parse_json3(), parse_vtt(), _vtt_ts_to_seconds(), test_merge_adjacent(), test_parse_json3_basic(), test_parse_json3_invalid_json() (+11 more)
|
||||||
|
|
||||||
|
### Community 18 - "test_blocked_videos.py"
|
||||||
|
Cohesion: 0.23
|
||||||
|
Nodes (18): Members-only / gated videos: identified from discovery, labelled, and kept out o, Buying the membership must not leave the video permanently stranded., `unlisted` downloads perfectly well — mislabelling it would hide videos the, Rows burned into `error` before availability was captured must still be iden, Bulk runs must not spend requests on videos that cannot be fetched., _ref(), _store(), test_a_blocked_video_is_still_reachable_by_id() (+10 more)
|
||||||
|
|
||||||
|
### Community 19 - "config.py"
|
||||||
|
Cohesion: 0.24
|
||||||
|
Nodes (14): _build_config(), DelayConfig, load_config(), Incremental channel sync — how far back a routine re-scan looks. The /video, SyncConfig, YtDlpConfig, Tests for the YAML config loader. Focuses on the ``languages`` and ``prefer_man, The default must fall back to auto captions. A manual-only default silently (+6 more)
|
||||||
|
|
||||||
|
### Community 20 - "setView"
|
||||||
|
Cohesion: 0.24
|
||||||
|
Nodes (10): cycleSort(), destroyCharts(), loadFolders(), loadFormat(), loadSearch(), loadVideos(), openChannel(), resetSearch() (+2 more)
|
||||||
|
|
||||||
|
### Community 21 - "Discovery-Only Scrape Design"
|
||||||
|
Cohesion: 0.22
|
||||||
|
Nodes (8): Architecture, Discovery-Only Scrape Design, Error Handling, Existing Context, Goal, Persistence and Results, Testing, UI
|
||||||
|
|
||||||
|
### Community 22 - "renderDashboardCharts"
|
||||||
|
Cohesion: 0.36
|
||||||
|
Nodes (9): axisOpts(), barOpts(), fmtMonth(), lineOpts(), loadAnalysis(), makeChart(), renderDashboardCharts(), renderTimelineChart() (+1 more)
|
||||||
|
|
||||||
|
### Community 23 - "Global Constraints"
|
||||||
|
Cohesion: 0.25
|
||||||
|
Nodes (7): Discovery-Only Scrape Implementation Plan, Global Constraints, Task 1: Make catalog insertion counts accurate, Task 2: Add discovery-only job execution, Task 3: Expose the discovery mode through the API, Task 4: Add one-channel and all-channel UI controls, Task 5: Final verification and review
|
||||||
|
|
||||||
|
### Community 24 - "Config"
|
||||||
|
Cohesion: 0.23
|
||||||
|
Nodes (11): Config, Path, _catalog_channel(), The window is 30 wide but the channel has 200 videos — video_count must not, test_all_channel_discovery_continues_after_one_error(), test_audio_job_downloads_into_audio_directory(), test_audio_job_with_video_ids_uses_audio_runner(), test_batch_job_reports_videos_without_transcripts() (+3 more)
|
||||||
|
|
||||||
|
### Community 25 - "init"
|
||||||
|
Cohesion: 0.15
|
||||||
|
Nodes (18): addChannel(), checkHealth(), downloadAvatarsScope(), hydrateURL(), init(), jobActive(), loadChannelPending(), loadChannels() (+10 more)
|
||||||
|
|
||||||
|
### Community 26 - "render.py"
|
||||||
|
Cohesion: 0.25
|
||||||
|
Nodes (12): Environment, format_timestamp(), _make_env(), Path, quote_yaml(), render_markdown(), to_json(), test_format_timestamp_hours() (+4 more)
|
||||||
|
|
||||||
|
### Community 27 - "loadChannels"
|
||||||
|
Cohesion: 0.23
|
||||||
|
Nodes (8): Row, is_expired(), list_vault(), CookieRow, JobRow, _row_to_cookierow(), _row_to_jobrow(), _sanitize_fts()
|
||||||
|
|
||||||
|
### Community 28 - "jobs.py"
|
||||||
|
Cohesion: 0.23
|
||||||
|
Nodes (10): resolve_active_path(), _extract_handle(), _keep_ref(), Same shorts/live rules the store was populated under — see jobs._keep_ref., Discover + process pending videos for a channel, optionally on an interval., _run_once(), watch_loop(), polite_sleep() (+2 more)
|
||||||
|
|
||||||
|
### Community 29 - "Segment"
|
||||||
|
Cohesion: 0.37
|
||||||
|
Nodes (11): align_chapters(), Chapter, chapters_from_info(), Section, Segment, test_align_no_chapters(), test_align_pre_chapter_segments(), test_align_with_chapters() (+3 more)
|
||||||
|
|
||||||
|
### Community 30 - "cookies.py"
|
||||||
|
Cohesion: 0.35
|
||||||
|
Nodes (11): auto_import_dir(), cookies_dir(), delete(), import_file(), import_text(), parse_netscape(), parse_netscape_file(), Path (+3 more)
|
||||||
|
|
||||||
|
### Community 31 - "discover.py"
|
||||||
Cohesion: 0.27
|
Cohesion: 0.27
|
||||||
Nodes (10): _iter_texts(), Path, Return [(YYYY-MM, count)] of months where `term` appears in transcripts., render_timeline_chart(), render_top_words_chart(), render_wordcloud(), term_timeline(), _tokenize() (+2 more)
|
Nodes (10): _channel_root_url(), deep_channel_avatar(), discover_channel(), _flatten_entries(), _pick_channel_avatar(), Any, Best-effort channel avatar URL from a yt-dlp channel info dict. Only source, Strip a tab suffix from a YouTube channel URL so yt-dlp extracts the channel (+2 more)
|
||||||
|
|
||||||
|
### Community 32 - "export.py"
|
||||||
|
Cohesion: 0.38
|
||||||
|
Nodes (10): export_csv(), export_html(), export_json(), export_srt(), export_srt_video(), Path, _segments_to_srt(), _srt_ts() (+2 more)
|
||||||
|
|
||||||
|
### Community 33 - "pipeline.py"
|
||||||
|
Cohesion: 0.29
|
||||||
|
Nodes (10): _normalize_date(), process_video(), Regenerate .md for done videos that have stored segments. Returns count re-rende, Extract + parse + store + render a single video. Returns final status string., re_render_videos(), _safe_dirname(), build_filename_stem(), test_build_filename_stem() (+2 more)
|
||||||
|
|
||||||
|
### Community 34 - "_channel_targets"
|
||||||
|
Cohesion: 0.20
|
||||||
|
Nodes (10): analyze_cmd(), audio_cmd(), _channel_targets(), export_cmd(), Devolver videos atascados en error/no_subtitles a 'pending' para reintentarlos., Export multi-formato., Descarga audio MP3 (requiere ffmpeg)., Analisis estadistico de contenido (frecuencia, timeline). (+2 more)
|
||||||
|
|
||||||
|
### Community 35 - "test_features.py"
|
||||||
|
Cohesion: 0.20
|
||||||
|
Nodes (5): store(), test_export_json_csv(), test_export_srt(), test_term_timeline(), test_word_frequency()
|
||||||
|
|
||||||
|
### Community 36 - "VideoRow"
|
||||||
|
Cohesion: 0.33
|
||||||
|
Nodes (4): Why this video can never be fetched, or None if it can. Prefers the dis, Pending videos, excluding ones discovery already told us we cannot fetch, _row_to_videorow(), VideoRow
|
||||||
|
|
||||||
|
### Community 37 - "test_since_filter.py"
|
||||||
|
Cohesion: 0.52
|
||||||
|
Nodes (6): _apply_filters(), `--since` / the scrape form's date filter must not eat undated entries. yt-dlp', _ref(), test_since_keeps_entries_with_no_upload_date(), test_since_still_drops_entries_that_are_provably_older(), test_without_since_nothing_is_dropped_on_date_grounds()
|
||||||
|
|
||||||
|
### Community 40 - "re_render_cmd"
|
||||||
|
Cohesion: 0.50
|
||||||
|
Nodes (4): _parse_tags(), Regenerar Markdown desde segmentos almacenados., re_render_cmd(), _safe_dirname()
|
||||||
|
|
||||||
|
### Community 41 - "handleFiles"
|
||||||
|
Cohesion: 0.67
|
||||||
|
Nodes (3): handleDrop(), handleFiles(), uploadCookies()
|
||||||
|
|
||||||
## Knowledge Gaps
|
## Knowledge Gaps
|
||||||
- **10 isolated node(s):** `yt-channel-scraper`, `OPPORTUNITIES.md - extension audit`, `Design language: dark data command center`, `yt-dlp InnerTube mechanism (ANDROID/IOS/WEB clients)`, `Feature proposal: search (FTS5 transcript search)` (+5 more)
|
- **58 isolated node(s):** `yt-channel-scraper`, `Qué es esto`, `Comandos`, `Orientación antes de leer código`, `Tres superficies, un solo pipeline` (+53 more)
|
||||||
These have ≤1 connection - possible missing edges or undocumented components.
|
These have ≤1 connection - possible missing edges or undocumented components.
|
||||||
- **2 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
|
- **5 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
|
||||||
|
|
||||||
## Suggested Questions
|
## Suggested Questions
|
||||||
_Questions this graph is uniquely positioned to answer:_
|
_Questions this graph is uniquely positioned to answer:_
|
||||||
|
|
||||||
- **Why does `Store` connect `Store & Export Layer` to `Transcript Extraction & Parsing`, `Config & Discovery Pipeline`, `CLI Command Layer`, `Segments Search & Backfill`, `Store Platform Tests`, `Content Analysis`?**
|
- **Why does `Store` connect `Store & Export Layer` to `export.py`, `Transcript Extraction & Parsing`, `pipeline.py`, `Config & Discovery Pipeline`, `CLI Command Layer`, `VideoRow`, `.reset_videos`, `._connect`, `Content Analysis`, `Segments Search & Backfill`, `test_features.py`, `pipeline.py`, `test_blocked_videos.py`, `Config`, `loadChannels`, `jobs.py`, `cookies.py`?**
|
||||||
_High betweenness centrality (0.194) - this node is a cross-community bridge._
|
_High betweenness centrality (0.162) - this node is a cross-community bridge._
|
||||||
- **Why does `Segment` connect `Transcript Extraction & Parsing` to `CLI Command Layer`, `Segments Search & Backfill`, `Store Platform Tests`?**
|
- **Why does `VideoRef` connect `Content Analysis` to `Store & Export Layer`, `CLI Command Layer`, `test_features.py`, `test_since_filter.py`, `Store Platform Tests`, `pipeline.py`, `test_blocked_videos.py`, `Config`, `loadChannels`, `jobs.py`, `discover.py`?**
|
||||||
_High betweenness centrality (0.057) - this node is a cross-community bridge._
|
_High betweenness centrality (0.053) - this node is a cross-community bridge._
|
||||||
- **Why does `Config` connect `Config & Discovery Pipeline` to `Transcript Extraction & Parsing`, `CLI Command Layer`?**
|
- **Why does `Segment` connect `Segment` to `pipeline.py`, `test_features.py`, `CLI Command Layer`, `re_render_cmd`, `Content Analysis`, `test_extract.py`, `pipeline.py`?**
|
||||||
_High betweenness centrality (0.024) - this node is a cross-community bridge._
|
_High betweenness centrality (0.031) - this node is a cross-community bridge._
|
||||||
- **Are the 4 inferred relationships involving `Store` (e.g. with `ParsedMarkdown` and `store()`) actually correct?**
|
- **Are the 14 inferred relationships involving `Store` (e.g. with `ParsedMarkdown` and `_store()`) actually correct?**
|
||||||
_`Store` has 4 INFERRED edges - model-reasoned connections that need verification._
|
_`Store` has 14 INFERRED edges - model-reasoned connections that need verification._
|
||||||
- **Are the 17 inferred relationships involving `Segment` (e.g. with `Chapter` and `Section`) actually correct?**
|
- **Are the 31 inferred relationships involving `VideoRef` (e.g. with `IncrementalDiscovery` and `store()`) actually correct?**
|
||||||
_`Segment` has 17 INFERRED edges - model-reasoned connections that need verification._
|
_`VideoRef` has 31 INFERRED edges - model-reasoned connections that need verification._
|
||||||
- **What connects `yt-channel-scraper`, `OPPORTUNITIES.md - extension audit`, `Design language: dark data command center` to the rest of the system?**
|
- **Are the 11 inferred relationships involving `Config` (e.g. with `test_config_languages_defaults_to_dict()` and `test_config_prefer_manual_defaults_true()`) actually correct?**
|
||||||
_10 weakly-connected nodes found - possible documentation gaps or missing edges._
|
_`Config` has 11 INFERRED edges - model-reasoned connections that need verification._
|
||||||
- **Should `Webapp Frontend (JS)` be split into smaller, more focused modules?**
|
- **What connects `yt-channel-scraper`, `Qué es esto`, `Comandos` to the rest of the system?**
|
||||||
_Cohesion score 0.062317429406037 - nodes in this community are weakly interconnected._
|
_58 weakly-connected nodes found - possible documentation gaps or missing edges._
|
||||||
graphify-out/cache/ast/v0.9.22/01d12afbd0457e8960455bcdd1197204d149941168e4452768492f17e12057b7.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/055ece0cc9a7abdfdce29fc2fe1d94165bdcca87a236ba398c105e4f8641c676.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/138f2e349300764b2b5ca342295a12efbaa6de1ff8df332be58871b130570ed9.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/165fa89acf2585a34e4b07436faeb8f225d14bc1a699fdc4ee506d7619eb1d95.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/1b7bfeeef4294b3ebbdfafd1e791473299d1bc77ad5e2f6b8a7b9427ba76f5ec.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/1d0aa0749bad2a7c8c3670e18750ea9cc581a8ac338d00b6501bb7361cf24991.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/1e52d419bb392cebab1e9715bcb53d51cceaa71d68b9032227d4741695af6e28.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/2552d4d9a9cc2a1ed9040be2fef3d5c460f5df7680e5e25ae14ed6f75907918b.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/2b9367cdaf1ee8bcd08a752824a819d0139dadd53f59086cb4b19fd67f32b53b.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/3f8e7e91eb49c07fcdbe4286b8c15848443c4edc734e4418e3f481bf7482a008.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/45cf21d405c4221cb88b1147cf7e87d1e1e4cc3252111019deeaf27a6d7a8b6a.json
Vendored
+1
@@ -0,0 +1 @@
|
|||||||
|
{"nodes": [{"id": "d_yt_channel_scraper_scripts_stop_server_ps1", "label": "stop-server.ps1", "file_type": "code", "source_file": "scripts/stop-server.ps1", "source_location": "L1"}, {"id": "d_yt_channel_scraper_scripts_stop_server_out_line", "label": "Out-Line()", "file_type": "code", "source_file": "scripts/stop-server.ps1", "source_location": "L5"}, {"id": "d_yt_channel_scraper_scripts_stop_server_test_portopen", "label": "Test-PortOpen()", "file_type": "code", "source_file": "scripts/stop-server.ps1", "source_location": "L13"}, {"id": "d_yt_channel_scraper_scripts_stop_server_kill_portowner", "label": "Kill-PortOwner()", "file_type": "code", "source_file": "scripts/stop-server.ps1", "source_location": "L24"}], "edges": [{"source": "d_yt_channel_scraper_scripts_stop_server_ps1", "target": "d_yt_channel_scraper_scripts_stop_server_out_line", "relation": "contains", "confidence": "EXTRACTED", "source_file": "scripts/stop-server.ps1", "source_location": "L5", "weight": 1.0}, {"source": "d_yt_channel_scraper_scripts_stop_server_ps1", "target": "d_yt_channel_scraper_scripts_stop_server_test_portopen", "relation": "contains", "confidence": "EXTRACTED", "source_file": "scripts/stop-server.ps1", "source_location": "L13", "weight": 1.0}, {"source": "d_yt_channel_scraper_scripts_stop_server_ps1", "target": "d_yt_channel_scraper_scripts_stop_server_kill_portowner", "relation": "contains", "confidence": "EXTRACTED", "source_file": "scripts/stop-server.ps1", "source_location": "L24", "weight": 1.0}, {"source": "d_yt_channel_scraper_scripts_stop_server_kill_portowner", "target": "d_yt_channel_scraper_scripts_stop_server_out_line", "relation": "calls", "confidence": "EXTRACTED", "source_file": "scripts/stop-server.ps1", "source_location": "L27", "weight": 1.0}], "raw_calls": [{"caller_nid": "d_yt_channel_scraper_scripts_stop_server_out_line", "callee": "Write-Host", "is_member_call": false, "source_file": "scripts/stop-server.ps1", "source_location": "L6"}, {"caller_nid": "d_yt_channel_scraper_scripts_stop_server_out_line", "callee": "Write-Host", "is_member_call": false, "source_file": "scripts/stop-server.ps1", "source_location": "L7"}, {"caller_nid": "d_yt_channel_scraper_scripts_stop_server_test_portopen", "callee": "New-Object", "is_member_call": false, "source_file": "scripts/stop-server.ps1", "source_location": "L14"}, {"caller_nid": "d_yt_channel_scraper_scripts_stop_server_kill_portowner", "callee": "Get-NetTCPConnection", "is_member_call": false, "source_file": "scripts/stop-server.ps1", "source_location": "L25"}, {"caller_nid": "d_yt_channel_scraper_scripts_stop_server_kill_portowner", "callee": "ForEach-Object", "is_member_call": false, "source_file": "scripts/stop-server.ps1", "source_location": "L26"}, {"caller_nid": "d_yt_channel_scraper_scripts_stop_server_kill_portowner", "callee": "Stop-Process", "is_member_call": false, "source_file": "scripts/stop-server.ps1", "source_location": "L28"}]}
|
||||||
graphify-out/cache/ast/v0.9.22/46f59277042b7eba1430abe6f302416a2f14834a2a39de944e7919afd30ea732.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/4cd4f80326c3df60abc579fa3be19eca6e08b8bdbe0ad936c8d42ee241b71edb.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/55d28bdb9e808d6e2ab2717ef307faedff76d2202f0d029621489c82b8fe3aa1.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/5ad53cd8158245cd67b3ea45dc7dad107636666bd957a76041a6e36c20b1b4df.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/60f9653b556213f971442048ebc563d916c6c2c8e7e729d1eb3f1f1ef9a6f74e.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/61e93a83b219449f61955a14b4f15ecfb5be4950adc7a90568e2d4c1af13dc00.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/651f8538a30ddaeddb38795a3eef200b5d426c33eea3812e578b532e09c13f45.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/65dfd9ff5b8d3db19bf1ef2b4eb7ff596e9f5647d010481376266dc63f1b083e.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/68dada94168c2b33eb99461b5db0bcf4ce8f2ef6c256effd5a1efa9a2f537a35.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/6d4af4049f6b3a40f7e0356cc1993dcaf22af164a35b080c55b6aabc001fd18d.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/718bb43705ffbaaabace8c1eacedaa69d89d34729ff65328d3dfe20f2d1098d3.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/73fd8d6e6d80e5ffd3e34ba7cd37ea93f8972f02d09533f44710b9f1e90e0803.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/7dbbf6b8f41b3db71cb97e7f5e9ac6bbdd326f835b96aa8be9ce255a3e7e33b4.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/84d78dabe5dd523c6f7de3c05844ac207347e04eb57939fbc1186894fb8fc334.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/8984317652a50a4ac2e8de007380ce35790d306453dff7053f8089b007349da3.json
Vendored
+1
@@ -0,0 +1 @@
|
|||||||
|
{"nodes": [{"id": "d_yt_channel_scraper_src_yt_scraper_yt_http_py", "label": "_yt_http.py", "file_type": "code", "source_file": "src/yt_scraper/_yt_http.py", "source_location": "L1"}, {"id": "d_yt_channel_scraper_src_yt_scraper_yt_http_yt_headers", "label": "yt_headers()", "file_type": "code", "source_file": "src/yt_scraper/_yt_http.py", "source_location": "L33", "_callable": true}, {"id": "d_yt_channel_scraper_src_yt_scraper_yt_http_yt_get", "label": "yt_get()", "file_type": "code", "source_file": "src/yt_scraper/_yt_http.py", "source_location": "L40", "_callable": true}, {"id": "any", "label": "Any", "file_type": "code", "source_file": "", "source_location": "", "origin_file": "D:\\yt-channel-scraper\\src\\yt_scraper\\_yt_http.py"}, {"id": "response", "label": "Response", "file_type": "code", "source_file": "", "source_location": "", "origin_file": "D:\\yt-channel-scraper\\src\\yt_scraper\\_yt_http.py"}, {"id": "d_yt_channel_scraper_src_yt_scraper_yt_http_rationale_1", "label": "Shared HTTP helpers for talking to YouTube CDNs. Single source of truth for the", "file_type": "rationale", "source_file": "src/yt_scraper/_yt_http.py", "source_location": "L1"}, {"id": "d_yt_channel_scraper_src_yt_scraper_yt_http_rationale_34", "label": "Return a copy of the YouTube CDN headers merged with ``extra``.", "file_type": "rationale", "source_file": "src/yt_scraper/_yt_http.py", "source_location": "L34"}, {"id": "d_yt_channel_scraper_src_yt_scraper_yt_http_rationale_43", "label": "GET ``url`` with the shared YouTube CDN headers. Centralising this lets eve", "file_type": "rationale", "source_file": "src/yt_scraper/_yt_http.py", "source_location": "L43"}], "edges": [{"source": "d_yt_channel_scraper_src_yt_scraper_yt_http_py", "target": "typing", "relation": "imports_from", "context": "import", "confidence": "EXTRACTED", "source_file": "src/yt_scraper/_yt_http.py", "source_location": "L10", "weight": 1.0}, {"source": "d_yt_channel_scraper_src_yt_scraper_yt_http_py", "target": "requests", "relation": "imports", "context": "import", "confidence": "EXTRACTED", "source_file": "src/yt_scraper/_yt_http.py", "source_location": "L12", "weight": 1.0}, {"source": "d_yt_channel_scraper_src_yt_scraper_yt_http_py", "target": "d_yt_channel_scraper_src_yt_scraper_yt_http_yt_headers", "relation": "contains", "confidence": "EXTRACTED", "source_file": "src/yt_scraper/_yt_http.py", "source_location": "L33", "weight": 1.0}, {"source": "d_yt_channel_scraper_src_yt_scraper_yt_http_py", "target": "d_yt_channel_scraper_src_yt_scraper_yt_http_yt_get", "relation": "contains", "confidence": "EXTRACTED", "source_file": "src/yt_scraper/_yt_http.py", "source_location": "L40", "weight": 1.0}, {"source": "d_yt_channel_scraper_src_yt_scraper_yt_http_yt_get", "target": "any", "relation": "references", "context": "generic_arg", "confidence": "EXTRACTED", "source_file": "src/yt_scraper/_yt_http.py", "source_location": "L40", "weight": 1.0}, {"source": "d_yt_channel_scraper_src_yt_scraper_yt_http_yt_get", "target": "response", "relation": "references", "context": "return_type", "confidence": "EXTRACTED", "source_file": "src/yt_scraper/_yt_http.py", "source_location": "L40", "weight": 1.0}, {"source": "d_yt_channel_scraper_src_yt_scraper_yt_http_yt_get", "target": "d_yt_channel_scraper_src_yt_scraper_yt_http_yt_headers", "relation": "calls", "context": "call", "confidence": "EXTRACTED", "source_file": "src/yt_scraper/_yt_http.py", "source_location": "L49", "weight": 1.0}, {"source": "d_yt_channel_scraper_src_yt_scraper_yt_http_rationale_1", "target": "d_yt_channel_scraper_src_yt_scraper_yt_http_py", "relation": "rationale_for", "confidence": "EXTRACTED", "source_file": "src/yt_scraper/_yt_http.py", "source_location": "L1", "weight": 1.0}, {"source": "d_yt_channel_scraper_src_yt_scraper_yt_http_rationale_34", "target": "d_yt_channel_scraper_src_yt_scraper_yt_http_yt_headers", "relation": "rationale_for", "confidence": "EXTRACTED", "source_file": "src/yt_scraper/_yt_http.py", "source_location": "L34", "weight": 1.0}, {"source": "d_yt_channel_scraper_src_yt_scraper_yt_http_rationale_43", "target": "d_yt_channel_scraper_src_yt_scraper_yt_http_yt_get", "relation": "rationale_for", "confidence": "EXTRACTED", "source_file": "src/yt_scraper/_yt_http.py", "source_location": "L43", "weight": 1.0}], "raw_calls": [{"caller_nid": "d_yt_channel_scraper_src_yt_scraper_yt_http_yt_headers", "callee": "_YT_HEADERS", "is_member_call": false, "indirect": true, "context": "argument", "source_file": "src/yt_scraper/_yt_http.py", "source_location": "L37"}, {"caller_nid": "d_yt_channel_scraper_src_yt_scraper_yt_http_yt_get", "callee": "get", "is_member_call": true, "source_file": "src/yt_scraper/_yt_http.py", "source_location": "L49", "receiver": "requests"}]}
|
||||||
graphify-out/cache/ast/v0.9.22/91e63e7a4a9cd9af19ab798569730680f8172995ff8eb74803b64de66bac1295.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/9804897ae7075e54c30e3ff644c5fb455a7c00e00c7af74ec141cacfe8a68c35.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/981149ba3a0875f62a1191a903e769cea97f4692897fe547800a6aa1fbe81508.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/9b69dde6b0a1aa81eb3e1f3629f71bb3398b89d8ac28bddfd66df8736dd7f841.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/a188993a249d683cc548d0a49630ad50fef9b9e01f50888c7737eb5beb278ed0.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/a23f32306d296829d2d5a3997e63e9c4f54452dfebfe13bf34d9ee6eb9b29ec3.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/ada75965055538556230762161fb5829df2e2bda12de82e24c9d0aec4de63e15.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/b76c96064d3a1eef8dbf2ad95c10e220da08be4b2203dbe3bed5e063090a164a.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/c7152e16af45313d40e6cd32591417d727f7f781e34cbd96cd3d0716d1a5fc8c.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/c73479f2c48be4ec998ffd20c079c22cbe23a523d7ee0e75e475544bcec1c8bf.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/cfd689ff36a543753dbb4069417b91b6ffadd87c1f2f535e8a9ca6e8f14c75cb.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/e3898b3c1015fcdf14b2ded9745068645e784bd58b6bdd07d46665dffeeee335.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/e5392d2b2097b1b3abe5e0cf8abd8c578f8c2ce31b6b6cc2d94f33e4f77011de.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/e78706ff4fb38d94863ba4a677a8e0924bc13bcf4f1a9e785d7329d141c9f5fb.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/e94e0430b251839d294e187004902d1dbe489bdffe5abc8d37ddabf5057b622d.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/e9f4e6a982201fc90e95acc8d10cd3279ec1fd92f33dda68976dcb0c88db272f.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/ee1d3dfe150c6f3ba8318bbfb1c8cbd0747a36b7590090dfc94540fe225535d0.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/ee93e837819a3a34757f24506b85165b97901af37eccb161b585053e90379120.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/f323a0f9a76830ce23bc882ee9bcc08d1f3043ef1b793a1b93b289b8e92b513a.json
Vendored
+1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/f860e97e16ef806ce0f1473489eba58a4ab3df878e99456f492ccdd4c117fd76.json
Vendored
+1
@@ -0,0 +1 @@
|
|||||||
|
{"nodes": [{"id": "d_yt_channel_scraper_scripts_start_server_ps1", "label": "start-server.ps1", "file_type": "code", "source_file": "scripts/start-server.ps1", "source_location": "L1"}, {"id": "d_yt_channel_scraper_scripts_start_server_out_line", "label": "Out-Line()", "file_type": "code", "source_file": "scripts/start-server.ps1", "source_location": "L10"}, {"id": "d_yt_channel_scraper_scripts_start_server_open_browser", "label": "Open-Browser()", "file_type": "code", "source_file": "scripts/start-server.ps1", "source_location": "L17"}, {"id": "d_yt_channel_scraper_scripts_start_server_test_portopen", "label": "Test-PortOpen()", "file_type": "code", "source_file": "scripts/start-server.ps1", "source_location": "L31"}, {"id": "d_yt_channel_scraper_scripts_start_server_get_serverhealth", "label": "Get-ServerHealth()", "file_type": "code", "source_file": "scripts/start-server.ps1", "source_location": "L42"}, {"id": "d_yt_channel_scraper_scripts_start_server_is_pidalive", "label": "Is-PidAlive()", "file_type": "code", "source_file": "scripts/start-server.ps1", "source_location": "L50"}, {"id": "d_yt_channel_scraper_scripts_start_server_show_state", "label": "Show-State()", "file_type": "code", "source_file": "scripts/start-server.ps1", "source_location": "L56"}], "edges": [{"source": "d_yt_channel_scraper_scripts_start_server_ps1", "target": "d_yt_channel_scraper_scripts_start_server_out_line", "relation": "contains", "confidence": "EXTRACTED", "source_file": "scripts/start-server.ps1", "source_location": "L10", "weight": 1.0}, {"source": "d_yt_channel_scraper_scripts_start_server_ps1", "target": "d_yt_channel_scraper_scripts_start_server_open_browser", "relation": "contains", "confidence": "EXTRACTED", "source_file": "scripts/start-server.ps1", "source_location": "L17", "weight": 1.0}, {"source": "d_yt_channel_scraper_scripts_start_server_ps1", "target": "d_yt_channel_scraper_scripts_start_server_test_portopen", "relation": "contains", "confidence": "EXTRACTED", "source_file": "scripts/start-server.ps1", "source_location": "L31", "weight": 1.0}, {"source": "d_yt_channel_scraper_scripts_start_server_ps1", "target": "d_yt_channel_scraper_scripts_start_server_get_serverhealth", "relation": "contains", "confidence": "EXTRACTED", "source_file": "scripts/start-server.ps1", "source_location": "L42", "weight": 1.0}, {"source": "d_yt_channel_scraper_scripts_start_server_ps1", "target": "d_yt_channel_scraper_scripts_start_server_is_pidalive", "relation": "contains", "confidence": "EXTRACTED", "source_file": "scripts/start-server.ps1", "source_location": "L50", "weight": 1.0}, {"source": "d_yt_channel_scraper_scripts_start_server_ps1", "target": "d_yt_channel_scraper_scripts_start_server_show_state", "relation": "contains", "confidence": "EXTRACTED", "source_file": "scripts/start-server.ps1", "source_location": "L56", "weight": 1.0}, {"source": "d_yt_channel_scraper_scripts_start_server_open_browser", "target": "d_yt_channel_scraper_scripts_start_server_out_line", "relation": "calls", "confidence": "EXTRACTED", "source_file": "scripts/start-server.ps1", "source_location": "L22", "weight": 1.0}, {"source": "d_yt_channel_scraper_scripts_start_server_show_state", "target": "d_yt_channel_scraper_scripts_start_server_out_line", "relation": "calls", "confidence": "EXTRACTED", "source_file": "scripts/start-server.ps1", "source_location": "L57", "weight": 1.0}], "raw_calls": [{"caller_nid": "d_yt_channel_scraper_scripts_start_server_out_line", "callee": "Write-Host", "is_member_call": false, "source_file": "scripts/start-server.ps1", "source_location": "L11"}, {"caller_nid": "d_yt_channel_scraper_scripts_start_server_out_line", "callee": "Write-Host", "is_member_call": false, "source_file": "scripts/start-server.ps1", "source_location": "L12"}, {"caller_nid": "d_yt_channel_scraper_scripts_start_server_open_browser", "callee": "Start-Process", "is_member_call": false, "source_file": "scripts/start-server.ps1", "source_location": "L19"}, {"caller_nid": "d_yt_channel_scraper_scripts_start_server_test_portopen", "callee": "New-Object", "is_member_call": false, "source_file": "scripts/start-server.ps1", "source_location": "L32"}, {"caller_nid": "d_yt_channel_scraper_scripts_start_server_get_serverhealth", "callee": "Invoke-WebRequest", "is_member_call": false, "source_file": "scripts/start-server.ps1", "source_location": "L44"}, {"caller_nid": "d_yt_channel_scraper_scripts_start_server_is_pidalive", "callee": "Get-Process", "is_member_call": false, "source_file": "scripts/start-server.ps1", "source_location": "L52"}, {"caller_nid": "d_yt_channel_scraper_scripts_start_server_show_state", "callee": "Get-Date", "is_member_call": false, "source_file": "scripts/start-server.ps1", "source_location": "L57"}]}
|
||||||
Vendored
+1
@@ -0,0 +1 @@
|
|||||||
|
1787443350.5198534
|
||||||
Vendored
+1
-1
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
+12669
-1545
File diff suppressed because it is too large
Load Diff
+121
-51
@@ -5,14 +5,14 @@
|
|||||||
"semantic_hash": "f80dc007b3b67c7f7a78d9df01a81ead"
|
"semantic_hash": "f80dc007b3b67c7f7a78d9df01a81ead"
|
||||||
},
|
},
|
||||||
"scripts/start-server.ps1": {
|
"scripts/start-server.ps1": {
|
||||||
"mtime": 1785130613.533591,
|
"mtime": 1785132936.4055703,
|
||||||
"ast_hash": "9baf397df6c14b7917ca59fbe6df61cc",
|
"ast_hash": "7390a228e676bff301ae85c5a4e7449e",
|
||||||
"semantic_hash": "9baf397df6c14b7917ca59fbe6df61cc"
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"scripts/stop-server.ps1": {
|
"scripts/stop-server.ps1": {
|
||||||
"mtime": 1785127403.2329483,
|
"mtime": 1785132762.2959402,
|
||||||
"ast_hash": "4398501c202a3eb624fb439f6d03fcf5",
|
"ast_hash": "4a62db3fa3dca6b7109fcaadc1219226",
|
||||||
"semantic_hash": "4398501c202a3eb624fb439f6d03fcf5"
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"src/yt_scraper/__init__.py": {
|
"src/yt_scraper/__init__.py": {
|
||||||
"mtime": 1785120608.3336751,
|
"mtime": 1785120608.3336751,
|
||||||
@@ -30,14 +30,14 @@
|
|||||||
"semantic_hash": "1bf51b7c13486b9db9bcc3420fba2593"
|
"semantic_hash": "1bf51b7c13486b9db9bcc3420fba2593"
|
||||||
},
|
},
|
||||||
"src/yt_scraper/cli.py": {
|
"src/yt_scraper/cli.py": {
|
||||||
"mtime": 1785126572.0575886,
|
"mtime": 1785561823.2458293,
|
||||||
"ast_hash": "9734cedf7c1e045e1ab0e4a102860e29",
|
"ast_hash": "100604834f369047b55641bd327333d0",
|
||||||
"semantic_hash": "9734cedf7c1e045e1ab0e4a102860e29"
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"src/yt_scraper/config.py": {
|
"src/yt_scraper/config.py": {
|
||||||
"mtime": 1785120620.335521,
|
"mtime": 1785561416.2574291,
|
||||||
"ast_hash": "2f3c8f7768a942e12c0a9e41b54891be",
|
"ast_hash": "712dddce27149cb3a5ceb34f78fac2fd",
|
||||||
"semantic_hash": "2f3c8f7768a942e12c0a9e41b54891be"
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"src/yt_scraper/cookies.py": {
|
"src/yt_scraper/cookies.py": {
|
||||||
"mtime": 1785126417.4218228,
|
"mtime": 1785126417.4218228,
|
||||||
@@ -45,9 +45,9 @@
|
|||||||
"semantic_hash": "dd5ec44bb4c6e6375be80f21307fc73d"
|
"semantic_hash": "dd5ec44bb4c6e6375be80f21307fc73d"
|
||||||
},
|
},
|
||||||
"src/yt_scraper/discover.py": {
|
"src/yt_scraper/discover.py": {
|
||||||
"mtime": 1785120708.331971,
|
"mtime": 1785569683.6527412,
|
||||||
"ast_hash": "8db06d0b131575f709b55467eabdc2a9",
|
"ast_hash": "6b626a24ae81406ea549e9989c9f5d5b",
|
||||||
"semantic_hash": "8db06d0b131575f709b55467eabdc2a9"
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"src/yt_scraper/export.py": {
|
"src/yt_scraper/export.py": {
|
||||||
"mtime": 1785126439.6829402,
|
"mtime": 1785126439.6829402,
|
||||||
@@ -55,14 +55,14 @@
|
|||||||
"semantic_hash": "f1cff31fb1004f08e997655f838e3a58"
|
"semantic_hash": "f1cff31fb1004f08e997655f838e3a58"
|
||||||
},
|
},
|
||||||
"src/yt_scraper/extract.py": {
|
"src/yt_scraper/extract.py": {
|
||||||
"mtime": 1785122183.7294595,
|
"mtime": 1785561382.7161057,
|
||||||
"ast_hash": "3e2179db2f7fd316d2b3acedee536086",
|
"ast_hash": "a46b96123cfef40a10436e6e65b45d4b",
|
||||||
"semantic_hash": "3e2179db2f7fd316d2b3acedee536086"
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"src/yt_scraper/monitor.py": {
|
"src/yt_scraper/monitor.py": {
|
||||||
"mtime": 1785126496.288271,
|
"mtime": 1785555591.7611978,
|
||||||
"ast_hash": "a64188f938dc4f8917335ef41d333ead",
|
"ast_hash": "e7948e61b74b84b246708518f02dbcae",
|
||||||
"semantic_hash": "a64188f938dc4f8917335ef41d333ead"
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"src/yt_scraper/parse.py": {
|
"src/yt_scraper/parse.py": {
|
||||||
"mtime": 1785120665.3384771,
|
"mtime": 1785120665.3384771,
|
||||||
@@ -70,9 +70,9 @@
|
|||||||
"semantic_hash": "25354365ac1ab52f1561c97256573806"
|
"semantic_hash": "25354365ac1ab52f1561c97256573806"
|
||||||
},
|
},
|
||||||
"src/yt_scraper/pipeline.py": {
|
"src/yt_scraper/pipeline.py": {
|
||||||
"mtime": 1785131195.7311583,
|
"mtime": 1785569737.341973,
|
||||||
"ast_hash": "ebce1b294140bf40f809848ef280c290",
|
"ast_hash": "4e22e69083b15c4b937fca406134539c",
|
||||||
"semantic_hash": "ebce1b294140bf40f809848ef280c290"
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"src/yt_scraper/ratelimit.py": {
|
"src/yt_scraper/ratelimit.py": {
|
||||||
"mtime": 1785120650.3795395,
|
"mtime": 1785120650.3795395,
|
||||||
@@ -85,14 +85,14 @@
|
|||||||
"semantic_hash": "ff83764f3a8c8874557993e11eab1f89"
|
"semantic_hash": "ff83764f3a8c8874557993e11eab1f89"
|
||||||
},
|
},
|
||||||
"src/yt_scraper/segments.py": {
|
"src/yt_scraper/segments.py": {
|
||||||
"mtime": 1785128385.6652262,
|
"mtime": 1785562183.2376134,
|
||||||
"ast_hash": "09c7da35bc09a3f75fd0a627c0534746",
|
"ast_hash": "902d851e321b83c99d4cae30c1ae39fc",
|
||||||
"semantic_hash": "09c7da35bc09a3f75fd0a627c0534746"
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"src/yt_scraper/store.py": {
|
"src/yt_scraper/store.py": {
|
||||||
"mtime": 1785131583.548138,
|
"mtime": 1785569922.3510046,
|
||||||
"ast_hash": "9a3d4286214aa54ae4ac4df2097b643d",
|
"ast_hash": "ea9a69606929df7ff0508814f2733204",
|
||||||
"semantic_hash": "9a3d4286214aa54ae4ac4df2097b643d"
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"src/yt_scraper/webapp/__init__.py": {
|
"src/yt_scraper/webapp/__init__.py": {
|
||||||
"mtime": 1785126827.4957752,
|
"mtime": 1785126827.4957752,
|
||||||
@@ -100,24 +100,24 @@
|
|||||||
"semantic_hash": "d41d8cd98f00b204e9800998ecf8427e"
|
"semantic_hash": "d41d8cd98f00b204e9800998ecf8427e"
|
||||||
},
|
},
|
||||||
"src/yt_scraper/webapp/api.py": {
|
"src/yt_scraper/webapp/api.py": {
|
||||||
"mtime": 1785131595.9914868,
|
"mtime": 1785569846.188889,
|
||||||
"ast_hash": "e64174f04d6eab0747b1d2a7c2a96a01",
|
"ast_hash": "e63b87bc3804f92f9078cdf36d7181da",
|
||||||
"semantic_hash": "e64174f04d6eab0747b1d2a7c2a96a01"
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"src/yt_scraper/webapp/app.py": {
|
"src/yt_scraper/webapp/app.py": {
|
||||||
"mtime": 1785126888.8032887,
|
"mtime": 1785562107.3107448,
|
||||||
"ast_hash": "e0b1925b4318305dcf746d360373ff50",
|
"ast_hash": "80e375becb41e202f78239138252c5c8",
|
||||||
"semantic_hash": "e0b1925b4318305dcf746d360373ff50"
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"src/yt_scraper/webapp/jobs.py": {
|
"src/yt_scraper/webapp/jobs.py": {
|
||||||
"mtime": 1785131211.0397103,
|
"mtime": 1785556090.2023566,
|
||||||
"ast_hash": "8b6c80001ff8c95eaaa16f7b61d1f207",
|
"ast_hash": "45ba7f10b13b847c213f04441a5e741d",
|
||||||
"semantic_hash": "8b6c80001ff8c95eaaa16f7b61d1f207"
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"src/yt_scraper/webapp/static/app.js": {
|
"src/yt_scraper/webapp/static/app.js": {
|
||||||
"mtime": 1785131664.8792574,
|
"mtime": 1785569873.551512,
|
||||||
"ast_hash": "f6b081fadc5884d379152abbd1e59700",
|
"ast_hash": "0f4fc27cfdde0a10ef3c568774b199a2",
|
||||||
"semantic_hash": "f6b081fadc5884d379152abbd1e59700"
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"tests/test_chapters.py": {
|
"tests/test_chapters.py": {
|
||||||
"mtime": 1785120789.3300896,
|
"mtime": 1785120789.3300896,
|
||||||
@@ -140,9 +140,9 @@
|
|||||||
"semantic_hash": "27125960e235813724fd669001faab59"
|
"semantic_hash": "27125960e235813724fd669001faab59"
|
||||||
},
|
},
|
||||||
"tests/test_store_platform.py": {
|
"tests/test_store_platform.py": {
|
||||||
"mtime": 1785126662.4772756,
|
"mtime": 1785194588.4186954,
|
||||||
"ast_hash": "bad13dcf37be7a414ffae8018c29e55b",
|
"ast_hash": "44069db75a444906fb5001880dedf180",
|
||||||
"semantic_hash": "bad13dcf37be7a414ffae8018c29e55b"
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
".opencode/agent/webapp-builder.md": {
|
".opencode/agent/webapp-builder.md": {
|
||||||
"mtime": 1785126783.7152567,
|
"mtime": 1785126783.7152567,
|
||||||
@@ -165,9 +165,9 @@
|
|||||||
"semantic_hash": "a0a8e6b62a991fce5d6c58171f80075e"
|
"semantic_hash": "a0a8e6b62a991fce5d6c58171f80075e"
|
||||||
},
|
},
|
||||||
"config.example.yaml": {
|
"config.example.yaml": {
|
||||||
"mtime": 1785120602.4114335,
|
"mtime": 1785561451.814817,
|
||||||
"ast_hash": "415b0fc8e71ecc305d21e9c3c08d585a",
|
"ast_hash": "43661d5c00d05c922bcf884e8ac707c3",
|
||||||
"semantic_hash": "415b0fc8e71ecc305d21e9c3c08d585a"
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"docs/superpowers/plans/2026-07-26-platform-build.md": {
|
"docs/superpowers/plans/2026-07-26-platform-build.md": {
|
||||||
"mtime": 1785126240.4211702,
|
"mtime": 1785126240.4211702,
|
||||||
@@ -180,8 +180,78 @@
|
|||||||
"semantic_hash": "9b85f02fd45c93cdabfd2e1c6d1e0c9e"
|
"semantic_hash": "9b85f02fd45c93cdabfd2e1c6d1e0c9e"
|
||||||
},
|
},
|
||||||
"src/yt_scraper/webapp/static/index.html": {
|
"src/yt_scraper/webapp/static/index.html": {
|
||||||
"mtime": 1785131695.7653384,
|
"mtime": 1785570599.948749,
|
||||||
"ast_hash": "13591b7c8c662d8e7c2540f3032ae26d",
|
"ast_hash": "46b1cb57281f7683acc5b39714d74073",
|
||||||
"semantic_hash": "13591b7c8c662d8e7c2540f3032ae26d"
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"src/yt_scraper/_yt_http.py": {
|
||||||
|
"mtime": 1785134945.1512043,
|
||||||
|
"ast_hash": "8af224580322df5aeeba10062592d961",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_config.py": {
|
||||||
|
"mtime": 1785561477.3907576,
|
||||||
|
"ast_hash": "46f3eefb24bcf44535c17977ac175509",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_discover_incremental.py": {
|
||||||
|
"mtime": 1785555704.1061945,
|
||||||
|
"ast_hash": "e1c382a9b8f5c927a85d67814d3dbb09",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_extract.py": {
|
||||||
|
"mtime": 1785561483.422817,
|
||||||
|
"ast_hash": "9949ef4f6e7fe8cbdf12bc8d1dc71e3f",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_store_sync_watermark.py": {
|
||||||
|
"mtime": 1785555725.6495814,
|
||||||
|
"ast_hash": "9f9943f5b914078cee49bf4aaa060564",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_webapp_jobs.py": {
|
||||||
|
"mtime": 1785555746.4912646,
|
||||||
|
"ast_hash": "7688c701256ff0cdd2a7d78d7090c8b3",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"CLAUDE.md": {
|
||||||
|
"mtime": 1785570018.086969,
|
||||||
|
"ast_hash": "e478d29b66549d4a24cdd0f91ddfc2aa",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"YOUTUBE-TRANSCRIPT-AUDIT.md": {
|
||||||
|
"mtime": 1785116715.5118847,
|
||||||
|
"ast_hash": "67c5b7e3b4661ca8c8a7280e63b13ba9",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"docs/superpowers/plans/2026-07-27-discovery-only-scrape.md": {
|
||||||
|
"mtime": 1785192916.5353482,
|
||||||
|
"ast_hash": "3cf41d6331bd67e39ccd96b0e6138cf2",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"docs/superpowers/specs/2026-07-27-discovery-only-scrape-design.md": {
|
||||||
|
"mtime": 1785192898.2889626,
|
||||||
|
"ast_hash": "529b39e9bd2adcafd17a2dc30263281c",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_recovery.py": {
|
||||||
|
"mtime": 1785569307.786051,
|
||||||
|
"ast_hash": "53f25aeda88cbf3755f45f2377970b62",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_since_filter.py": {
|
||||||
|
"mtime": 1785556104.7475204,
|
||||||
|
"ast_hash": "1d827c61156ab488e838dac91261624f",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
".claude/settings.json": {
|
||||||
|
"mtime": 1785569155.4443069,
|
||||||
|
"ast_hash": "dc072413449915d630bc773c9fb57dc7",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_blocked_videos.py": {
|
||||||
|
"mtime": 1785569898.4683661,
|
||||||
|
"ast_hash": "674b087c46f1ee8de0e650295f0cb8cd",
|
||||||
|
"semantic_hash": ""
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
+1
-1
@@ -8,7 +8,7 @@ version = "0.1.0"
|
|||||||
description = "Scraper de canales de YouTube hacia notas Markdown para Obsidian"
|
description = "Scraper de canales de YouTube hacia notas Markdown para Obsidian"
|
||||||
requires-python = ">=3.10"
|
requires-python = ">=3.10"
|
||||||
dependencies = [
|
dependencies = [
|
||||||
"yt-dlp>=2024.10.7",
|
"yt-dlp[default]>=2026.8.19",
|
||||||
"click>=8.1",
|
"click>=8.1",
|
||||||
"rich>=13.7",
|
"rich>=13.7",
|
||||||
"jinja2>=3.1",
|
"jinja2>=3.1",
|
||||||
|
|||||||
@@ -0,0 +1,50 @@
|
|||||||
|
"""Shared HTTP helpers for talking to YouTube CDNs.
|
||||||
|
|
||||||
|
Single source of truth for the headers we send to ``*.youtube.com`` endpoints
|
||||||
|
(``ytimg.com``, ``ggpht.com``, ``youtube.com/api/timedtext``). Centralising
|
||||||
|
them prevents the failures we hit when ``Referer: https://www.youtube.com/``
|
||||||
|
was missing on avatar/thumbnail/subtitle fetches.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from typing import Any, Mapping
|
||||||
|
|
||||||
|
import requests
|
||||||
|
|
||||||
|
|
||||||
|
_YT_HEADERS: dict[str, str] = {
|
||||||
|
# Chrome 124 on Windows. Mimics a real browser request; some Google CDN
|
||||||
|
# endpoints (notably ``yt3.ggpht.com`` avatars and ``timedtext`` captions)
|
||||||
|
# reject ``python-requests`` style UA-only calls with HTTP 403.
|
||||||
|
"User-Agent": (
|
||||||
|
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
|
||||||
|
"(KHTML, like Gecko) Chrome/124.0 Safari/537.36"
|
||||||
|
),
|
||||||
|
"Referer": "https://www.youtube.com/",
|
||||||
|
"Accept-Language": "en-US,en;q=0.9",
|
||||||
|
# Image fetches want an Accept that lists image/* so the CDN can pick
|
||||||
|
# the right encoded variant (avif/webp/etc).
|
||||||
|
"Accept": (
|
||||||
|
"image/avif,image/webp,image/png,image/jpeg,image/*,*/*;q=0.8"
|
||||||
|
),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def yt_headers(extra: Mapping[str, str] | None = None) -> dict[str, str]:
|
||||||
|
"""Return a copy of the YouTube CDN headers merged with ``extra``."""
|
||||||
|
if extra:
|
||||||
|
return {**_YT_HEADERS, **dict(extra)}
|
||||||
|
return dict(_YT_HEADERS)
|
||||||
|
|
||||||
|
|
||||||
|
def yt_get(url: str, *, timeout: float = 20.0,
|
||||||
|
headers: Mapping[str, str] | None = None,
|
||||||
|
params: Mapping[str, Any] | None = None) -> requests.Response:
|
||||||
|
"""GET ``url`` with the shared YouTube CDN headers.
|
||||||
|
|
||||||
|
Centralising this lets every call site (subtitle download, video
|
||||||
|
thumbnail cache, channel avatar cache) inherit header updates from a
|
||||||
|
single place.
|
||||||
|
"""
|
||||||
|
return requests.get(url, params=params, headers=yt_headers(headers),
|
||||||
|
timeout=timeout)
|
||||||
+154
-14
@@ -3,6 +3,7 @@ from __future__ import annotations
|
|||||||
import logging
|
import logging
|
||||||
import shutil
|
import shutil
|
||||||
import sys
|
import sys
|
||||||
|
import time
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from types import SimpleNamespace
|
from types import SimpleNamespace
|
||||||
|
|
||||||
@@ -11,13 +12,13 @@ from rich.console import Console
|
|||||||
from rich.progress import Progress, SpinnerColumn, TextColumn, BarColumn, TaskProgressColumn, TimeRemainingColumn
|
from rich.progress import Progress, SpinnerColumn, TextColumn, BarColumn, TaskProgressColumn, TimeRemainingColumn
|
||||||
from rich.table import Table
|
from rich.table import Table
|
||||||
|
|
||||||
from .config import Config, load_config
|
from .config import Config, load_config, parse_languages
|
||||||
from .store import Store, VideoRef, VideoRow
|
from .store import Store, VideoRef, VideoRow
|
||||||
from .discover import discover_channel
|
from .discover import discover_channel, discover_incremental
|
||||||
from .chapters import align_chapters, chapters_from_info, Chapter, Section
|
from .chapters import align_chapters, chapters_from_info, Chapter, Section
|
||||||
from .parse import Segment
|
from .parse import Segment
|
||||||
from .render import build_filename_stem, render_markdown
|
from .render import build_filename_stem, render_markdown
|
||||||
from .ratelimit import polite_sleep
|
from .ratelimit import ThrottleGuard, configure_global_pacer, polite_sleep
|
||||||
from .pipeline import process_video
|
from .pipeline import process_video
|
||||||
from .cookies import auto_import_dir, resolve_active_path
|
from .cookies import auto_import_dir, resolve_active_path
|
||||||
|
|
||||||
@@ -51,6 +52,7 @@ def cli(ctx, config_path, channel, all_channels, cookies, cookies_from_browser,
|
|||||||
cfg = load_config(cfg_path) if cfg_path.exists() else Config()
|
cfg = load_config(cfg_path) if cfg_path.exists() else Config()
|
||||||
if channel:
|
if channel:
|
||||||
cfg.channel_url = channel
|
cfg.channel_url = channel
|
||||||
|
configure_global_pacer(cfg.delay.min_request_interval)
|
||||||
store = Store(cfg.database_path_resolved)
|
store = Store(cfg.database_path_resolved)
|
||||||
# import any loose cookie files into the vault (idempotent)
|
# import any loose cookie files into the vault (idempotent)
|
||||||
try:
|
try:
|
||||||
@@ -72,7 +74,10 @@ def cli(ctx, config_path, channel, all_channels, cookies, cookies_from_browser,
|
|||||||
def _apply_filters(refs, since, no_shorts, no_live, min_duration, limit):
|
def _apply_filters(refs, since, no_shorts, no_live, min_duration, limit):
|
||||||
filtered = list(refs)
|
filtered = list(refs)
|
||||||
if since:
|
if since:
|
||||||
filtered = [r for r in filtered if (r.upload_date or "") >= since.replace("-", "")]
|
# Undated entries survive: yt-dlp's flat listing omits upload_date, so a
|
||||||
|
# plain `>= since` would silently drop every video discovery found.
|
||||||
|
cutoff = since.replace("-", "")
|
||||||
|
filtered = [r for r in filtered if not r.upload_date or r.upload_date >= cutoff]
|
||||||
if no_shorts:
|
if no_shorts:
|
||||||
filtered = [r for r in filtered if "/shorts/" not in (r.url or "")]
|
filtered = [r for r in filtered if "/shorts/" not in (r.url or "")]
|
||||||
if no_live:
|
if no_live:
|
||||||
@@ -95,20 +100,24 @@ def _apply_filters(refs, since, no_shorts, no_live, min_duration, limit):
|
|||||||
@click.option("--resume/--no-resume", default=True)
|
@click.option("--resume/--no-resume", default=True)
|
||||||
@click.option("--dry-run", is_flag=True, default=False)
|
@click.option("--dry-run", is_flag=True, default=False)
|
||||||
@click.option("--reset-errors", is_flag=True, default=False)
|
@click.option("--reset-errors", is_flag=True, default=False)
|
||||||
|
@click.option("--full", is_flag=True, default=False,
|
||||||
|
help="Recorrer el canal entero en vez de solo los videos nuevos")
|
||||||
@click.pass_obj
|
@click.pass_obj
|
||||||
def scrape(obj, limit, since, languages, no_auto, no_shorts, include_shorts, no_live, resume, dry_run, reset_errors):
|
def scrape(obj, limit, since, languages, no_auto, no_shorts, include_shorts, no_live, resume, dry_run, reset_errors, full):
|
||||||
"""Scrape completo: discovery + extraccion + markdown."""
|
"""Scrape completo: discovery + extraccion + markdown."""
|
||||||
_run_scrape(obj, limit, since, languages, no_auto, no_shorts, include_shorts, no_live, resume, dry_run, reset_errors)
|
_run_scrape(obj, limit, since, languages, no_auto, no_shorts, include_shorts, no_live, resume, dry_run, reset_errors, full)
|
||||||
|
|
||||||
|
|
||||||
def _run_scrape(obj, limit, since, languages, no_auto, no_shorts, include_shorts, no_live, resume, dry_run, reset_errors):
|
def _run_scrape(obj, limit, since, languages, no_auto, no_shorts, include_shorts, no_live, resume, dry_run, reset_errors, full=False):
|
||||||
cfg: Config = obj.cfg
|
cfg: Config = obj.cfg
|
||||||
store: Store = obj.store
|
store: Store = obj.store
|
||||||
if not cfg.channel_url:
|
if not cfg.channel_url:
|
||||||
console.print("[red]Error:[/red] falta la URL del canal. Usa --channel o config.yaml")
|
console.print("[red]Error:[/red] falta la URL del canal. Usa --channel o config.yaml")
|
||||||
sys.exit(1)
|
sys.exit(1)
|
||||||
if languages:
|
if languages:
|
||||||
cfg.languages = [l.strip() for l in languages.split(",") if l.strip()]
|
# CLI flag accepts legacy CSV; per-language mode requires editing YAML.
|
||||||
|
list_value = [l.strip() for l in languages.split(",") if l.strip()]
|
||||||
|
cfg.languages = parse_languages(list_value, cfg.prefer_manual)
|
||||||
if no_auto:
|
if no_auto:
|
||||||
cfg.prefer_manual = True
|
cfg.prefer_manual = True
|
||||||
if include_shorts:
|
if include_shorts:
|
||||||
@@ -119,22 +128,52 @@ def _run_scrape(obj, limit, since, languages, no_auto, no_shorts, include_shorts
|
|||||||
cfg.include_live = False
|
cfg.include_live = False
|
||||||
|
|
||||||
console.print(f"[cyan]Canal:[/cyan] {cfg.channel_url}\n[cyan]DB:[/cyan] {cfg.database_path_resolved}")
|
console.print(f"[cyan]Canal:[/cyan] {cfg.channel_url}\n[cyan]DB:[/cyan] {cfg.database_path_resolved}")
|
||||||
console.print("\n[bold blue]Paso 1:[/bold blue] Discovery")
|
|
||||||
|
# Incremental by default: only pull pages newer than the last video we have.
|
||||||
|
known_channel = _known_channel_for(store, cfg.channel_url)
|
||||||
|
incremental = cfg.sync.incremental and not full and bool(known_channel)
|
||||||
|
known = store.known_video_ids(known_channel) if incremental else set()
|
||||||
|
watermark = store.latest_upload_date(known_channel) if incremental else None
|
||||||
|
|
||||||
|
mode = (
|
||||||
|
f"incremental (desde {_fmt_date(watermark)} en adelante)" if watermark
|
||||||
|
else "incremental (solo videos nuevos)" if incremental
|
||||||
|
else "completo"
|
||||||
|
)
|
||||||
|
console.print(f"\n[bold blue]Paso 1:[/bold blue] Discovery — [dim]{mode}[/dim]")
|
||||||
try:
|
try:
|
||||||
with Progress(SpinnerColumn(), TextColumn("[progress.description]{task.description}"), transient=True) as prog:
|
with Progress(SpinnerColumn(), TextColumn("[progress.description]{task.description}"), transient=True) as prog:
|
||||||
task = prog.add_task("Descubriendo videos...", total=None)
|
task = prog.add_task("Descubriendo videos...", total=None)
|
||||||
channel_id, channel_name, refs = discover_channel(cfg.channel_url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests)
|
result = discover_incremental(
|
||||||
|
cfg.channel_url, known,
|
||||||
|
sleep_subrequests=cfg.yt_dlp.sleep_subrequests,
|
||||||
|
window=cfg.sync.window, max_window=cfg.sync.max_window,
|
||||||
|
overlap=cfg.sync.overlap, since=watermark,
|
||||||
|
)
|
||||||
prog.update(task, completed=1, total=1)
|
prog.update(task, completed=1, total=1)
|
||||||
except Exception as exc:
|
except Exception as exc:
|
||||||
console.print(f"[red]Error en discovery:[/red] {exc}")
|
console.print(f"[red]Error en discovery:[/red] {exc}")
|
||||||
sys.exit(1)
|
sys.exit(1)
|
||||||
|
|
||||||
|
channel_id, channel_name, refs = result.channel_id, result.channel_name, result.refs
|
||||||
|
if result.full_scan:
|
||||||
store.upsert_channel(channel_id, _extract_handle(cfg.channel_url), channel_name, len(refs))
|
store.upsert_channel(channel_id, _extract_handle(cfg.channel_url), channel_name, len(refs))
|
||||||
console.print(f"[green]Canal:[/green] {channel_name} ({channel_id}) — {len(refs)} videos")
|
else:
|
||||||
|
store.update_channel_meta(channel_id, name=channel_name)
|
||||||
|
console.print(
|
||||||
|
f"[green]Canal:[/green] {channel_name} ({channel_id}) — "
|
||||||
|
f"{result.fetched} videos leidos en {result.passes} pasada(s), {result.new_count} nuevos"
|
||||||
|
)
|
||||||
|
if not result.caught_up:
|
||||||
|
console.print(
|
||||||
|
"[yellow]Aviso:[/yellow] se alcanzo el limite de ventana "
|
||||||
|
f"({cfg.sync.max_window}). Usa --full si faltan videos antiguos."
|
||||||
|
)
|
||||||
|
|
||||||
refs = _apply_filters(refs, since, not cfg.include_shorts, not cfg.include_live, cfg.min_duration_sec, limit)
|
refs = _apply_filters(refs, since, not cfg.include_shorts, not cfg.include_live, cfg.min_duration_sec, limit)
|
||||||
console.print(f"[yellow]Tras filtros:[/yellow] {len(refs)} videos")
|
console.print(f"[yellow]Tras filtros:[/yellow] {len(refs)} videos")
|
||||||
store.upsert_videos(refs)
|
store.upsert_videos(refs)
|
||||||
|
store.mark_channel_synced(channel_id)
|
||||||
|
|
||||||
if reset_errors:
|
if reset_errors:
|
||||||
n = store.reset_errors(channel_id)
|
n = store.reset_errors(channel_id)
|
||||||
@@ -145,8 +184,13 @@ def _run_scrape(obj, limit, since, languages, no_auto, no_shorts, include_shorts
|
|||||||
return
|
return
|
||||||
|
|
||||||
pending = store.get_pending(channel_id) if resume else store.get_all(channel_id)
|
pending = store.get_pending(channel_id) if resume else store.get_all(channel_id)
|
||||||
|
if result.full_scan:
|
||||||
ref_ids = {r.video_id for r in refs}
|
ref_ids = {r.video_id for r in refs}
|
||||||
pending = [p for p in pending if p.video_id in ref_ids] if ref_ids else pending
|
pending = [p for p in pending if p.video_id in ref_ids] if ref_ids else pending
|
||||||
|
else:
|
||||||
|
# Discovery only saw the newest slice; keep the older backlog reachable
|
||||||
|
# but put this run's videos first so --limit still means "los mas nuevos".
|
||||||
|
pending = _order_pending(pending, refs)
|
||||||
if limit:
|
if limit:
|
||||||
pending = pending[:limit]
|
pending = pending[:limit]
|
||||||
if not pending:
|
if not pending:
|
||||||
@@ -162,17 +206,64 @@ def _run_scrape(obj, limit, since, languages, no_auto, no_shorts, include_shorts
|
|||||||
with Progress(SpinnerColumn(), TextColumn("[progress.description]{task.description}"),
|
with Progress(SpinnerColumn(), TextColumn("[progress.description]{task.description}"),
|
||||||
BarColumn(), TaskProgressColumn(), TimeRemainingColumn()) as progress:
|
BarColumn(), TaskProgressColumn(), TimeRemainingColumn()) as progress:
|
||||||
task = progress.add_task("Procesando", total=len(pending))
|
task = progress.add_task("Procesando", total=len(pending))
|
||||||
|
guard = ThrottleGuard(
|
||||||
|
threshold=cfg.delay.throttle_threshold,
|
||||||
|
base=cfg.delay.backoff_base,
|
||||||
|
cap=cfg.delay.backoff_cap,
|
||||||
|
)
|
||||||
for i, row in enumerate(pending):
|
for i, row in enumerate(pending):
|
||||||
progress.update(task, description=f"{row.video_id} {(row.title or '')[:30]}", completed=i)
|
progress.update(task, description=f"{row.video_id} {(row.title or '')[:30]}", completed=i)
|
||||||
process_video(row, cfg, store, channel_name, channel_id, cfg.channel_url,
|
status = process_video(row, cfg, store, channel_name, channel_id, cfg.channel_url,
|
||||||
cookies_file=cookie_path, cookies_from_browser=obj.cookies_from_browser)
|
cookies_file=cookie_path, cookies_from_browser=obj.cookies_from_browser)
|
||||||
progress.advance(task)
|
progress.advance(task)
|
||||||
|
if status == "done":
|
||||||
|
guard.note_success()
|
||||||
|
else:
|
||||||
|
failed = store.get_video(row.video_id)
|
||||||
|
wait = guard.note_failure(getattr(failed, "error_msg", None) or status)
|
||||||
|
if guard.tripped:
|
||||||
|
# Everything still pending stays pending: that is what makes
|
||||||
|
# this recoverable instead of 500 rows marked failed.
|
||||||
|
console.print(f"\n[bold red]Detenido:[/bold red] {guard.tripped_reason}")
|
||||||
|
console.print(f"[yellow]{len(pending) - i - 1} videos sin tocar, siguen pendientes.[/yellow]")
|
||||||
|
break
|
||||||
|
if wait > 0:
|
||||||
|
console.print(f"[yellow]YouTube nos esta limitando; esperando {wait:.1f}s[/yellow]")
|
||||||
|
time.sleep(wait)
|
||||||
if i < len(pending) - 1:
|
if i < len(pending) - 1:
|
||||||
polite_sleep(cfg.delay.min_seconds, cfg.delay.max_seconds)
|
polite_sleep(cfg.delay.min_seconds, cfg.delay.max_seconds)
|
||||||
console.print()
|
console.print()
|
||||||
_print_stats(store, channel_id)
|
_print_stats(store, channel_id)
|
||||||
|
|
||||||
|
|
||||||
|
def _known_channel_for(store: Store, channel_url: str) -> str | None:
|
||||||
|
"""Match a channel URL against an already-tracked channel, by handle or id.
|
||||||
|
|
||||||
|
Without a hit there is no local history to stop at, so discovery has to walk
|
||||||
|
the whole channel — which is correct for a first run.
|
||||||
|
"""
|
||||||
|
if not channel_url:
|
||||||
|
return None
|
||||||
|
handle = _extract_handle(channel_url).lstrip("@").lower()
|
||||||
|
tail = channel_url.rstrip("/").split("/")[-1]
|
||||||
|
for ch in store.list_channels():
|
||||||
|
cid = ch.get("channel_id") or ""
|
||||||
|
ch_handle = (ch.get("handle") or "").lstrip("@").lower()
|
||||||
|
if handle and ch_handle == handle:
|
||||||
|
return cid
|
||||||
|
if cid and (cid == tail or f"/channel/{cid}" in channel_url):
|
||||||
|
return cid
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def _order_pending(rows: list, refs: list) -> list:
|
||||||
|
"""Videos from this run's window first, then the rest of the backlog."""
|
||||||
|
by_id = {r.video_id: r for r in rows}
|
||||||
|
ordered = [by_id.pop(r.video_id) for r in refs if r.video_id in by_id]
|
||||||
|
ordered.extend(by_id.values())
|
||||||
|
return ordered
|
||||||
|
|
||||||
|
|
||||||
def _print_dry_run(refs):
|
def _print_dry_run(refs):
|
||||||
table = Table(show_lines=False)
|
table = Table(show_lines=False)
|
||||||
table.add_column("Fecha", style="dim")
|
table.add_column("Fecha", style="dim")
|
||||||
@@ -186,6 +277,55 @@ def _print_dry_run(refs):
|
|||||||
console.print(f"... y {len(refs) - 50} mas")
|
console.print(f"... y {len(refs) - 50} mas")
|
||||||
|
|
||||||
|
|
||||||
|
# --------------------------------------------------------------------------- recovery
|
||||||
|
|
||||||
|
@cli.command("reset")
|
||||||
|
@click.option("--status", "statuses", multiple=True,
|
||||||
|
type=click.Choice(Store.RETRYABLE_STATUSES),
|
||||||
|
help="Estados a reintentar (por defecto: error y no_subtitles)")
|
||||||
|
@click.option("--yes", is_flag=True, default=False, help="No preguntar")
|
||||||
|
@click.pass_obj
|
||||||
|
def reset_cmd(obj, statuses, yes):
|
||||||
|
"""Devolver videos atascados en error/no_subtitles a 'pending' para reintentarlos."""
|
||||||
|
store: Store = obj.store
|
||||||
|
targets = tuple(statuses) if statuses else Store.RETRYABLE_STATUSES
|
||||||
|
channels = _channel_targets(obj)
|
||||||
|
channel_id = channels[0] if len(channels) == 1 else None
|
||||||
|
|
||||||
|
counts = store.retryable_counts(channel_id)
|
||||||
|
affected = sum(counts.get(s, 0) for s in targets)
|
||||||
|
if not affected:
|
||||||
|
console.print("[green]No hay videos que reintentar.[/green]")
|
||||||
|
return
|
||||||
|
detail = ", ".join(f"{s}={counts.get(s, 0)}" for s in targets)
|
||||||
|
scope = channel_id or "todos los canales"
|
||||||
|
if not yes and not click.confirm(f"Reintentar {affected} videos ({detail}) en {scope}?"):
|
||||||
|
return
|
||||||
|
n = store.reset_videos(channel_id, targets)
|
||||||
|
console.print(f"[yellow]{n}[/yellow] videos vueltos a 'pending'.")
|
||||||
|
|
||||||
|
|
||||||
|
@cli.command("reconcile")
|
||||||
|
@click.option("--prune", is_flag=True, default=False,
|
||||||
|
help="Ademas, devolver a 'pending' los 'done' cuyo .md ya no existe")
|
||||||
|
@click.pass_obj
|
||||||
|
def reconcile_cmd(obj, prune):
|
||||||
|
"""Re-escanear data/markdown y hacer que la DB coincida con el disco."""
|
||||||
|
from .segments import reconcile_markdown
|
||||||
|
cfg: Config = obj.cfg
|
||||||
|
store: Store = obj.store
|
||||||
|
result = reconcile_markdown(
|
||||||
|
store, Path(cfg.output_dir_resolved),
|
||||||
|
log=lambda m: console.print(f"[dim]{m}[/dim]"), prune=prune,
|
||||||
|
)
|
||||||
|
table = Table(title="Reconciliacion disco <-> DB")
|
||||||
|
table.add_column("Metrica", style="bold")
|
||||||
|
table.add_column("Valor", justify="right")
|
||||||
|
for k, v in result.items():
|
||||||
|
table.add_row(k, str(v))
|
||||||
|
console.print(table)
|
||||||
|
|
||||||
|
|
||||||
# --------------------------------------------------------------------------- search
|
# --------------------------------------------------------------------------- search
|
||||||
|
|
||||||
@cli.command("search")
|
@cli.command("search")
|
||||||
@@ -313,8 +453,8 @@ def channels_add(obj, url):
|
|||||||
cfg: Config = obj.cfg
|
cfg: Config = obj.cfg
|
||||||
cfg.channel_url = url
|
cfg.channel_url = url
|
||||||
console.print("[cyan]Resolviendo canal...[/cyan]")
|
console.print("[cyan]Resolviendo canal...[/cyan]")
|
||||||
channel_id, name, refs = discover_channel(url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests)
|
channel_id, name, avatar, refs = discover_channel(url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests)
|
||||||
obj.store.upsert_channel(channel_id, _extract_handle(url), name, len(refs))
|
obj.store.upsert_channel(channel_id, _extract_handle(url), name, len(refs), avatar=avatar)
|
||||||
obj.store.upsert_videos(refs)
|
obj.store.upsert_videos(refs)
|
||||||
console.print(f"[green]Added:[/green] {name} ({channel_id}) — {len(refs)} videos")
|
console.print(f"[green]Added:[/green] {name} ({channel_id}) — {len(refs)} videos")
|
||||||
|
|
||||||
|
|||||||
@@ -9,28 +9,79 @@ import yaml
|
|||||||
|
|
||||||
@dataclass
|
@dataclass
|
||||||
class DelayConfig:
|
class DelayConfig:
|
||||||
|
"""Pacing and failure handling for everything that touches YouTube.
|
||||||
|
|
||||||
|
`min_seconds`/`max_seconds` are the randomised gap between units of work
|
||||||
|
(one video, one channel). `backoff_*` and `throttle_threshold` only come
|
||||||
|
into play once YouTube starts refusing: the backoff pair feeds the
|
||||||
|
truncated-exponential-with-jitter formula Google documents for its own
|
||||||
|
APIs, and the threshold is how many consecutive rate-limit responses end
|
||||||
|
the run instead of burning through the rest of the queue.
|
||||||
|
"""
|
||||||
|
|
||||||
min_seconds: float = 1.5
|
min_seconds: float = 1.5
|
||||||
max_seconds: float = 3.5
|
max_seconds: float = 3.5
|
||||||
backoff_base: float = 2.0
|
backoff_base: float = 2.0
|
||||||
backoff_cap: float = 60.0
|
backoff_cap: float = 60.0
|
||||||
|
throttle_threshold: int = 3
|
||||||
|
# Minimum gap between *any* two requests to YouTube from this process,
|
||||||
|
# including the ones the `/api/tools/*` endpoints make outside the job
|
||||||
|
# runner. 0 disables the shared pacer and leaves only the per-unit sleeps.
|
||||||
|
min_request_interval: float = 0.0
|
||||||
|
# Bytes/sec ceiling for audio downloads (yt-dlp `ratelimit`). 0 = unlimited.
|
||||||
|
audio_rate_limit: int = 0
|
||||||
|
|
||||||
|
|
||||||
@dataclass
|
@dataclass
|
||||||
class YtDlpConfig:
|
class YtDlpConfig:
|
||||||
retries: int = 10
|
retries: int = 10
|
||||||
|
# Seconds between the individual HTTP requests inside one extraction.
|
||||||
|
# Passed to yt-dlp as `sleep_interval_requests`; the name here is the
|
||||||
|
# project's own and predates the discovery that the option this used to be
|
||||||
|
# forwarded under (`sleep_subrequests`) does not exist in yt-dlp at all.
|
||||||
sleep_subrequests: float = 2.0
|
sleep_subrequests: float = 2.0
|
||||||
|
extractor_retries: int = 3
|
||||||
|
socket_timeout: float = 30.0
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class SyncConfig:
|
||||||
|
"""Incremental channel sync — how far back a routine re-scan looks.
|
||||||
|
|
||||||
|
The /videos tab is reverse-chronological and every extra page is another
|
||||||
|
request to YouTube, so a sync fetches `window` entries and stops as soon as
|
||||||
|
it has seen `overlap` consecutive videos already in the DB. Only if the
|
||||||
|
whole window turns out to be new does it widen (doubling up to `max_window`),
|
||||||
|
which is the case where the channel really did publish a lot since last time.
|
||||||
|
"""
|
||||||
|
|
||||||
|
incremental: bool = True
|
||||||
|
window: int = 30
|
||||||
|
max_window: int = 300
|
||||||
|
overlap: int = 3
|
||||||
|
|
||||||
|
|
||||||
@dataclass
|
@dataclass
|
||||||
class Config:
|
class Config:
|
||||||
channel_url: str = ""
|
channel_url: str = ""
|
||||||
languages: list[str] = field(default_factory=lambda: ["es", "en"])
|
# Per-language subtitle preference. Each value is one of:
|
||||||
|
# "manual" - only manually uploaded captions
|
||||||
|
# "auto" - only YouTube-auto-generated captions
|
||||||
|
# "any" - defer to the legacy `prefer_manual` flag
|
||||||
|
# Legacy form: ``["es", "en"]`` - treated as ``{lang: <prefer_manual_default>}``.
|
||||||
|
# Default to "any" (manual first, auto as fallback). A manual-only default
|
||||||
|
# silently yields nothing on the many channels that publish only
|
||||||
|
# auto-generated captions, and stores that as `no_subtitles`.
|
||||||
|
languages: dict[str, str] = field(
|
||||||
|
default_factory=lambda: {"es": "any", "en": "any"}
|
||||||
|
)
|
||||||
prefer_manual: bool = True
|
prefer_manual: bool = True
|
||||||
include_shorts: bool = False
|
include_shorts: bool = False
|
||||||
include_live: bool = True
|
include_live: bool = True
|
||||||
min_duration_sec: int = 0
|
min_duration_sec: int = 0
|
||||||
delay: DelayConfig = field(default_factory=DelayConfig)
|
delay: DelayConfig = field(default_factory=DelayConfig)
|
||||||
yt_dlp: YtDlpConfig = field(default_factory=YtDlpConfig)
|
yt_dlp: YtDlpConfig = field(default_factory=YtDlpConfig)
|
||||||
|
sync: SyncConfig = field(default_factory=SyncConfig)
|
||||||
database_path: str = "data/state.db"
|
database_path: str = "data/state.db"
|
||||||
output_dir: str = "data/markdown"
|
output_dir: str = "data/markdown"
|
||||||
template_path: str = "templates/video.md.j2"
|
template_path: str = "templates/video.md.j2"
|
||||||
@@ -49,6 +100,27 @@ class Config:
|
|||||||
return Path(self.template_path).resolve()
|
return Path(self.template_path).resolve()
|
||||||
|
|
||||||
|
|
||||||
|
# Modes accepted in the ``languages`` dict. Kept here so tests and CLI
|
||||||
|
# share a single definition without importing the private extract constant.
|
||||||
|
LANGUAGE_MODES = ("manual", "auto", "any")
|
||||||
|
|
||||||
|
|
||||||
|
def parse_languages(raw: Any, prefer_manual: bool) -> dict[str, str]:
|
||||||
|
"""Normalise legacy list / new dict / None into ``{lang: mode}``."""
|
||||||
|
default = "manual" if prefer_manual else "auto"
|
||||||
|
if raw is None:
|
||||||
|
return {}
|
||||||
|
if isinstance(raw, dict):
|
||||||
|
out: dict[str, str] = {}
|
||||||
|
for lang, mode in raw.items():
|
||||||
|
m = str(mode).lower().strip()
|
||||||
|
out[str(lang)] = m if m in LANGUAGE_MODES else "any"
|
||||||
|
return out
|
||||||
|
if isinstance(raw, (list, tuple)):
|
||||||
|
return {str(l): default for l in raw}
|
||||||
|
return {}
|
||||||
|
|
||||||
|
|
||||||
def load_config(path: str | Path) -> Config:
|
def load_config(path: str | Path) -> Config:
|
||||||
p = Path(path)
|
p = Path(path)
|
||||||
if not p.exists():
|
if not p.exists():
|
||||||
@@ -61,10 +133,15 @@ def load_config(path: str | Path) -> Config:
|
|||||||
def _build_config(raw: dict[str, Any]) -> Config:
|
def _build_config(raw: dict[str, Any]) -> Config:
|
||||||
delay_raw = raw.get("delay") or {}
|
delay_raw = raw.get("delay") or {}
|
||||||
ydl_raw = raw.get("yt_dlp") or {}
|
ydl_raw = raw.get("yt_dlp") or {}
|
||||||
|
sync_raw = raw.get("sync") or {}
|
||||||
|
prefer_manual = bool(raw.get("prefer_manual", True))
|
||||||
|
languages = parse_languages(raw.get("languages"), prefer_manual)
|
||||||
|
if not languages:
|
||||||
|
languages = {"es": "any", "en": "any"}
|
||||||
return Config(
|
return Config(
|
||||||
channel_url=raw.get("channel_url", ""),
|
channel_url=raw.get("channel_url", ""),
|
||||||
languages=list(raw.get("languages", ["es", "en"])),
|
languages=languages,
|
||||||
prefer_manual=bool(raw.get("prefer_manual", True)),
|
prefer_manual=prefer_manual,
|
||||||
include_shorts=bool(raw.get("include_shorts", False)),
|
include_shorts=bool(raw.get("include_shorts", False)),
|
||||||
include_live=bool(raw.get("include_live", True)),
|
include_live=bool(raw.get("include_live", True)),
|
||||||
min_duration_sec=int(raw.get("min_duration_sec", 0)),
|
min_duration_sec=int(raw.get("min_duration_sec", 0)),
|
||||||
@@ -73,10 +150,21 @@ def _build_config(raw: dict[str, Any]) -> Config:
|
|||||||
max_seconds=float(delay_raw.get("max_seconds", 3.5)),
|
max_seconds=float(delay_raw.get("max_seconds", 3.5)),
|
||||||
backoff_base=float(delay_raw.get("backoff_base", 2.0)),
|
backoff_base=float(delay_raw.get("backoff_base", 2.0)),
|
||||||
backoff_cap=float(delay_raw.get("backoff_cap", 60.0)),
|
backoff_cap=float(delay_raw.get("backoff_cap", 60.0)),
|
||||||
|
throttle_threshold=max(1, int(delay_raw.get("throttle_threshold", 3))),
|
||||||
|
min_request_interval=max(0.0, float(delay_raw.get("min_request_interval", 0.0))),
|
||||||
|
audio_rate_limit=max(0, int(delay_raw.get("audio_rate_limit", 0))),
|
||||||
),
|
),
|
||||||
yt_dlp=YtDlpConfig(
|
yt_dlp=YtDlpConfig(
|
||||||
retries=int(ydl_raw.get("retries", 10)),
|
retries=int(ydl_raw.get("retries", 10)),
|
||||||
sleep_subrequests=float(ydl_raw.get("sleep_subrequests", 2.0)),
|
sleep_subrequests=float(ydl_raw.get("sleep_subrequests", 2.0)),
|
||||||
|
extractor_retries=max(0, int(ydl_raw.get("extractor_retries", 3))),
|
||||||
|
socket_timeout=float(ydl_raw.get("socket_timeout", 30.0)),
|
||||||
|
),
|
||||||
|
sync=SyncConfig(
|
||||||
|
incremental=bool(sync_raw.get("incremental", True)),
|
||||||
|
window=max(1, int(sync_raw.get("window", 30))),
|
||||||
|
max_window=max(1, int(sync_raw.get("max_window", 300))),
|
||||||
|
overlap=max(1, int(sync_raw.get("overlap", 3))),
|
||||||
),
|
),
|
||||||
database_path=raw.get("database_path", "data/state.db"),
|
database_path=raw.get("database_path", "data/state.db"),
|
||||||
output_dir=raw.get("output_dir", "data/markdown"),
|
output_dir=raw.get("output_dir", "data/markdown"),
|
||||||
|
|||||||
+224
-28
@@ -2,12 +2,13 @@ from __future__ import annotations
|
|||||||
|
|
||||||
import logging
|
import logging
|
||||||
from dataclasses import dataclass
|
from dataclasses import dataclass
|
||||||
from typing import Any
|
from typing import Any, Mapping
|
||||||
|
|
||||||
import requests
|
|
||||||
import yt_dlp
|
import yt_dlp
|
||||||
|
|
||||||
|
from ._yt_http import yt_get
|
||||||
from .parse import Segment, parse_auto_dump
|
from .parse import Segment, parse_auto_dump
|
||||||
|
from .ratelimit import GLOBAL_PACER, ydl_throttle_opts
|
||||||
|
|
||||||
log = logging.getLogger(__name__)
|
log = logging.getLogger(__name__)
|
||||||
|
|
||||||
@@ -26,66 +27,190 @@ class VideoData:
|
|||||||
segments: list[Segment]
|
segments: list[Segment]
|
||||||
subtitle: SubtitlePick | None
|
subtitle: SubtitlePick | None
|
||||||
has_chapters: bool
|
has_chapters: bool
|
||||||
|
# Why `segments` came back empty. "No transcript" has several very
|
||||||
|
# different causes — the video genuinely has no captions, the language
|
||||||
|
# policy rejected the tracks that do exist, or the download was throttled —
|
||||||
|
# and collapsing them into one terminal status hides recoverable failures.
|
||||||
|
skip_reason: str | None = None
|
||||||
|
|
||||||
|
|
||||||
|
_LANGUAGE_MODES = ("manual", "auto", "any")
|
||||||
|
|
||||||
|
|
||||||
|
def _sources_for(mode: str, prefer_manual: bool, manual: dict, auto: dict) -> list[tuple[str, dict]]:
|
||||||
|
"""Return ordered list of (label, tracks_dict) to try for ``mode``.
|
||||||
|
|
||||||
|
``any`` defers to the legacy :data:`prefer_manual` global default.
|
||||||
|
``manual`` / ``auto`` force the track family even if the other has
|
||||||
|
a higher-priority language elsewhere in the iteration.
|
||||||
|
"""
|
||||||
|
if mode == "manual":
|
||||||
|
return [("manual", manual)]
|
||||||
|
if mode == "auto":
|
||||||
|
return [("auto", auto)]
|
||||||
|
if prefer_manual:
|
||||||
|
return [("manual", manual), ("auto", auto)]
|
||||||
|
return [("auto", auto), ("manual", manual)]
|
||||||
|
|
||||||
|
|
||||||
def extract_video(
|
def extract_video(
|
||||||
video_url: str,
|
video_url: str,
|
||||||
languages: list[str],
|
languages: Mapping[str, str] | list[str],
|
||||||
retries: int = 10,
|
retries: int = 10,
|
||||||
sleep_subrequests: float = 2.0,
|
sleep_subrequests: float = 2.0,
|
||||||
prefer_manual: bool = True,
|
prefer_manual: bool = True,
|
||||||
cookies_file: str | None = None,
|
cookies_file: str | None = None,
|
||||||
cookies_from_browser: str | None = None,
|
cookies_from_browser: str | None = None,
|
||||||
|
extractor_retries: int = 3,
|
||||||
|
socket_timeout: float = 30.0,
|
||||||
) -> VideoData:
|
) -> VideoData:
|
||||||
|
"""Run yt-dlp on ``video_url`` and pull the preferred subtitle track.
|
||||||
|
|
||||||
|
``languages`` may be either a list (legacy, every entry uses
|
||||||
|
``prefer_manual``) or a mapping ``{lang: mode}`` where ``mode`` is one
|
||||||
|
of ``"manual"``, ``"auto"`` or ``"any"``. The mapping form is the
|
||||||
|
preferred interface because it lets you mix per-language policies such
|
||||||
|
as ``{"en": "manual", "es": "auto", "pt": "any"}``.
|
||||||
|
"""
|
||||||
|
languages_dict = _coerce_languages(languages, prefer_manual)
|
||||||
|
|
||||||
ydl_opts: dict[str, Any] = {
|
ydl_opts: dict[str, Any] = {
|
||||||
"writesubtitles": True,
|
"writesubtitles": True,
|
||||||
"writeautomaticsub": True,
|
"writeautomaticsub": True,
|
||||||
"subtitleslangs": languages,
|
"subtitleslangs": list(languages_dict.keys()),
|
||||||
"skip_download": True,
|
"skip_download": True,
|
||||||
"quiet": True,
|
"quiet": True,
|
||||||
"no_warnings": True,
|
"no_warnings": True,
|
||||||
"retries": retries,
|
"retries": retries,
|
||||||
"sleep_subrequests": sleep_subrequests,
|
|
||||||
"noprogress": True,
|
"noprogress": True,
|
||||||
|
# Never let yt-dlp probe formats: it costs one HTTP request per format
|
||||||
|
# and we only ever want captions and metadata.
|
||||||
|
"check_formats": None,
|
||||||
|
**ydl_throttle_opts(
|
||||||
|
sleep_subrequests,
|
||||||
|
extractor_retries=extractor_retries,
|
||||||
|
socket_timeout=socket_timeout,
|
||||||
|
),
|
||||||
}
|
}
|
||||||
if cookies_file:
|
if cookies_file:
|
||||||
ydl_opts["cookiefile"] = cookies_file
|
ydl_opts["cookiefile"] = cookies_file
|
||||||
if cookies_from_browser:
|
if cookies_from_browser:
|
||||||
ydl_opts["cookiesfrombrowser"] = (cookies_from_browser,)
|
ydl_opts["cookiesfrombrowser"] = (cookies_from_browser,)
|
||||||
|
|
||||||
|
# One video extraction is two requests: the watch page and the InnerTube
|
||||||
|
# player call. yt-dlp spaces them itself; the pacer needs to know they exist.
|
||||||
|
GLOBAL_PACER.wait(cost=2)
|
||||||
with yt_dlp.YoutubeDL(ydl_opts) as ydl:
|
with yt_dlp.YoutubeDL(ydl_opts) as ydl:
|
||||||
info = ydl.extract_info(video_url, download=False)
|
info = ydl.extract_info(video_url, download=False)
|
||||||
|
|
||||||
pick = pick_subtitle(info, languages, prefer_manual)
|
pick = pick_subtitle(info, languages_dict, prefer_manual)
|
||||||
segments: list[Segment] = []
|
segments: list[Segment] = []
|
||||||
|
skip_reason: str | None = None
|
||||||
if pick:
|
if pick:
|
||||||
raw = _download_subtitle(pick.url)
|
raw, dl_error = _download_subtitle(pick.url)
|
||||||
if raw:
|
if raw:
|
||||||
segments = parse_auto_dump(raw)
|
segments = parse_auto_dump(raw)
|
||||||
if not segments:
|
if not segments:
|
||||||
log.warning("Could not parse subtitle for %s (format=%s)", video_url, pick.ext)
|
log.warning("Could not parse subtitle for %s (format=%s)", video_url, pick.ext)
|
||||||
|
skip_reason = f"subtitle downloaded but parsed empty (lang={pick.lang}, format={pick.ext})"
|
||||||
|
else:
|
||||||
|
skip_reason = (
|
||||||
|
f"subtitle track found (lang={pick.lang}, {pick.source}) but the download failed "
|
||||||
|
f"— usually throttling; retry later [{dl_error}]"
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
skip_reason = describe_missing_subtitle(info, languages_dict)
|
||||||
|
|
||||||
has_chapters = bool(info.get("chapters"))
|
has_chapters = bool(info.get("chapters"))
|
||||||
return VideoData(info=info, segments=segments, subtitle=pick, has_chapters=has_chapters)
|
return VideoData(
|
||||||
|
info=info, segments=segments, subtitle=pick,
|
||||||
|
has_chapters=has_chapters, skip_reason=skip_reason,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def pick_subtitle(info: dict[str, Any], languages: list[str], prefer_manual: bool = True) -> SubtitlePick | None:
|
def describe_missing_subtitle(info: dict[str, Any], languages: Mapping[str, str]) -> str:
|
||||||
|
"""Explain why no track matched, distinguishing 'none exist' from 'policy rejected them'.
|
||||||
|
|
||||||
|
A channel that only publishes auto-generated captions scanned under a
|
||||||
|
manual-only policy yields nothing — which is a config problem, not a
|
||||||
|
property of the video, and the message has to say so.
|
||||||
|
"""
|
||||||
|
manual = {k: v for k, v in (info.get("subtitles") or {}).items() if v}
|
||||||
|
auto = {k: v for k, v in (info.get("automatic_captions") or {}).items() if v}
|
||||||
|
if not manual and not auto:
|
||||||
|
return "no caption tracks published for this video"
|
||||||
|
|
||||||
|
wanted = ", ".join(f"{lang}={mode}" for lang, mode in languages.items()) or "(none configured)"
|
||||||
|
modes = {str(m).lower() for m in languages.values()}
|
||||||
|
parts = [f"no track matched the language policy ({wanted})"]
|
||||||
|
parts.append(f"available: {len(manual)} manual, {len(auto)} auto")
|
||||||
|
if auto and not manual and modes == {"manual"}:
|
||||||
|
parts.append(
|
||||||
|
"this video has ONLY auto-generated captions — set the language mode "
|
||||||
|
"to 'any' or 'auto' to use them"
|
||||||
|
)
|
||||||
|
return "; ".join(parts)
|
||||||
|
|
||||||
|
|
||||||
|
def _coerce_languages(languages: Mapping[str, str] | list[str] | None,
|
||||||
|
prefer_manual: bool) -> dict[str, str]:
|
||||||
|
"""Normalise legacy list / new dict / None into ``{lang: mode}``."""
|
||||||
|
default = "manual" if prefer_manual else "auto"
|
||||||
|
if languages is None:
|
||||||
|
return {}
|
||||||
|
if isinstance(languages, Mapping):
|
||||||
|
out: dict[str, str] = {}
|
||||||
|
for lang, mode in languages.items():
|
||||||
|
m = str(mode).lower().strip()
|
||||||
|
if m not in _LANGUAGE_MODES:
|
||||||
|
m = "any"
|
||||||
|
out[str(lang)] = m
|
||||||
|
return out
|
||||||
|
if isinstance(languages, (list, tuple)):
|
||||||
|
return {str(l): default for l in languages}
|
||||||
|
return {}
|
||||||
|
|
||||||
|
|
||||||
|
def pick_subtitle(info: dict[str, Any],
|
||||||
|
languages: Mapping[str, str],
|
||||||
|
prefer_manual: bool = True) -> SubtitlePick | None:
|
||||||
|
"""Pick the best subtitle track for ``info`` honouring per-language mode.
|
||||||
|
|
||||||
|
See :func:`extract_video` for the ``languages`` schema. ``prefer_manual``
|
||||||
|
is only consulted for entries whose mode is ``"any"``.
|
||||||
|
"""
|
||||||
manual = info.get("subtitles") or {}
|
manual = info.get("subtitles") or {}
|
||||||
auto = info.get("automatic_captions") or {}
|
auto = info.get("automatic_captions") or {}
|
||||||
|
|
||||||
ordered_sources: list[tuple[str, dict[str, Any]]]
|
if isinstance(languages, (list, tuple)):
|
||||||
if prefer_manual:
|
# legacy path: convert on the fly
|
||||||
ordered_sources = [("manual", manual), ("auto", auto)]
|
default = "manual" if prefer_manual else "auto"
|
||||||
else:
|
languages = {l: default for l in languages}
|
||||||
ordered_sources = [("auto", auto), ("manual", manual)]
|
|
||||||
|
|
||||||
for source_label, tracks in ordered_sources:
|
# Config order is a preference between languages we can read, not an
|
||||||
for lang in languages:
|
# instruction to accept a machine translation when the real transcript is
|
||||||
|
# sitting right there. An English channel scanned under {es, es-419, en}
|
||||||
|
# was yielding Spanish auto-translations of English speech.
|
||||||
|
#
|
||||||
|
# Two passes rather than a reorder. The reorder alone needed to know the
|
||||||
|
# spoken language, and when neither an `-orig` key nor `info["language"]`
|
||||||
|
# was present it silently fell back to config order and reintroduced the
|
||||||
|
# bug. Rejecting translations outright in the first pass needs no such
|
||||||
|
# knowledge: whatever language it lands on, it is the one actually spoken.
|
||||||
|
ordered = list(languages.items())
|
||||||
|
spoken = original_language(info)
|
||||||
|
if spoken and any(_normalize_lang(l) == spoken for l, _ in ordered):
|
||||||
|
ordered.sort(key=lambda kv: _normalize_lang(kv[0]) != spoken)
|
||||||
|
|
||||||
|
for allow_translations in (False, True):
|
||||||
|
for lang, mode in ordered:
|
||||||
|
if mode not in _LANGUAGE_MODES:
|
||||||
|
mode = "any"
|
||||||
|
sources = _sources_for(mode, prefer_manual, manual, auto)
|
||||||
normalized = _normalize_lang(lang)
|
normalized = _normalize_lang(lang)
|
||||||
for track_lang, formats in tracks.items():
|
for source_label, tracks in sources:
|
||||||
if _normalize_lang(track_lang) != normalized:
|
for track_lang, formats in _ordered_tracks(tracks, normalized):
|
||||||
continue
|
if not allow_translations and _is_translation(track_lang, formats):
|
||||||
if not formats:
|
|
||||||
continue
|
continue
|
||||||
pick = _pick_best_format(formats)
|
pick = _pick_best_format(formats)
|
||||||
if pick:
|
if pick:
|
||||||
@@ -93,7 +218,7 @@ def pick_subtitle(info: dict[str, Any], languages: list[str], prefer_manual: boo
|
|||||||
url=pick["url"],
|
url=pick["url"],
|
||||||
ext=pick["ext"],
|
ext=pick["ext"],
|
||||||
lang=track_lang,
|
lang=track_lang,
|
||||||
source="manual" if source_label == "manual" else "auto",
|
source=source_label,
|
||||||
)
|
)
|
||||||
return None
|
return None
|
||||||
|
|
||||||
@@ -115,11 +240,82 @@ def _normalize_lang(code: str) -> str:
|
|||||||
return base
|
return base
|
||||||
|
|
||||||
|
|
||||||
def _download_subtitle(url: str) -> str | None:
|
def _is_original_track(code: str) -> bool:
|
||||||
try:
|
"""True for YouTube's original-ASR track, which it suffixes with ``-orig``.
|
||||||
resp = requests.get(url, timeout=15, headers={"User-Agent": "Mozilla/5.0"})
|
|
||||||
resp.raise_for_status()
|
YouTube publishes the speech-recognised track as ``<lang>-orig`` and then a
|
||||||
return resp.text
|
long tail of machine translations keyed by bare language code — including a
|
||||||
except requests.RequestException as exc:
|
translation *into the video's own language*. So on a Spanish video both
|
||||||
log.error("Failed to download subtitle from %s: %s", url, exc)
|
``es-orig`` and ``es`` exist, and only the first is the real transcript.
|
||||||
|
"""
|
||||||
|
return code.replace("_", "-").lower().endswith("-orig")
|
||||||
|
|
||||||
|
|
||||||
|
def original_language(info: dict[str, Any]) -> str | None:
|
||||||
|
"""The language actually spoken in the video, normalised, or None.
|
||||||
|
|
||||||
|
Prefers the ``-orig`` track that YouTube itself publishes over ``info`` keys,
|
||||||
|
because the ``-orig`` suffix is direct evidence from the caption list while
|
||||||
|
``language`` is metadata that YouTube localises along with the title.
|
||||||
|
"""
|
||||||
|
for code in (info.get("automatic_captions") or {}):
|
||||||
|
if _is_original_track(code):
|
||||||
|
return _normalize_lang(code)
|
||||||
|
for key in ("language", "original_language"):
|
||||||
|
val = info.get(key)
|
||||||
|
if isinstance(val, str) and val:
|
||||||
|
return _normalize_lang(val)
|
||||||
return None
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def _is_translation(code: str, formats: list[dict[str, Any]] | None) -> bool:
|
||||||
|
"""True when this track is YouTube machine-translating some other track.
|
||||||
|
|
||||||
|
The caption URL says so outright: yt-dlp builds a translated track by
|
||||||
|
appending ``tlang=`` to the base track's URL and omits it when the target
|
||||||
|
equals the source language. That is direct evidence, unlike the ``-orig``
|
||||||
|
naming convention, and it is what lets us reject a translation even for a
|
||||||
|
video whose spoken language we could not otherwise determine.
|
||||||
|
"""
|
||||||
|
if _is_original_track(code):
|
||||||
|
return False
|
||||||
|
for fmt in formats or []:
|
||||||
|
url = fmt.get("url") or ""
|
||||||
|
if "tlang=" in url:
|
||||||
|
return True
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
def _ordered_tracks(tracks: dict, normalized: str) -> list[tuple[str, Any]]:
|
||||||
|
"""Tracks matching `normalized`, original-ASR first.
|
||||||
|
|
||||||
|
Without this the picker took whichever key yt-dlp happened to list first.
|
||||||
|
That silently returned the right thing on Spanish channels (``es-orig``
|
||||||
|
sorts before ``es``) and the wrong thing everywhere else.
|
||||||
|
"""
|
||||||
|
matches = [(c, f) for c, f in tracks.items() if _normalize_lang(c) == normalized and f]
|
||||||
|
matches.sort(key=lambda kv: not _is_original_track(kv[0]))
|
||||||
|
return matches
|
||||||
|
|
||||||
|
|
||||||
|
def _download_subtitle(url: str, *, timeout: float = 15.0) -> tuple[str | None, str | None]:
|
||||||
|
"""Fetch a caption track. Returns (text, error_description).
|
||||||
|
|
||||||
|
The error text is returned rather than only logged because a 429 here is
|
||||||
|
how YouTube throttling most often shows up on this path, and the circuit
|
||||||
|
breaker upstream can only see it if it survives into the stored reason.
|
||||||
|
"""
|
||||||
|
try:
|
||||||
|
GLOBAL_PACER.wait()
|
||||||
|
resp = yt_get(url, timeout=timeout)
|
||||||
|
resp.raise_for_status()
|
||||||
|
return resp.text, None
|
||||||
|
except Exception as exc: # pylint: disable=broad-except
|
||||||
|
log.error("Failed to download subtitle from %s: %s", url, exc)
|
||||||
|
# Lead with a normalised "HTTP Error <status>" token. Downstream, the
|
||||||
|
# throttle detector has to recognise a 429 here, and depending on the
|
||||||
|
# prose is fragile: a 429 served without a reason phrase (routine over
|
||||||
|
# HTTP/2) says nothing about "too many requests".
|
||||||
|
status = getattr(getattr(exc, "response", None), "status_code", None)
|
||||||
|
prefix = f"HTTP Error {status}: " if status else ""
|
||||||
|
return None, f"{prefix}{type(exc).__name__}: {exc}"
|
||||||
|
|||||||
@@ -6,9 +6,9 @@ from typing import Callable
|
|||||||
|
|
||||||
from .config import Config
|
from .config import Config
|
||||||
from .cookies import resolve_active_path
|
from .cookies import resolve_active_path
|
||||||
from .discover import discover_channel
|
from .discover import discover_incremental
|
||||||
from .pipeline import process_video
|
from .pipeline import process_video
|
||||||
from .ratelimit import polite_sleep
|
from .ratelimit import ThrottleGuard, polite_sleep
|
||||||
from .store import Store
|
from .store import Store
|
||||||
|
|
||||||
|
|
||||||
@@ -65,13 +65,34 @@ def _run_once(
|
|||||||
if on_log:
|
if on_log:
|
||||||
on_log(msg)
|
on_log(msg)
|
||||||
|
|
||||||
channel_id_found, channel_name, refs = discover_channel(
|
# A watch loop re-runs on an interval; walking the whole channel every tick
|
||||||
cfg.channel_url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests
|
# is exactly what burns through YouTube's tolerance. Only look at what is
|
||||||
|
# newer than the last video we already have.
|
||||||
|
known = store.known_video_ids(channel_id) if (channel_id and cfg.sync.incremental) else set()
|
||||||
|
result = discover_incremental(
|
||||||
|
cfg.channel_url,
|
||||||
|
known,
|
||||||
|
sleep_subrequests=cfg.yt_dlp.sleep_subrequests,
|
||||||
|
window=cfg.sync.window,
|
||||||
|
max_window=cfg.sync.max_window,
|
||||||
|
overlap=cfg.sync.overlap,
|
||||||
|
since=store.latest_upload_date(channel_id) if known else None,
|
||||||
|
keep=_keep_ref(cfg),
|
||||||
)
|
)
|
||||||
target_channel = channel_id or channel_id_found
|
channel_name, refs = result.channel_name, result.refs
|
||||||
store.upsert_channel(target_channel, _extract_handle(cfg.channel_url), channel_name, len(refs))
|
target_channel = channel_id or result.channel_id
|
||||||
|
if result.full_scan:
|
||||||
|
store.upsert_channel(
|
||||||
|
target_channel, _extract_handle(cfg.channel_url), channel_name, len(refs), avatar=result.avatar
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
store.update_channel_meta(target_channel, name=channel_name, avatar=result.avatar)
|
||||||
store.upsert_videos(refs)
|
store.upsert_videos(refs)
|
||||||
_emit(f"watch: discovered {len(refs)} videos on {channel_name}")
|
store.mark_channel_synced(target_channel)
|
||||||
|
_emit(
|
||||||
|
f"watch: scanned {result.fetched} newest video(s) on {channel_name} — "
|
||||||
|
f"{result.new_count} new"
|
||||||
|
)
|
||||||
|
|
||||||
pending = store.get_pending(target_channel)
|
pending = store.get_pending(target_channel)
|
||||||
if not pending:
|
if not pending:
|
||||||
@@ -86,16 +107,48 @@ def _run_once(
|
|||||||
cookie_path = vault
|
cookie_path = vault
|
||||||
_emit(f"watch: using active vault cookie {vault}")
|
_emit(f"watch: using active vault cookie {vault}")
|
||||||
|
|
||||||
|
guard = ThrottleGuard(
|
||||||
|
threshold=cfg.delay.throttle_threshold,
|
||||||
|
base=cfg.delay.backoff_base,
|
||||||
|
cap=cfg.delay.backoff_cap,
|
||||||
|
)
|
||||||
for row in pending:
|
for row in pending:
|
||||||
process_video(
|
status = process_video(
|
||||||
row, cfg, store, channel_name, target_channel, cfg.channel_url,
|
row, cfg, store, channel_name, target_channel, cfg.channel_url,
|
||||||
cookies_file=cookie_path,
|
cookies_file=cookie_path,
|
||||||
cookies_from_browser=cookies_from_browser,
|
cookies_from_browser=cookies_from_browser,
|
||||||
on_log=on_log,
|
on_log=on_log,
|
||||||
)
|
)
|
||||||
|
if status == "done":
|
||||||
|
guard.note_success()
|
||||||
|
else:
|
||||||
|
failed = store.get_video(row.video_id)
|
||||||
|
wait = guard.note_failure(getattr(failed, "error_msg", None) or status)
|
||||||
|
if guard.tripped:
|
||||||
|
# Unattended loop: stop this pass rather than spend the whole
|
||||||
|
# backlog against a throttled session. The next tick retries.
|
||||||
|
_emit(f"watch: {guard.tripped_reason} — stopping this pass")
|
||||||
|
return
|
||||||
|
if wait > 0:
|
||||||
|
_emit(f"watch: throttled, waiting {wait:.1f}s")
|
||||||
|
time.sleep(wait)
|
||||||
polite_sleep(cfg.delay.min_seconds, cfg.delay.max_seconds)
|
polite_sleep(cfg.delay.min_seconds, cfg.delay.max_seconds)
|
||||||
|
|
||||||
|
|
||||||
|
def _keep_ref(cfg: Config):
|
||||||
|
"""Same shorts/live rules the store was populated under — see jobs._keep_ref."""
|
||||||
|
|
||||||
|
def keep(r) -> bool:
|
||||||
|
url = r.url or ""
|
||||||
|
if not cfg.include_shorts and "/shorts/" in url:
|
||||||
|
return False
|
||||||
|
if not cfg.include_live and url.startswith("https://www.youtube.com/live/"):
|
||||||
|
return False
|
||||||
|
return True
|
||||||
|
|
||||||
|
return keep
|
||||||
|
|
||||||
|
|
||||||
def _extract_handle(url: str) -> str:
|
def _extract_handle(url: str) -> str:
|
||||||
if "@" in url:
|
if "@" in url:
|
||||||
return "@" + url.split("@", 1)[1].split("/", 1)[0]
|
return "@" + url.split("@", 1)[1].split("/", 1)[0]
|
||||||
|
|||||||
+127
-34
@@ -8,13 +8,84 @@ from typing import Callable
|
|||||||
from .chapters import align_chapters, chapters_from_info
|
from .chapters import align_chapters, chapters_from_info
|
||||||
from .config import Config
|
from .config import Config
|
||||||
from .extract import extract_video
|
from .extract import extract_video
|
||||||
|
from .ratelimit import is_rate_limited
|
||||||
from .render import build_filename_stem, render_markdown
|
from .render import build_filename_stem, render_markdown
|
||||||
from .store import Store, VideoRow
|
from .store import Store, VideoRow
|
||||||
|
from ._yt_http import yt_get
|
||||||
|
|
||||||
|
|
||||||
log = logging.getLogger(__name__)
|
log = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
|
||||||
|
def _render_and_retire(
|
||||||
|
cfg: Config,
|
||||||
|
video_id: str,
|
||||||
|
channel_name: str,
|
||||||
|
context: dict,
|
||||||
|
previous: str | None,
|
||||||
|
) -> tuple[Path, str]:
|
||||||
|
"""Write the note for `video_id` and delete whatever file it used to own.
|
||||||
|
|
||||||
|
Identity is the video id; the *filename* is derived from the title, and
|
||||||
|
titles are not stable. YouTube serves them localised — the same video came
|
||||||
|
back as "La controversia de Claude Fable 5" on one pass and "The Claude
|
||||||
|
Fable controversy 5" on the next — and creators rename videos outright.
|
||||||
|
Re-rendering under a new stem without retiring the old path leaves an
|
||||||
|
orphan the database no longer references, which is how a library ends up
|
||||||
|
with more notes than videos.
|
||||||
|
|
||||||
|
Shared by `process_video` and `re_render_videos` because they previously
|
||||||
|
each built the stem their own way and drifted: one used the raw compact
|
||||||
|
date, the other the normalised one, and the result was 94 files for 61 rows.
|
||||||
|
"""
|
||||||
|
stem = build_filename_stem(
|
||||||
|
upload_date=context["upload_date"],
|
||||||
|
title=context["title"],
|
||||||
|
template=cfg.filename_template,
|
||||||
|
video_id=video_id,
|
||||||
|
)
|
||||||
|
out_subdir = Path(cfg.output_dir_resolved) / _safe_dirname(channel_name)
|
||||||
|
md_path = render_markdown(cfg.template_path_resolved, out_subdir, stem, context)
|
||||||
|
|
||||||
|
out_root = Path(cfg.output_dir_resolved).parent
|
||||||
|
try:
|
||||||
|
rel = md_path.relative_to(out_root) if md_path.is_relative_to(out_root) else md_path
|
||||||
|
except ValueError:
|
||||||
|
rel = md_path
|
||||||
|
rel_str = str(rel)
|
||||||
|
|
||||||
|
if previous and previous != rel_str:
|
||||||
|
old = out_root / previous
|
||||||
|
try:
|
||||||
|
if old.exists() and old.resolve() != md_path.resolve():
|
||||||
|
old.unlink()
|
||||||
|
log.info("retired superseded note %s -> %s", previous, rel_str)
|
||||||
|
except OSError as exc:
|
||||||
|
log.warning("could not remove superseded note %s: %s", old, exc)
|
||||||
|
return md_path, rel_str
|
||||||
|
|
||||||
|
|
||||||
|
def _store_metadata(store: Store, row: VideoRow, info: dict) -> None:
|
||||||
|
"""Persist everything the extraction learned that is not the transcript.
|
||||||
|
|
||||||
|
Split out of the render path so it can run before the no-transcript exit:
|
||||||
|
these fields are already paid for by the time we know whether captions
|
||||||
|
downloaded, and `update_video_metadata` only writes the keys it is given.
|
||||||
|
"""
|
||||||
|
store.update_video_metadata(
|
||||||
|
row.video_id,
|
||||||
|
view_count=info.get("view_count"),
|
||||||
|
like_count=info.get("like_count"),
|
||||||
|
tags=info.get("tags") or None,
|
||||||
|
thumbnail=info.get("thumbnail"),
|
||||||
|
description=(info.get("description") or "").strip() or None,
|
||||||
|
)
|
||||||
|
# discovery in flat mode reports no upload_date, so this is where it lands
|
||||||
|
ud = info.get("upload_date") or row.upload_date
|
||||||
|
if ud:
|
||||||
|
store.set_upload_date(row.video_id, ud.replace("-", "") if "-" in str(ud) else str(ud))
|
||||||
|
|
||||||
|
|
||||||
def process_video(
|
def process_video(
|
||||||
row: VideoRow,
|
row: VideoRow,
|
||||||
cfg: Config,
|
cfg: Config,
|
||||||
@@ -42,22 +113,41 @@ def process_video(
|
|||||||
prefer_manual=cfg.prefer_manual,
|
prefer_manual=cfg.prefer_manual,
|
||||||
cookies_file=cookies_file,
|
cookies_file=cookies_file,
|
||||||
cookies_from_browser=cookies_from_browser,
|
cookies_from_browser=cookies_from_browser,
|
||||||
|
extractor_retries=cfg.yt_dlp.extractor_retries,
|
||||||
|
socket_timeout=cfg.yt_dlp.socket_timeout,
|
||||||
)
|
)
|
||||||
except Exception as exc:
|
except Exception as exc:
|
||||||
log.error("Error extrayendo %s: %s", row.video_id, exc)
|
log.error("Error extrayendo %s: %s", row.video_id, exc)
|
||||||
store.mark_error(row.video_id, str(exc))
|
store.mark_error(row.video_id, str(exc))
|
||||||
return "error"
|
return "error"
|
||||||
|
|
||||||
|
info = data.info
|
||||||
|
store.set_availability(row.video_id, info.get("availability"))
|
||||||
|
|
||||||
|
# Metadata is persisted BEFORE the no-transcript exit. The extraction already
|
||||||
|
# cost its requests and the info dict is in hand; discarding it because the
|
||||||
|
# separate caption fetch failed means a retry re-spends them for data we
|
||||||
|
# already had. Measured after a throttling incident: five rows left with
|
||||||
|
# upload_date, view_count, description and thumbnail all NULL.
|
||||||
|
_store_metadata(store, row, info)
|
||||||
|
|
||||||
if not data.segments:
|
if not data.segments:
|
||||||
log.warning("Sin transcripcion para %s", row.video_id)
|
reason = data.skip_reason or "no transcript"
|
||||||
store.mark_status(row.video_id, "no_subtitles")
|
log.warning("Sin transcripcion para %s: %s", row.video_id, reason)
|
||||||
|
_emit(f"{row.video_id}: {reason}")
|
||||||
|
# A throttled caption fetch is not "this video has no captions". Both
|
||||||
|
# states are retryable, but only `error` is honest about the cause, and
|
||||||
|
# `no_subtitles` counts are what tell you a channel publishes none.
|
||||||
|
# Which of the three requests YouTube refused should not decide this.
|
||||||
|
if is_rate_limited(reason):
|
||||||
|
store.mark_error(row.video_id, reason)
|
||||||
|
return "error"
|
||||||
|
store.mark_status(row.video_id, "no_subtitles", reason)
|
||||||
return "no_subtitles"
|
return "no_subtitles"
|
||||||
|
|
||||||
chapters = chapters_from_info(data.info)
|
chapters = chapters_from_info(data.info)
|
||||||
sections = align_chapters(data.segments, chapters)
|
sections = align_chapters(data.segments, chapters)
|
||||||
|
|
||||||
info = data.info
|
|
||||||
|
|
||||||
# persist segments + rich metadata to DB (for search, stats, webapp)
|
# persist segments + rich metadata to DB (for search, stats, webapp)
|
||||||
store.store_segments(row.video_id, data.segments)
|
store.store_segments(row.video_id, data.segments)
|
||||||
seg_json = json.dumps(
|
seg_json = json.dumps(
|
||||||
@@ -70,20 +160,10 @@ def process_video(
|
|||||||
)
|
)
|
||||||
store.update_video_metadata(
|
store.update_video_metadata(
|
||||||
row.video_id,
|
row.video_id,
|
||||||
view_count=info.get("view_count"),
|
|
||||||
like_count=info.get("like_count"),
|
|
||||||
tags=info.get("tags") or None,
|
|
||||||
thumbnail=info.get("thumbnail"),
|
|
||||||
description=(info.get("description") or "").strip() or None,
|
|
||||||
chapters_json=ch_json,
|
chapters_json=ch_json,
|
||||||
segments_json=seg_json,
|
segments_json=seg_json,
|
||||||
)
|
)
|
||||||
|
|
||||||
# ensure upload_date is populated (discovery sometimes lacks it)
|
|
||||||
ud = info.get("upload_date") or row.upload_date
|
|
||||||
if ud:
|
|
||||||
store.set_upload_date(row.video_id, ud.replace("-", "") if "-" in str(ud) else str(ud))
|
|
||||||
|
|
||||||
context = {
|
context = {
|
||||||
"video_id": row.video_id,
|
"video_id": row.video_id,
|
||||||
"title": info.get("title") or row.title or row.video_id,
|
"title": info.get("title") or row.title or row.video_id,
|
||||||
@@ -103,22 +183,12 @@ def process_video(
|
|||||||
"sections": sections,
|
"sections": sections,
|
||||||
}
|
}
|
||||||
|
|
||||||
stem = build_filename_stem(
|
md_path, rel_str = _render_and_retire(
|
||||||
upload_date=context["upload_date"],
|
cfg, row.video_id, channel_name, context, row.markdown_path
|
||||||
title=context["title"],
|
|
||||||
template=cfg.filename_template,
|
|
||||||
)
|
)
|
||||||
out_subdir = Path(cfg.output_dir_resolved) / _safe_dirname(channel_name)
|
|
||||||
md_path = render_markdown(cfg.template_path_resolved, out_subdir, stem, context)
|
|
||||||
|
|
||||||
out_root = Path(cfg.output_dir_resolved).parent
|
|
||||||
try:
|
|
||||||
rel = md_path.relative_to(out_root) if md_path.is_relative_to(out_root) else md_path
|
|
||||||
except ValueError:
|
|
||||||
rel = md_path
|
|
||||||
store.mark_done(
|
store.mark_done(
|
||||||
row.video_id,
|
row.video_id,
|
||||||
str(rel),
|
rel_str,
|
||||||
data.subtitle.lang if data.subtitle else None,
|
data.subtitle.lang if data.subtitle else None,
|
||||||
data.subtitle.source if data.subtitle else None,
|
data.subtitle.source if data.subtitle else None,
|
||||||
data.has_chapters,
|
data.has_chapters,
|
||||||
@@ -150,20 +220,38 @@ def thumbnail_url_for(video_row) -> str:
|
|||||||
|
|
||||||
def cache_thumbnail(store, video_id: str, thumbnails_dir) -> bool:
|
def cache_thumbnail(store, video_id: str, thumbnails_dir) -> bool:
|
||||||
"""Download + cache a video's thumbnail to <thumbnails_dir>/<video_id>.jpg. Returns True on success/existing."""
|
"""Download + cache a video's thumbnail to <thumbnails_dir>/<video_id>.jpg. Returns True on success/existing."""
|
||||||
import requests as _requests
|
|
||||||
out = Path(thumbnails_dir) / f"{video_id}.jpg"
|
out = Path(thumbnails_dir) / f"{video_id}.jpg"
|
||||||
if out.exists():
|
if out.exists() and out.stat().st_size > 0:
|
||||||
return True
|
return True
|
||||||
v = store.get_video(video_id)
|
v = store.get_video(video_id)
|
||||||
url = thumbnail_url_for(v) if v else f"https://i.ytimg.com/vi/{video_id}/hqdefault.jpg"
|
url = thumbnail_url_for(v) if v else f"https://i.ytimg.com/vi/{video_id}/hqdefault.jpg"
|
||||||
try:
|
try:
|
||||||
r = _requests.get(url, timeout=15, headers={"User-Agent": "Mozilla/5.0"})
|
r = yt_get(url, timeout=20.0)
|
||||||
r.raise_for_status()
|
r.raise_for_status()
|
||||||
|
if r.content:
|
||||||
out.parent.mkdir(parents=True, exist_ok=True)
|
out.parent.mkdir(parents=True, exist_ok=True)
|
||||||
out.write_bytes(r.content)
|
out.write_bytes(r.content)
|
||||||
return True
|
return True
|
||||||
except Exception:
|
except Exception:
|
||||||
return False
|
return False
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
def cache_channel_avatar(channel_id: str, url: str, avatars_dir) -> bool:
|
||||||
|
"""Download + cache a channel avatar to <avatars_dir>/<channel_id>.jpg. Returns True on success/existing."""
|
||||||
|
out = Path(avatars_dir) / f"{channel_id}.jpg"
|
||||||
|
out.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
if out.exists() and out.stat().st_size > 0:
|
||||||
|
return True
|
||||||
|
try:
|
||||||
|
r = yt_get(url, timeout=20.0)
|
||||||
|
r.raise_for_status()
|
||||||
|
if r.content:
|
||||||
|
out.write_bytes(r.content)
|
||||||
|
return True
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
def re_render_videos(store: Store, cfg: Config, channel_id: str | None = None) -> int:
|
def re_render_videos(store: Store, cfg: Config, channel_id: str | None = None) -> int:
|
||||||
@@ -193,14 +281,19 @@ def re_render_videos(store: Store, cfg: Config, channel_id: str | None = None) -
|
|||||||
context = {
|
context = {
|
||||||
"video_id": v.video_id, "title": v.title or v.video_id,
|
"video_id": v.video_id, "title": v.title or v.video_id,
|
||||||
"channel_name": ch.get("name") or "", "channel_id": v.channel_id, "channel_url": "",
|
"channel_name": ch.get("name") or "", "channel_id": v.channel_id, "channel_url": "",
|
||||||
"upload_date": v.upload_date or "", "duration": v.duration or 0, "url": v.url,
|
# Must match process_video's normalisation (pipeline.py:96): the
|
||||||
|
# frontmatter is a bidirectional contract that backfill re-reads.
|
||||||
|
"upload_date": _normalize_date(v.upload_date), "duration": v.duration or 0, "url": v.url,
|
||||||
"transcript_lang": v.transcript_lang or "", "transcript_src": v.transcript_src or "",
|
"transcript_lang": v.transcript_lang or "", "transcript_src": v.transcript_src or "",
|
||||||
"view_count": v.view_count, "like_count": v.like_count, "tags": tags,
|
"view_count": v.view_count, "like_count": v.like_count, "tags": tags,
|
||||||
"thumbnail": v.thumbnail or "", "description": v.description or "", "sections": sections,
|
"thumbnail": v.thumbnail or "", "description": v.description or "", "sections": sections,
|
||||||
}
|
}
|
||||||
stem = build_filename_stem(v.upload_date, v.title or v.video_id, cfg.filename_template)
|
md_path, rel_str = _render_and_retire(
|
||||||
out_subdir = Path(cfg.output_dir_resolved) / _safe_dirname(ch.get("name") or "unknown")
|
cfg, v.video_id, ch.get("name") or "unknown", context, v.markdown_path
|
||||||
render_markdown(cfg.template_path_resolved, out_subdir, stem, context)
|
)
|
||||||
|
# mark_done is the only writer of markdown_path; without it the DB
|
||||||
|
# keeps pointing at the pre-render file.
|
||||||
|
store.mark_done(v.video_id, rel_str, v.transcript_lang, v.transcript_src, bool(chapters))
|
||||||
n += 1
|
n += 1
|
||||||
except Exception as exc:
|
except Exception as exc:
|
||||||
log.warning("re-render failed for %s: %s", v.video_id, exc)
|
log.warning("re-render failed for %s: %s", v.video_id, exc)
|
||||||
|
|||||||
+278
-4
@@ -1,14 +1,288 @@
|
|||||||
|
"""Politeness primitives shared by every path that reaches YouTube.
|
||||||
|
|
||||||
|
Three separate concerns live here, and they are not interchangeable:
|
||||||
|
|
||||||
|
* `Pacer` spaces requests out so a burst never leaves this process. It is
|
||||||
|
process-global on purpose — the JobManager serialises *jobs*, but the
|
||||||
|
`/api/tools/*` endpoints run outside it, so without a shared pacer two
|
||||||
|
consumers can hammer YouTube while each believes it is being polite.
|
||||||
|
* `backoff_delay` is what to wait *after* a failure. It follows the algorithm
|
||||||
|
Google documents for its own APIs: `min(base * 2**n + jitter, cap)`, with the
|
||||||
|
jitter redrawn on every attempt so retries from concurrent clients do not
|
||||||
|
re-synchronise into waves.
|
||||||
|
* `ThrottleGuard` decides when to stop trying. YouTube's throttle is a session
|
||||||
|
ban of up to an hour; once it lands, every further request is both useless
|
||||||
|
and harmful. The guard is what turns "511 videos marked failed" into "stopped
|
||||||
|
after 3, kept them retryable".
|
||||||
|
|
||||||
|
Deliberately absent: any retry of the request itself. yt-dlp's YouTube
|
||||||
|
extractor explicitly excludes 403/429 from its own RetryManager, so a throttled
|
||||||
|
call fails once and comes back here — retrying it in a tight loop is the exact
|
||||||
|
behaviour that earns the ban in the first place.
|
||||||
|
"""
|
||||||
|
|
||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
import random
|
import random
|
||||||
|
import re
|
||||||
|
import threading
|
||||||
import time
|
import time
|
||||||
|
from dataclasses import dataclass, field
|
||||||
|
from typing import Callable
|
||||||
|
|
||||||
|
# Substrings that mean "YouTube is refusing because of *rate*, not content".
|
||||||
|
#
|
||||||
|
# Kept distinct from `Store.PERMANENT_ERROR_PATTERNS`: those describe videos we
|
||||||
|
# will never get (members-only, deleted). These describe videos we could get if
|
||||||
|
# we asked more slowly, so they must stay retryable and must trip the breaker.
|
||||||
|
_THROTTLE_SIGNS = (
|
||||||
|
"rate-limited by youtube",
|
||||||
|
"this content isn't available, try again later",
|
||||||
|
"sign in to confirm you're not a bot",
|
||||||
|
"sign in to confirm your age", # same bot-wall, different copy
|
||||||
|
"http error 429",
|
||||||
|
"too many requests",
|
||||||
|
"ratelimitexceeded",
|
||||||
|
"userratelimitexceeded",
|
||||||
|
"temporarily blocked",
|
||||||
|
)
|
||||||
|
|
||||||
|
# 403 with a quota reason is NOT the same failure: Google documents it as
|
||||||
|
# daily and explicitly warns retries may not work for hours (AIP-194). Treated
|
||||||
|
# as fatal-for-now rather than as something to back off and retry into.
|
||||||
|
_QUOTA_SIGNS = (
|
||||||
|
"quotaexceeded",
|
||||||
|
"dailylimitexceeded",
|
||||||
|
)
|
||||||
|
|
||||||
|
# Two spellings reach us and they are not the same string:
|
||||||
|
# yt-dlp -> "HTTP Error 429: Too Many Requests"
|
||||||
|
# requests -> "HTTPError: 429 Client Error: Too Many Requests for url: ..."
|
||||||
|
# The original pattern only matched the first, so the second was detected purely
|
||||||
|
# by its "too many requests" prose — and a 429 served without a reason phrase
|
||||||
|
# (routine over HTTP/2) carries no such prose and slipped through the breaker.
|
||||||
|
_HTTP_STATUS = re.compile(
|
||||||
|
r"HTTP\s*Error[:\s]+(\d{3})"
|
||||||
|
r"|(\d{3})\s+(?:Client|Server)\s+Error",
|
||||||
|
re.IGNORECASE,
|
||||||
|
)
|
||||||
|
|
||||||
|
#: Statuses Google documents as retryable, and which mean "slow down" here.
|
||||||
|
_RETRYABLE_STATUSES = frozenset({"408", "429"})
|
||||||
|
|
||||||
|
|
||||||
|
def is_rate_limited(error: object) -> bool:
|
||||||
|
"""True when `error` looks like YouTube throttling us rather than a bad video."""
|
||||||
|
text = str(error or "").lower()
|
||||||
|
if any(sign in text for sign in _THROTTLE_SIGNS):
|
||||||
|
return True
|
||||||
|
for m in _HTTP_STATUS.finditer(text):
|
||||||
|
status = m.group(1) or m.group(2)
|
||||||
|
if status in _RETRYABLE_STATUSES:
|
||||||
|
return True
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
def is_quota_exhausted(error: object) -> bool:
|
||||||
|
"""True for Google's daily-quota refusals, which backing off will not fix."""
|
||||||
|
text = str(error or "").lower()
|
||||||
|
return any(sign in text for sign in _QUOTA_SIGNS)
|
||||||
|
|
||||||
|
|
||||||
def polite_sleep(min_s: float = 1.5, max_s: float = 3.5) -> None:
|
def polite_sleep(min_s: float = 1.5, max_s: float = 3.5) -> None:
|
||||||
delay = random.uniform(min_s, max_s)
|
"""Randomised pause between units of work (one video, one channel)."""
|
||||||
time.sleep(delay)
|
if max_s < min_s:
|
||||||
|
max_s = min_s
|
||||||
|
time.sleep(random.uniform(min_s, max_s))
|
||||||
|
|
||||||
|
|
||||||
def backoff_sleep(attempt: int, base: float = 2.0, cap: float = 60.0) -> None:
|
def backoff_delay(attempt: int, base: float = 2.0, cap: float = 60.0) -> float:
|
||||||
delay = min(cap, base * (2 ** attempt) + random.uniform(0, 1))
|
"""Truncated exponential backoff with full jitter, per Google's retry guidance.
|
||||||
|
|
||||||
|
`attempt` is 0-indexed, so the first wait after a failure is ~`base`.
|
||||||
|
Returns the delay instead of sleeping so callers can log it and tests can
|
||||||
|
assert on it without spending real seconds.
|
||||||
|
"""
|
||||||
|
if attempt < 0:
|
||||||
|
attempt = 0
|
||||||
|
# 2**attempt overflows into absurd floats long before it matters; clamp the
|
||||||
|
# exponent so a runaway counter can't turn into an OverflowError.
|
||||||
|
exponent = min(attempt, 32)
|
||||||
|
return min(cap, base * (2**exponent) + random.uniform(0, 1))
|
||||||
|
|
||||||
|
|
||||||
|
def backoff_sleep(attempt: int, base: float = 2.0, cap: float = 60.0) -> float:
|
||||||
|
"""Sleep `backoff_delay(...)` and return how long it waited."""
|
||||||
|
delay = backoff_delay(attempt, base, cap)
|
||||||
time.sleep(delay)
|
time.sleep(delay)
|
||||||
|
return delay
|
||||||
|
|
||||||
|
|
||||||
|
class Pacer:
|
||||||
|
"""Enforces a minimum gap between requests, process-wide and thread-safe.
|
||||||
|
|
||||||
|
`wait()` blocks only for the remainder of the gap, so a slow caller never
|
||||||
|
pays twice: if the previous request already took longer than `min_interval`
|
||||||
|
there is nothing left to wait for.
|
||||||
|
"""
|
||||||
|
|
||||||
|
def __init__(
|
||||||
|
self,
|
||||||
|
min_interval: float = 0.0,
|
||||||
|
*,
|
||||||
|
clock: Callable[[], float] = time.monotonic,
|
||||||
|
sleeper: Callable[[float], None] = time.sleep,
|
||||||
|
) -> None:
|
||||||
|
self.min_interval = max(0.0, float(min_interval))
|
||||||
|
self._clock = clock
|
||||||
|
self._sleeper = sleeper
|
||||||
|
self._lock = threading.Lock()
|
||||||
|
self._next_at = 0.0
|
||||||
|
|
||||||
|
def configure(self, min_interval: float) -> None:
|
||||||
|
with self._lock:
|
||||||
|
self.min_interval = max(0.0, float(min_interval))
|
||||||
|
|
||||||
|
def wait(self, cost: int = 1) -> float:
|
||||||
|
"""Block until the next request is allowed. Returns seconds actually slept.
|
||||||
|
|
||||||
|
`cost` is how many HTTP requests the caller is about to make. It matters
|
||||||
|
because the caller is usually yt-dlp: one `extract_info` on a video is
|
||||||
|
two requests (watch page + InnerTube player), and charging it as one
|
||||||
|
made the pacer under-count by half. yt-dlp spaces those two internally
|
||||||
|
via `sleep_interval_requests`; `cost` is what keeps this pacer's idea of
|
||||||
|
the budget honest about them.
|
||||||
|
"""
|
||||||
|
cost = max(1, int(cost))
|
||||||
|
with self._lock:
|
||||||
|
if self.min_interval <= 0:
|
||||||
|
return 0.0
|
||||||
|
now = self._clock()
|
||||||
|
delay = self._next_at - now
|
||||||
|
if delay < 0:
|
||||||
|
delay = 0.0
|
||||||
|
# Reserve the slots before releasing the lock so concurrent callers
|
||||||
|
# queue up behind each other instead of all reading the same `now`.
|
||||||
|
self._next_at = now + delay + self.min_interval * cost
|
||||||
|
if delay > 0:
|
||||||
|
self._sleeper(delay)
|
||||||
|
return delay
|
||||||
|
|
||||||
|
def penalise(self, seconds: float) -> None:
|
||||||
|
"""Push the next allowed request out by `seconds` (used after a 429)."""
|
||||||
|
if seconds <= 0:
|
||||||
|
return
|
||||||
|
with self._lock:
|
||||||
|
self._next_at = max(self._next_at, self._clock() + seconds)
|
||||||
|
|
||||||
|
|
||||||
|
#: Shared by `_yt_http.yt_get` and every yt-dlp entry point. Idle (0.0) until
|
||||||
|
#: something calls `configure_global_pacer`, so importing this module never
|
||||||
|
#: slows down a test suite that does not opt in.
|
||||||
|
GLOBAL_PACER = Pacer(0.0)
|
||||||
|
|
||||||
|
|
||||||
|
def configure_global_pacer(min_interval: float) -> None:
|
||||||
|
GLOBAL_PACER.configure(min_interval)
|
||||||
|
|
||||||
|
|
||||||
|
def ydl_throttle_opts(
|
||||||
|
sleep_requests: float = 2.0,
|
||||||
|
*,
|
||||||
|
extractor_retries: int = 3,
|
||||||
|
socket_timeout: float = 30.0,
|
||||||
|
) -> dict[str, object]:
|
||||||
|
"""The politeness half of every `ydl_opts` dict in this project.
|
||||||
|
|
||||||
|
Exists because the option names are easy to get subtly wrong, and yt-dlp
|
||||||
|
silently ignores keys it does not recognise — this project shipped
|
||||||
|
`sleep_subrequests` (not a real option) for its entire history, so nothing
|
||||||
|
ever slept between the sub-requests of an extraction.
|
||||||
|
|
||||||
|
Only `sleep_interval_requests` throttles *extraction*; `sleep_interval` and
|
||||||
|
`max_sleep_interval` fire in the file downloader and never trigger under
|
||||||
|
`skip_download`, which is every metadata path here.
|
||||||
|
"""
|
||||||
|
return {
|
||||||
|
# Sleeps inside InfoExtractor._request_webpage, i.e. before each HTTP
|
||||||
|
# call an extractor makes: watch page, InnerTube player/browse, and the
|
||||||
|
# continuation pages of a channel tab.
|
||||||
|
"sleep_interval_requests": max(0.0, float(sleep_requests)),
|
||||||
|
# `retries` governs the downloader; extraction retries are this one.
|
||||||
|
# Note yt-dlp's YouTube extractor refuses to retry 403/429 at all, so
|
||||||
|
# this only covers transient 5xx and network errors.
|
||||||
|
"extractor_retries": int(extractor_retries),
|
||||||
|
"socket_timeout": float(socket_timeout),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class ThrottleGuard:
|
||||||
|
"""Circuit breaker for a batch of YouTube work.
|
||||||
|
|
||||||
|
Feed it every outcome. It answers one question — "should this run keep
|
||||||
|
going?" — and, while it still says yes, how long to wait first.
|
||||||
|
|
||||||
|
The counter is *consecutive*: an isolated throttled video between successes
|
||||||
|
is noise, three in a row means the session is banned and everything after
|
||||||
|
it will fail too. Only the consecutive form distinguishes those.
|
||||||
|
"""
|
||||||
|
|
||||||
|
threshold: int = 3
|
||||||
|
base: float = 2.0
|
||||||
|
cap: float = 60.0
|
||||||
|
|
||||||
|
consecutive: int = 0
|
||||||
|
throttled_total: int = 0
|
||||||
|
tripped_reason: str | None = None
|
||||||
|
_waits: list[float] = field(default_factory=list)
|
||||||
|
|
||||||
|
@property
|
||||||
|
def tripped(self) -> bool:
|
||||||
|
return self.tripped_reason is not None
|
||||||
|
|
||||||
|
def note_success(self) -> None:
|
||||||
|
"""A request went through: the session is healthy again."""
|
||||||
|
self.consecutive = 0
|
||||||
|
|
||||||
|
def note_failure(self, error: object) -> float:
|
||||||
|
"""Record a failed unit of work. Returns seconds the caller should wait.
|
||||||
|
|
||||||
|
Non-throttle failures (a private video, a parse error) reset the
|
||||||
|
consecutive counter: they say nothing about our request rate, and
|
||||||
|
letting them accumulate would trip the breaker on a channel that simply
|
||||||
|
has a few dead videos.
|
||||||
|
"""
|
||||||
|
if is_quota_exhausted(error):
|
||||||
|
self.throttled_total += 1
|
||||||
|
self.tripped_reason = (
|
||||||
|
"YouTube reports the quota is exhausted; Google documents this as "
|
||||||
|
"daily, so retrying now cannot succeed"
|
||||||
|
)
|
||||||
|
return 0.0
|
||||||
|
|
||||||
|
if not is_rate_limited(error):
|
||||||
|
self.consecutive = 0
|
||||||
|
return 0.0
|
||||||
|
|
||||||
|
self.throttled_total += 1
|
||||||
|
self.consecutive += 1
|
||||||
|
if self.consecutive >= self.threshold:
|
||||||
|
self.tripped_reason = (
|
||||||
|
f"{self.consecutive} consecutive rate-limit responses from YouTube; "
|
||||||
|
"the session is throttled (YouTube states up to an hour) and further "
|
||||||
|
"requests would only mark healthy videos as failed"
|
||||||
|
)
|
||||||
|
return 0.0
|
||||||
|
|
||||||
|
delay = backoff_delay(self.consecutive - 1, self.base, self.cap)
|
||||||
|
self._waits.append(delay)
|
||||||
|
GLOBAL_PACER.penalise(delay)
|
||||||
|
return delay
|
||||||
|
|
||||||
|
def summary(self) -> str:
|
||||||
|
if self.tripped_reason:
|
||||||
|
return f"stopped: {self.tripped_reason}"
|
||||||
|
if self.throttled_total:
|
||||||
|
return f"{self.throttled_total} throttled response(s) absorbed by backoff"
|
||||||
|
return "no throttling seen"
|
||||||
|
|||||||
@@ -63,10 +63,26 @@ def render_markdown(
|
|||||||
return out_file
|
return out_file
|
||||||
|
|
||||||
|
|
||||||
def build_filename_stem(upload_date: str | None, title: str, template: str = "{upload_date}_{slug}") -> str:
|
def build_filename_stem(
|
||||||
|
upload_date: str | None,
|
||||||
|
title: str,
|
||||||
|
template: str = "{upload_date}_{slug}",
|
||||||
|
video_id: str | None = None,
|
||||||
|
) -> str:
|
||||||
|
"""Filename for a video's note, from `template`.
|
||||||
|
|
||||||
|
`{video_id}` is offered because the other two variables are unstable:
|
||||||
|
YouTube serves titles localised (the same video came back Spanish on one
|
||||||
|
pass and English on the next) and creators rename things. A template
|
||||||
|
including the id makes the file identifiable from disk alone, without
|
||||||
|
consulting the database. It is opt-in — the default is unchanged so
|
||||||
|
existing libraries keep their filenames.
|
||||||
|
"""
|
||||||
from slugify import slugify
|
from slugify import slugify
|
||||||
date_part = upload_date or "unknown-date"
|
date_part = upload_date or "unknown-date"
|
||||||
slug = slugify(title, max_length=60) or "untitled"
|
slug = slugify(title, max_length=60) or "untitled"
|
||||||
stem = template.format(upload_date=date_part, slug=slug, title=title)
|
stem = template.format(
|
||||||
|
upload_date=date_part, slug=slug, title=title, video_id=video_id or "",
|
||||||
|
)
|
||||||
safe = "".join(c for c in stem if c not in r'\/:*?"<>|')
|
safe = "".join(c for c in stem if c not in r'\/:*?"<>|')
|
||||||
return safe.strip().strip(".")
|
return safe.strip().strip(".")
|
||||||
|
|||||||
+113
-1
@@ -122,7 +122,9 @@ def backfill_from_markdown(
|
|||||||
for md_path in md_files:
|
for md_path in md_files:
|
||||||
try:
|
try:
|
||||||
text = md_path.read_text(encoding="utf-8")
|
text = md_path.read_text(encoding="utf-8")
|
||||||
except OSError as exc:
|
except (OSError, UnicodeDecodeError) as exc:
|
||||||
|
# UnicodeDecodeError is a ValueError, not an OSError — letting it
|
||||||
|
# escape aborted the loop and silently skipped every later file.
|
||||||
_log(f"backfill: skip unreadable {md_path}: {exc}")
|
_log(f"backfill: skip unreadable {md_path}: {exc}")
|
||||||
continue
|
continue
|
||||||
parsed = parse_markdown(text)
|
parsed = parse_markdown(text)
|
||||||
@@ -170,6 +172,116 @@ def backfill_from_markdown(
|
|||||||
return n
|
return n
|
||||||
|
|
||||||
|
|
||||||
|
def reconcile_markdown(
|
||||||
|
store: Store,
|
||||||
|
md_root: Path,
|
||||||
|
log: Callable[[str], None] | None = None,
|
||||||
|
*,
|
||||||
|
prune: bool = False,
|
||||||
|
) -> dict[str, int]:
|
||||||
|
"""Make the DB agree with what is actually on disk.
|
||||||
|
|
||||||
|
`backfill_from_markdown` fills in segments and metadata but never touches
|
||||||
|
`status` or `markdown_path`, so a video whose .md exists can sit at
|
||||||
|
`error`/`no_subtitles`/`pending` forever and the UI keeps showing a failure
|
||||||
|
for work that is already done. This walks the markdown tree and repairs:
|
||||||
|
|
||||||
|
- a row with a real .md but a non-done status -> marked done
|
||||||
|
|
||||||
|
`prune=True` additionally sends `done` rows whose .md has disappeared back
|
||||||
|
to pending. That direction is opt-in because it is destructive when aimed
|
||||||
|
at the wrong root: pointed at an empty or unrelated markdown tree it would
|
||||||
|
demote every finished video in the database. It is also skipped outright
|
||||||
|
when the tree contains no .md at all, which is never a real "everything was
|
||||||
|
deleted" state — it means the root is wrong.
|
||||||
|
|
||||||
|
Returns counts so the caller can report what changed. Idempotent.
|
||||||
|
"""
|
||||||
|
def _log(msg: str) -> None:
|
||||||
|
if log:
|
||||||
|
log(msg)
|
||||||
|
else:
|
||||||
|
logging.getLogger(__name__).info(msg)
|
||||||
|
|
||||||
|
md_root = Path(md_root)
|
||||||
|
data_root = md_root.parent
|
||||||
|
out = {
|
||||||
|
"scanned": 0, "repaired_done": 0, "orphan_md": 0,
|
||||||
|
"missing_md": 0, "backfilled": 0, "stale_dupe": 0,
|
||||||
|
}
|
||||||
|
if not md_root.exists():
|
||||||
|
_log(f"reconcile: markdown root not found: {md_root}")
|
||||||
|
return out
|
||||||
|
|
||||||
|
out["backfilled"] = backfill_from_markdown(store, md_root, log=log)
|
||||||
|
|
||||||
|
seen: dict[str, Path] = {}
|
||||||
|
for md_path in sorted(md_root.rglob("*.md")):
|
||||||
|
out["scanned"] += 1
|
||||||
|
try:
|
||||||
|
text = md_path.read_text(encoding="utf-8")
|
||||||
|
except (OSError, UnicodeDecodeError) as exc:
|
||||||
|
_log(f"reconcile: skip unreadable {md_path}: {exc}")
|
||||||
|
continue
|
||||||
|
meta = parse_markdown(text).metadata
|
||||||
|
video_id = meta.get("video_id")
|
||||||
|
if not video_id:
|
||||||
|
continue
|
||||||
|
row = store.get_video(video_id)
|
||||||
|
if not row:
|
||||||
|
out["orphan_md"] += 1
|
||||||
|
continue
|
||||||
|
rel = md_path.relative_to(data_root).as_posix()
|
||||||
|
|
||||||
|
# A second .md for a video the DB already resolves elsewhere. Older
|
||||||
|
# re-renders built the filename from a differently-formatted date, so
|
||||||
|
# they wrote a sibling file the DB never learned about; it is dead
|
||||||
|
# weight that every later scan has to wade through.
|
||||||
|
canonical = (row.markdown_path or "").replace("\\", "/")
|
||||||
|
if video_id in seen or (row.status == "done" and canonical and canonical != rel):
|
||||||
|
out["stale_dupe"] += 1
|
||||||
|
if prune:
|
||||||
|
try:
|
||||||
|
md_path.unlink()
|
||||||
|
_log(f"reconcile: removed stale duplicate {rel}")
|
||||||
|
except OSError as exc:
|
||||||
|
_log(f"reconcile: could not remove {rel}: {exc}")
|
||||||
|
else:
|
||||||
|
_log(f"reconcile: stale duplicate (use prune to delete): {rel}")
|
||||||
|
continue
|
||||||
|
|
||||||
|
seen[video_id] = md_path
|
||||||
|
if row.status != "done" or not row.markdown_path:
|
||||||
|
store.mark_done(
|
||||||
|
video_id, rel,
|
||||||
|
meta.get("transcript_lang") or row.transcript_lang,
|
||||||
|
meta.get("transcript_src") or row.transcript_src,
|
||||||
|
bool(meta.get("has_chapters")) or bool(row.has_chapters),
|
||||||
|
)
|
||||||
|
out["repaired_done"] += 1
|
||||||
|
_log(f"reconcile: {video_id} had a .md on disk but status={row.status} -> done")
|
||||||
|
|
||||||
|
# The other direction is destructive, so it needs both an explicit opt-in
|
||||||
|
# and evidence that we are looking at a real markdown tree.
|
||||||
|
if prune and out["scanned"]:
|
||||||
|
for row in store.get_all():
|
||||||
|
if row.status != "done" or row.video_id in seen:
|
||||||
|
continue
|
||||||
|
path = data_root / row.markdown_path if row.markdown_path else None
|
||||||
|
if path is None or not path.exists():
|
||||||
|
store.mark_status(row.video_id, "pending", "markdown file missing on disk")
|
||||||
|
out["missing_md"] += 1
|
||||||
|
_log(f"reconcile: {row.video_id} marked done but .md is gone -> pending")
|
||||||
|
elif prune:
|
||||||
|
_log("reconcile: markdown tree is empty — refusing to prune (wrong root?)")
|
||||||
|
|
||||||
|
_log(
|
||||||
|
"reconcile: scanned {scanned} .md, repaired {repaired_done}, "
|
||||||
|
"re-queued {missing_md}, orphans {orphan_md}, stale duplicates {stale_dupe}".format(**out)
|
||||||
|
)
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
def _to_int(value: str | None) -> int | None:
|
def _to_int(value: str | None) -> int | None:
|
||||||
if value is None:
|
if value is None:
|
||||||
return None
|
return None
|
||||||
|
|||||||
@@ -383,6 +383,24 @@ def build_router(store: Store, cfg: Config, jobs) -> APIRouter:
|
|||||||
raise HTTPException(404, "audio not downloaded yet")
|
raise HTTPException(404, "audio not downloaded yet")
|
||||||
return FileResponse(str(p), filename=f"{video_id}.mp3", media_type="audio/mpeg")
|
return FileResponse(str(p), filename=f"{video_id}.mp3", media_type="audio/mpeg")
|
||||||
|
|
||||||
|
@r.api_route("/videos/{video_id}/media", methods=["GET", "HEAD"])
|
||||||
|
def video_media(video_id: str, request: Request, inline: bool = False):
|
||||||
|
v = store.get_video(video_id)
|
||||||
|
if not v or v.video_download_status != "done" or not v.video_path:
|
||||||
|
raise HTTPException(404, "video not downloaded yet")
|
||||||
|
root = Path(cfg.output_dir_resolved).parent.resolve()
|
||||||
|
p = (root / v.video_path).resolve()
|
||||||
|
try:
|
||||||
|
p.relative_to(root)
|
||||||
|
except ValueError:
|
||||||
|
raise HTTPException(500, "invalid stored video path")
|
||||||
|
if not p.exists():
|
||||||
|
raise HTTPException(404, "video file missing on disk")
|
||||||
|
return FileResponse(
|
||||||
|
str(p), filename=v.video_filename or p.name, media_type="video/webm",
|
||||||
|
content_disposition_type="inline" if inline else "attachment",
|
||||||
|
)
|
||||||
|
|
||||||
# -------------------------------------------------- thumbnails (local cache)
|
# -------------------------------------------------- thumbnails (local cache)
|
||||||
@r.post("/tools/thumbnails")
|
@r.post("/tools/thumbnails")
|
||||||
def download_thumbnails(payload: dict):
|
def download_thumbnails(payload: dict):
|
||||||
@@ -614,6 +632,22 @@ def build_router(store: Store, cfg: Config, jobs) -> APIRouter:
|
|||||||
job_id = jobs.enqueue(channel_id, opts)
|
job_id = jobs.enqueue(channel_id, opts)
|
||||||
return {"job_id": job_id}
|
return {"job_id": job_id}
|
||||||
|
|
||||||
|
@r.post("/tools/video")
|
||||||
|
async def tools_video(payload: dict):
|
||||||
|
video_ids = (payload or {}).get("video_ids") or []
|
||||||
|
if len(video_ids) != 1:
|
||||||
|
raise HTTPException(400, "exactly one video_id required")
|
||||||
|
video_id = str(video_ids[0])
|
||||||
|
v = store.get_video(video_id)
|
||||||
|
if not v:
|
||||||
|
raise HTTPException(404, "video not found")
|
||||||
|
md_path = Path(cfg.output_dir_resolved).parent / v.markdown_path if v.markdown_path else None
|
||||||
|
if v.status != "done" or not v.markdown_path or not md_path or not md_path.exists():
|
||||||
|
raise HTTPException(409, "markdown must be generated and present before downloading video")
|
||||||
|
opts = {"mode": "video", "video_ids": [video_id]}
|
||||||
|
job_id = jobs.enqueue(v.channel_id, opts)
|
||||||
|
return {"job_id": job_id}
|
||||||
|
|
||||||
return r
|
return r
|
||||||
|
|
||||||
|
|
||||||
@@ -633,6 +667,10 @@ def _video_dict(v) -> dict:
|
|||||||
"transcript_lang": v.transcript_lang, "has_chapters": bool(v.has_chapters),
|
"transcript_lang": v.transcript_lang, "has_chapters": bool(v.has_chapters),
|
||||||
"view_count": v.view_count, "like_count": v.like_count, "tags": tags,
|
"view_count": v.view_count, "like_count": v.like_count, "tags": tags,
|
||||||
"thumbnail": v.thumbnail, "markdown_path": v.markdown_path,
|
"thumbnail": v.thumbnail, "markdown_path": v.markdown_path,
|
||||||
|
"video_download_status": v.video_download_status or "not_downloaded",
|
||||||
|
"video_path": v.video_path, "video_filename": v.video_filename,
|
||||||
|
"video_size": v.video_size, "video_downloaded_at": v.video_downloaded_at,
|
||||||
|
"video_error": v.video_error,
|
||||||
# The date the row was ordered by. For a video discovery has not
|
# The date the row was ordered by. For a video discovery has not
|
||||||
# extracted yet there is no real upload_date, so this is inferred from
|
# extracted yet there is no real upload_date, so this is inferred from
|
||||||
# its position in the channel listing — flagged, never passed off as
|
# its position in the channel listing — flagged, never passed off as
|
||||||
|
|||||||
@@ -1,5 +1,6 @@
|
|||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import logging
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
from fastapi import FastAPI
|
from fastapi import FastAPI
|
||||||
@@ -8,7 +9,8 @@ from fastapi.staticfiles import StaticFiles
|
|||||||
|
|
||||||
from ..config import Config, load_config
|
from ..config import Config, load_config
|
||||||
from ..cookies import auto_import_dir
|
from ..cookies import auto_import_dir
|
||||||
from ..segments import backfill_from_markdown
|
from ..ratelimit import configure_global_pacer
|
||||||
|
from ..segments import reconcile_markdown
|
||||||
from ..store import Store
|
from ..store import Store
|
||||||
from .api import build_router
|
from .api import build_router
|
||||||
from .jobs import JobManager
|
from .jobs import JobManager
|
||||||
@@ -23,18 +25,34 @@ def create_app(cfg: Config | None = None, db_path: str | Path | None = None) ->
|
|||||||
cfg.database_path = str(db_path)
|
cfg.database_path = str(db_path)
|
||||||
store = Store(cfg.database_path_resolved)
|
store = Store(cfg.database_path_resolved)
|
||||||
|
|
||||||
|
# The job runner serialises jobs, but /api/tools/* endpoints run outside it
|
||||||
|
# on the threadpool. The pacer is the only thing that stops those two from
|
||||||
|
# hitting YouTube simultaneously, so it has to be armed before any router.
|
||||||
|
configure_global_pacer(cfg.delay.min_request_interval)
|
||||||
|
|
||||||
# auto-import any loose cookies into the vault
|
# auto-import any loose cookies into the vault
|
||||||
try:
|
try:
|
||||||
auto_import_dir(store)
|
auto_import_dir(store)
|
||||||
except Exception:
|
except Exception:
|
||||||
pass
|
pass
|
||||||
# auto-backfill segments from existing markdown (idempotent)
|
# uvicorn only configures its own loggers, so the root logger has no
|
||||||
|
# handler and everything yt_scraper logs at startup vanishes.
|
||||||
|
pkg_log = logging.getLogger("yt_scraper")
|
||||||
|
if not pkg_log.handlers:
|
||||||
|
handler = logging.StreamHandler()
|
||||||
|
handler.setFormatter(logging.Formatter("%(levelname)s %(name)s: %(message)s"))
|
||||||
|
pkg_log.addHandler(handler)
|
||||||
|
pkg_log.setLevel(logging.INFO)
|
||||||
|
|
||||||
|
# Reconcile, not just backfill: backfill_from_markdown populates segments
|
||||||
|
# and metadata but never touches `status`/`markdown_path`, so a video whose
|
||||||
|
# .md is already on disk would keep showing a failure after every restart.
|
||||||
try:
|
try:
|
||||||
md_root = cfg.output_dir_resolved
|
md_root = Path(cfg.output_dir_resolved)
|
||||||
if Path(md_root).exists():
|
if md_root.exists():
|
||||||
backfill_from_markdown(store, Path(md_root))
|
reconcile_markdown(store, md_root, log=pkg_log.info)
|
||||||
except Exception:
|
except Exception:
|
||||||
pass
|
pkg_log.exception("startup reconcile failed")
|
||||||
|
|
||||||
app = FastAPI(title="yt-scraper platform", version="1.0.0")
|
app = FastAPI(title="yt-scraper platform", version="1.0.0")
|
||||||
|
|
||||||
|
|||||||
+488
-36
@@ -6,14 +6,15 @@ import threading
|
|||||||
import time
|
import time
|
||||||
import uuid
|
import uuid
|
||||||
from collections import deque
|
from collections import deque
|
||||||
|
from pathlib import Path
|
||||||
from typing import Any
|
from typing import Any
|
||||||
|
|
||||||
from ..config import Config, load_config
|
from ..config import Config, load_config, parse_languages
|
||||||
from ..cookies import resolve_active_path
|
from ..cookies import resolve_active_path
|
||||||
from ..discover import discover_channel
|
from ..discover import discover_incremental
|
||||||
from ..pipeline import process_video
|
from ..pipeline import process_video
|
||||||
from ..ratelimit import polite_sleep
|
from ..ratelimit import ThrottleGuard, polite_sleep
|
||||||
from ..store import Store
|
from ..store import Store, VideoRef
|
||||||
|
|
||||||
|
|
||||||
class JobManager:
|
class JobManager:
|
||||||
@@ -89,12 +90,18 @@ class JobManager:
|
|||||||
if not job:
|
if not job:
|
||||||
return
|
return
|
||||||
opts = json.loads(job.opts_json) if job.opts_json else {}
|
opts = json.loads(job.opts_json) if job.opts_json else {}
|
||||||
if opts.get("video_ids"):
|
if opts.get("mode") == "discover":
|
||||||
self._run_batch(job_id, opts)
|
self._run_discovery(job_id, opts)
|
||||||
return
|
return
|
||||||
if opts.get("mode") == "audio":
|
if opts.get("mode") == "audio":
|
||||||
self._run_audio(job_id, opts)
|
self._run_audio(job_id, opts)
|
||||||
return
|
return
|
||||||
|
if opts.get("mode") == "video":
|
||||||
|
self._run_video(job_id, opts)
|
||||||
|
return
|
||||||
|
if opts.get("video_ids"):
|
||||||
|
self._run_batch(job_id, opts)
|
||||||
|
return
|
||||||
self._run_channel(job_id, opts)
|
self._run_channel(job_id, opts)
|
||||||
|
|
||||||
def _resolve_channel_for_video(self, video_id: str) -> tuple[str, str, str]:
|
def _resolve_channel_for_video(self, video_id: str) -> tuple[str, str, str]:
|
||||||
@@ -107,19 +114,85 @@ class JobManager:
|
|||||||
url = f"https://www.youtube.com/@{handle}/videos" if handle else row.url
|
url = f"https://www.youtube.com/@{handle}/videos" if handle else row.url
|
||||||
return (name, row.channel_id, url)
|
return (name, row.channel_id, url)
|
||||||
|
|
||||||
|
# ---- throttling ------------------------------------------------------
|
||||||
|
#
|
||||||
|
# Every loop that touches YouTube shares these three helpers. Before them,
|
||||||
|
# a throttled session was invisible to the runner: it kept going and turned
|
||||||
|
# one rate-limit into hundreds of videos marked failed (the incident that
|
||||||
|
# left 511 rows in `no_subtitles` and 347 in `error`, all of them saying
|
||||||
|
# "rate-limited by YouTube"). Now the run stops and everything it has not
|
||||||
|
# reached stays `pending`, which is the retryable state.
|
||||||
|
|
||||||
|
def _new_guard(self) -> ThrottleGuard:
|
||||||
|
d = self.cfg.delay
|
||||||
|
return ThrottleGuard(
|
||||||
|
threshold=d.throttle_threshold,
|
||||||
|
base=d.backoff_base,
|
||||||
|
cap=d.backoff_cap,
|
||||||
|
)
|
||||||
|
|
||||||
|
def _note_outcome(self, job_id: str, guard: ThrottleGuard, video_id: str, status: str) -> bool:
|
||||||
|
"""Feed one video's result to the breaker. Returns False to stop the run.
|
||||||
|
|
||||||
|
`process_video` returns only a status string, so the reason is read back
|
||||||
|
from the row: the message it stored is the one place the throttling
|
||||||
|
signature survives.
|
||||||
|
"""
|
||||||
|
if status == "done":
|
||||||
|
guard.note_success()
|
||||||
|
return True
|
||||||
|
row = self.store.get_video(video_id)
|
||||||
|
wait = guard.note_failure(getattr(row, "error_msg", None) or status)
|
||||||
|
if guard.tripped:
|
||||||
|
return False
|
||||||
|
if wait > 0:
|
||||||
|
self._emit(job_id, "log", {
|
||||||
|
"msg": f"YouTube is throttling us — waiting {wait:.1f}s before the next video"
|
||||||
|
})
|
||||||
|
time.sleep(wait)
|
||||||
|
return True
|
||||||
|
|
||||||
|
def _stop_throttled(
|
||||||
|
self, job_id: str, guard: ThrottleGuard, completed: int, total: int,
|
||||||
|
extra: dict[str, Any] | None = None,
|
||||||
|
) -> None:
|
||||||
|
msg = guard.tripped_reason or "stopped by the throttling circuit breaker"
|
||||||
|
remaining = max(0, total - completed)
|
||||||
|
detail = (
|
||||||
|
f"{msg}. Stopped after {completed}/{total}; the remaining {remaining} "
|
||||||
|
"were left untouched and are still pending."
|
||||||
|
)
|
||||||
|
self.store.update_job(
|
||||||
|
job_id, status="error", completed=completed, last_error=detail, finished=True
|
||||||
|
)
|
||||||
|
self._emit(job_id, "log", {"msg": f"ABORTED — {detail}"})
|
||||||
|
self._emit(job_id, "error", {
|
||||||
|
"message": detail, "throttled": True,
|
||||||
|
"completed": completed, "total": total, **(extra or {}),
|
||||||
|
})
|
||||||
|
|
||||||
def _run_batch(self, job_id: str, opts: dict[str, Any]) -> None:
|
def _run_batch(self, job_id: str, opts: dict[str, Any]) -> None:
|
||||||
from ..pipeline import process_video as _process_video, cache_thumbnail
|
"""Generate the .md for an explicit list of videos — and nothing else.
|
||||||
|
|
||||||
|
Thumbnails are deliberately NOT fetched here. They are already cached by
|
||||||
|
the channel-level paths (add-channel, the Thumbnails tool), and
|
||||||
|
/api/thumbnails/{id} redirects to the CDN for anything missing, so
|
||||||
|
piggybacking them on a .md batch only spent extra requests per video for
|
||||||
|
an image the UI could already display.
|
||||||
|
"""
|
||||||
|
from ..pipeline import process_video as _process_video
|
||||||
video_ids = opts.get("video_ids") or []
|
video_ids = opts.get("video_ids") or []
|
||||||
cookie_path = opts.get("cookies_file") or resolve_active_path(self.store)
|
cookie_path = opts.get("cookies_file") or resolve_active_path(self.store)
|
||||||
thumb_dir = Path(self.cfg.output_dir_resolved).parent / "thumbnails"
|
|
||||||
total = len(video_ids)
|
total = len(video_ids)
|
||||||
self.store.update_job(job_id, status="running", total=total, completed=0)
|
self.store.update_job(job_id, status="running", total=total, completed=0)
|
||||||
self._emit(job_id, "progress", {"completed": 0, "total": total})
|
self._emit(job_id, "progress", {"completed": 0, "total": total})
|
||||||
self._emit(job_id, "log", {"msg": f"processing {total} videos (.md + thumbnails)"})
|
self._emit(job_id, "log", {"msg": f"generating .md for {total} videos"})
|
||||||
cfg = _clone_config(self.cfg)
|
cfg = _clone_config(self.cfg)
|
||||||
if opts.get("languages"):
|
if opts.get("languages"):
|
||||||
cfg.languages = opts["languages"]
|
cfg.languages = parse_languages(opts["languages"], cfg.prefer_manual)
|
||||||
completed = 0
|
completed = 0
|
||||||
|
outcomes = {"processed": 0, "no_subtitles": 0, "errors": 0}
|
||||||
|
guard = self._new_guard()
|
||||||
for vid in video_ids:
|
for vid in video_ids:
|
||||||
if job_id in self._cancel:
|
if job_id in self._cancel:
|
||||||
self.store.update_job(job_id, status="cancelled", completed=completed, finished=True)
|
self.store.update_job(job_id, status="cancelled", completed=completed, finished=True)
|
||||||
@@ -128,15 +201,17 @@ class JobManager:
|
|||||||
row = self.store.get_video(vid)
|
row = self.store.get_video(vid)
|
||||||
if not row:
|
if not row:
|
||||||
self._emit(job_id, "log", {"msg": f"skip unknown {vid}"})
|
self._emit(job_id, "log", {"msg": f"skip unknown {vid}"})
|
||||||
|
outcomes["errors"] += 1
|
||||||
completed += 1
|
completed += 1
|
||||||
self.store.update_job(job_id, completed=completed)
|
self.store.update_job(job_id, completed=completed)
|
||||||
continue
|
continue
|
||||||
# UX optimization: skip re-extracting videos whose .md already exists — just ensure thumbnail cached.
|
# UX optimization: a video whose .md is already on disk costs nothing
|
||||||
|
# to "download" — skip the extraction entirely.
|
||||||
md_path = Path(cfg.output_dir_resolved).parent / row.markdown_path if row.markdown_path else None
|
md_path = Path(cfg.output_dir_resolved).parent / row.markdown_path if row.markdown_path else None
|
||||||
if row.status == "done" and row.markdown_path and md_path and md_path.exists():
|
if row.status == "done" and row.markdown_path and md_path and md_path.exists():
|
||||||
cache_thumbnail(self.store, vid, thumb_dir)
|
|
||||||
status = "done"
|
status = "done"
|
||||||
self._emit(job_id, "log", {"msg": f"cached {vid} (md already present)"})
|
extracted = False
|
||||||
|
self._emit(job_id, "log", {"msg": f"skip {vid} (.md already present)"})
|
||||||
else:
|
else:
|
||||||
channel_name, channel_id, channel_url = self._resolve_channel_for_video(vid)
|
channel_name, channel_id, channel_url = self._resolve_channel_for_video(vid)
|
||||||
status = _process_video(
|
status = _process_video(
|
||||||
@@ -144,12 +219,28 @@ class JobManager:
|
|||||||
cookies_file=cookie_path,
|
cookies_file=cookie_path,
|
||||||
on_log=lambda m: self._emit(job_id, "log", {"msg": m}),
|
on_log=lambda m: self._emit(job_id, "log", {"msg": m}),
|
||||||
)
|
)
|
||||||
cache_thumbnail(self.store, vid, thumb_dir)
|
extracted = True
|
||||||
|
if status == "done":
|
||||||
|
outcomes["processed"] += 1
|
||||||
|
elif status == "no_subtitles":
|
||||||
|
outcomes["no_subtitles"] += 1
|
||||||
|
self._emit(job_id, "log", {"msg": f"{vid}: no transcript; .md not generated"})
|
||||||
|
else:
|
||||||
|
outcomes["errors"] += 1
|
||||||
|
self._emit(job_id, "log", {"msg": f"{vid}: processing failed; .md not generated"})
|
||||||
completed += 1
|
completed += 1
|
||||||
self.store.update_job(job_id, completed=completed)
|
self.store.update_job(job_id, completed=completed)
|
||||||
self._emit(job_id, "progress", {"completed": completed, "total": total, "video_id": vid, "status": status})
|
self._emit(job_id, "progress", {"completed": completed, "total": total, "video_id": vid, "status": status})
|
||||||
|
if not self._note_outcome(job_id, guard, vid, status):
|
||||||
|
self._stop_throttled(job_id, guard, completed, total, outcomes)
|
||||||
|
return
|
||||||
|
# Only pace when we actually hit YouTube. A batch of already-rendered
|
||||||
|
# videos costs nothing but a disk check, and sleeping through it
|
||||||
|
# would make multi-select feel broken for no benefit.
|
||||||
|
if extracted and completed < total:
|
||||||
|
polite_sleep(cfg.delay.min_seconds, cfg.delay.max_seconds)
|
||||||
self.store.update_job(job_id, status="done", completed=completed, finished=True)
|
self.store.update_job(job_id, status="done", completed=completed, finished=True)
|
||||||
self._emit(job_id, "done", {"completed": completed, "total": total})
|
self._emit(job_id, "done", {"completed": completed, "total": total, **outcomes, "throttling": guard.summary()})
|
||||||
|
|
||||||
def _run_audio(self, job_id: str, opts: dict[str, Any]) -> None:
|
def _run_audio(self, job_id: str, opts: dict[str, Any]) -> None:
|
||||||
import yt_dlp
|
import yt_dlp
|
||||||
@@ -165,15 +256,28 @@ class JobManager:
|
|||||||
total = len(videos)
|
total = len(videos)
|
||||||
self.store.update_job(job_id, status="running", total=total, completed=0)
|
self.store.update_job(job_id, status="running", total=total, completed=0)
|
||||||
self._emit(job_id, "log", {"msg": f"downloading {total} audio tracks -> {out_dir}"})
|
self._emit(job_id, "log", {"msg": f"downloading {total} audio tracks -> {out_dir}"})
|
||||||
|
cfg = self.cfg
|
||||||
ydl_opts = {
|
ydl_opts = {
|
||||||
"format": "bestaudio/best",
|
"format": "bestaudio/best",
|
||||||
"outtmpl": str(out_dir / "%(id)s.%(ext)s"),
|
"outtmpl": str(out_dir / "%(id)s.%(ext)s"),
|
||||||
"postprocessors": [{"key": "FFmpegExtractAudio", "preferredcodec": "mp3", "preferredquality": "128"}],
|
"postprocessors": [{"key": "FFmpegExtractAudio", "preferredcodec": "mp3", "preferredquality": "128"}],
|
||||||
"quiet": True, "no_warnings": True, "noprogress": True,
|
"quiet": True, "no_warnings": True, "noprogress": True,
|
||||||
|
# This is the only path that downloads media, so it is also the only
|
||||||
|
# one where yt-dlp's own download-side throttles actually fire.
|
||||||
|
"sleep_interval": cfg.delay.min_seconds,
|
||||||
|
"max_sleep_interval": cfg.delay.max_seconds,
|
||||||
|
"sleep_interval_requests": cfg.yt_dlp.sleep_subrequests,
|
||||||
|
"retries": cfg.yt_dlp.retries,
|
||||||
|
"socket_timeout": 30.0,
|
||||||
}
|
}
|
||||||
|
if cfg.delay.audio_rate_limit:
|
||||||
|
ydl_opts["ratelimit"] = cfg.delay.audio_rate_limit
|
||||||
if cookie_path:
|
if cookie_path:
|
||||||
ydl_opts["cookiefile"] = cookie_path
|
ydl_opts["cookiefile"] = cookie_path
|
||||||
completed = 0
|
completed = 0
|
||||||
|
guard = self._new_guard()
|
||||||
|
# One YoutubeDL for the whole batch: keeps the connection pool, the
|
||||||
|
# cookie jar and the resolved player JS alive across videos.
|
||||||
with yt_dlp.YoutubeDL(ydl_opts) as ydl:
|
with yt_dlp.YoutubeDL(ydl_opts) as ydl:
|
||||||
for v in videos:
|
for v in videos:
|
||||||
if job_id in self._cancel:
|
if job_id in self._cancel:
|
||||||
@@ -183,13 +287,132 @@ class JobManager:
|
|||||||
try:
|
try:
|
||||||
ydl.download([v.url])
|
ydl.download([v.url])
|
||||||
self._emit(job_id, "log", {"msg": f"OK {v.video_id}"})
|
self._emit(job_id, "log", {"msg": f"OK {v.video_id}"})
|
||||||
|
guard.note_success()
|
||||||
except Exception as exc:
|
except Exception as exc:
|
||||||
self._emit(job_id, "log", {"msg": f"FAIL {v.video_id}: {exc}"})
|
self._emit(job_id, "log", {"msg": f"FAIL {v.video_id}: {exc}"})
|
||||||
|
wait = guard.note_failure(exc)
|
||||||
|
if guard.tripped:
|
||||||
|
self._stop_throttled(job_id, guard, completed, total)
|
||||||
|
return
|
||||||
|
if wait > 0:
|
||||||
|
self._emit(job_id, "log", {
|
||||||
|
"msg": f"YouTube is throttling us — waiting {wait:.1f}s before the next track"
|
||||||
|
})
|
||||||
|
time.sleep(wait)
|
||||||
completed += 1
|
completed += 1
|
||||||
self.store.update_job(job_id, completed=completed)
|
self.store.update_job(job_id, completed=completed)
|
||||||
self._emit(job_id, "progress", {"completed": completed, "total": total, "video_id": v.video_id})
|
self._emit(job_id, "progress", {"completed": completed, "total": total, "video_id": v.video_id})
|
||||||
self.store.update_job(job_id, status="done", completed=completed, finished=True)
|
self.store.update_job(job_id, status="done", completed=completed, finished=True)
|
||||||
self._emit(job_id, "done", {"completed": completed, "total": total})
|
self._emit(job_id, "done", {"completed": completed, "total": total, "throttling": guard.summary()})
|
||||||
|
|
||||||
|
def _run_video(self, job_id: str, opts: dict[str, Any]) -> None:
|
||||||
|
"""Download one validated video as a Chromium-friendly WebM."""
|
||||||
|
import yt_dlp
|
||||||
|
|
||||||
|
video_id = (opts.get("video_ids") or [None])[0]
|
||||||
|
row = self.store.get_video(video_id) if video_id else None
|
||||||
|
if not row:
|
||||||
|
self._fail_video_job(job_id, "video not found")
|
||||||
|
return
|
||||||
|
|
||||||
|
root = Path(self.cfg.output_dir_resolved).parent
|
||||||
|
md_path = root / row.markdown_path if row.markdown_path else None
|
||||||
|
if row.status != "done" or not row.markdown_path or not md_path or not md_path.exists():
|
||||||
|
self._fail_video_job(job_id, "markdown must be generated and present before downloading video", video_id)
|
||||||
|
return
|
||||||
|
|
||||||
|
out_dir = root / "videos" / video_id
|
||||||
|
out_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
filename = row.video_filename or _safe_video_filename(row.title or video_id)
|
||||||
|
target = out_dir / filename
|
||||||
|
if target.suffix.lower() != ".webm":
|
||||||
|
target = target.with_suffix(".webm")
|
||||||
|
rel_path = target.relative_to(root).as_posix()
|
||||||
|
|
||||||
|
if target.exists() and target.stat().st_size > 0:
|
||||||
|
self.store.update_video_download(video_id, "done", path=rel_path, filename=target.name, size=target.stat().st_size)
|
||||||
|
self.store.update_job(job_id, status="done", total=1, completed=1, finished=True)
|
||||||
|
self._emit(job_id, "progress", {"completed": 1, "total": 1, "video_id": video_id, "status": "done", "skipped": True})
|
||||||
|
self._emit(job_id, "done", {"completed": 1, "total": 1, "skipped": True})
|
||||||
|
return
|
||||||
|
|
||||||
|
cookie_path = opts.get("cookies_file") or resolve_active_path(self.store)
|
||||||
|
restricted = bool(row.block_reason or row.availability in {
|
||||||
|
"subscriber_only", "premium_only", "private", "needs_auth",
|
||||||
|
})
|
||||||
|
self.store.update_video_download(video_id, "queued", filename=target.name, error=None)
|
||||||
|
self.store.update_job(job_id, status="running", total=1, completed=0)
|
||||||
|
self._emit(job_id, "log", {"msg": f"downloading {video_id} -> {target.name}"})
|
||||||
|
|
||||||
|
def hook(data: dict[str, Any]) -> None:
|
||||||
|
if data.get("status") not in ("downloading", "finished"):
|
||||||
|
return
|
||||||
|
downloaded = int(data.get("downloaded_bytes") or 0)
|
||||||
|
total = int(data.get("total_bytes") or data.get("total_bytes_estimate") or 0)
|
||||||
|
percent = round(downloaded * 100 / total, 1) if total else None
|
||||||
|
self._emit(job_id, "progress", {
|
||||||
|
"completed": 0, "total": 1, "video_id": video_id,
|
||||||
|
"status": data.get("status"), "downloaded_bytes": downloaded,
|
||||||
|
"total_bytes": total, "percent": percent,
|
||||||
|
"speed": data.get("speed"), "eta": data.get("eta"),
|
||||||
|
})
|
||||||
|
|
||||||
|
self.store.update_video_download(video_id, "downloading", filename=target.name, error=None)
|
||||||
|
ydl_opts = {
|
||||||
|
# Prefer YouTube's HLS VP9 + Opus pair. Chromium can play the
|
||||||
|
# resulting WebM directly, and HLS avoids the recurring mid-range
|
||||||
|
# 403s seen on long HTTPS media requests.
|
||||||
|
"format": "bestvideo[height<=1080][protocol^=m3u8_native][vcodec^=vp09]+bestaudio[acodec^=opus]/bestvideo[height<=1080][protocol^=m3u8_native]+bestaudio[acodec^=opus]/bestvideo[height<=1080][ext=webm]+bestaudio[ext=webm]",
|
||||||
|
"merge_output_format": "webm",
|
||||||
|
"outtmpl": str(out_dir / f"{filename.rsplit('.', 1)[0]}.%(ext)s"),
|
||||||
|
"quiet": True, "no_warnings": True, "noprogress": True,
|
||||||
|
"progress_hooks": [hook],
|
||||||
|
"continuedl": True,
|
||||||
|
# Current YouTube extraction requires an external JS runtime and
|
||||||
|
# the EJS challenge scripts. Node is available in the supported
|
||||||
|
# local browser setup and is installed by yt-dlp[default].
|
||||||
|
"js_runtimes": {"node": {}},
|
||||||
|
"fragment_retries": 10,
|
||||||
|
"retries": self.cfg.yt_dlp.retries,
|
||||||
|
"extractor_retries": self.cfg.yt_dlp.extractor_retries,
|
||||||
|
"sleep_interval": self.cfg.delay.min_seconds,
|
||||||
|
"max_sleep_interval": self.cfg.delay.max_seconds,
|
||||||
|
"sleep_interval_requests": self.cfg.yt_dlp.sleep_subrequests,
|
||||||
|
"socket_timeout": self.cfg.yt_dlp.socket_timeout,
|
||||||
|
"extractor_args": {"youtube": {"player_client": ["default", "web"]}},
|
||||||
|
}
|
||||||
|
if cookie_path and restricted:
|
||||||
|
ydl_opts["cookiefile"] = cookie_path
|
||||||
|
|
||||||
|
try:
|
||||||
|
with yt_dlp.YoutubeDL(ydl_opts) as ydl:
|
||||||
|
ydl.download([row.url])
|
||||||
|
if not target.exists():
|
||||||
|
# yt-dlp may retain the requested stem but choose a different
|
||||||
|
# extension when the merge was skipped; locate only this ID's
|
||||||
|
# directory and accept the generated WebM as the canonical file.
|
||||||
|
candidates = sorted(out_dir.glob("*.webm"), key=lambda p: p.stat().st_mtime, reverse=True)
|
||||||
|
if candidates:
|
||||||
|
target = candidates[0]
|
||||||
|
if not target.exists() or target.stat().st_size <= 0:
|
||||||
|
raise RuntimeError("yt-dlp finished without producing a WebM file")
|
||||||
|
rel_path = target.relative_to(root).as_posix()
|
||||||
|
size = target.stat().st_size
|
||||||
|
self.store.update_video_download(video_id, "done", path=rel_path, filename=target.name, size=size, error=None)
|
||||||
|
self.store.update_job(job_id, status="done", total=1, completed=1, finished=True)
|
||||||
|
self._emit(job_id, "progress", {"completed": 1, "total": 1, "video_id": video_id, "status": "done", "percent": 100, "total_bytes": size})
|
||||||
|
self._emit(job_id, "done", {"completed": 1, "total": 1, "video_id": video_id, "size": size})
|
||||||
|
except Exception as exc:
|
||||||
|
msg = str(exc)
|
||||||
|
self.store.update_video_download(video_id, "error", filename=target.name, error=msg)
|
||||||
|
self.store.update_job(job_id, status="error", total=1, completed=0, last_error=msg, finished=True)
|
||||||
|
self._emit(job_id, "error", {"message": msg, "video_id": video_id})
|
||||||
|
|
||||||
|
def _fail_video_job(self, job_id: str, message: str, video_id: str | None = None) -> None:
|
||||||
|
if video_id:
|
||||||
|
self.store.update_video_download(video_id, "error", error=message)
|
||||||
|
self.store.update_job(job_id, status="error", total=1, completed=0, last_error=message, finished=True)
|
||||||
|
self._emit(job_id, "error", {"message": message, **({"video_id": video_id} if video_id else {})})
|
||||||
|
|
||||||
def _run_channel(self, job_id: str, opts: dict[str, Any]) -> None:
|
def _run_channel(self, job_id: str, opts: dict[str, Any]) -> None:
|
||||||
job = self.store.get_job(job_id)
|
job = self.store.get_job(job_id)
|
||||||
@@ -204,37 +427,75 @@ class JobManager:
|
|||||||
return
|
return
|
||||||
cfg.channel_url = channel_url
|
cfg.channel_url = channel_url
|
||||||
|
|
||||||
self.store.update_job(job_id, status="running")
|
|
||||||
self._emit(job_id, "log", {"msg": f"discovering {channel_url}"})
|
|
||||||
try:
|
|
||||||
_channel_id, channel_name, refs = discover_channel(channel_url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests)
|
|
||||||
except Exception as exc:
|
|
||||||
self.store.update_job(job_id, status="error", last_error=str(exc), finished=True)
|
|
||||||
self._emit(job_id, "error", {"message": f"discovery failed: {exc}"})
|
|
||||||
return
|
|
||||||
|
|
||||||
limit = opts.get("limit")
|
limit = opts.get("limit")
|
||||||
since = opts.get("since")
|
since = opts.get("since")
|
||||||
languages = opts.get("languages")
|
languages = opts.get("languages")
|
||||||
include_shorts = opts.get("include_shorts", cfg.include_shorts)
|
include_shorts = opts.get("include_shorts", cfg.include_shorts)
|
||||||
no_live = opts.get("no_live", not cfg.include_live)
|
no_live = opts.get("no_live", not cfg.include_live)
|
||||||
cookie_override = opts.get("cookies_file")
|
cookie_override = opts.get("cookies_file")
|
||||||
|
full = bool(opts.get("full"))
|
||||||
|
|
||||||
|
sync = cfg.sync
|
||||||
|
incremental = sync.incremental and not full and bool(channel_id)
|
||||||
|
known = self.store.known_video_ids(channel_id) if incremental else set()
|
||||||
|
|
||||||
|
self.store.update_job(job_id, status="running")
|
||||||
|
self._emit(job_id, "log", {
|
||||||
|
"msg": f"discovering {channel_url}"
|
||||||
|
+ (f" (incremental — newest {sync.window} first)" if known else " (full)")
|
||||||
|
})
|
||||||
|
try:
|
||||||
|
result = discover_incremental(
|
||||||
|
channel_url,
|
||||||
|
known,
|
||||||
|
sleep_subrequests=cfg.yt_dlp.sleep_subrequests,
|
||||||
|
window=sync.window,
|
||||||
|
max_window=sync.max_window,
|
||||||
|
overlap=sync.overlap,
|
||||||
|
since=self.store.latest_upload_date(channel_id) if incremental else None,
|
||||||
|
keep=_keep_ref(cfg, include_shorts, no_live),
|
||||||
|
)
|
||||||
|
except Exception as exc:
|
||||||
|
self.store.update_job(job_id, status="error", last_error=str(exc), finished=True)
|
||||||
|
self._emit(job_id, "error", {"message": f"discovery failed: {exc}"})
|
||||||
|
return
|
||||||
|
|
||||||
|
_channel_id, channel_name, avatar = result.channel_id, result.channel_name, result.avatar
|
||||||
|
refs = result.refs
|
||||||
|
self._emit(job_id, "log", {
|
||||||
|
"msg": f"fetched {result.fetched} entries in {result.passes} pass(es), {result.new_count} new"
|
||||||
|
})
|
||||||
|
|
||||||
if languages:
|
if languages:
|
||||||
cfg.languages = languages
|
cfg.languages = parse_languages(languages, cfg.prefer_manual)
|
||||||
if since:
|
if since:
|
||||||
refs = [r for r in refs if (r.upload_date or "") >= since.replace("-", "")]
|
# Keep undated entries: flat discovery does not report upload_date,
|
||||||
if not include_shorts:
|
# so `>= since` on a missing date would discard the whole channel.
|
||||||
refs = [r for r in refs if "/shorts/" not in (r.url or "")]
|
cutoff = since.replace("-", "")
|
||||||
if no_live:
|
refs = [r for r in refs if not r.upload_date or r.upload_date >= cutoff]
|
||||||
refs = [r for r in refs if not (r.url or "").startswith("https://www.youtube.com/live/")]
|
|
||||||
if limit:
|
|
||||||
refs = refs[: int(limit)]
|
|
||||||
|
|
||||||
self.store.upsert_channel(_channel_id, _handle(channel_url), channel_name, len(refs))
|
if not avatar:
|
||||||
|
from ..discover import deep_channel_avatar
|
||||||
|
avatar = deep_channel_avatar(channel_url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests)
|
||||||
|
if result.full_scan:
|
||||||
|
self.store.upsert_channel(_channel_id, _handle(channel_url), channel_name, len(refs), avatar=avatar)
|
||||||
|
else:
|
||||||
|
self.store.update_channel_meta(_channel_id, name=channel_name, avatar=avatar)
|
||||||
|
if avatar:
|
||||||
|
from ..pipeline import cache_channel_avatar
|
||||||
|
cache_channel_avatar(_channel_id, avatar, Path(self.cfg.output_dir_resolved).parent / "avatars")
|
||||||
self.store.upsert_videos(refs)
|
self.store.upsert_videos(refs)
|
||||||
|
self.store.mark_channel_synced(_channel_id)
|
||||||
|
|
||||||
pending = [r for r in self.store.get_pending(_channel_id) if r.video_id in {x.video_id for x in refs}]
|
# Process the freshly-seen window first, then the older backlog, so a
|
||||||
|
# `limit` still means "the newest N" now that discovery stops early.
|
||||||
|
pending = _order_pending(self.store.get_pending(_channel_id), refs)
|
||||||
|
pending = [r for r in pending if _keep_ref(cfg, include_shorts, no_live)(r)]
|
||||||
|
if since:
|
||||||
|
cutoff = since.replace("-", "")
|
||||||
|
pending = [r for r in pending if not r.upload_date or r.upload_date >= cutoff]
|
||||||
|
if limit:
|
||||||
|
pending = pending[: int(limit)]
|
||||||
total = len(pending)
|
total = len(pending)
|
||||||
self.store.update_job(job_id, total=total)
|
self.store.update_job(job_id, total=total)
|
||||||
self._emit(job_id, "progress", {"completed": 0, "total": total})
|
self._emit(job_id, "progress", {"completed": 0, "total": total})
|
||||||
@@ -242,6 +503,8 @@ class JobManager:
|
|||||||
|
|
||||||
cookie_path = cookie_override or resolve_active_path(self.store)
|
cookie_path = cookie_override or resolve_active_path(self.store)
|
||||||
completed = 0
|
completed = 0
|
||||||
|
outcomes = {"processed": 0, "no_subtitles": 0, "errors": 0}
|
||||||
|
guard = self._new_guard()
|
||||||
for row in pending:
|
for row in pending:
|
||||||
if job_id in self._cancel:
|
if job_id in self._cancel:
|
||||||
self.store.update_job(job_id, status="cancelled", completed=completed, finished=True)
|
self.store.update_job(job_id, status="cancelled", completed=completed, finished=True)
|
||||||
@@ -252,14 +515,193 @@ class JobManager:
|
|||||||
cookies_file=cookie_path,
|
cookies_file=cookie_path,
|
||||||
on_log=lambda m: self._emit(job_id, "log", {"msg": m}),
|
on_log=lambda m: self._emit(job_id, "log", {"msg": m}),
|
||||||
)
|
)
|
||||||
|
if status == "done":
|
||||||
|
outcomes["processed"] += 1
|
||||||
|
elif status == "no_subtitles":
|
||||||
|
outcomes["no_subtitles"] += 1
|
||||||
|
self._emit(job_id, "log", {"msg": f"{row.video_id}: no transcript; .md not generated"})
|
||||||
|
else:
|
||||||
|
outcomes["errors"] += 1
|
||||||
|
self._emit(job_id, "log", {"msg": f"{row.video_id}: processing failed; .md not generated"})
|
||||||
completed += 1
|
completed += 1
|
||||||
self.store.update_job(job_id, completed=completed)
|
self.store.update_job(job_id, completed=completed)
|
||||||
self._emit(job_id, "progress", {"completed": completed, "total": total, "video_id": row.video_id, "status": status})
|
self._emit(job_id, "progress", {"completed": completed, "total": total, "video_id": row.video_id, "status": status})
|
||||||
|
if not self._note_outcome(job_id, guard, row.video_id, status):
|
||||||
|
self._stop_throttled(job_id, guard, completed, total, outcomes)
|
||||||
|
return
|
||||||
if completed < total:
|
if completed < total:
|
||||||
polite_sleep(cfg.delay.min_seconds, cfg.delay.max_seconds)
|
polite_sleep(cfg.delay.min_seconds, cfg.delay.max_seconds)
|
||||||
|
|
||||||
self.store.update_job(job_id, status="done", completed=completed, finished=True)
|
self.store.update_job(job_id, status="done", completed=completed, finished=True)
|
||||||
self._emit(job_id, "done", {"completed": completed, "total": total})
|
self._emit(job_id, "done", {"completed": completed, "total": total, **outcomes, "throttling": guard.summary()})
|
||||||
|
|
||||||
|
def _run_discovery(self, job_id: str, opts: dict[str, Any] | None = None) -> None:
|
||||||
|
"""Refresh the catalog without extracting or downloading video content."""
|
||||||
|
job = self.store.get_job(job_id)
|
||||||
|
if not job:
|
||||||
|
return
|
||||||
|
opts = opts or {}
|
||||||
|
full = bool(opts.get("full"))
|
||||||
|
|
||||||
|
if job.channel_id:
|
||||||
|
channel = self.store.get_channel(job.channel_id)
|
||||||
|
channels = [channel] if channel else []
|
||||||
|
else:
|
||||||
|
channels = self.store.list_channels()
|
||||||
|
channels = [c for c in channels if c]
|
||||||
|
total = len(channels)
|
||||||
|
self.store.update_job(job_id, status="running", total=total, completed=0)
|
||||||
|
mode = "full rescan" if full else "incremental (recent uploads only)"
|
||||||
|
self._emit(job_id, "log", {"msg": f"investigating {total} channel(s) — {mode}"})
|
||||||
|
|
||||||
|
totals = {"new_videos": 0, "known_videos": 0, "errors": 0, "fetched": 0}
|
||||||
|
completed = 0
|
||||||
|
guard = self._new_guard()
|
||||||
|
for channel in channels:
|
||||||
|
if job_id in self._cancel:
|
||||||
|
self.store.update_job(job_id, status="cancelled", completed=completed, finished=True)
|
||||||
|
self._emit(job_id, "cancelled", {"completed": completed, "total": total})
|
||||||
|
return
|
||||||
|
|
||||||
|
# Pace between channels too: each one is a fresh burst of
|
||||||
|
# continuation requests, and a catalog refresh over every channel
|
||||||
|
# used to fire them back to back.
|
||||||
|
if completed:
|
||||||
|
polite_sleep(self.cfg.delay.min_seconds, self.cfg.delay.max_seconds)
|
||||||
|
|
||||||
|
channel_id = channel["channel_id"]
|
||||||
|
try:
|
||||||
|
result = self._discover_catalog_channel(channel, full=full)
|
||||||
|
except Exception as exc:
|
||||||
|
wait = guard.note_failure(exc)
|
||||||
|
if guard.tripped:
|
||||||
|
self._stop_throttled(job_id, guard, completed, total, totals)
|
||||||
|
return
|
||||||
|
if wait > 0:
|
||||||
|
self._emit(job_id, "log", {
|
||||||
|
"msg": f"YouTube is throttling us — waiting {wait:.1f}s before the next channel"
|
||||||
|
})
|
||||||
|
time.sleep(wait)
|
||||||
|
if job.channel_id:
|
||||||
|
self.store.update_job(job_id, status="error", completed=completed, last_error=str(exc), finished=True)
|
||||||
|
self._emit(job_id, "error", {"message": f"discovery failed: {exc}"})
|
||||||
|
return
|
||||||
|
totals["errors"] += 1
|
||||||
|
self._emit(job_id, "log", {"msg": f"discovery failed for {channel.get('name') or channel_id}: {exc}"})
|
||||||
|
else:
|
||||||
|
guard.note_success()
|
||||||
|
totals["new_videos"] += result["new_videos"]
|
||||||
|
totals["known_videos"] += result["known_videos"]
|
||||||
|
totals["fetched"] += result["fetched"]
|
||||||
|
name = channel.get("name") or channel_id
|
||||||
|
scope = (
|
||||||
|
f"scanned {result['fetched']} newest"
|
||||||
|
if result["incremental"] else f"scanned all {result['fetched']}"
|
||||||
|
)
|
||||||
|
if not result["caught_up"]:
|
||||||
|
scope += " (hit window ceiling — run a full rescan if videos are missing)"
|
||||||
|
self._emit(job_id, "log", {"msg": f"{name}: {scope}, {result['new_videos']} new"})
|
||||||
|
self._emit(
|
||||||
|
job_id,
|
||||||
|
"progress",
|
||||||
|
{
|
||||||
|
"channel_id": channel_id,
|
||||||
|
"new_videos": result["new_videos"],
|
||||||
|
"known_videos": result["known_videos"],
|
||||||
|
"fetched": result["fetched"],
|
||||||
|
"incremental": result["incremental"],
|
||||||
|
"caught_up": result["caught_up"],
|
||||||
|
"last_video_date": result["last_video_date"],
|
||||||
|
"completed": completed + 1,
|
||||||
|
"total": total,
|
||||||
|
},
|
||||||
|
)
|
||||||
|
completed += 1
|
||||||
|
self.store.update_job(job_id, completed=completed)
|
||||||
|
|
||||||
|
self.store.update_job(job_id, status="done", completed=completed, finished=True)
|
||||||
|
self._emit(job_id, "done", {"completed": completed, "total": total, **totals})
|
||||||
|
|
||||||
|
def _discover_catalog_channel(self, channel: dict, *, full: bool = False) -> dict[str, Any]:
|
||||||
|
"""Refresh one channel's catalog.
|
||||||
|
|
||||||
|
Incremental by default: only the newest slice of the channel is fetched,
|
||||||
|
stopping at the first run of videos we already have. `full` forces the
|
||||||
|
old behaviour (walk every page) for when the local catalog is suspect.
|
||||||
|
"""
|
||||||
|
channel_id = channel["channel_id"]
|
||||||
|
channel_url = _resolve_channel_url(self.store, self.cfg, channel_id)
|
||||||
|
if not channel_url:
|
||||||
|
raise RuntimeError("no channel url")
|
||||||
|
|
||||||
|
sync = self.cfg.sync
|
||||||
|
incremental = sync.incremental and not full
|
||||||
|
known = self.store.known_video_ids(channel_id) if incremental else set()
|
||||||
|
|
||||||
|
result = discover_incremental(
|
||||||
|
channel_url,
|
||||||
|
known,
|
||||||
|
sleep_subrequests=self.cfg.yt_dlp.sleep_subrequests,
|
||||||
|
window=sync.window,
|
||||||
|
max_window=sync.max_window,
|
||||||
|
overlap=sync.overlap,
|
||||||
|
since=self.store.latest_upload_date(channel_id) if incremental else None,
|
||||||
|
keep=_keep_ref(self.cfg),
|
||||||
|
)
|
||||||
|
refs = result.refs
|
||||||
|
|
||||||
|
if result.full_scan:
|
||||||
|
self.store.upsert_channel(result.channel_id, _handle(channel_url), result.channel_name, len(refs))
|
||||||
|
new_videos = self.store.upsert_videos(refs)
|
||||||
|
else:
|
||||||
|
self.store.update_channel_meta(result.channel_id, name=result.channel_name)
|
||||||
|
new_videos = self.store.upsert_videos(refs)
|
||||||
|
marks = self.store.mark_channel_synced(result.channel_id)
|
||||||
|
|
||||||
|
return {
|
||||||
|
"new_videos": new_videos,
|
||||||
|
"known_videos": len({r.video_id for r in refs}) - new_videos,
|
||||||
|
"fetched": result.fetched,
|
||||||
|
"passes": result.passes,
|
||||||
|
"incremental": not result.full_scan,
|
||||||
|
"caught_up": result.caught_up,
|
||||||
|
"last_video_date": marks["last_video_date"],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _keep_ref(cfg: Config, include_shorts: bool | None = None, no_live: bool | None = None):
|
||||||
|
"""Predicate matching the shorts/live rules that decide what reaches the DB.
|
||||||
|
|
||||||
|
Incremental discovery needs the same filter its stored ids were created
|
||||||
|
under, otherwise the tail of a window is full of entries that can never be
|
||||||
|
recognised as known and the window keeps widening for nothing.
|
||||||
|
"""
|
||||||
|
shorts = cfg.include_shorts if include_shorts is None else include_shorts
|
||||||
|
skip_live = (not cfg.include_live) if no_live is None else no_live
|
||||||
|
|
||||||
|
def keep(r: Any) -> bool: # VideoRef or VideoRow — both carry `.url`
|
||||||
|
url = r.url or ""
|
||||||
|
if not shorts and "/shorts/" in url:
|
||||||
|
return False
|
||||||
|
if skip_live and url.startswith("https://www.youtube.com/live/"):
|
||||||
|
return False
|
||||||
|
return True
|
||||||
|
|
||||||
|
return keep
|
||||||
|
|
||||||
|
|
||||||
|
def _order_pending(rows: list, refs: list[VideoRef]) -> list:
|
||||||
|
"""Newest-window-first ordering for the pending queue.
|
||||||
|
|
||||||
|
`get_pending` is ordered by discovery time, which used to coincide with
|
||||||
|
newest-first because discovery saw the whole channel at once. With windowed
|
||||||
|
sync that no longer holds, so put the videos from this run's window at the
|
||||||
|
front and keep the rest of the backlog behind them.
|
||||||
|
"""
|
||||||
|
by_id = {r.video_id: r for r in rows}
|
||||||
|
ordered = [by_id.pop(x.video_id) for x in refs if x.video_id in by_id]
|
||||||
|
ordered.extend(by_id.values())
|
||||||
|
return ordered
|
||||||
|
|
||||||
|
|
||||||
def _clone_config(cfg: Config) -> Config:
|
def _clone_config(cfg: Config) -> Config:
|
||||||
@@ -283,3 +725,13 @@ def _handle(url: str) -> str:
|
|||||||
if "@" in url:
|
if "@" in url:
|
||||||
return "@" + url.split("@", 1)[1].split("/", 1)[0]
|
return "@" + url.split("@", 1)[1].split("/", 1)[0]
|
||||||
return ""
|
return ""
|
||||||
|
|
||||||
|
|
||||||
|
def _safe_video_filename(title: str) -> str:
|
||||||
|
"""Keep the displayed title while making a valid, bounded Windows name."""
|
||||||
|
invalid = set(r'\\/:*?"<>|')
|
||||||
|
safe = "".join("_" if c in invalid or ord(c) < 32 else c for c in title)
|
||||||
|
safe = " ".join(safe.strip().split())
|
||||||
|
safe = safe.rstrip(" .") or "video"
|
||||||
|
# Leave room for the id directory and yt-dlp's temporary suffixes.
|
||||||
|
return safe[:180] + ".webm"
|
||||||
|
|||||||
@@ -38,7 +38,7 @@
|
|||||||
channels: { items: [], pending: {} },
|
channels: { items: [], pending: {} },
|
||||||
videos: { items: [], total: 0, page: 1, size: 25, selected: [] },
|
videos: { items: [], total: 0, page: 1, size: 25, selected: [] },
|
||||||
sidebarOpen: false,
|
sidebarOpen: false,
|
||||||
detail: { video: null, transcript: [], hasAudio: false, audioPlaying: false, activeSeg: -1 },
|
detail: { video: null, transcript: [], hasAudio: false, audioPlaying: false, activeSeg: -1, videoError: false },
|
||||||
search: { q: "", channel: "", items: [], ran: false },
|
search: { q: "", channel: "", items: [], ran: false },
|
||||||
analysis: { tab: "wordcloud", channel: "", term: "", words: [], timeline: [] },
|
analysis: { tab: "wordcloud", channel: "", term: "", words: [], timeline: [] },
|
||||||
|
|
||||||
@@ -330,7 +330,7 @@
|
|||||||
this.view = "detail";
|
this.view = "detail";
|
||||||
this.loading.detail = true;
|
this.loading.detail = true;
|
||||||
this.loading.transcript = true;
|
this.loading.transcript = true;
|
||||||
this.detail = { video: null, transcript: [], hasAudio: false, audioPlaying: false, activeSeg: -1 };
|
this.detail = { video: null, transcript: [], hasAudio: false, audioPlaying: false, activeSeg: -1, videoError: false };
|
||||||
try {
|
try {
|
||||||
const v = await this.api("/api/videos/" + encodeURIComponent(id));
|
const v = await this.api("/api/videos/" + encodeURIComponent(id));
|
||||||
// chapters may come embedded or be absent
|
// chapters may come embedded or be absent
|
||||||
@@ -365,6 +365,34 @@
|
|||||||
} catch (e) { this.toast("Audio failed: " + e.message, "error"); }
|
} catch (e) { this.toast("Audio failed: " + e.message, "error"); }
|
||||||
},
|
},
|
||||||
|
|
||||||
|
async downloadVideoOne(id) {
|
||||||
|
if (!this.detail.video || this.detail.video.video_id !== id) return;
|
||||||
|
if (this.detail.video.status !== "done") {
|
||||||
|
this.toast("Generate the .md before downloading the video", "error");
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
try {
|
||||||
|
const d = await this.api("/api/tools/video", { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify({ video_ids: [id] }) });
|
||||||
|
this.detail.video.video_download_status = "queued";
|
||||||
|
this.detail.video.video_error = null;
|
||||||
|
this.detail.videoError = false;
|
||||||
|
this.toast("Video download started");
|
||||||
|
this.subscribeJob(d.job_id, "video");
|
||||||
|
} catch (e) { this.toast("Video download failed: " + e.message, "error"); }
|
||||||
|
},
|
||||||
|
|
||||||
|
openVideoPlayer() {
|
||||||
|
this.detail.videoError = false;
|
||||||
|
this.$nextTick(() => {
|
||||||
|
const player = this.$refs.videoPlayer;
|
||||||
|
if (player) { player.load(); player.play().catch(() => {}); }
|
||||||
|
});
|
||||||
|
},
|
||||||
|
|
||||||
|
onVideoError() {
|
||||||
|
this.detail.videoError = true;
|
||||||
|
},
|
||||||
|
|
||||||
_pollAudio(id) {
|
_pollAudio(id) {
|
||||||
if (this._audioPoll) clearInterval(this._audioPoll);
|
if (this._audioPoll) clearInterval(this._audioPoll);
|
||||||
let attempts = 0;
|
let attempts = 0;
|
||||||
@@ -928,7 +956,16 @@
|
|||||||
const es = new EventSource("/api/scrape/" + encodeURIComponent(jobId) + "/stream");
|
const es = new EventSource("/api/scrape/" + encodeURIComponent(jobId) + "/stream");
|
||||||
this.scrape.es = es;
|
this.scrape.es = es;
|
||||||
es.addEventListener("log", e => { try { const d = JSON.parse(e.data); this.scrape.log.push(d.msg || ""); } catch (_) {} });
|
es.addEventListener("log", e => { try { const d = JSON.parse(e.data); this.scrape.log.push(d.msg || ""); } catch (_) {} });
|
||||||
es.addEventListener("progress", e => { try { const d = JSON.parse(e.data); this.scrape.progress = d; } catch (_) {} });
|
es.addEventListener("progress", e => {
|
||||||
|
try {
|
||||||
|
const d = JSON.parse(e.data);
|
||||||
|
this.scrape.progress = d;
|
||||||
|
if (kind === "video" && this.detail.video && this.detail.video.video_id === d.video_id) {
|
||||||
|
this.detail.video.video_download_status = d.status === "finished" || d.status === "done" ? "downloading" : "downloading";
|
||||||
|
this.detail.video.video_progress = d;
|
||||||
|
}
|
||||||
|
} catch (_) {}
|
||||||
|
});
|
||||||
es.addEventListener("done", e => {
|
es.addEventListener("done", e => {
|
||||||
let result = {};
|
let result = {};
|
||||||
try { result = e.data ? JSON.parse(e.data) : {}; } catch (_) {}
|
try { result = e.data ? JSON.parse(e.data) : {}; } catch (_) {}
|
||||||
@@ -937,7 +974,9 @@
|
|||||||
this.scrape.log.push("[done] job finished");
|
this.scrape.log.push("[done] job finished");
|
||||||
this.closeStream();
|
this.closeStream();
|
||||||
this.loadJobs();
|
this.loadJobs();
|
||||||
if (this.scrape.kind === "audio") {
|
if (this.scrape.kind === "video") {
|
||||||
|
this.toast("Video download complete", "success");
|
||||||
|
} else if (this.scrape.kind === "audio") {
|
||||||
this.toast("Audio job complete", "success");
|
this.toast("Audio job complete", "success");
|
||||||
} else if (this.scrape.kind === "discovery") {
|
} else if (this.scrape.kind === "discovery") {
|
||||||
const n = Number(result.new_videos || 0);
|
const n = Number(result.new_videos || 0);
|
||||||
@@ -1238,6 +1277,19 @@
|
|||||||
return m.toFixed(0) + "m";
|
return m.toFixed(0) + "m";
|
||||||
},
|
},
|
||||||
|
|
||||||
|
fmtBytes(n) {
|
||||||
|
n = Number(n || 0);
|
||||||
|
if (!n) return "—";
|
||||||
|
const units = ["B", "KB", "MB", "GB", "TB"];
|
||||||
|
let i = 0;
|
||||||
|
while (n >= 1024 && i < units.length - 1) { n /= 1024; i++; }
|
||||||
|
return (i ? n.toFixed(n >= 10 ? 0 : 1) : Math.round(n)) + " " + units[i];
|
||||||
|
},
|
||||||
|
|
||||||
|
fmtSpeed(n) {
|
||||||
|
return n ? this.fmtBytes(n) + "/s" : "—";
|
||||||
|
},
|
||||||
|
|
||||||
fmtNum(n) {
|
fmtNum(n) {
|
||||||
n = Number(n || 0);
|
n = Number(n || 0);
|
||||||
if (!isFinite(n)) return "0";
|
if (!isFinite(n)) return "0";
|
||||||
|
|||||||
@@ -406,9 +406,9 @@
|
|||||||
</button>
|
</button>
|
||||||
<button class="btn accent-grad btn-primary !py-1.5 !px-3 text-xs" @click="downloadVideoOne(detail.video.video_id)" x-show="detail.video.status==='done' && detail.video.video_download_status!=='done'" :disabled="detail.video.video_download_status==='queued' || detail.video.video_download_status==='downloading'">
|
<button class="btn accent-grad btn-primary !py-1.5 !px-3 text-xs" @click="downloadVideoOne(detail.video.video_id)" x-show="detail.video.status==='done' && detail.video.video_download_status!=='done'" :disabled="detail.video.video_download_status==='queued' || detail.video.video_download_status==='downloading'">
|
||||||
<svg x-show="detail.video.video_download_status==='queued' || detail.video.video_download_status==='downloading'" class="spin w-3.5 h-3.5" viewBox="0 0 24 24" fill="none"><circle cx="12" cy="12" r="9" stroke="currentColor" stroke-width="3" stroke-dasharray="40 20"/></svg>
|
<svg x-show="detail.video.video_download_status==='queued' || detail.video.video_download_status==='downloading'" class="spin w-3.5 h-3.5" viewBox="0 0 24 24" fill="none"><circle cx="12" cy="12" r="9" stroke="currentColor" stroke-width="3" stroke-dasharray="40 20"/></svg>
|
||||||
<span x-text="detail.video.video_download_status==='queued' || detail.video.video_download_status==='downloading' ? 'Downloading video…' : 'Download video · 1080p MKV'"></span>
|
<span x-text="detail.video.video_download_status==='queued' || detail.video.video_download_status==='downloading' ? 'Downloading video…' : 'Download video · 1080p WebM'"></span>
|
||||||
</button>
|
</button>
|
||||||
<a class="btn btn-ghost !py-1.5 !px-3 text-xs" x-show="detail.video.video_download_status==='done'" :href="'/api/videos/'+detail.video.video_id+'/media'" :download="detail.video.video_filename || ''">Download MKV</a>
|
<a class="btn btn-ghost !py-1.5 !px-3 text-xs" x-show="detail.video.video_download_status==='done'" :href="'/api/videos/'+detail.video.video_id+'/media'" :download="detail.video.video_filename || ''">Download WebM</a>
|
||||||
<button class="btn btn-ghost !py-1.5 !px-3 text-xs" @click="openClip(detail.video.video_id)">Clip</button>
|
<button class="btn btn-ghost !py-1.5 !px-3 text-xs" @click="openClip(detail.video.video_id)">Clip</button>
|
||||||
<template x-if="detail.video.url"><a class="btn btn-ghost !py-1.5 !px-3 text-xs" :href="detail.video.url" target="_blank" rel="noopener">Open on YouTube ↗</a></template>
|
<template x-if="detail.video.url"><a class="btn btn-ghost !py-1.5 !px-3 text-xs" :href="detail.video.url" target="_blank" rel="noopener">Open on YouTube ↗</a></template>
|
||||||
</div>
|
</div>
|
||||||
@@ -429,18 +429,18 @@
|
|||||||
<div class="cinema-card" x-show="detail.video && detail.video.video_download_status==='done'">
|
<div class="cinema-card" x-show="detail.video && detail.video.video_download_status==='done'">
|
||||||
<div class="cinema-stage">
|
<div class="cinema-stage">
|
||||||
<video x-ref="videoPlayer" class="cinema-video" controls playsinline preload="metadata"
|
<video x-ref="videoPlayer" class="cinema-video" controls playsinline preload="metadata"
|
||||||
:src="'/api/videos/'+detail.video.video_id+'/media'"
|
:src="'/api/videos/'+detail.video.video_id+'/media?inline=1'"
|
||||||
@error="onVideoError()"></video>
|
@error="onVideoError()"></video>
|
||||||
<div class="cinema-error" x-show="detail.videoError">
|
<div class="cinema-error" x-show="detail.videoError">
|
||||||
<div class="text-white font-semibold">Este navegador no pudo reproducir el MKV directamente.</div>
|
<div class="text-white font-semibold">Este navegador no pudo reproducir el WebM directamente.</div>
|
||||||
<div class="text-sm text-zinc-400 mt-1">Puedes descargar el archivo y abrirlo con VLC u otro reproductor local.</div>
|
<div class="text-sm text-zinc-400 mt-1">Puedes descargar el archivo y abrirlo con VLC u otro reproductor local.</div>
|
||||||
<a class="btn btn-ghost mt-3" :href="'/api/videos/'+detail.video.video_id+'/media'" :download="detail.video.video_filename || ''">Descargar MKV</a>
|
<a class="btn btn-ghost mt-3" :href="'/api/videos/'+detail.video.video_id+'/media'" :download="detail.video.video_filename || ''">Descargar WebM</a>
|
||||||
</div>
|
</div>
|
||||||
</div>
|
</div>
|
||||||
<div class="cinema-meta">
|
<div class="cinema-meta">
|
||||||
<div>
|
<div>
|
||||||
<div class="text-white font-semibold" x-text="detail.video.title"></div>
|
<div class="text-white font-semibold" x-text="detail.video.title"></div>
|
||||||
<div class="text-xs text-zinc-500 mt-1">1080p · MKV · <span x-text="fmtBytes(detail.video.video_size)"></span></div>
|
<div class="text-xs text-zinc-500 mt-1">1080p · WebM · <span x-text="fmtBytes(detail.video.video_size)"></span></div>
|
||||||
</div>
|
</div>
|
||||||
<button class="btn btn-ghost !py-1.5 !px-3 text-xs" @click="openVideoPlayer()">Play</button>
|
<button class="btn btn-ghost !py-1.5 !px-3 text-xs" @click="openVideoPlayer()">Play</button>
|
||||||
</div>
|
</div>
|
||||||
|
|||||||
@@ -0,0 +1,143 @@
|
|||||||
|
"""Members-only / gated videos: identified from discovery, labelled, and kept
|
||||||
|
out of bulk work without becoming unreachable.
|
||||||
|
|
||||||
|
yt-dlp's flat listing reports availability="subscriber_only" for members-only
|
||||||
|
videos, so they are knowable before an extraction attempt is ever spent on them.
|
||||||
|
Measured on a real channel: 120 flat entries in one request, exactly 4 flagged,
|
||||||
|
matching exactly the 4 the DB had learned about the expensive way.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from yt_scraper.store import BLOCKING_AVAILABILITY, Store, VideoRef
|
||||||
|
|
||||||
|
|
||||||
|
def _store(tmp_path) -> Store:
|
||||||
|
store = Store(tmp_path / "state.db")
|
||||||
|
store.upsert_channel("UC1", "@alpha", "Alpha", 0)
|
||||||
|
return store
|
||||||
|
|
||||||
|
|
||||||
|
def _ref(vid: str, availability: str | None = None) -> VideoRef:
|
||||||
|
return VideoRef(vid, "UC1", vid, f"https://www.youtube.com/watch?v={vid}", availability=availability)
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------- classification
|
||||||
|
|
||||||
|
def test_availability_from_discovery_is_stored_and_classified(tmp_path):
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([_ref("gated", "subscriber_only"), _ref("open", "public")])
|
||||||
|
|
||||||
|
assert store.get_video("gated").availability == "subscriber_only"
|
||||||
|
assert store.get_video("gated").block_reason == "members_only"
|
||||||
|
assert store.get_video("open").block_reason is None
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("availability,expected", [
|
||||||
|
("subscriber_only", "members_only"),
|
||||||
|
("premium_only", "premium_only"),
|
||||||
|
("private", "private"),
|
||||||
|
("needs_auth", "needs_auth"),
|
||||||
|
])
|
||||||
|
def test_every_blocking_availability_maps_to_a_reason(tmp_path, availability, expected):
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([_ref("v", availability)])
|
||||||
|
assert store.get_video("v").block_reason == expected
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("availability", ["public", "unlisted", None])
|
||||||
|
def test_fetchable_availability_is_never_blocked(tmp_path, availability):
|
||||||
|
"""`unlisted` downloads perfectly well — mislabelling it would hide videos
|
||||||
|
the user can actually have."""
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([_ref("v", availability)])
|
||||||
|
assert store.get_video("v").block_reason is None
|
||||||
|
assert "unlisted" not in BLOCKING_AVAILABILITY
|
||||||
|
|
||||||
|
|
||||||
|
def test_block_reason_falls_back_to_the_recorded_error(tmp_path):
|
||||||
|
"""Rows burned into `error` before availability was captured must still be
|
||||||
|
identifiable without re-fetching them."""
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([_ref("old")])
|
||||||
|
store.mark_error("old", "ERROR: [youtube] x: Join this channel to get access to members-only content")
|
||||||
|
|
||||||
|
assert store.get_video("old").availability is None
|
||||||
|
assert store.get_video("old").block_reason == "members_only"
|
||||||
|
|
||||||
|
|
||||||
|
def test_rate_limit_error_is_not_a_block_reason(tmp_path):
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([_ref("t")])
|
||||||
|
store.mark_error("t", "ERROR: Video unavailable. The current session has been rate-limited by YouTube")
|
||||||
|
|
||||||
|
assert store.get_video("t").block_reason is None
|
||||||
|
|
||||||
|
|
||||||
|
def test_rediscovery_does_not_wipe_a_known_availability(tmp_path):
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([_ref("v", "subscriber_only")])
|
||||||
|
store.upsert_videos([_ref("v", None)]) # a later flat pass omitted the field
|
||||||
|
|
||||||
|
assert store.get_video("v").availability == "subscriber_only"
|
||||||
|
|
||||||
|
|
||||||
|
def test_set_availability_records_what_extraction_learned(tmp_path):
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([_ref("v")])
|
||||||
|
store.set_availability("v", "subscriber_only")
|
||||||
|
assert store.get_video("v").block_reason == "members_only"
|
||||||
|
|
||||||
|
store.set_availability("v", None) # must not clear it
|
||||||
|
assert store.get_video("v").availability == "subscriber_only"
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------- queue behaviour
|
||||||
|
|
||||||
|
def test_blocked_videos_are_kept_out_of_the_pending_queue(tmp_path):
|
||||||
|
"""Bulk runs must not spend requests on videos that cannot be fetched."""
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([_ref("gated", "subscriber_only"), _ref("ok"), _ref("unlisted", "unlisted")])
|
||||||
|
|
||||||
|
ids = {v.video_id for v in store.get_pending("UC1")}
|
||||||
|
assert ids == {"ok", "unlisted"}
|
||||||
|
|
||||||
|
all_ids = {v.video_id for v in store.get_pending("UC1", include_blocked=True)}
|
||||||
|
assert all_ids == {"ok", "unlisted", "gated"}
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_blocked_video_is_still_reachable_by_id(tmp_path):
|
||||||
|
"""Buying the membership must not leave the video permanently stranded."""
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([_ref("gated", "subscriber_only")])
|
||||||
|
|
||||||
|
assert store.get_video("gated") is not None, "explicit per-video processing still works"
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------- filtering
|
||||||
|
|
||||||
|
def test_blocked_filter_finds_them_across_statuses(tmp_path):
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([
|
||||||
|
_ref("gated", "subscriber_only"),
|
||||||
|
_ref("legacy"),
|
||||||
|
_ref("fine"),
|
||||||
|
])
|
||||||
|
store.mark_error("legacy", "ERROR: Join this channel to get access to members-only content")
|
||||||
|
store.mark_status("fine", "no_subtitles")
|
||||||
|
|
||||||
|
rows, total = store.query_videos(blocked=True)
|
||||||
|
assert {r.video_id for r in rows} == {"gated", "legacy"}
|
||||||
|
assert total == 2
|
||||||
|
|
||||||
|
rows, total = store.query_videos(blocked=False)
|
||||||
|
assert {r.video_id for r in rows} == {"fine"}
|
||||||
|
|
||||||
|
|
||||||
|
def test_blocked_filter_absent_means_everything(tmp_path):
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([_ref("gated", "subscriber_only"), _ref("fine")])
|
||||||
|
|
||||||
|
_rows, total = store.query_videos()
|
||||||
|
assert total == 2
|
||||||
@@ -0,0 +1,75 @@
|
|||||||
|
"""Tests for the YAML config loader.
|
||||||
|
|
||||||
|
Focuses on the ``languages`` and ``prefer_manual`` fields, including the
|
||||||
|
legacy/compat behaviour. Network-free, DB-free.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import textwrap
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from yt_scraper.config import load_config, parse_languages
|
||||||
|
|
||||||
|
|
||||||
|
def _write_config(tmp_path, body: str) -> str:
|
||||||
|
p = tmp_path / "config.yaml"
|
||||||
|
p.write_text(textwrap.dedent(body), encoding="utf-8")
|
||||||
|
return str(p)
|
||||||
|
|
||||||
|
|
||||||
|
def test_load_legacy_languages_list(tmp_path) -> None:
|
||||||
|
p = _write_config(tmp_path, """
|
||||||
|
channel_url: "https://example.com/@x/videos"
|
||||||
|
languages: ["es", "en"]
|
||||||
|
prefer_manual: true
|
||||||
|
""")
|
||||||
|
cfg = load_config(p)
|
||||||
|
assert cfg.languages == {"es": "manual", "en": "manual"}
|
||||||
|
|
||||||
|
|
||||||
|
def test_load_legacy_languages_list_with_prefer_manual_false(tmp_path) -> None:
|
||||||
|
p = _write_config(tmp_path, """
|
||||||
|
languages: ["es", "en"]
|
||||||
|
prefer_manual: false
|
||||||
|
""")
|
||||||
|
cfg = load_config(p)
|
||||||
|
assert cfg.languages == {"es": "auto", "en": "auto"}
|
||||||
|
|
||||||
|
|
||||||
|
def test_load_new_dict_languages(tmp_path) -> None:
|
||||||
|
p = _write_config(tmp_path, """
|
||||||
|
languages:
|
||||||
|
en: manual
|
||||||
|
es: auto
|
||||||
|
pt: any
|
||||||
|
prefer_manual: false
|
||||||
|
""")
|
||||||
|
cfg = load_config(p)
|
||||||
|
assert cfg.languages == {"en": "manual", "es": "auto", "pt": "any"}
|
||||||
|
# prefer_manual remains accessible for fallback on `"any"` entries
|
||||||
|
assert cfg.prefer_manual is False
|
||||||
|
|
||||||
|
|
||||||
|
def test_load_unknown_mode_normalises_to_any(tmp_path) -> None:
|
||||||
|
p = _write_config(tmp_path, """
|
||||||
|
languages:
|
||||||
|
en: garbage
|
||||||
|
""")
|
||||||
|
cfg = load_config(p)
|
||||||
|
assert cfg.languages == {"en": "any"}
|
||||||
|
|
||||||
|
|
||||||
|
def test_load_missing_languages_defaults_to_any_not_manual_only(tmp_path) -> None:
|
||||||
|
"""The default must fall back to auto captions.
|
||||||
|
|
||||||
|
A manual-only default silently produces zero transcripts on the many
|
||||||
|
channels that publish only auto-generated captions, and records them as
|
||||||
|
`no_subtitles` — a terminal status that hides a purely configural failure.
|
||||||
|
"""
|
||||||
|
p = _write_config(tmp_path, """
|
||||||
|
channel_url: "https://example.com/@x/videos"
|
||||||
|
""")
|
||||||
|
cfg = load_config(p)
|
||||||
|
assert "es" in cfg.languages and "en" in cfg.languages
|
||||||
|
assert cfg.languages["es"] == "any"
|
||||||
@@ -0,0 +1,157 @@
|
|||||||
|
"""Windowed channel sync: only fetch what is newer than what we already have.
|
||||||
|
|
||||||
|
The whole point is request economy against YouTube, so these tests assert on the
|
||||||
|
`limit` values handed to yt-dlp (which become `playlistend`, i.e. how many
|
||||||
|
continuation pages get requested), not just on the refs that come back.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from yt_scraper.discover import discover_incremental
|
||||||
|
from yt_scraper.store import VideoRef
|
||||||
|
|
||||||
|
|
||||||
|
def _ref(vid: str, upload_date: str | None = None, url: str | None = None) -> VideoRef:
|
||||||
|
return VideoRef(
|
||||||
|
video_id=vid,
|
||||||
|
channel_id="UC1",
|
||||||
|
title=vid,
|
||||||
|
url=url or f"https://www.youtube.com/watch?v={vid}",
|
||||||
|
upload_date=upload_date,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _fake_channel(monkeypatch, catalog: list[VideoRef], calls: list | None = None):
|
||||||
|
"""Patch discover_channel with a channel whose /videos tab is `catalog`
|
||||||
|
(newest first) and that honours `limit` the way playlistend does."""
|
||||||
|
|
||||||
|
def fake(url, sleep_subrequests=2.0, limit=None):
|
||||||
|
if calls is not None:
|
||||||
|
calls.append(limit)
|
||||||
|
return ("UC1", "Alpha", "http://avatar", catalog[:limit] if limit else list(catalog))
|
||||||
|
|
||||||
|
monkeypatch.setattr("yt_scraper.discover.discover_channel", fake)
|
||||||
|
|
||||||
|
|
||||||
|
def test_stops_at_first_run_of_known_videos(monkeypatch):
|
||||||
|
catalog = [_ref(f"v{i:03d}") for i in range(500)]
|
||||||
|
known = {r.video_id for r in catalog[5:]} # everything except the 5 newest
|
||||||
|
calls: list = []
|
||||||
|
_fake_channel(monkeypatch, catalog, calls)
|
||||||
|
|
||||||
|
result = discover_incremental("https://y/@alpha/videos", known, window=30, overlap=3)
|
||||||
|
|
||||||
|
assert calls == [30], "one pass only — must not paginate the whole channel"
|
||||||
|
assert result.fetched == 30
|
||||||
|
assert [r.video_id for r in result.new_refs] == [f"v{i:03d}" for i in range(5)]
|
||||||
|
assert result.caught_up is True
|
||||||
|
assert result.full_scan is False
|
||||||
|
|
||||||
|
|
||||||
|
def test_widens_window_when_the_whole_window_is_new(monkeypatch):
|
||||||
|
catalog = [_ref(f"v{i:03d}") for i in range(500)]
|
||||||
|
known = {r.video_id for r in catalog[100:]} # 100 new uploads since last sync
|
||||||
|
calls: list = []
|
||||||
|
_fake_channel(monkeypatch, catalog, calls)
|
||||||
|
|
||||||
|
result = discover_incremental("https://y/@alpha/videos", known, window=30, overlap=3)
|
||||||
|
|
||||||
|
# 30 -> 60 -> 120: doubles only as far as needed, never the full 500.
|
||||||
|
assert calls == [30, 60, 120]
|
||||||
|
assert result.new_count == 100
|
||||||
|
assert result.caught_up is True
|
||||||
|
|
||||||
|
|
||||||
|
def test_gives_up_at_max_window_and_says_so(monkeypatch):
|
||||||
|
catalog = [_ref(f"v{i:03d}") for i in range(500)]
|
||||||
|
calls: list = []
|
||||||
|
_fake_channel(monkeypatch, catalog, calls)
|
||||||
|
|
||||||
|
result = discover_incremental(
|
||||||
|
"https://y/@alpha/videos", {"not-in-this-channel"}, window=10, max_window=40, overlap=3
|
||||||
|
)
|
||||||
|
|
||||||
|
assert calls == [10, 20, 40]
|
||||||
|
assert result.caught_up is False, "caller must be able to tell the scan was truncated"
|
||||||
|
assert result.fetched == 40
|
||||||
|
|
||||||
|
|
||||||
|
def test_no_local_history_walks_the_whole_channel(monkeypatch):
|
||||||
|
catalog = [_ref(f"v{i:03d}") for i in range(120)]
|
||||||
|
calls: list = []
|
||||||
|
_fake_channel(monkeypatch, catalog, calls)
|
||||||
|
|
||||||
|
result = discover_incremental("https://y/@alpha/videos", set(), window=30)
|
||||||
|
|
||||||
|
assert calls == [None], "a first sync has no boundary to stop at"
|
||||||
|
assert result.full_scan is True
|
||||||
|
assert result.new_count == 120
|
||||||
|
|
||||||
|
|
||||||
|
def test_short_channel_is_exhausted_in_one_pass(monkeypatch):
|
||||||
|
catalog = [_ref("a"), _ref("b")]
|
||||||
|
calls: list = []
|
||||||
|
_fake_channel(monkeypatch, catalog, calls)
|
||||||
|
|
||||||
|
result = discover_incremental("https://y/@alpha/videos", {"b"}, window=30)
|
||||||
|
|
||||||
|
assert calls == [30]
|
||||||
|
assert result.exhausted is True
|
||||||
|
assert result.caught_up is True
|
||||||
|
assert [r.video_id for r in result.new_refs] == ["a"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_keep_filter_applies_before_the_overlap_check(monkeypatch):
|
||||||
|
"""Shorts the store never recorded must not keep the window widening."""
|
||||||
|
catalog = [
|
||||||
|
_ref("s1", url="https://www.youtube.com/shorts/s1"),
|
||||||
|
_ref("s2", url="https://www.youtube.com/shorts/s2"),
|
||||||
|
_ref("s3", url="https://www.youtube.com/shorts/s3"),
|
||||||
|
_ref("known1"),
|
||||||
|
_ref("known2"),
|
||||||
|
_ref("known3"),
|
||||||
|
] + [_ref(f"v{i}") for i in range(50)]
|
||||||
|
calls: list = []
|
||||||
|
_fake_channel(monkeypatch, catalog, calls)
|
||||||
|
|
||||||
|
result = discover_incremental(
|
||||||
|
"https://y/@alpha/videos",
|
||||||
|
{"known1", "known2", "known3"},
|
||||||
|
window=6,
|
||||||
|
overlap=3,
|
||||||
|
keep=lambda r: "/shorts/" not in (r.url or ""),
|
||||||
|
)
|
||||||
|
|
||||||
|
assert calls == [6], "the three known long-form videos end the scan"
|
||||||
|
assert result.new_refs == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_upload_date_cutoff_stops_the_scan_when_dates_are_available(monkeypatch):
|
||||||
|
catalog = [
|
||||||
|
_ref("n1", "20260701"),
|
||||||
|
_ref("n2", "20260630"),
|
||||||
|
_ref("o1", "20250101"),
|
||||||
|
_ref("o2", "20241231"),
|
||||||
|
] + [_ref(f"old{i}", "20200101") for i in range(50)]
|
||||||
|
calls: list = []
|
||||||
|
_fake_channel(monkeypatch, catalog, calls)
|
||||||
|
|
||||||
|
result = discover_incremental(
|
||||||
|
"https://y/@alpha/videos", {"unrelated"}, window=4, overlap=2, since="20260601"
|
||||||
|
)
|
||||||
|
|
||||||
|
assert calls == [4], "entries older than the watermark end the scan"
|
||||||
|
assert result.caught_up is True
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("window,overlap", [(0, 0), (-5, -1)])
|
||||||
|
def test_degenerate_settings_are_clamped(monkeypatch, window, overlap):
|
||||||
|
catalog = [_ref("a"), _ref("b")]
|
||||||
|
calls: list = []
|
||||||
|
_fake_channel(monkeypatch, catalog, calls)
|
||||||
|
|
||||||
|
result = discover_incremental("https://y/@alpha/videos", {"b"}, window=window, overlap=overlap)
|
||||||
|
|
||||||
|
assert calls and calls[0] >= 1
|
||||||
|
assert result.fetched >= 1
|
||||||
@@ -0,0 +1,154 @@
|
|||||||
|
"""Tests for per-language subtitle selection and dict-typed ``Config.languages``.
|
||||||
|
|
||||||
|
These tests are pure: no network, no yt-dlp, no DB. They cover the policy
|
||||||
|
logic that decides which subtitle track to pick for a given
|
||||||
|
``info`` dict produced by yt-dlp, plus the legacy/back-compat shims that
|
||||||
|
let older YAML configs still load.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from yt_scraper.config import Config, parse_languages
|
||||||
|
from yt_scraper.extract import pick_subtitle
|
||||||
|
|
||||||
|
|
||||||
|
# ---------- fixtures ----------
|
||||||
|
|
||||||
|
def _track(ext: str = "json3", url: str = "u") -> dict:
|
||||||
|
return {"ext": ext, "url": url}
|
||||||
|
|
||||||
|
|
||||||
|
def _info(manual: dict | None = None, auto: dict | None = None) -> dict:
|
||||||
|
return {"subtitles": manual or {}, "automatic_captions": auto or {}}
|
||||||
|
|
||||||
|
|
||||||
|
# ---------- pick_subtitle: mixed per-language mode ----------
|
||||||
|
|
||||||
|
def test_mixed_manual_and_auto_per_language() -> None:
|
||||||
|
"""`es: auto` only takes auto; iterating `es` first means it wins.
|
||||||
|
|
||||||
|
With ``{"es": "auto", "en": "manual"}`` the iteration matches `es`
|
||||||
|
against the auto dict and picks the auto track; for `en` only the
|
||||||
|
manual dict is consulted and the manual track is returned. The two
|
||||||
|
policies do NOT cross-pollinate between languages.
|
||||||
|
"""
|
||||||
|
info = _info(
|
||||||
|
manual={"en": [_track(url="man-en")]},
|
||||||
|
auto={"en": [_track(url="auto-en")], "es": [_track(url="auto-es")]},
|
||||||
|
)
|
||||||
|
pick = pick_subtitle(info, {"es": "auto", "en": "manual"}, prefer_manual=True)
|
||||||
|
assert pick is not None
|
||||||
|
assert pick.lang == "es"
|
||||||
|
assert pick.source == "auto"
|
||||||
|
|
||||||
|
|
||||||
|
def test_manual_only_skips_auto_even_when_present() -> None:
|
||||||
|
info = _info(manual={}, auto={"es": [_track(url="auto")]})
|
||||||
|
pick = pick_subtitle(info, {"es": "manual"}, prefer_manual=True)
|
||||||
|
assert pick is None
|
||||||
|
|
||||||
|
|
||||||
|
def test_auto_only_skips_manual_even_when_present() -> None:
|
||||||
|
info = _info(manual={"es": [_track(url="man")]}, auto={})
|
||||||
|
pick = pick_subtitle(info, {"es": "auto"}, prefer_manual=True)
|
||||||
|
assert pick is None
|
||||||
|
|
||||||
|
|
||||||
|
def test_any_defer_to_prefer_manual_default_true() -> None:
|
||||||
|
"""`any` honours the legacy prefer_manual=True default."""
|
||||||
|
info = _info(
|
||||||
|
manual={"es": [_track(url="man")]},
|
||||||
|
auto={"es": [_track(url="auto")]},
|
||||||
|
)
|
||||||
|
pick = pick_subtitle(info, {"es": "any"}, prefer_manual=True)
|
||||||
|
assert pick is not None
|
||||||
|
assert pick.source == "manual"
|
||||||
|
|
||||||
|
|
||||||
|
def test_any_defer_to_prefer_manual_default_false() -> None:
|
||||||
|
info = _info(
|
||||||
|
manual={"es": [_track(url="man")]},
|
||||||
|
auto={"es": [_track(url="auto")]},
|
||||||
|
)
|
||||||
|
pick = pick_subtitle(info, {"es": "any"}, prefer_manual=False)
|
||||||
|
assert pick is not None
|
||||||
|
assert pick.source == "auto"
|
||||||
|
|
||||||
|
|
||||||
|
def test_picks_best_format_within_track() -> None:
|
||||||
|
info = _info(manual={"en": [_track(ext="ttml", url="t"), _track(ext="json3", url="j")]})
|
||||||
|
pick = pick_subtitle(info, {"en": "manual"}, prefer_manual=True)
|
||||||
|
assert pick.url == "j"
|
||||||
|
|
||||||
|
|
||||||
|
# ---------- language normalisation ----------
|
||||||
|
|
||||||
|
def test_language_base_match() -> None:
|
||||||
|
"""`es-419` preference matches the `es` caption track."""
|
||||||
|
info = _info(manual={"es": [_track(url="u")]})
|
||||||
|
pick = pick_subtitle(info, {"es-419": "manual"}, prefer_manual=True)
|
||||||
|
assert pick is not None
|
||||||
|
assert pick.lang == "es"
|
||||||
|
|
||||||
|
|
||||||
|
# ---------- legacy list input ----------
|
||||||
|
|
||||||
|
def test_legacy_list_input_uses_prefer_manual() -> None:
|
||||||
|
info = _info(
|
||||||
|
manual={"es": [_track(url="m")]},
|
||||||
|
auto={"es": [_track(url="a")]},
|
||||||
|
)
|
||||||
|
pick = pick_subtitle(info, ["es"], prefer_manual=True)
|
||||||
|
assert pick.source == "manual"
|
||||||
|
pick = pick_subtitle(info, ["es"], prefer_manual=False)
|
||||||
|
assert pick.source == "auto"
|
||||||
|
|
||||||
|
|
||||||
|
def test_unknown_mode_falls_back_to_any() -> None:
|
||||||
|
info = _info(manual={"es": [_track()]}, auto={"es": [_track()]})
|
||||||
|
pick = pick_subtitle(info, {"es": "garbage"}, prefer_manual=True)
|
||||||
|
assert pick is not None
|
||||||
|
assert pick.source == "manual"
|
||||||
|
|
||||||
|
|
||||||
|
# ---------- parse_languages config helper ----------
|
||||||
|
|
||||||
|
def test_parse_languages_dict_pass_through() -> None:
|
||||||
|
assert parse_languages({"en": "manual", "es": "auto"}, True) == {
|
||||||
|
"en": "manual", "es": "auto",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_languages_dict_unknown_mode_to_any() -> None:
|
||||||
|
assert parse_languages({"en": "garbage"}, True) == {"en": "any"}
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_languages_legacy_list_to_manual_by_default() -> None:
|
||||||
|
assert parse_languages(["es", "en"], prefer_manual=True) == {
|
||||||
|
"es": "manual", "en": "manual",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_languages_legacy_list_to_auto_when_prefer_manual_false() -> None:
|
||||||
|
assert parse_languages(["es", "en"], prefer_manual=False) == {
|
||||||
|
"es": "auto", "en": "auto",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_languages_none_returns_empty() -> None:
|
||||||
|
assert parse_languages(None, True) == {}
|
||||||
|
|
||||||
|
|
||||||
|
# ---------- Config dataclass default shape ----------
|
||||||
|
|
||||||
|
def test_config_languages_defaults_to_dict() -> None:
|
||||||
|
cfg = Config()
|
||||||
|
assert isinstance(cfg.languages, dict)
|
||||||
|
# "any", not "manual": the default must not exclude auto-generated captions.
|
||||||
|
assert cfg.languages == {"es": "any", "en": "any"}
|
||||||
|
|
||||||
|
|
||||||
|
def test_config_prefer_manual_defaults_true() -> None:
|
||||||
|
cfg = Config()
|
||||||
|
assert cfg.prefer_manual is True
|
||||||
@@ -0,0 +1,118 @@
|
|||||||
|
"""One video must own exactly one .md, whatever happens to its title.
|
||||||
|
|
||||||
|
The filename is derived from the title, and titles are not stable: YouTube
|
||||||
|
serves them localised, so the same video came back as "La controversia de
|
||||||
|
Claude Fable 5" on one pass and "The Claude Fable controversy 5" on the next.
|
||||||
|
Creators also simply rename videos.
|
||||||
|
|
||||||
|
`re_render_videos` already deleted the superseded file; `process_video` did not,
|
||||||
|
so a re-scrape after a title change left the old file orphaned on disk. The DB
|
||||||
|
repointed, the stale file stayed, and every later scan had to wade through it —
|
||||||
|
the same shape as the incident that left 94 files for 61 rows.
|
||||||
|
|
||||||
|
The identity that matters is the video id, which never changes. These tests pin
|
||||||
|
that: the row's `markdown_path` is authoritative, and anything it used to point
|
||||||
|
at gets cleaned up.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from yt_scraper.config import Config
|
||||||
|
from yt_scraper.extract import SubtitlePick, VideoData
|
||||||
|
from yt_scraper.parse import Segment
|
||||||
|
from yt_scraper.pipeline import process_video
|
||||||
|
from yt_scraper.render import build_filename_stem
|
||||||
|
from yt_scraper.store import Store, VideoRef
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture
|
||||||
|
def env(tmp_path):
|
||||||
|
cfg = Config(
|
||||||
|
database_path=str(tmp_path / "state.db"),
|
||||||
|
output_dir=str(tmp_path / "markdown"),
|
||||||
|
template_path="templates/video.md.j2",
|
||||||
|
)
|
||||||
|
store = Store(cfg.database_path_resolved)
|
||||||
|
store.upsert_channel("UC1", "@alpha", "Alpha", 1)
|
||||||
|
store.upsert_videos([VideoRef("vid123", "UC1", "t", "https://y/watch?v=vid123", "20260101", 60)])
|
||||||
|
return cfg, store, tmp_path
|
||||||
|
|
||||||
|
|
||||||
|
def _data(title: str) -> VideoData:
|
||||||
|
return VideoData(
|
||||||
|
info={"id": "vid123", "title": title, "upload_date": "20260101", "channel": "Alpha"},
|
||||||
|
segments=[Segment(start=0.0, end=2.0, text="hello")],
|
||||||
|
subtitle=SubtitlePick(url="u", ext="json3", lang="en-orig", source="auto"),
|
||||||
|
has_chapters=False,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _md_files(root: Path) -> list[str]:
|
||||||
|
return sorted(p.name for p in root.rglob("*.md"))
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_retitled_video_does_not_leave_a_second_file(env, monkeypatch):
|
||||||
|
"""The production case: the same video, title localised differently."""
|
||||||
|
cfg, store, tmp_path = env
|
||||||
|
md_root = Path(cfg.output_dir_resolved)
|
||||||
|
|
||||||
|
monkeypatch.setattr("yt_scraper.pipeline.extract_video",
|
||||||
|
lambda *a, **k: _data("La controversia de Claude Fable 5"))
|
||||||
|
assert process_video(store.get_video("vid123"), cfg, store, "Alpha", "UC1", "u") == "done"
|
||||||
|
first = _md_files(md_root)
|
||||||
|
assert len(first) == 1
|
||||||
|
|
||||||
|
monkeypatch.setattr("yt_scraper.pipeline.extract_video",
|
||||||
|
lambda *a, **k: _data("The Claude Fable controversy 5"))
|
||||||
|
assert process_video(store.get_video("vid123"), cfg, store, "Alpha", "UC1", "u") == "done"
|
||||||
|
|
||||||
|
after = _md_files(md_root)
|
||||||
|
assert len(after) == 1, f"one video, {len(after)} files on disk: {after}"
|
||||||
|
# And the DB points at the one that exists.
|
||||||
|
row = store.get_video("vid123")
|
||||||
|
assert (Path(cfg.output_dir_resolved).parent / row.markdown_path).exists()
|
||||||
|
assert Path(row.markdown_path).name == after[0]
|
||||||
|
|
||||||
|
|
||||||
|
def test_rescraping_an_unchanged_video_is_idempotent(env, monkeypatch):
|
||||||
|
cfg, store, tmp_path = env
|
||||||
|
md_root = Path(cfg.output_dir_resolved)
|
||||||
|
monkeypatch.setattr("yt_scraper.pipeline.extract_video", lambda *a, **k: _data("Same Title"))
|
||||||
|
for _ in range(3):
|
||||||
|
process_video(store.get_video("vid123"), cfg, store, "Alpha", "UC1", "u")
|
||||||
|
assert len(_md_files(md_root)) == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_changing_the_filename_template_relocates_rather_than_duplicates(env, monkeypatch):
|
||||||
|
"""Opting into ids in the filename must not strand the old files."""
|
||||||
|
cfg, store, tmp_path = env
|
||||||
|
md_root = Path(cfg.output_dir_resolved)
|
||||||
|
monkeypatch.setattr("yt_scraper.pipeline.extract_video", lambda *a, **k: _data("A Title"))
|
||||||
|
process_video(store.get_video("vid123"), cfg, store, "Alpha", "UC1", "u")
|
||||||
|
|
||||||
|
cfg.filename_template = "{upload_date}_{slug}_{video_id}"
|
||||||
|
process_video(store.get_video("vid123"), cfg, store, "Alpha", "UC1", "u")
|
||||||
|
|
||||||
|
files = _md_files(md_root)
|
||||||
|
assert len(files) == 1, f"template change duplicated the file: {files}"
|
||||||
|
assert "vid123" in files[0]
|
||||||
|
|
||||||
|
|
||||||
|
# ------------------------------------------------------------ template vars
|
||||||
|
|
||||||
|
|
||||||
|
def test_video_id_is_available_to_the_filename_template():
|
||||||
|
"""Lets an operator make the file self-identifying without the DB."""
|
||||||
|
stem = build_filename_stem(
|
||||||
|
"20260101", "Some Title", template="{upload_date}_{slug}_{video_id}", video_id="abc123XYZ_-"
|
||||||
|
)
|
||||||
|
assert stem == "20260101_some-title_abc123XYZ_-"
|
||||||
|
|
||||||
|
|
||||||
|
def test_default_template_is_unchanged():
|
||||||
|
"""Existing libraries keep their filenames; adding the variable is opt-in."""
|
||||||
|
assert build_filename_stem("20260101", "Some Title", video_id="abc123") == "20260101_some-title"
|
||||||
@@ -0,0 +1,246 @@
|
|||||||
|
"""Unit tests for the politeness primitives.
|
||||||
|
|
||||||
|
Nothing here sleeps for real: `Pacer` takes an injectable clock and sleeper, and
|
||||||
|
`backoff_delay` returns the delay instead of consuming it. A test suite that
|
||||||
|
actually waited would be the first thing anyone deleted.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from yt_scraper.ratelimit import (
|
||||||
|
Pacer,
|
||||||
|
ThrottleGuard,
|
||||||
|
backoff_delay,
|
||||||
|
is_quota_exhausted,
|
||||||
|
is_rate_limited,
|
||||||
|
ydl_throttle_opts,
|
||||||
|
)
|
||||||
|
|
||||||
|
# Verbatim from the production database, where 343 of 350 `error` rows carried
|
||||||
|
# one of these. If the detector stops matching them the circuit breaker becomes
|
||||||
|
# decorative, so they are pinned here rather than paraphrased.
|
||||||
|
REAL_THROTTLE_MESSAGES = [
|
||||||
|
"ERROR: [youtube] abcdefghijk: Video unavailable. This content isn't available, "
|
||||||
|
"try again later. The current session has been rate-limited by YouTube for up to an hour.",
|
||||||
|
"ERROR: [youtube] abcdefghijk: This content isn't available, try again later. "
|
||||||
|
"The current session has been rate-limited by YouTube for up to an hour. It is recommended t",
|
||||||
|
"HTTPError: 429 Client Error: Too Many Requests for url: https://www.youtube.com/api/timedtext",
|
||||||
|
"ERROR: [youtube] xyz: Sign in to confirm you're not a bot",
|
||||||
|
"HTTP Error 429: Too Many Requests",
|
||||||
|
]
|
||||||
|
|
||||||
|
# Equally verbatim: these are permanent, must NOT trip the breaker, and must
|
||||||
|
# stay distinguishable from throttling.
|
||||||
|
REAL_PERMANENT_MESSAGES = [
|
||||||
|
"ERROR: [youtube] abcdefghijk: Join this channel to get access to members-only "
|
||||||
|
"content like this video, and other exclusive perks.",
|
||||||
|
"ERROR: [youtube] abcdefghijk: Private video. Sign in if you've been granted access to this video",
|
||||||
|
"ERROR: [youtube] abcdefghijk: This video has been removed by the uploader",
|
||||||
|
"no caption tracks published for this video",
|
||||||
|
"subtitle downloaded but parsed empty (lang=es, format=json3)",
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("msg", REAL_THROTTLE_MESSAGES)
|
||||||
|
def test_detects_real_throttle_messages(msg):
|
||||||
|
assert is_rate_limited(msg) is True
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("msg", REAL_PERMANENT_MESSAGES)
|
||||||
|
def test_ignores_permanent_failures(msg):
|
||||||
|
assert is_rate_limited(msg) is False
|
||||||
|
|
||||||
|
|
||||||
|
def test_quota_is_not_treated_as_plain_throttling():
|
||||||
|
"""Google documents quota exhaustion as daily; backing off cannot fix it."""
|
||||||
|
assert is_quota_exhausted("403 quotaExceeded") is True
|
||||||
|
assert is_quota_exhausted("dailyLimitExceeded") is True
|
||||||
|
assert is_quota_exhausted("rate-limited by YouTube") is False
|
||||||
|
|
||||||
|
|
||||||
|
def test_rate_limited_accepts_exception_objects():
|
||||||
|
assert is_rate_limited(RuntimeError("HTTP Error 429: Too Many Requests")) is True
|
||||||
|
|
||||||
|
|
||||||
|
# Both spellings occur, and they are different strings. yt-dlp raises
|
||||||
|
# "HTTP Error 429: ..."; `requests` raises "429 Client Error: ... for url: ...",
|
||||||
|
# which the project wraps as "HTTPError: 429 ...". Detection used to rely on the
|
||||||
|
# prose for the second form, so a 429 with no reason phrase — routine over
|
||||||
|
# HTTP/2 — went unnoticed and the breaker never counted it.
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
"msg",
|
||||||
|
[
|
||||||
|
"HTTP Error 429: Too Many Requests",
|
||||||
|
"HTTPError: 429 Client Error: Too Many Requests for url: https://youtube.com/api/timedtext",
|
||||||
|
"HTTP Error 429: HTTPError: 429 Client Error: for url: https://youtube.com/api/timedtext",
|
||||||
|
"429 Client Error: for url: https://www.youtube.com/api/timedtext?v=x",
|
||||||
|
"HTTP Error 408: Request Timeout",
|
||||||
|
"HTTPError: 408 Client Error: Request Timeout for url: https://youtube.com/",
|
||||||
|
],
|
||||||
|
)
|
||||||
|
def test_detects_the_status_code_without_relying_on_the_reason_phrase(msg):
|
||||||
|
assert is_rate_limited(msg) is True, f"undetected throttle: {msg}"
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
"msg",
|
||||||
|
[
|
||||||
|
"HTTP Error 404: Not Found",
|
||||||
|
"HTTPError: 403 Client Error: Forbidden for url: https://youtube.com/",
|
||||||
|
"HTTP Error 500: Internal Server Error",
|
||||||
|
"no caption tracks published for this video",
|
||||||
|
# A bare number must not be read as a status code.
|
||||||
|
"video 429 seconds long with 408 segments",
|
||||||
|
],
|
||||||
|
)
|
||||||
|
def test_does_not_treat_other_statuses_as_throttling(msg):
|
||||||
|
assert is_rate_limited(msg) is False, f"false positive: {msg}"
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------- backoff
|
||||||
|
|
||||||
|
|
||||||
|
def test_backoff_grows_and_is_capped():
|
||||||
|
delays = [backoff_delay(n, base=2.0, cap=60.0) for n in range(8)]
|
||||||
|
# Jitter is < 1s so successive doublings still order strictly until the cap.
|
||||||
|
assert delays[0] < delays[1] < delays[2] < delays[3]
|
||||||
|
assert all(d <= 60.0 for d in delays)
|
||||||
|
assert delays[-1] == 60.0
|
||||||
|
|
||||||
|
|
||||||
|
def test_backoff_jitters():
|
||||||
|
"""Same attempt must not produce the same delay twice, or concurrent
|
||||||
|
clients would re-synchronise into waves — the reason Google mandates it."""
|
||||||
|
seen = {backoff_delay(2, base=2.0, cap=60.0) for _ in range(30)}
|
||||||
|
assert len(seen) > 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_backoff_survives_a_runaway_counter():
|
||||||
|
assert backoff_delay(10_000, base=2.0, cap=60.0) == 60.0
|
||||||
|
assert backoff_delay(-5, base=2.0, cap=60.0) <= 3.0
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------- pacer
|
||||||
|
|
||||||
|
|
||||||
|
class FakeClock:
|
||||||
|
def __init__(self):
|
||||||
|
self.now = 1000.0
|
||||||
|
self.slept: list[float] = []
|
||||||
|
|
||||||
|
def time(self) -> float:
|
||||||
|
return self.now
|
||||||
|
|
||||||
|
def sleep(self, seconds: float) -> None:
|
||||||
|
self.slept.append(seconds)
|
||||||
|
self.now += seconds
|
||||||
|
|
||||||
|
|
||||||
|
def test_pacer_spaces_calls():
|
||||||
|
clock = FakeClock()
|
||||||
|
pacer = Pacer(2.0, clock=clock.time, sleeper=clock.sleep)
|
||||||
|
assert pacer.wait() == 0.0 # first call is free
|
||||||
|
assert pacer.wait() == pytest.approx(2.0)
|
||||||
|
assert pacer.wait() == pytest.approx(2.0)
|
||||||
|
assert clock.slept == [2.0, 2.0]
|
||||||
|
|
||||||
|
|
||||||
|
def test_pacer_does_not_charge_for_time_already_spent():
|
||||||
|
"""A caller slower than the interval should never wait on top of its own work."""
|
||||||
|
clock = FakeClock()
|
||||||
|
pacer = Pacer(2.0, clock=clock.time, sleeper=clock.sleep)
|
||||||
|
pacer.wait()
|
||||||
|
clock.now += 10.0 # the request itself took 10s
|
||||||
|
assert pacer.wait() == 0.0
|
||||||
|
assert clock.slept == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_pacer_disabled_by_default_interval():
|
||||||
|
clock = FakeClock()
|
||||||
|
pacer = Pacer(0.0, clock=clock.time, sleeper=clock.sleep)
|
||||||
|
assert [pacer.wait() for _ in range(5)] == [0.0] * 5
|
||||||
|
assert clock.slept == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_pacer_charges_for_multi_request_callers():
|
||||||
|
"""One extract_info is two HTTP requests; billing it as one halves the budget."""
|
||||||
|
clock = FakeClock()
|
||||||
|
pacer = Pacer(2.0, clock=clock.time, sleeper=clock.sleep)
|
||||||
|
pacer.wait(cost=2) # first call still free...
|
||||||
|
assert pacer.wait() == pytest.approx(4.0) # ...but it reserved two slots
|
||||||
|
|
||||||
|
|
||||||
|
def test_penalise_pushes_the_next_slot_out():
|
||||||
|
clock = FakeClock()
|
||||||
|
pacer = Pacer(1.0, clock=clock.time, sleeper=clock.sleep)
|
||||||
|
pacer.wait()
|
||||||
|
pacer.penalise(30.0)
|
||||||
|
assert pacer.wait() == pytest.approx(30.0)
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------- guard
|
||||||
|
|
||||||
|
|
||||||
|
def _guard() -> ThrottleGuard:
|
||||||
|
return ThrottleGuard(threshold=3, base=0.01, cap=0.05)
|
||||||
|
|
||||||
|
|
||||||
|
def test_guard_trips_after_consecutive_throttling():
|
||||||
|
g = _guard()
|
||||||
|
assert g.note_failure(REAL_THROTTLE_MESSAGES[0]) > 0
|
||||||
|
assert not g.tripped
|
||||||
|
g.note_failure(REAL_THROTTLE_MESSAGES[0])
|
||||||
|
assert not g.tripped
|
||||||
|
g.note_failure(REAL_THROTTLE_MESSAGES[0])
|
||||||
|
assert g.tripped
|
||||||
|
assert "consecutive" in (g.tripped_reason or "")
|
||||||
|
|
||||||
|
|
||||||
|
def test_success_resets_the_streak():
|
||||||
|
"""Isolated throttled videos between successes are noise, not a banned session."""
|
||||||
|
g = _guard()
|
||||||
|
for _ in range(10):
|
||||||
|
g.note_failure(REAL_THROTTLE_MESSAGES[0])
|
||||||
|
g.note_success()
|
||||||
|
assert not g.tripped
|
||||||
|
assert g.throttled_total == 10
|
||||||
|
|
||||||
|
|
||||||
|
def test_permanent_failures_never_trip_the_breaker():
|
||||||
|
"""A channel with a few members-only videos must not look like a ban."""
|
||||||
|
g = _guard()
|
||||||
|
for msg in REAL_PERMANENT_MESSAGES * 5:
|
||||||
|
g.note_failure(msg)
|
||||||
|
assert not g.tripped
|
||||||
|
assert g.throttled_total == 0
|
||||||
|
|
||||||
|
|
||||||
|
def test_mixed_failures_do_not_accumulate_into_a_trip():
|
||||||
|
g = _guard()
|
||||||
|
g.note_failure(REAL_THROTTLE_MESSAGES[0])
|
||||||
|
g.note_failure("Private video")
|
||||||
|
g.note_failure(REAL_THROTTLE_MESSAGES[0])
|
||||||
|
g.note_failure("no caption tracks published for this video")
|
||||||
|
g.note_failure(REAL_THROTTLE_MESSAGES[0])
|
||||||
|
assert not g.tripped
|
||||||
|
|
||||||
|
|
||||||
|
def test_quota_trips_immediately_without_backoff():
|
||||||
|
g = _guard()
|
||||||
|
assert g.note_failure("403 quotaExceeded") == 0.0
|
||||||
|
assert g.tripped
|
||||||
|
assert "quota" in (g.tripped_reason or "").lower()
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------- ydl opts
|
||||||
|
|
||||||
|
|
||||||
|
def test_throttle_opts_use_names_yt_dlp_actually_reads():
|
||||||
|
opts = ydl_throttle_opts(2.5, extractor_retries=4, socket_timeout=15.0)
|
||||||
|
assert opts["sleep_interval_requests"] == 2.5
|
||||||
|
assert opts["extractor_retries"] == 4
|
||||||
|
assert opts["socket_timeout"] == 15.0
|
||||||
|
# The bug this whole module exists to prevent.
|
||||||
|
assert "sleep_subrequests" not in opts
|
||||||
@@ -0,0 +1,356 @@
|
|||||||
|
"""Recovery paths: retryable statuses, recorded skip reasons, disk<->DB reconcile.
|
||||||
|
|
||||||
|
The motivating incident: 511 videos were stored as `no_subtitles` because the
|
||||||
|
language policy was manual-only while the channel publishes only auto-generated
|
||||||
|
captions. `reset_errors` could not reach them, and nothing recorded why they
|
||||||
|
were skipped, so the failure was both invisible and irreversible.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from yt_scraper.extract import describe_missing_subtitle
|
||||||
|
from yt_scraper.segments import reconcile_markdown
|
||||||
|
from yt_scraper.store import Store, VideoRef
|
||||||
|
|
||||||
|
|
||||||
|
def _store(tmp_path) -> Store:
|
||||||
|
store = Store(tmp_path / "state.db")
|
||||||
|
store.upsert_channel("UC1", "@alpha", "Alpha", 0)
|
||||||
|
return store
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------- reset
|
||||||
|
|
||||||
|
def test_reset_reaches_no_subtitles_not_just_error(tmp_path):
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([
|
||||||
|
VideoRef("a", "UC1", "A", "https://y/watch?v=a"),
|
||||||
|
VideoRef("b", "UC1", "B", "https://y/watch?v=b"),
|
||||||
|
VideoRef("c", "UC1", "C", "https://y/watch?v=c"),
|
||||||
|
])
|
||||||
|
store.mark_status("a", "no_subtitles", "policy rejected auto captions")
|
||||||
|
store.mark_error("b", "rate limited")
|
||||||
|
store.mark_done("c", "markdown/c.md", "es", "auto", False)
|
||||||
|
|
||||||
|
n = store.reset_videos("UC1", ("error", "no_subtitles"))
|
||||||
|
|
||||||
|
assert n == 2
|
||||||
|
assert store.get_video("a").status == "pending"
|
||||||
|
assert store.get_video("b").status == "pending"
|
||||||
|
assert store.get_video("c").status == "done", "finished work must not be re-queued"
|
||||||
|
|
||||||
|
|
||||||
|
def test_reset_never_touches_done_even_if_asked(tmp_path):
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([VideoRef("c", "UC1", "C", "https://y/watch?v=c")])
|
||||||
|
store.mark_done("c", "markdown/c.md", "es", "auto", False)
|
||||||
|
|
||||||
|
assert store.reset_videos("UC1", ("done",)) == 0
|
||||||
|
assert store.get_video("c").status == "done"
|
||||||
|
|
||||||
|
|
||||||
|
def test_reset_errors_still_only_resets_errors(tmp_path):
|
||||||
|
"""The narrower legacy helper keeps its old meaning."""
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([
|
||||||
|
VideoRef("a", "UC1", "A", "https://y/watch?v=a"),
|
||||||
|
VideoRef("b", "UC1", "B", "https://y/watch?v=b"),
|
||||||
|
])
|
||||||
|
store.mark_status("a", "no_subtitles")
|
||||||
|
store.mark_error("b", "boom")
|
||||||
|
|
||||||
|
assert store.reset_errors("UC1") == 1
|
||||||
|
assert store.get_video("a").status == "no_subtitles"
|
||||||
|
assert store.get_video("b").status == "pending"
|
||||||
|
|
||||||
|
|
||||||
|
def test_permanent_failures_are_excluded_from_retries(tmp_path):
|
||||||
|
"""Members-only videos cannot be fixed by retrying; re-running them only
|
||||||
|
spends requests the recoverable videos need."""
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([
|
||||||
|
VideoRef(v, "UC1", v, f"https://y/watch?v={v}") for v in ("members", "throttled", "priv")
|
||||||
|
])
|
||||||
|
store.mark_error("members", "ERROR: [youtube] x: Join this channel to get access to members-only content")
|
||||||
|
store.mark_error("throttled", "ERROR: Video unavailable. The current session has been rate-limited by YouTube")
|
||||||
|
store.mark_error("priv", "ERROR: Private video. Sign in if you've been granted access")
|
||||||
|
|
||||||
|
counts = store.retryable_counts("UC1")
|
||||||
|
assert counts["error"] == 1, "only the throttled one is worth retrying"
|
||||||
|
assert counts["permanent"] == 2
|
||||||
|
|
||||||
|
assert store.reset_videos("UC1", ("error",)) == 1
|
||||||
|
assert store.get_video("throttled").status == "pending"
|
||||||
|
assert store.get_video("members").status == "error"
|
||||||
|
assert store.get_video("priv").status == "error"
|
||||||
|
|
||||||
|
|
||||||
|
def test_permanent_failures_can_be_reset_when_explicitly_asked(tmp_path):
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([VideoRef("members", "UC1", "M", "https://y/watch?v=members")])
|
||||||
|
store.mark_error("members", "ERROR: members-only content")
|
||||||
|
|
||||||
|
assert store.reset_videos("UC1", ("error",)) == 0
|
||||||
|
assert store.reset_videos("UC1", ("error",), include_permanent=True) == 1
|
||||||
|
assert store.get_video("members").status == "pending"
|
||||||
|
|
||||||
|
|
||||||
|
def test_rate_limit_wording_is_never_treated_as_permanent(tmp_path):
|
||||||
|
"""The throttling message is the one that must stay retryable."""
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([VideoRef("t", "UC1", "T", "https://y/watch?v=t")])
|
||||||
|
store.mark_error(
|
||||||
|
"t",
|
||||||
|
"ERROR: [youtube] t: Video unavailable. This content isn't available, try again later. "
|
||||||
|
"The current session has been rate-limited by YouTube for up to an hour.",
|
||||||
|
)
|
||||||
|
|
||||||
|
assert store.retryable_counts("UC1")["permanent"] == 0
|
||||||
|
assert store.reset_videos("UC1", ("error",)) == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_retryable_counts_reports_both_statuses(tmp_path):
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([
|
||||||
|
VideoRef(v, "UC1", v, f"https://y/watch?v={v}") for v in ("a", "b", "c")
|
||||||
|
])
|
||||||
|
store.mark_status("a", "no_subtitles")
|
||||||
|
store.mark_status("b", "no_subtitles")
|
||||||
|
store.mark_error("c", "boom")
|
||||||
|
|
||||||
|
assert store.retryable_counts("UC1") == {"error": 1, "no_subtitles": 2, "permanent": 0}
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------- skip reasons
|
||||||
|
|
||||||
|
def test_mark_status_records_the_reason(tmp_path):
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([VideoRef("a", "UC1", "A", "https://y/watch?v=a")])
|
||||||
|
|
||||||
|
store.mark_status("a", "no_subtitles", "no track matched the language policy")
|
||||||
|
|
||||||
|
assert "language policy" in store.get_video("a").error_msg
|
||||||
|
|
||||||
|
|
||||||
|
def test_describe_distinguishes_no_captions_from_policy_rejection():
|
||||||
|
none_at_all = describe_missing_subtitle({"subtitles": {}, "automatic_captions": {}}, {"es": "manual"})
|
||||||
|
assert "no caption tracks published" in none_at_all
|
||||||
|
|
||||||
|
auto_only = describe_missing_subtitle(
|
||||||
|
{"subtitles": {}, "automatic_captions": {"es": [{"url": "u", "ext": "json3"}]}},
|
||||||
|
{"es": "manual"},
|
||||||
|
)
|
||||||
|
assert "ONLY auto-generated" in auto_only
|
||||||
|
assert "'any' or 'auto'" in auto_only
|
||||||
|
|
||||||
|
|
||||||
|
def test_describe_does_not_blame_config_when_mode_already_allows_auto():
|
||||||
|
msg = describe_missing_subtitle(
|
||||||
|
{"subtitles": {}, "automatic_captions": {"de": [{"url": "u", "ext": "json3"}]}},
|
||||||
|
{"es": "any"},
|
||||||
|
)
|
||||||
|
assert "ONLY auto-generated" not in msg
|
||||||
|
assert "no track matched the language policy" in msg
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------- reconcile
|
||||||
|
|
||||||
|
_MD = """---
|
||||||
|
video_id: "{vid}"
|
||||||
|
title: "T"
|
||||||
|
upload_date: "2026-01-02"
|
||||||
|
---
|
||||||
|
|
||||||
|
## Transcript
|
||||||
|
|
||||||
|
**00:00** · hola mundo
|
||||||
|
"""
|
||||||
|
|
||||||
|
|
||||||
|
def test_reconcile_marks_done_when_the_md_is_already_on_disk(tmp_path):
|
||||||
|
"""The exact symptom the user reported: a .md exists but the row still
|
||||||
|
shows a failure, and nothing reconciles it without a server restart."""
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([VideoRef("vid1", "UC1", "T", "https://y/watch?v=vid1")])
|
||||||
|
store.mark_status("vid1", "no_subtitles", "stale failure")
|
||||||
|
|
||||||
|
md_root = tmp_path / "markdown" / "Alpha"
|
||||||
|
md_root.mkdir(parents=True)
|
||||||
|
(md_root / "2026-01-02_t.md").write_text(_MD.format(vid="vid1"), encoding="utf-8")
|
||||||
|
|
||||||
|
result = reconcile_markdown(store, tmp_path / "markdown")
|
||||||
|
|
||||||
|
row = store.get_video("vid1")
|
||||||
|
assert row.status == "done"
|
||||||
|
assert row.markdown_path == "markdown/Alpha/2026-01-02_t.md"
|
||||||
|
assert result["repaired_done"] == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_reconcile_requeues_rows_whose_md_vanished_only_with_prune(tmp_path):
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([
|
||||||
|
VideoRef("gone", "UC1", "G", "https://y/watch?v=gone"),
|
||||||
|
VideoRef("vid1", "UC1", "T", "https://y/watch?v=vid1"),
|
||||||
|
])
|
||||||
|
store.mark_done("gone", "markdown/Alpha/nope.md", "es", "auto", False)
|
||||||
|
md_root = tmp_path / "markdown" / "Alpha"
|
||||||
|
md_root.mkdir(parents=True)
|
||||||
|
(md_root / "2026-01-02_t.md").write_text(_MD.format(vid="vid1"), encoding="utf-8")
|
||||||
|
|
||||||
|
# default is non-destructive
|
||||||
|
assert reconcile_markdown(store, tmp_path / "markdown")["missing_md"] == 0
|
||||||
|
assert store.get_video("gone").status == "done"
|
||||||
|
|
||||||
|
result = reconcile_markdown(store, tmp_path / "markdown", prune=True)
|
||||||
|
assert result["missing_md"] == 1
|
||||||
|
assert store.get_video("gone").status == "pending"
|
||||||
|
|
||||||
|
|
||||||
|
def test_prune_refuses_to_demote_everything_when_the_root_is_empty(tmp_path):
|
||||||
|
"""Pointed at a wrong or not-yet-populated markdown root, prune must be a
|
||||||
|
no-op rather than wiping every finished video in the database."""
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([VideoRef("c", "UC1", "C", "https://y/watch?v=c")])
|
||||||
|
store.mark_done("c", "markdown/Alpha/c.md", "es", "auto", False)
|
||||||
|
(tmp_path / "markdown").mkdir()
|
||||||
|
|
||||||
|
result = reconcile_markdown(store, tmp_path / "markdown", prune=True)
|
||||||
|
|
||||||
|
assert result["missing_md"] == 0
|
||||||
|
assert store.get_video("c").status == "done"
|
||||||
|
|
||||||
|
|
||||||
|
def test_reconcile_is_idempotent(tmp_path):
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([VideoRef("vid1", "UC1", "T", "https://y/watch?v=vid1")])
|
||||||
|
md_root = tmp_path / "markdown" / "Alpha"
|
||||||
|
md_root.mkdir(parents=True)
|
||||||
|
(md_root / "2026-01-02_t.md").write_text(_MD.format(vid="vid1"), encoding="utf-8")
|
||||||
|
|
||||||
|
first = reconcile_markdown(store, tmp_path / "markdown")
|
||||||
|
second = reconcile_markdown(store, tmp_path / "markdown")
|
||||||
|
|
||||||
|
assert first["repaired_done"] == 1
|
||||||
|
assert second["repaired_done"] == 0
|
||||||
|
assert second["missing_md"] == 0
|
||||||
|
assert store.get_video("vid1").status == "done"
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_bad_encoding_does_not_abort_the_whole_scan(tmp_path):
|
||||||
|
"""UnicodeDecodeError is a ValueError, not an OSError. Letting it escape
|
||||||
|
aborted the loop, so one bad file silently hid every later one."""
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([
|
||||||
|
VideoRef("v1", "UC1", "A", "https://y/watch?v=v1"),
|
||||||
|
VideoRef("v3", "UC1", "C", "https://y/watch?v=v3"),
|
||||||
|
])
|
||||||
|
md_root = tmp_path / "markdown" / "Alpha"
|
||||||
|
md_root.mkdir(parents=True)
|
||||||
|
(md_root / "a.md").write_text(_MD.format(vid="v1"), encoding="utf-8")
|
||||||
|
(md_root / "b.md").write_bytes(b"---\nvideo_id: \xe9\xe9\xe9\n---\n")
|
||||||
|
(md_root / "c.md").write_text(_MD.format(vid="v3"), encoding="utf-8")
|
||||||
|
|
||||||
|
result = reconcile_markdown(store, tmp_path / "markdown")
|
||||||
|
|
||||||
|
assert store.get_video("v3").status == "done", "the file after the bad one must still import"
|
||||||
|
assert result["repaired_done"] == 2
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------- re-render
|
||||||
|
|
||||||
|
def test_re_render_updates_markdown_path_and_removes_the_old_file(tmp_path):
|
||||||
|
"""re-render used a raw compact upload_date while process_video uses the
|
||||||
|
hyphenated form, so it wrote a SECOND .md and never told the DB — leaving
|
||||||
|
the app serving the older file. Measured on real data: 94 files, 61 rows."""
|
||||||
|
from yt_scraper.config import Config
|
||||||
|
from yt_scraper.pipeline import re_render_videos
|
||||||
|
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([VideoRef("v1", "UC1", "Mi Video", "https://y/watch?v=v1", "20240519", 60)])
|
||||||
|
store.update_video_metadata(
|
||||||
|
"v1", view_count=1, like_count=1, tags=None, thumbnail=None, description=None,
|
||||||
|
chapters_json="[]",
|
||||||
|
segments_json='[{"start": 0.0, "end": 2.0, "text": "hola"}]',
|
||||||
|
)
|
||||||
|
md_root = tmp_path / "markdown"
|
||||||
|
old_dir = md_root / "Alpha"
|
||||||
|
old_dir.mkdir(parents=True)
|
||||||
|
(old_dir / "2024-05-19_mi-video.md").write_text("stale", encoding="utf-8")
|
||||||
|
store.mark_done("v1", "markdown/Alpha/2024-05-19_mi-video.md", "es", "auto", False)
|
||||||
|
|
||||||
|
cfg = Config(
|
||||||
|
database_path=str(tmp_path / "state.db"),
|
||||||
|
output_dir=str(md_root),
|
||||||
|
template_path=str(Path("templates/video.md.j2").resolve()),
|
||||||
|
)
|
||||||
|
assert re_render_videos(store, cfg) == 1
|
||||||
|
|
||||||
|
files = sorted(p.name for p in md_root.rglob("*.md"))
|
||||||
|
assert len(files) == 1, f"re-render must not leave an orphan beside it: {files}"
|
||||||
|
row = store.get_video("v1")
|
||||||
|
assert row.markdown_path.replace("\\", "/").endswith(files[0])
|
||||||
|
assert (md_root.parent / row.markdown_path).exists()
|
||||||
|
assert "stale" not in (md_root.parent / row.markdown_path).read_text(encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def test_re_render_is_idempotent(tmp_path):
|
||||||
|
from yt_scraper.config import Config
|
||||||
|
from yt_scraper.pipeline import re_render_videos
|
||||||
|
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([VideoRef("v1", "UC1", "Mi Video", "https://y/watch?v=v1", "20240519", 60)])
|
||||||
|
store.update_video_metadata(
|
||||||
|
"v1", view_count=None, like_count=None, tags=None, thumbnail=None, description=None,
|
||||||
|
chapters_json="[]", segments_json='[{"start": 0.0, "end": 2.0, "text": "hola"}]',
|
||||||
|
)
|
||||||
|
store.mark_done("v1", "markdown/Alpha/whatever.md", "es", "auto", False)
|
||||||
|
cfg = Config(
|
||||||
|
database_path=str(tmp_path / "state.db"),
|
||||||
|
output_dir=str(tmp_path / "markdown"),
|
||||||
|
template_path=str(Path("templates/video.md.j2").resolve()),
|
||||||
|
)
|
||||||
|
|
||||||
|
re_render_videos(store, cfg)
|
||||||
|
first = store.get_video("v1").markdown_path
|
||||||
|
re_render_videos(store, cfg)
|
||||||
|
|
||||||
|
assert store.get_video("v1").markdown_path == first
|
||||||
|
assert len(list((tmp_path / "markdown").rglob("*.md"))) == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_reconcile_reports_stale_duplicates_and_deletes_them_only_with_prune(tmp_path):
|
||||||
|
"""The 33 leftover files the old re-render wrote under a second filename:
|
||||||
|
the DB points at one, the other is dead weight."""
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([VideoRef("vid1", "UC1", "T", "https://y/watch?v=vid1")])
|
||||||
|
md_root = tmp_path / "markdown" / "Alpha"
|
||||||
|
md_root.mkdir(parents=True)
|
||||||
|
canonical = md_root / "2026-01-02_t.md"
|
||||||
|
duplicate = md_root / "20260102_t.md"
|
||||||
|
canonical.write_text(_MD.format(vid="vid1"), encoding="utf-8")
|
||||||
|
duplicate.write_text(_MD.format(vid="vid1"), encoding="utf-8")
|
||||||
|
store.mark_done("vid1", "markdown/Alpha/2026-01-02_t.md", "es", "auto", False)
|
||||||
|
|
||||||
|
result = reconcile_markdown(store, tmp_path / "markdown")
|
||||||
|
assert result["stale_dupe"] == 1
|
||||||
|
assert duplicate.exists(), "reporting only by default"
|
||||||
|
assert store.get_video("vid1").markdown_path == "markdown/Alpha/2026-01-02_t.md"
|
||||||
|
|
||||||
|
result = reconcile_markdown(store, tmp_path / "markdown", prune=True)
|
||||||
|
assert result["stale_dupe"] == 1
|
||||||
|
assert not duplicate.exists()
|
||||||
|
assert canonical.exists(), "the file the DB points at must survive"
|
||||||
|
assert store.get_video("vid1").status == "done"
|
||||||
|
|
||||||
|
|
||||||
|
def test_reconcile_counts_orphan_markdown(tmp_path):
|
||||||
|
store = _store(tmp_path)
|
||||||
|
md_root = tmp_path / "markdown" / "Alpha"
|
||||||
|
md_root.mkdir(parents=True)
|
||||||
|
(md_root / "ghost.md").write_text(_MD.format(vid="not-in-db"), encoding="utf-8")
|
||||||
|
|
||||||
|
result = reconcile_markdown(store, tmp_path / "markdown")
|
||||||
|
|
||||||
|
assert result["orphan_md"] == 1
|
||||||
|
assert result["repaired_done"] == 0
|
||||||
@@ -0,0 +1,171 @@
|
|||||||
|
"""Guards on how many requests we spend and under what option names.
|
||||||
|
|
||||||
|
Two classes of regression live here, both of which actually happened:
|
||||||
|
|
||||||
|
1. An option name yt-dlp does not recognise. yt-dlp ignores unknown keys
|
||||||
|
silently, so `sleep_subrequests` looked configured for the project's whole
|
||||||
|
history while nothing ever slept between requests. `test_*_options_are_real`
|
||||||
|
checks every key against yt-dlp's own list instead of trusting review.
|
||||||
|
|
||||||
|
2. An extraction that walks far more of a channel than it needs.
|
||||||
|
`deep_channel_avatar` read one avatar URL by fully extracting every video the
|
||||||
|
channel had ever published — 735 requests and still going when a measurement
|
||||||
|
aborted it. The bound is asserted here because nothing else would notice.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
import yt_dlp
|
||||||
|
|
||||||
|
from yt_scraper import discover, extract
|
||||||
|
|
||||||
|
#: Every option name yt-dlp's own CLI parser produces. Anything outside this is
|
||||||
|
#: either a typo or something yt-dlp will silently drop.
|
||||||
|
KNOWN_YDL_OPTIONS = set(yt_dlp.parse_options([]).ydl_opts)
|
||||||
|
|
||||||
|
|
||||||
|
class CapturingYDL:
|
||||||
|
"""Stands in for yt_dlp.YoutubeDL and records the options it was built with."""
|
||||||
|
|
||||||
|
captured: list[dict] = []
|
||||||
|
info: dict = {}
|
||||||
|
|
||||||
|
def __init__(self, options=None, *args, **kwargs):
|
||||||
|
type(self).captured.append(dict(options or {}))
|
||||||
|
self.options = options or {}
|
||||||
|
|
||||||
|
def __enter__(self):
|
||||||
|
return self
|
||||||
|
|
||||||
|
def __exit__(self, *exc):
|
||||||
|
return False
|
||||||
|
|
||||||
|
def extract_info(self, url, download=False, process=True):
|
||||||
|
return dict(type(self).info)
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture
|
||||||
|
def capture(monkeypatch):
|
||||||
|
CapturingYDL.captured = []
|
||||||
|
CapturingYDL.info = {
|
||||||
|
"id": "UC123",
|
||||||
|
"channel_id": "UC123",
|
||||||
|
"channel": "Test Channel",
|
||||||
|
"thumbnails": [{"url": "https://yt3.ggpht.com/avatar.jpg"}],
|
||||||
|
"entries": [],
|
||||||
|
}
|
||||||
|
monkeypatch.setattr(yt_dlp, "YoutubeDL", CapturingYDL)
|
||||||
|
return CapturingYDL
|
||||||
|
|
||||||
|
|
||||||
|
def assert_options_are_real(opts: dict, where: str) -> None:
|
||||||
|
unknown = sorted(set(opts) - KNOWN_YDL_OPTIONS)
|
||||||
|
assert not unknown, (
|
||||||
|
f"{where} passes option(s) yt-dlp does not recognise and will silently "
|
||||||
|
f"ignore: {unknown}"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
# ------------------------------------------------------------------ names
|
||||||
|
|
||||||
|
|
||||||
|
def test_discover_channel_options_are_real(capture):
|
||||||
|
discover.discover_channel("https://www.youtube.com/@x/videos", sleep_subrequests=2.0)
|
||||||
|
assert_options_are_real(capture.captured[0], "discover_channel")
|
||||||
|
|
||||||
|
|
||||||
|
def test_deep_channel_avatar_options_are_real(capture):
|
||||||
|
discover.deep_channel_avatar("https://www.youtube.com/@x/videos", sleep_subrequests=2.0)
|
||||||
|
assert_options_are_real(capture.captured[0], "deep_channel_avatar")
|
||||||
|
|
||||||
|
|
||||||
|
def test_extract_video_options_are_real(capture):
|
||||||
|
capture.info = {"id": "vid", "title": "t", "subtitles": {}, "automatic_captions": {}}
|
||||||
|
extract.extract_video("https://www.youtube.com/watch?v=vid", {"es": "any"})
|
||||||
|
assert_options_are_real(capture.captured[0], "extract_video")
|
||||||
|
|
||||||
|
|
||||||
|
# ------------------------------------------------------------------ throttles wired
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
"call",
|
||||||
|
[
|
||||||
|
pytest.param(
|
||||||
|
lambda: discover.discover_channel("https://www.youtube.com/@x/videos",
|
||||||
|
sleep_subrequests=3.25),
|
||||||
|
id="discover_channel",
|
||||||
|
),
|
||||||
|
pytest.param(
|
||||||
|
lambda: discover.deep_channel_avatar("https://www.youtube.com/@x/videos",
|
||||||
|
sleep_subrequests=3.25),
|
||||||
|
id="deep_channel_avatar",
|
||||||
|
),
|
||||||
|
pytest.param(
|
||||||
|
lambda: extract.extract_video("https://www.youtube.com/watch?v=vid",
|
||||||
|
{"es": "any"}, sleep_subrequests=3.25),
|
||||||
|
id="extract_video",
|
||||||
|
),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
def test_every_entry_point_forwards_the_real_sleep_option(capture, call):
|
||||||
|
capture.info = {"id": "vid", "title": "t", "subtitles": {}, "automatic_captions": {},
|
||||||
|
"channel_id": "UC1", "entries": []}
|
||||||
|
call()
|
||||||
|
opts = capture.captured[0]
|
||||||
|
assert opts.get("sleep_interval_requests") == 3.25
|
||||||
|
assert opts.get("socket_timeout"), "a hung connection must not block the worker forever"
|
||||||
|
assert "extractor_retries" in opts
|
||||||
|
|
||||||
|
|
||||||
|
# ------------------------------------------------------------------ request bounds
|
||||||
|
|
||||||
|
|
||||||
|
def test_deep_avatar_does_not_walk_the_channel(capture):
|
||||||
|
"""The 735-request bug. Both halves of the fix are asserted.
|
||||||
|
|
||||||
|
`extract_flat` stops yt-dlp expanding each entry into a full extraction, and
|
||||||
|
`playlistend` stops it paginating past the first page. Either one missing
|
||||||
|
puts the whole channel back on the wire.
|
||||||
|
"""
|
||||||
|
discover.deep_channel_avatar("https://www.youtube.com/@x/videos")
|
||||||
|
opts = capture.captured[0]
|
||||||
|
assert opts.get("extract_flat"), "must not fully extract every video"
|
||||||
|
assert opts.get("playlistend") == 1, "must not paginate beyond the first page"
|
||||||
|
|
||||||
|
|
||||||
|
def test_discover_channel_limit_becomes_playlistend(capture):
|
||||||
|
discover.discover_channel("https://www.youtube.com/@x/videos", limit=30)
|
||||||
|
assert capture.captured[0].get("playlistend") == 30
|
||||||
|
|
||||||
|
|
||||||
|
def test_discover_channel_without_limit_has_no_ceiling(capture):
|
||||||
|
"""A brand-new channel legitimately walks everything; that must stay possible."""
|
||||||
|
discover.discover_channel("https://www.youtube.com/@x/videos")
|
||||||
|
assert "playlistend" not in capture.captured[0]
|
||||||
|
|
||||||
|
|
||||||
|
def test_extraction_never_probes_formats(capture):
|
||||||
|
"""`check_formats` costs one HTTP request per format and we only want captions."""
|
||||||
|
capture.info = {"id": "vid", "title": "t", "subtitles": {}, "automatic_captions": {}}
|
||||||
|
extract.extract_video("https://www.youtube.com/watch?v=vid", {"es": "any"})
|
||||||
|
assert capture.captured[0].get("check_formats") is None
|
||||||
|
|
||||||
|
|
||||||
|
def test_discovery_processes_the_result_so_playlistend_applies(capture, monkeypatch):
|
||||||
|
"""`playlistend` is silently ignored when extract_info runs with process=False.
|
||||||
|
|
||||||
|
Measured on a 2564-video channel: processed + playlistend=60 costs 2
|
||||||
|
requests; unprocessed, the lazy generator ignores the limit and walking it
|
||||||
|
costs 86. Nothing else in the codebase would catch that flip.
|
||||||
|
"""
|
||||||
|
seen = {}
|
||||||
|
|
||||||
|
def extract_info(self, url, download=False, process=True):
|
||||||
|
seen["process"] = process
|
||||||
|
return dict(CapturingYDL.info)
|
||||||
|
|
||||||
|
monkeypatch.setattr(CapturingYDL, "extract_info", extract_info)
|
||||||
|
discover.discover_channel("https://www.youtube.com/@x/videos", limit=30)
|
||||||
|
assert seen["process"] is not False
|
||||||
@@ -0,0 +1,33 @@
|
|||||||
|
"""`--since` / the scrape form's date filter must not eat undated entries.
|
||||||
|
|
||||||
|
yt-dlp's flat channel listing does not report upload_date, so comparing a
|
||||||
|
missing date against the cutoff used to discard everything discovery found.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from yt_scraper.cli import _apply_filters
|
||||||
|
from yt_scraper.store import VideoRef
|
||||||
|
|
||||||
|
|
||||||
|
def _ref(vid: str, upload_date: str | None) -> VideoRef:
|
||||||
|
return VideoRef(vid, "UC1", vid, f"https://www.youtube.com/watch?v={vid}", upload_date)
|
||||||
|
|
||||||
|
|
||||||
|
def test_since_keeps_entries_with_no_upload_date():
|
||||||
|
refs = [_ref("undated", None), _ref("new", "20260715"), _ref("old", "20200101")]
|
||||||
|
|
||||||
|
kept = _apply_filters(refs, "2026-07-01", False, False, 0, None)
|
||||||
|
|
||||||
|
assert [r.video_id for r in kept] == ["undated", "new"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_since_still_drops_entries_that_are_provably_older():
|
||||||
|
refs = [_ref("old", "20200101"), _ref("older", "20190101")]
|
||||||
|
|
||||||
|
assert _apply_filters(refs, "2026-07-01", False, False, 0, None) == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_without_since_nothing_is_dropped_on_date_grounds():
|
||||||
|
refs = [_ref("undated", None), _ref("old", "20200101")]
|
||||||
|
|
||||||
|
assert len(_apply_filters(refs, None, False, False, 0, None)) == 2
|
||||||
@@ -100,6 +100,37 @@ def test_query_videos_filters(seeded_store):
|
|||||||
assert rows3[0].video_id == "v2"
|
assert rows3[0].video_id == "v2"
|
||||||
|
|
||||||
|
|
||||||
|
def test_upsert_videos_reports_only_new_rows(seeded_store):
|
||||||
|
inserted = seeded_store.upsert_videos([
|
||||||
|
VideoRef("v1", "UC1", "Updated title", "https://y/watch?v=v1", "20240101", 120),
|
||||||
|
VideoRef("v4", "UC1", "New video", "https://y/watch?v=v4", "20240501", 180),
|
||||||
|
])
|
||||||
|
|
||||||
|
assert inserted == 1
|
||||||
|
assert seeded_store.get_video("v1").title == "Updated title"
|
||||||
|
assert seeded_store.get_video("v4").status == "pending"
|
||||||
|
|
||||||
|
|
||||||
|
def test_upload_sort_uses_discovered_time_when_date_is_missing(seeded_store):
|
||||||
|
seeded_store.upsert_videos([
|
||||||
|
VideoRef("recent", "UC1", "Recent discovered", "https://y/watch?v=recent"),
|
||||||
|
])
|
||||||
|
|
||||||
|
rows, _ = seeded_store.query_videos(channel_id="UC1", sort="upload_date", page=1, size=1)
|
||||||
|
|
||||||
|
assert rows[0].video_id == "recent"
|
||||||
|
|
||||||
|
|
||||||
|
def test_mark_status_clears_stale_error_message(seeded_store):
|
||||||
|
seeded_store.mark_error("v1", "temporary extraction failure")
|
||||||
|
|
||||||
|
seeded_store.mark_status("v1", "no_subtitles")
|
||||||
|
|
||||||
|
video = seeded_store.get_video("v1")
|
||||||
|
assert video.status == "no_subtitles"
|
||||||
|
assert video.error_msg is None
|
||||||
|
|
||||||
|
|
||||||
def test_cookie_vault(store):
|
def test_cookie_vault(store):
|
||||||
from yt_scraper import cookies
|
from yt_scraper import cookies
|
||||||
sample = (
|
sample = (
|
||||||
|
|||||||
@@ -0,0 +1,80 @@
|
|||||||
|
"""Store-side support for incremental sync: the known-id boundary and the
|
||||||
|
watermark that survives a re-discovery."""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from yt_scraper.store import Store, VideoRef
|
||||||
|
|
||||||
|
|
||||||
|
def _store(tmp_path) -> Store:
|
||||||
|
store = Store(tmp_path / "state.db")
|
||||||
|
store.upsert_channel("UC1", "@alpha", "Alpha", 0)
|
||||||
|
return store
|
||||||
|
|
||||||
|
|
||||||
|
def test_known_video_ids_is_scoped_to_the_channel(tmp_path):
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_channel("UC2", "@beta", "Beta", 0)
|
||||||
|
store.upsert_videos([
|
||||||
|
VideoRef("a", "UC1", "A", "https://y/watch?v=a"),
|
||||||
|
VideoRef("b", "UC1", "B", "https://y/watch?v=b"),
|
||||||
|
VideoRef("c", "UC2", "C", "https://y/watch?v=c"),
|
||||||
|
])
|
||||||
|
|
||||||
|
assert store.known_video_ids("UC1") == {"a", "b"}
|
||||||
|
assert store.known_video_ids("UC2") == {"c"}
|
||||||
|
assert store.known_video_ids("nope") == set()
|
||||||
|
|
||||||
|
|
||||||
|
def test_rediscovery_does_not_wipe_a_known_upload_date(tmp_path):
|
||||||
|
"""Flat discovery reports upload_date=None; without COALESCE a routine sync
|
||||||
|
would erase the dates learned during extraction — and with them the very
|
||||||
|
watermark this feature is built on."""
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([VideoRef("a", "UC1", "A", "https://y/watch?v=a", "20260715", 120)])
|
||||||
|
|
||||||
|
store.upsert_videos([VideoRef("a", "UC1", "A (renamed)", "https://y/watch?v=a", None, None)])
|
||||||
|
|
||||||
|
row = store.get_video("a")
|
||||||
|
assert row.upload_date == "20260715"
|
||||||
|
assert row.duration == 120
|
||||||
|
assert row.title == "A (renamed)", "titles should still refresh"
|
||||||
|
|
||||||
|
|
||||||
|
def test_latest_upload_date_ignores_undated_rows(tmp_path):
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([
|
||||||
|
VideoRef("a", "UC1", "A", "https://y/watch?v=a", "20260101"),
|
||||||
|
VideoRef("b", "UC1", "B", "https://y/watch?v=b", None),
|
||||||
|
VideoRef("c", "UC1", "C", "https://y/watch?v=c", "20260720"),
|
||||||
|
])
|
||||||
|
|
||||||
|
assert store.latest_upload_date("UC1") == "20260720"
|
||||||
|
|
||||||
|
|
||||||
|
def test_mark_channel_synced_recounts_instead_of_trusting_the_window(tmp_path):
|
||||||
|
"""An incremental pass only sees the newest slice, so video_count must come
|
||||||
|
from the DB — otherwise an 848-video channel shrinks to the window size."""
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([
|
||||||
|
VideoRef(f"v{i}", "UC1", f"V{i}", f"https://y/watch?v=v{i}", "20260101")
|
||||||
|
for i in range(40)
|
||||||
|
])
|
||||||
|
|
||||||
|
marks = store.mark_channel_synced("UC1")
|
||||||
|
|
||||||
|
assert marks["video_count"] == 40
|
||||||
|
channel = store.get_channel("UC1")
|
||||||
|
assert channel["video_count"] == 40
|
||||||
|
assert channel["last_video_date"] == "20260101"
|
||||||
|
assert channel["last_synced_at"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_mark_channel_synced_keeps_the_last_date_when_nothing_is_dated(tmp_path):
|
||||||
|
store = _store(tmp_path)
|
||||||
|
store.upsert_videos([VideoRef("a", "UC1", "A", "https://y/watch?v=a", "20260101")])
|
||||||
|
store.mark_channel_synced("UC1")
|
||||||
|
|
||||||
|
store.upsert_videos([VideoRef("b", "UC1", "B", "https://y/watch?v=b", None)])
|
||||||
|
store.mark_channel_synced("UC1")
|
||||||
|
|
||||||
|
assert store.get_channel("UC1")["last_video_date"] == "20260101"
|
||||||
@@ -0,0 +1,219 @@
|
|||||||
|
"""The transcript must be the language actually spoken, not a machine translation.
|
||||||
|
|
||||||
|
Found in production on the Alex Hormozi channel (English). Under
|
||||||
|
`languages: {es: any, es-419: any, en: any}` the picker walked the config in
|
||||||
|
order, matched `es` first, and stored a Spanish auto-translation of English
|
||||||
|
speech — "Soy Nim Jenkinson y enseño a los aficionados a las manualidades" for
|
||||||
|
a video whose speaker says it in English.
|
||||||
|
|
||||||
|
Two things conspired:
|
||||||
|
|
||||||
|
1. `_normalize_lang("es-orig") == "es"`, so the `-orig` suffix — the one piece
|
||||||
|
of evidence distinguishing YouTube's real ASR track from a translation *into
|
||||||
|
the same language* — was thrown away before comparison.
|
||||||
|
2. Config order was treated as absolute preference, so a translation into a
|
||||||
|
preferred language beat the original.
|
||||||
|
|
||||||
|
Spanish channels were unaffected only by luck: yt-dlp happens to list `es-orig`
|
||||||
|
before `es`, so dict order gave the right answer. Every `done` row on the four
|
||||||
|
Spanish channels recorded `transcript_lang=es-orig`; the three Hormozi rows
|
||||||
|
recorded `es`. That asymmetry is what these tests pin down.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from yt_scraper.extract import original_language, pick_subtitle
|
||||||
|
|
||||||
|
CFG = {"es": "any", "es-419": "any", "en": "any"}
|
||||||
|
|
||||||
|
|
||||||
|
def _track(url: str) -> list[dict]:
|
||||||
|
return [{"ext": "json3", "url": url}]
|
||||||
|
|
||||||
|
|
||||||
|
def _english_video() -> dict:
|
||||||
|
"""An English video as YouTube actually presents it: the original ASR under
|
||||||
|
`en-orig`, plus a long tail of translations keyed by bare language code."""
|
||||||
|
return {
|
||||||
|
"id": "aRVv5NLVRwE",
|
||||||
|
"title": "My honest advice to someone who wants to get rich.",
|
||||||
|
"subtitles": {},
|
||||||
|
"automatic_captions": {
|
||||||
|
"en-orig": _track("https://timedtext/en-orig"),
|
||||||
|
"en": _track("https://timedtext/en-translated"),
|
||||||
|
"es": _track("https://timedtext/es"),
|
||||||
|
"es-419": _track("https://timedtext/es-419"),
|
||||||
|
"fr": _track("https://timedtext/fr"),
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _spanish_video() -> dict:
|
||||||
|
return {
|
||||||
|
"id": "uZH3FKH_yNw",
|
||||||
|
"title": "Escribir codigo a mano sera irresponsable",
|
||||||
|
"subtitles": {},
|
||||||
|
"automatic_captions": {
|
||||||
|
"es-orig": _track("https://timedtext/es-orig"),
|
||||||
|
"es": _track("https://timedtext/es"),
|
||||||
|
"en": _track("https://timedtext/en"),
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def test_english_video_yields_english_not_a_spanish_translation():
|
||||||
|
"""The production bug, verbatim."""
|
||||||
|
pick = pick_subtitle(_english_video(), CFG, prefer_manual=True)
|
||||||
|
assert pick is not None
|
||||||
|
assert pick.lang == "en-orig", f"picked {pick.lang}: a translation, not the spoken language"
|
||||||
|
|
||||||
|
|
||||||
|
def test_spanish_video_still_yields_the_spanish_original():
|
||||||
|
"""The fix must not regress the four Spanish channels already in the DB."""
|
||||||
|
pick = pick_subtitle(_spanish_video(), CFG, prefer_manual=True)
|
||||||
|
assert pick is not None
|
||||||
|
assert pick.lang == "es-orig"
|
||||||
|
|
||||||
|
|
||||||
|
def test_original_beats_a_same_language_translation():
|
||||||
|
"""YouTube publishes a translation *into the video's own language* too.
|
||||||
|
|
||||||
|
`en-orig` and `en` both normalise to "en"; only the suffix says which one is
|
||||||
|
the real transcript, and dict order must not decide it.
|
||||||
|
"""
|
||||||
|
info = {
|
||||||
|
"subtitles": {},
|
||||||
|
# Deliberately listed translation-first to defeat insertion order.
|
||||||
|
"automatic_captions": {
|
||||||
|
"en": _track("https://timedtext/en-translated"),
|
||||||
|
"en-orig": _track("https://timedtext/en-orig"),
|
||||||
|
},
|
||||||
|
}
|
||||||
|
pick = pick_subtitle(info, {"en": "any"}, prefer_manual=True)
|
||||||
|
assert pick.lang == "en-orig"
|
||||||
|
assert pick.url.endswith("en-orig")
|
||||||
|
|
||||||
|
|
||||||
|
def test_translation_is_a_documented_last_resort_not_a_silent_default():
|
||||||
|
"""A German channel under a Spanish/English policy.
|
||||||
|
|
||||||
|
The spoken language is not one the operator asked for, so there is no
|
||||||
|
original to give them and a translation is the only thing on offer. We do
|
||||||
|
take it — but only on the second pass, after every untranslated option has
|
||||||
|
been rejected, and `transcript_lang` records the bare code so the row is
|
||||||
|
distinguishable from an `-orig` one afterwards.
|
||||||
|
|
||||||
|
This case is a deliberate fallback. It is NOT the behaviour that caused the
|
||||||
|
Hormozi bug: there, `en` *was* configured and was being skipped.
|
||||||
|
"""
|
||||||
|
info = {
|
||||||
|
"subtitles": {},
|
||||||
|
"automatic_captions": {
|
||||||
|
"de-orig": _track("https://timedtext/de-orig"),
|
||||||
|
"es": _track("https://timedtext/de?tlang=es"),
|
||||||
|
"en": _track("https://timedtext/de?tlang=en"),
|
||||||
|
},
|
||||||
|
}
|
||||||
|
pick = pick_subtitle(info, CFG, prefer_manual=True)
|
||||||
|
assert pick.lang == "es"
|
||||||
|
assert "tlang=" in pick.url, "the fallback really is a translation; nothing else was available"
|
||||||
|
|
||||||
|
|
||||||
|
# ------------------------------------------------- the tlang= guard
|
||||||
|
#
|
||||||
|
# A translated caption URL is the base track's URL with `tlang=` appended, and
|
||||||
|
# yt-dlp omits it when target == source. That is direct evidence, unlike the
|
||||||
|
# `-orig` naming convention, so it catches videos whose spoken language cannot
|
||||||
|
# be determined any other way.
|
||||||
|
|
||||||
|
|
||||||
|
def test_untranslated_track_wins_even_with_no_orig_key_and_no_language_field():
|
||||||
|
"""The gap the first version of this fix left open.
|
||||||
|
|
||||||
|
Without an `-orig` key and without `info["language"]`, the spoken language
|
||||||
|
is unknown, the reorder cannot fire, and config order used to hand back the
|
||||||
|
Spanish translation. Rejecting `tlang=` needs no such knowledge.
|
||||||
|
"""
|
||||||
|
info = {
|
||||||
|
"subtitles": {},
|
||||||
|
"automatic_captions": {
|
||||||
|
"es": _track("https://timedtext/base?lang=en&kind=asr&tlang=es"),
|
||||||
|
"en": _track("https://timedtext/base?lang=en&kind=asr"),
|
||||||
|
},
|
||||||
|
}
|
||||||
|
assert original_language(info) is None, "precondition: spoken language is undeterminable"
|
||||||
|
pick = pick_subtitle(info, CFG, prefer_manual=True)
|
||||||
|
assert pick.lang == "en"
|
||||||
|
assert "tlang=" not in pick.url
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_real_hormozi_url_shape_is_recognised_as_a_translation():
|
||||||
|
"""Verbatim parameters from the production timedtext URLs in .run/server.log."""
|
||||||
|
info = {
|
||||||
|
"subtitles": {},
|
||||||
|
"automatic_captions": {
|
||||||
|
"es": _track(
|
||||||
|
"https://www.youtube.com/api/timedtext?v=aRVv5NLVRwE&caps=asr&opi=112496729"
|
||||||
|
"&lang=en&kind=asr&variant=gemini&fmt=json3&tlang=es"
|
||||||
|
),
|
||||||
|
"en-orig": _track(
|
||||||
|
"https://www.youtube.com/api/timedtext?v=aRVv5NLVRwE&caps=asr&opi=112496729"
|
||||||
|
"&lang=en&kind=asr&variant=gemini&fmt=json3"
|
||||||
|
),
|
||||||
|
},
|
||||||
|
}
|
||||||
|
pick = pick_subtitle(info, CFG, prefer_manual=True)
|
||||||
|
assert pick.lang == "en-orig"
|
||||||
|
assert "tlang=" not in pick.url
|
||||||
|
|
||||||
|
|
||||||
|
def test_manual_captions_still_outrank_auto_for_the_same_language():
|
||||||
|
"""Preferring the original must not override the manual/auto policy."""
|
||||||
|
info = {
|
||||||
|
"subtitles": {"en": _track("https://timedtext/en-manual")},
|
||||||
|
"automatic_captions": {"en-orig": _track("https://timedtext/en-orig")},
|
||||||
|
}
|
||||||
|
pick = pick_subtitle(info, {"en": "any"}, prefer_manual=True)
|
||||||
|
assert pick.source == "manual"
|
||||||
|
assert pick.url.endswith("en-manual")
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_manual_translation_does_not_beat_the_spoken_language():
|
||||||
|
"""Config lists es first, but the video is English with English manual subs."""
|
||||||
|
info = {
|
||||||
|
"subtitles": {"en": _track("https://timedtext/en-manual")},
|
||||||
|
"automatic_captions": {
|
||||||
|
"en-orig": _track("https://timedtext/en-orig"),
|
||||||
|
"es": _track("https://timedtext/es"),
|
||||||
|
},
|
||||||
|
}
|
||||||
|
pick = pick_subtitle(info, CFG, prefer_manual=True)
|
||||||
|
assert pick.lang == "en"
|
||||||
|
assert pick.source == "manual"
|
||||||
|
|
||||||
|
|
||||||
|
# ------------------------------------------------------------- detection
|
||||||
|
|
||||||
|
|
||||||
|
def test_original_language_read_from_the_orig_suffix():
|
||||||
|
assert original_language(_english_video()) == "en"
|
||||||
|
assert original_language(_spanish_video()) == "es"
|
||||||
|
|
||||||
|
|
||||||
|
def test_original_language_falls_back_to_the_info_key():
|
||||||
|
assert original_language({"language": "pt-BR", "automatic_captions": {}}) == "pt"
|
||||||
|
|
||||||
|
|
||||||
|
def test_orig_suffix_wins_over_the_info_key():
|
||||||
|
"""`language` is metadata YouTube localises; the caption list is evidence.
|
||||||
|
|
||||||
|
Production showed YouTube returning English titles for Spanish videos, so
|
||||||
|
localised metadata is not trustworthy for this decision.
|
||||||
|
"""
|
||||||
|
info = {"language": "es", "automatic_captions": {"en-orig": _track("u")}}
|
||||||
|
assert original_language(info) == "en"
|
||||||
|
|
||||||
|
|
||||||
|
def test_original_language_is_none_when_unknowable():
|
||||||
|
assert original_language({"automatic_captions": {"es": _track("u")}}) is None
|
||||||
|
assert original_language({}) is None
|
||||||
@@ -0,0 +1,246 @@
|
|||||||
|
"""The circuit breaker, tested through the webapp job runner.
|
||||||
|
|
||||||
|
This is the regression test for the incident the whole change exists to
|
||||||
|
prevent: a session gets rate-limited, the runner keeps going anyway, and every
|
||||||
|
remaining video is marked failed. In production that turned one throttling
|
||||||
|
event into 511 `no_subtitles` and 347 `error` rows on a single channel.
|
||||||
|
|
||||||
|
The invariant asserted everywhere below is the same: **videos the run never
|
||||||
|
reached must still be `pending`**, because `pending` is what a later run picks
|
||||||
|
up. A video wrongly marked `error` needs a manual reset first.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from yt_scraper.config import Config, DelayConfig
|
||||||
|
from yt_scraper.store import Store, VideoRef
|
||||||
|
from yt_scraper.webapp.jobs import JobManager
|
||||||
|
|
||||||
|
THROTTLED = (
|
||||||
|
"ERROR: [youtube] {vid}: Video unavailable. This content isn't available, try "
|
||||||
|
"again later. The current session has been rate-limited by YouTube for up to an hour."
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _cfg(tmp_path) -> Config:
|
||||||
|
return Config(
|
||||||
|
database_path=str(tmp_path / "state.db"),
|
||||||
|
output_dir=str(tmp_path / "markdown"),
|
||||||
|
# No real waiting in tests; the breaker's arithmetic is unit-tested
|
||||||
|
# separately in test_ratelimit.py.
|
||||||
|
delay=DelayConfig(
|
||||||
|
min_seconds=0.0, max_seconds=0.0,
|
||||||
|
backoff_base=0.001, backoff_cap=0.002, throttle_threshold=3,
|
||||||
|
),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _seed(store: Store, n: int) -> list[str]:
|
||||||
|
store.upsert_channel("UC1", "@alpha", "Alpha", n)
|
||||||
|
ids = [f"v{i:02d}" for i in range(n)]
|
||||||
|
store.upsert_videos([
|
||||||
|
VideoRef(v, "UC1", f"Video {i}", f"https://y/watch?v={v}") for i, v in enumerate(ids)
|
||||||
|
])
|
||||||
|
return ids
|
||||||
|
|
||||||
|
|
||||||
|
def _always_throttled(store: Store):
|
||||||
|
"""Stand-in for process_video that fails the way a throttled session does."""
|
||||||
|
|
||||||
|
def fake(row, cfg, st, *args, **kwargs):
|
||||||
|
st.mark_error(row.video_id, THROTTLED.format(vid=row.video_id))
|
||||||
|
return "error"
|
||||||
|
|
||||||
|
return fake
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture
|
||||||
|
def manager(tmp_path):
|
||||||
|
store = Store(tmp_path / "state.db")
|
||||||
|
return JobManager(store, _cfg(tmp_path)), store
|
||||||
|
|
||||||
|
|
||||||
|
def test_batch_stops_and_leaves_the_rest_pending(manager, monkeypatch):
|
||||||
|
mgr, store = manager
|
||||||
|
ids = _seed(store, 20)
|
||||||
|
monkeypatch.setattr("yt_scraper.pipeline.process_video", _always_throttled(store))
|
||||||
|
monkeypatch.setattr("yt_scraper.pipeline.cache_thumbnail", lambda *a, **k: False)
|
||||||
|
store.create_job("job", None, {"video_ids": ids})
|
||||||
|
|
||||||
|
mgr._run_batch("job", {"video_ids": ids})
|
||||||
|
|
||||||
|
touched = [v for v in ids if store.get_video(v).status != "pending"]
|
||||||
|
assert len(touched) == 3, "must stop at the threshold, not walk all 20"
|
||||||
|
assert all(store.get_video(v).status == "pending" for v in ids[3:])
|
||||||
|
|
||||||
|
job = store.get_job("job")
|
||||||
|
assert job.status == "error"
|
||||||
|
assert "rate-limit" in (job.last_error or "")
|
||||||
|
|
||||||
|
events = mgr.events_since("job", 0)
|
||||||
|
err = [e for e in events if e["event"] == "error"]
|
||||||
|
assert err and err[-1]["data"]["throttled"] is True
|
||||||
|
|
||||||
|
|
||||||
|
def test_channel_run_stops_and_leaves_the_rest_pending(manager, monkeypatch):
|
||||||
|
mgr, store = manager
|
||||||
|
ids = _seed(store, 20)
|
||||||
|
monkeypatch.setattr("yt_scraper.webapp.jobs.process_video", _always_throttled(store))
|
||||||
|
monkeypatch.setattr(
|
||||||
|
"yt_scraper.discover.discover_channel",
|
||||||
|
lambda url, sleep_subrequests=2.0, limit=None: ("UC1", "Alpha", None, []),
|
||||||
|
)
|
||||||
|
store.create_job("job", "UC1", {})
|
||||||
|
|
||||||
|
mgr._run_channel("job", {})
|
||||||
|
|
||||||
|
assert len([v for v in ids if store.get_video(v).status != "pending"]) == 3
|
||||||
|
assert store.get_job("job").status == "error"
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_healthy_run_is_untouched_by_the_breaker(manager, monkeypatch):
|
||||||
|
"""The breaker must be invisible when nothing is throttling."""
|
||||||
|
mgr, store = manager
|
||||||
|
ids = _seed(store, 8)
|
||||||
|
monkeypatch.setattr("yt_scraper.pipeline.process_video",
|
||||||
|
lambda row, *a, **k: store.mark_status(row.video_id, "done") or "done")
|
||||||
|
monkeypatch.setattr("yt_scraper.pipeline.cache_thumbnail", lambda *a, **k: False)
|
||||||
|
store.create_job("job", None, {"video_ids": ids})
|
||||||
|
|
||||||
|
mgr._run_batch("job", {"video_ids": ids})
|
||||||
|
|
||||||
|
assert store.get_job("job").status == "done"
|
||||||
|
assert all(store.get_video(v).status == "done" for v in ids)
|
||||||
|
|
||||||
|
|
||||||
|
def test_ordinary_failures_do_not_stop_the_run(manager, monkeypatch):
|
||||||
|
"""A channel with dead videos must still be processed to the end.
|
||||||
|
|
||||||
|
Without the throttle/permanent distinction, three members-only videos in a
|
||||||
|
row would abort a perfectly healthy scrape.
|
||||||
|
"""
|
||||||
|
mgr, store = manager
|
||||||
|
ids = _seed(store, 10)
|
||||||
|
|
||||||
|
def fake(row, cfg, st, *args, **kwargs):
|
||||||
|
st.mark_error(row.video_id, "ERROR: [youtube] x: Private video. Sign in if you've been granted access")
|
||||||
|
return "error"
|
||||||
|
|
||||||
|
monkeypatch.setattr("yt_scraper.pipeline.process_video", fake)
|
||||||
|
monkeypatch.setattr("yt_scraper.pipeline.cache_thumbnail", lambda *a, **k: False)
|
||||||
|
store.create_job("job", None, {"video_ids": ids})
|
||||||
|
|
||||||
|
mgr._run_batch("job", {"video_ids": ids})
|
||||||
|
|
||||||
|
assert store.get_job("job").status == "done"
|
||||||
|
assert all(store.get_video(v).status == "error" for v in ids)
|
||||||
|
|
||||||
|
|
||||||
|
def test_isolated_throttling_between_successes_does_not_stop_the_run(manager, monkeypatch):
|
||||||
|
"""Only *consecutive* throttling means the session is banned."""
|
||||||
|
mgr, store = manager
|
||||||
|
ids = _seed(store, 12)
|
||||||
|
calls = {"n": 0}
|
||||||
|
|
||||||
|
def fake(row, cfg, st, *args, **kwargs):
|
||||||
|
calls["n"] += 1
|
||||||
|
if calls["n"] % 2:
|
||||||
|
st.mark_error(row.video_id, THROTTLED.format(vid=row.video_id))
|
||||||
|
return "error"
|
||||||
|
st.mark_status(row.video_id, "done")
|
||||||
|
return "done"
|
||||||
|
|
||||||
|
monkeypatch.setattr("yt_scraper.pipeline.process_video", fake)
|
||||||
|
monkeypatch.setattr("yt_scraper.pipeline.cache_thumbnail", lambda *a, **k: False)
|
||||||
|
store.create_job("job", None, {"video_ids": ids})
|
||||||
|
|
||||||
|
mgr._run_batch("job", {"video_ids": ids})
|
||||||
|
|
||||||
|
assert store.get_job("job").status == "done"
|
||||||
|
assert calls["n"] == 12
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_throttled_caption_fetch_is_an_error_not_no_subtitles(tmp_path, monkeypatch):
|
||||||
|
"""Which of the three requests YouTube refused must not decide the status.
|
||||||
|
|
||||||
|
A 429 on `extract_info` produced `error`; a 429 on the separate caption
|
||||||
|
download produced `no_subtitles` — the state that means "this video
|
||||||
|
publishes no captions", which is what the `no_subtitles` counts are read as.
|
||||||
|
Both are retryable, but only one is honest.
|
||||||
|
"""
|
||||||
|
from yt_scraper.extract import VideoData
|
||||||
|
from yt_scraper.pipeline import process_video
|
||||||
|
|
||||||
|
store = Store(tmp_path / "state.db")
|
||||||
|
cfg = _cfg(tmp_path)
|
||||||
|
_seed(store, 1)
|
||||||
|
row = store.get_video("v00")
|
||||||
|
|
||||||
|
throttled = VideoData(
|
||||||
|
info={"id": "v00", "title": "t"},
|
||||||
|
segments=[],
|
||||||
|
subtitle=None,
|
||||||
|
has_chapters=False,
|
||||||
|
skip_reason=(
|
||||||
|
"subtitle track found (lang=en, auto) but the download failed "
|
||||||
|
"[HTTP Error 429: HTTPError: 429 Client Error: Too Many Requests for url: ...]"
|
||||||
|
),
|
||||||
|
)
|
||||||
|
monkeypatch.setattr("yt_scraper.pipeline.extract_video", lambda *a, **k: throttled)
|
||||||
|
assert process_video(row, cfg, store, "Alpha", "UC1", "u") == "error"
|
||||||
|
|
||||||
|
genuinely_absent = VideoData(
|
||||||
|
info={"id": "v00", "title": "t"}, segments=[], subtitle=None,
|
||||||
|
has_chapters=False, skip_reason="no caption tracks published for this video",
|
||||||
|
)
|
||||||
|
monkeypatch.setattr("yt_scraper.pipeline.extract_video", lambda *a, **k: genuinely_absent)
|
||||||
|
assert process_video(row, cfg, store, "Alpha", "UC1", "u") == "no_subtitles"
|
||||||
|
|
||||||
|
|
||||||
|
def test_metadata_survives_a_failed_caption_fetch(tmp_path, monkeypatch):
|
||||||
|
"""The extraction already paid for this data; a retry must not re-buy it."""
|
||||||
|
from yt_scraper.extract import VideoData
|
||||||
|
from yt_scraper.pipeline import process_video
|
||||||
|
|
||||||
|
store = Store(tmp_path / "state.db")
|
||||||
|
cfg = _cfg(tmp_path)
|
||||||
|
_seed(store, 1)
|
||||||
|
row = store.get_video("v00")
|
||||||
|
|
||||||
|
monkeypatch.setattr(
|
||||||
|
"yt_scraper.pipeline.extract_video",
|
||||||
|
lambda *a, **k: VideoData(
|
||||||
|
info={"id": "v00", "title": "t", "view_count": 4321, "upload_date": "20260101",
|
||||||
|
"thumbnail": "https://i.ytimg.com/x.jpg", "description": "hola"},
|
||||||
|
segments=[], subtitle=None, has_chapters=False,
|
||||||
|
skip_reason="no caption tracks published for this video",
|
||||||
|
),
|
||||||
|
)
|
||||||
|
process_video(row, cfg, store, "Alpha", "UC1", "u")
|
||||||
|
|
||||||
|
after = store.get_video("v00")
|
||||||
|
assert after.status == "no_subtitles"
|
||||||
|
assert after.view_count == 4321, "metadata was discarded with the failed transcript"
|
||||||
|
assert after.upload_date == "20260101"
|
||||||
|
|
||||||
|
|
||||||
|
def test_throttled_videos_stay_retryable(manager, monkeypatch):
|
||||||
|
"""The three videos that did fail must not be classified as permanent.
|
||||||
|
|
||||||
|
`Store.PERMANENT_ERROR_PATTERNS` deliberately excludes the throttling
|
||||||
|
message; if that ever changed, a rate-limit incident would poison rows that
|
||||||
|
a later run could have recovered.
|
||||||
|
"""
|
||||||
|
mgr, store = manager
|
||||||
|
ids = _seed(store, 10)
|
||||||
|
monkeypatch.setattr("yt_scraper.pipeline.process_video", _always_throttled(store))
|
||||||
|
monkeypatch.setattr("yt_scraper.pipeline.cache_thumbnail", lambda *a, **k: False)
|
||||||
|
store.create_job("job", None, {"video_ids": ids})
|
||||||
|
|
||||||
|
mgr._run_batch("job", {"video_ids": ids})
|
||||||
|
|
||||||
|
counts = store.retryable_counts("UC1")
|
||||||
|
assert counts.get("permanent", 0) == 0
|
||||||
|
assert store.reset_videos("UC1", ("error",)) == 3
|
||||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user