Compare commits
17
Commits
06497299ab
...
c693b20e19
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
c693b20e19 | ||
|
|
450aee216d | ||
|
|
3c49f73fa2 | ||
|
|
b3b27ce883 | ||
|
|
1190228a81 | ||
|
|
8a59b39c98 | ||
|
|
4f5a68b572 | ||
|
|
ca2ae23f70 | ||
|
|
e219cfcbfa | ||
|
|
8940c325ea | ||
|
|
b04a2710cf | ||
|
|
cd9f16f329 | ||
|
|
fae9496201 | ||
|
|
f038790caa | ||
|
|
a5fd39a3e6 | ||
|
|
a71401666c | ||
|
|
8fea202c21 |
@@ -0,0 +1,33 @@
|
||||
# Normalizar todo el texto a LF en el repo (checkout nativo por OS controlado abajo)
|
||||
* text=auto eol=lf
|
||||
|
||||
# Extensiones con LF forzado
|
||||
*.py text eol=lf
|
||||
*.js text eol=lf
|
||||
*.css text eol=lf
|
||||
*.html text eol=lf
|
||||
*.md text eol=lf
|
||||
*.yaml text eol=lf
|
||||
*.yml text eol=lf
|
||||
*.toml text eol=lf
|
||||
*.json text eol=lf
|
||||
*.sh text eol=lf
|
||||
*.j2 text eol=lf
|
||||
*.bat text eol=crlf
|
||||
*.ps1 text eol=crlf
|
||||
.gitattributes text eol=lf
|
||||
.gitignore text eol=lf
|
||||
Makefile text eol=lf
|
||||
|
||||
# Binarios
|
||||
*.png binary
|
||||
*.jpg binary
|
||||
*.jpeg binary
|
||||
*.gif binary
|
||||
*.ico binary
|
||||
*.db binary
|
||||
*.mp3 binary
|
||||
*.mp4 binary
|
||||
*.webm binary
|
||||
*.woff binary
|
||||
*.woff2 binary
|
||||
+35
-4
@@ -1,3 +1,4 @@
|
||||
# --- Python ---
|
||||
__pycache__/
|
||||
*.py[cod]
|
||||
*.egg-info/
|
||||
@@ -7,14 +8,44 @@ dist/
|
||||
.pytest_cache/
|
||||
.coverage
|
||||
htmlcov/
|
||||
data/
|
||||
cookies/
|
||||
.run/
|
||||
|
||||
# --- Entornos virtuales ---
|
||||
.venv/
|
||||
.venv-*/
|
||||
venv/
|
||||
|
||||
# --- Datos del usuario: NUNCA versionar (canal(es), contenido scrapeado, BD, media) ---
|
||||
data/*
|
||||
!data/.gitkeep
|
||||
|
||||
# --- Cookies de sesión de YouTube: sensibles, nunca versionar ---
|
||||
cookies/
|
||||
*.cookies.txt
|
||||
cookies.txt
|
||||
|
||||
# --- Config real del usuario (usar config.example.yaml como plantilla) ---
|
||||
config.yaml
|
||||
.env
|
||||
.env.*
|
||||
|
||||
# --- Estado de runtime del servidor ---
|
||||
.run/
|
||||
*.log
|
||||
|
||||
# opencode runtime state (keep agent/ + goal definitions, exclude loop/session state)
|
||||
# --- Salida regenerable de herramientas de análisis (graphify) ---
|
||||
graphify-out/
|
||||
|
||||
# --- Estado local de agentes IA ---
|
||||
.superpowers/
|
||||
.claude/
|
||||
.opencode/goals/state.json
|
||||
.opencode/goals/state.json
|
||||
.opencode/goals/state.json.ledger.jsonl
|
||||
.opencode/opencode-loop/
|
||||
.opencode/node_modules/
|
||||
|
||||
# --- macOS / editors ---
|
||||
.DS_Store
|
||||
Thumbs.db
|
||||
.idea/
|
||||
.vscode/
|
||||
|
||||
@@ -0,0 +1,61 @@
|
||||
# AGENTS.md
|
||||
|
||||
Guía de entrada para cualquier agente de código (Claude Code, Cursor, opencode, ZCode, ...) que vaya a trabajar en este repo. Léela antes de tocar nada.
|
||||
|
||||
## Qué es (en 5 líneas)
|
||||
|
||||
Plataforma 100% local de minería de contenido de creadores de YouTube. `yt-dlp` extrae metadatos + transcripciones de canales completos; todo se guarda en SQLite (con FTS5) y se renderiza a notas Markdown estilo Obsidian. Dos superficies: CLI `yt-scraper` (Click) y webapp FastAPI con SPA vanilla (Alpine). Sin auth, sin deploy remoto, sin features de IA (regla de alcance explícita en `docs/superpowers/specs/2026-07-26-platform-design.md` §14).
|
||||
|
||||
Arquitectura profunda, decisiones y trampas: [CLAUDE.md](CLAUDE.md) — es la referencia principal; no la dupliques aquí.
|
||||
|
||||
## Mapa de módulos (`src/yt_scraper/`)
|
||||
|
||||
| Módulo | Responsabilidad |
|
||||
|---|---|
|
||||
| `cli.py` | CLI Click (`yt-scraper`); fachada del pipeline. Sin subcomando = scrape completo |
|
||||
| `config.py` | Dataclasses `Config`/`DelayConfig`/`YtDlpConfig`/`SyncConfig` + `load_config` |
|
||||
| `discover.py` | yt-dlp `--flat-playlist` → `VideoRef[]`; `discover_incremental` (ventana + overlap) |
|
||||
| `extract.py` | `extract_info` → metadatos + subtítulos; `pick_subtitle` (json3 > srv1 > vtt) |
|
||||
| `parse.py` | JSON3/VTT → `Segment{start, end, text}` |
|
||||
| `chapters.py` | `align_chapters`: capítulos ↔ segmentos → secciones |
|
||||
| `pipeline.py` | `process_video()`: la unidad de trabajo por vídeo (extract→parse→store→render) |
|
||||
| `store.py` | Única capa de datos: SQLite a mano (WAL, FTS5, migraciones idempotentes) |
|
||||
| `segments.py` | Camino inverso `.md` → DB (`backfill_from_markdown`, `reconcile_markdown`) |
|
||||
| `render.py` | Jinja2 → Markdown con frontmatter YAML (`templates/video.md.j2`) |
|
||||
| `ratelimit.py` | Políticas de red: `Pacer` global, backoff exponencial, `ThrottleGuard` |
|
||||
| `cookies.py` | Vault de cookies Netscape (import, activación, caducidad) |
|
||||
| `monitor.py` | Watch loop (discovery periódico) |
|
||||
| `export.py` | Export json/csv/srt/html |
|
||||
| `analysis.py` | Top words, wordcloud, timeline |
|
||||
| `_yt_http.py` | `yt_get()`: toda petición HTTP directa a CDNs de YouTube (UA + Referer) |
|
||||
| `webapp/` | FastAPI: `app.py` (`create_app`), `api.py` (routers), `jobs.py` (`JobManager`), `static/` (SPA Alpine) |
|
||||
|
||||
## Comandos esenciales
|
||||
|
||||
```bash
|
||||
uv sync --all-extras # o: pip install -e ".[dev,web,analysis]"
|
||||
uv run python -m pytest -q # suite hermética (~240 tests, 24 archivos, sin red)
|
||||
python -m uvicorn yt_scraper.webapp.app:app --port 8000 # webapp manual
|
||||
yt-scraper --help # CLI; ej: yt-scraper scrape --dry-run --limit 10
|
||||
make serve # macOS/Linux; Windows: start-server.bat / stop-server.bat
|
||||
```
|
||||
|
||||
Todo se ejecuta **desde la raíz del repo**: las rutas de `config.yaml` se resuelven relativas al CWD.
|
||||
|
||||
## Reglas de oro
|
||||
|
||||
1. **No toques `data/` ni `cookies/`**: contienen la biblioteca real y sesiones de YouTube (equivalente a contraseñas). Están gitignored; nunca las commitees ni pegues su contenido.
|
||||
2. **Tests herméticos**: solo `tmp_path` + `monkeypatch`, cero red. `yt_dlp.YoutubeDL` se fakea (ver `tests/test_webapp_jobs.py`).
|
||||
3. **Handlers con I/O de red en la webapp: `def`, no `async def`** (FastAPI manda los `async def` al event loop y congelan el servidor; los `def` van al threadpool). Excepción: `stream_job`, que devuelve el SSE.
|
||||
4. **Mantén CLAUDE.md al día** cuando cambies decisiones de arquitectura o contratos (plantilla ↔ regex, endpoints, flags). Hay checklist al final de ese archivo.
|
||||
5. **LF, no CRLF**: `.gitattributes` fuerza LF para código y docs (`*.bat`/`*.ps1` salen en CRLF). No lo deshagas; no añades `^M`.
|
||||
6. Una feature se implementa en `pipeline`/`store` y se expone dos veces (subcomando CLI + endpoint webapp). No dupliques lógica en la capa web.
|
||||
7. `yt-dlp` es la única interfaz con YouTube; el resto de peticiones HTTP van por `_yt_http.yt_get()`.
|
||||
|
||||
## Punteros
|
||||
|
||||
- [CLAUDE.md](CLAUDE.md) — arquitectura profunda, invariantes, incidentes y porqués.
|
||||
- [docs/GETTING-STARTED.md](docs/GETTING-STARTED.md) — máquina nueva desde cero por OS.
|
||||
- [docs/CONFIG.md](docs/CONFIG.md) — referencia completa de `config.yaml`.
|
||||
- [docs/COOKIES.md](docs/COOKIES.md) — guía de cookies (export, import, rotación, seguridad).
|
||||
- [docs/superpowers/specs/](docs/superpowers/specs/) — specs de diseño históricos (fuente de verdad de decisiones; donde el código difiere, gana el código).
|
||||
@@ -0,0 +1,294 @@
|
||||
# CLAUDE.md
|
||||
|
||||
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
|
||||
|
||||
## Qué es esto
|
||||
|
||||
Plataforma **100% local** de minería de contenido de creadores de YouTube: `yt-dlp` → transcripciones + metadatos → SQLite (con FTS5) → notas Markdown estilo Obsidian, más una webapp FastAPI/Alpine que expone todo eso. Sin auth, sin deploy remoto, sin features de IA (regla de alcance explícita en `docs/superpowers/specs/2026-07-26-platform-design.md` §14).
|
||||
|
||||
## Comandos
|
||||
|
||||
```bash
|
||||
pip install -e ".[dev,web,analysis]" # Python >= 3.10; ffmpeg es requisito del sistema para audio
|
||||
|
||||
python -m pytest tests/ -q # suite completa (~240 tests en 24 archivos, sin red)
|
||||
uv run python -m pytest -q # equivalente si usas uv
|
||||
python -m pytest tests/test_store_platform.py -v # un archivo
|
||||
python -m pytest tests/test_store_platform.py::test_dashboard_aggregates -v # un test
|
||||
python -m pytest -m "not integration" # marker declarado en pyproject (aún sin uso)
|
||||
|
||||
yt-scraper --help # CLI (equivale a: python -m yt_scraper.cli)
|
||||
yt-scraper scrape --dry-run --limit 10 # discovery sin descargar
|
||||
yt-scraper search "vaporwave" # FTS5 sobre transcripciones
|
||||
yt-scraper re-render --backfill # reimportar .md a la DB y regenerar Markdown
|
||||
|
||||
start-server.bat # webapp: busca puerto libre 8000-8100, arranca uvicorn, abre browser
|
||||
stop-server.bat # mata el proceso registrado en .run/server.info y libera el puerto
|
||||
scripts/doctor.ps1 # (primera vez, lo llama start-server.bat) instala Python 3.12 + deps si faltan
|
||||
make setup / make serve / make stop / make test / make clean # macOS/Linux vía Makefile
|
||||
scripts/bootstrap.sh # macOS/Linux: prepara python/ffmpeg/node (brew) + venv + extras
|
||||
scripts/start-server.sh / scripts/stop-server.sh # equivalentes bash de los .ps1
|
||||
python -m uvicorn yt_scraper.webapp.app:app --port 8765 # arranque manual (debug)
|
||||
```
|
||||
|
||||
No hay linter ni formatter configurado. Los `.bat` de la raíz solo envuelven `scripts/*.ps1` (`start-server.bat` pasa primero por `scripts/doctor.ps1`); en macOS/Linux el equivalente son `scripts/bootstrap.sh` + `scripts/start-server.sh` / `stop-server.sh`, expuestos en el `Makefile` raíz.
|
||||
|
||||
**Ojo:** `yt-scraper` sin subcomando ejecuta un scrape completo del canal de `config.yaml` (`invoke_without_command=True` en [cli.py:40](src/yt_scraper/cli.py#L40)).
|
||||
|
||||
## Orientación antes de leer código
|
||||
|
||||
Existe un grafo de conocimiento en [graphify-out/](graphify-out/) (gitignored, regenerable). Recomendado —no obligatorio, no es un hook del repo— consultarlo antes de leer/grepear fuentes: `graphify query "<pregunta>"`, `graphify explain "<concepto>"`, `graphify path "<A>" "<B>"`. Si el grafo está desactualizado respecto a los archivos, `graphify update`. [graphify-out/GRAPH_REPORT.md](graphify-out/GRAPH_REPORT.md) resume comunidades y god nodes (`Store` es el más conectado con diferencia).
|
||||
|
||||
## Arquitectura
|
||||
|
||||
### Tres superficies, un solo pipeline
|
||||
|
||||
`cli.py` (Click+Rich), `webapp/` (FastAPI+SSE) y `monitor.py` (watch loop) son **fachadas**; las tres convergen en `pipeline.process_video(row, cfg, store, ...)`, la unidad de trabajo por vídeo: extract → parse → align chapters → persistir en DB → renderizar `.md` → `mark_done`. Devuelve `"done" | "no_subtitles" | "error"` y nunca lanza por fallo de un vídeo.
|
||||
|
||||
Consecuencia práctica: una feature nueva se implementa en `pipeline`/`store` y **se expone dos veces** (subcomando en `cli.py` + endpoint en `webapp/api.py`). No dupliques lógica en la capa web.
|
||||
|
||||
```
|
||||
discover.py yt-dlp --flat-playlist → VideoRef[] (+ avatar; deep_channel_avatar como fallback)
|
||||
extract.py yt-dlp extract_info → info dict + pick_subtitle (prioridad json3 > srv1 > srv3 > vtt > ttml)
|
||||
parse.py JSON3/VTT → Segment{start,end,text}, con merge_adjacent
|
||||
chapters.py align_chapters: capítulos ↔ segmentos → Section[]
|
||||
render.py Jinja2 → .md (filtros format_timestamp / quote_yaml / to_json)
|
||||
store.py SQLite: todo el SQL a mano, filas como dataclasses
|
||||
segments.py camino inverso: .md → DB (backfill)
|
||||
```
|
||||
|
||||
### `Store` es la única capa de datos
|
||||
|
||||
SQLite sin ORM, `sqlite3.Row`, WAL, `foreign_keys=ON`. Cada método abre y cierra su conexión vía el contextmanager `_cursor()` (commit al salir), así que es seguro desde el worker thread de jobs.
|
||||
|
||||
**Migraciones:** `_init_schema()` corre en **cada** instanciación de `Store`. Es `SCHEMA` base + `ALTER TABLE ADD COLUMN` idempotente guiado por los dicts `_VIDEO_COLUMNS` / `_CHANNEL_COLUMNS` + `_EXTRA_SCHEMA`. Para añadir una columna: agrégala al dict, no escribas script de migración.
|
||||
|
||||
**FTS5:** `transcript_fts` es una tabla **standalone** (columnas duplicadas como `UNINDEXED`), no external-content — pese a lo que dice el spec. Por eso `store_segments()` borra e inserta en `transcript_segments` **y** en `transcript_fts`; si escribes segmentos por otra vía, replica ambas. Las queries de usuario pasan por `_sanitize_fts()` (tokens entre comillas unidos con `AND`).
|
||||
|
||||
**Doble persistencia de transcripciones**, intencional: filas en `transcript_segments` (búsqueda, análisis, reader) *y* `segments_json`/`chapters_json` en `videos` (permite `re-render` sin volver a descargar). `pipeline.process_video()` escribe las dos.
|
||||
|
||||
### El Markdown es fuente de datos, no solo salida
|
||||
|
||||
`segments.backfill_from_markdown()` parsea `data/markdown/**/*.md` de vuelta a la DB (idempotente, salta vídeos que ya tienen segmentos). Se invoca al arrancar la webapp (el reconcile de arranque corre en un hilo de fondo, [webapp/app.py:47-53](src/yt_scraper/webapp/app.py#L47)) y con `re-render --backfill`.
|
||||
|
||||
Eso convierte el formato del `.md` en un **contrato bidireccional**: [templates/video.md.j2](templates/video.md.j2) escribe `**MM:SS** · texto` y `### Título (MM:SS)`; los regex `_SEG_LINE` / `_CHAPTER` / `_FRONTMATTER_KEY` de [segments.py:17](src/yt_scraper/segments.py#L17) los leen. Si tocas la plantilla, actualiza los regex en el mismo cambio o el backfill se rompe en silencio.
|
||||
|
||||
### Sincronización incremental (no recorras el canal entero)
|
||||
|
||||
Re-escanear un canal ya trackeado **no** pagina todo el canal. `discover.discover_incremental()` lee la pestaña `/videos` (cronológica inversa) con `playlistend` y corta en cuanto ve `sync.overlap` vídeos consecutivos que ya están en la DB. Si toda la ventana resulta nueva, la duplica (`window` → `max_window`) en vez de perderse subidas. Medido sobre los 4 canales reales: **21.3 s / 1661 entradas → 3.3 s / 120 entradas**, mismos vídeos nuevos detectados.
|
||||
|
||||
Detalle que condiciona el diseño: en modo flat, yt-dlp **no** devuelve `upload_date` ni `timestamp` para los entries de YouTube (verificado). Por eso el corte real se hace por **solapamiento de IDs**, no por fecha; `since=` existe como condición secundaria por si el extractor sí las trae. La fecha del último vídeo se guarda como marca de agua en `channels.last_video_date` (poblada desde `videos.upload_date`, que sí se conoce tras la extracción) y es lo que se muestra en la UI y en el CLI.
|
||||
|
||||
- El `keep=` que se pasa a `discover_incremental` **debe** replicar el filtro shorts/live bajo el que se pobló la DB (`jobs._keep_ref`). Si no, la cola de la ventana se llena de entradas que nunca podrán ser "conocidas" y la ventana crece sin motivo.
|
||||
- `upsert_videos` usa `COALESCE` en `upload_date`/`duration`: el discovery flat manda `None` y sin eso cada sync borraría las fechas aprendidas en la extracción — justo la marca de agua.
|
||||
- `video_count` se recalcula con `store.mark_channel_synced()`, nunca con `len(refs)`: la ventana son 30 y el canal puede tener 861.
|
||||
- Sin historial local (canal nuevo) se recorre el canal completo, que es lo correcto en la primera pasada.
|
||||
- Escotilla de escape: `--full` en el CLI, `opts["full"]` en los jobs, botones **Full rescan** / **Full** y el checkbox *Full channel rescan* en la webapp. Config en `sync:` de `config.yaml`.
|
||||
- `/tools/sync-channels` y `/tools/avatars` solo quieren metadatos+avatar del objeto canal, así que llaman a `discover_channel(..., limit=1)`: una petición en vez de paginar el canal.
|
||||
|
||||
### El orden de la lista es cronológico, y la fecha casi siempre es inferida
|
||||
|
||||
La sección de vídeos ordena por defecto como YouTube: subida más reciente primero, **haya o no `.md`**. El problema es que la mayoría de las filas no tienen fecha — medido sobre la biblioteca real, **4541 de 4959** — porque el discovery flat no la trae y solo aparece con la extracción.
|
||||
|
||||
`videos.channel_seq` es la posición del vídeo en la pestaña `/videos` (mayor = más nuevo) y es lo único que sitúa a un vídeo sin extraer. `store.SORT_DATE_SQL` deriva de ahí la fecha de orden, en cascada: (1) su `upload_date`; (2) la del vídeo con fecha **inmediatamente superior** en su canal — no es anterior a esa, y el rank desempata; (3) la fecha más nueva conocida en el canal, para la racha que está por encima de todo lo datado; (4) `NO_DATE_SENTINEL` (`"00000000"`) cuando el canal entero está sin extraer. El resultado se expone como `sort_date` + `date_estimated`, y la UI lo pinta con `~`. El sentinel no sale nunca al cliente: es un rank, no una fecha.
|
||||
|
||||
- **El fallback anterior era `discovered_at`**, que fechaba "hoy" todo lo no scrapeado y clavaba el backlog entero encima de los vídeos realmente recientes. Es la razón de ser de todo esto.
|
||||
- El paso (3) también es deliberado: con `'99999999'` un canal sin un solo vídeo extraído (Víctor Pérez, 212 vídeos) se adueñaba de la página 1 sobre canales con fechas reales. Un vídeo sin evidencia no adelanta a uno con evidencia; dentro de su canal el rank lo sigue ordenando bien.
|
||||
- `upsert_videos` asigna `channel_seq` **por encima del máximo actual del canal**, no desde cero: un sync solo ve la ventana más nueva, y todo lo que no trajo es más viejo por construcción. Así ambas mitades quedan ordenadas sin repaginar el canal, y sincronizar repetidamente no reordena nada.
|
||||
- Los dos lookups de `SORT_DATE_SQL` corren una vez por fila sin fecha, así que van sobre **índices parciales** (`idx_videos_dated_seq`, `idx_videos_dated`) que solo indexan las filas datadas. Sin el `WHERE` parcial, buscar "el vídeo datado que tengo encima" recorre todas las filas intermedias mirando la tabla: 2500 filas dentro de un canal cuyos 8 vídeos datados están arriba del todo es cuadrático. Medido: **576 ms → 22 ms por página**. El `WHERE` de la query tiene que estar escrito igual que el del índice o SQLite lo descarta en silencio; hay un test sobre `EXPLAIN QUERY PLAN` que lo fija.
|
||||
- El orden lleva desempate total (`channel_seq`, `video_id`). Sin él, `LIMIT/OFFSET` repite o se salta filas entre páginas.
|
||||
- `_order_clause` acepta las dos ortografías de cada clave. La webapp mandaba `view_count`/`like_count`/`duration` y el store solo conocía `views_desc`/`duration_desc`: **todo** sort que no fuera fecha caía en un `upload_date DESC` crudo que ignoraba la inferencia.
|
||||
- Bases anteriores a la columna se siembran una sola vez en `_seed_channel_seq()`, rankeando por `discovered_at DESC, rowid ASC` — el orden en que se aprendieron: un batch de discovery comparte timestamp y se inserta de nuevo a viejo, y un sync incremental solo añade ids más nuevos que todo lo guardado.
|
||||
|
||||
**Verificado contra la lista real de YouTube** (`discover_channel(limit=N)` sobre los 7 canales, 1 petición por ventana), no contra la intuición:
|
||||
|
||||
| Canal | Filas comparadas | Pares invertidos |
|
||||
|---|---|---|
|
||||
| Alex Hormozi | 28 | 0 / 378 |
|
||||
| Benjamín Cordero | 29 | 0 / 406 |
|
||||
| Fazt | 26 | 0 / 325 |
|
||||
| Nostal Vlad | 30 | 0 / 435 |
|
||||
| Platzi | 29 | 0 / 406 |
|
||||
| Víctor Pérez | **212 (canal completo, 0 fechas reales)** | **0 / 22 366** |
|
||||
| Código Espinoza | 30 | 9 / 435 → **0 tras un sync** |
|
||||
| Código Espinoza (fondo) | 200 | 0 / 19 900 |
|
||||
|
||||
- Víctor Pérez es la prueba fuerte del seed: 212 vídeos **sin una sola fecha real**, orden derivado enteramente del rank reconstruido, y **cada fila en el índice exacto** de YouTube.
|
||||
- El único desvío (Espinoza, 9 pares) venía de un seed erróneo en la ventana más nueva, no del mecanismo: los rangos ahí no se habían observado. **Un pase normal de discovery (1 petición, 30 entradas) lo dejó en coincidencia exacta**, porque `upsert_videos` re-rankea toda la ventana, no solo lo nuevo. Para el fondo de un canal hace falta un **Full rescan**.
|
||||
- Ese caso también decide el diseño: con rank puro (`channel_seq DESC` solo) Espinoza salía **peor**. La fecha real va primero justo para que lo que costó una extracción corrija un rank reconstruido, y el rank solo desempata cuando la fecha —que es solo día, sin hora— no puede.
|
||||
- Scripts de la medición: comparan `query_videos` contra `discover_channel` y cuentan pares invertidos. Repetible con ~1 petición por canal.
|
||||
|
||||
**Lo que la coincidencia exacta NO demuestra.** Una auditoría posterior sobre la base viva encontró que **`channel_seq` no contiene ni una sola posición observada de YouTube**: el 100 % es la reconstrucción de `_rank_unranked`, y los rangos son perfectamente contiguos por canal (`MIN=1`, `MAX=n`, sin huecos) — la firma del seed, no de un sync, que deja huecos al solapar ventanas. Que acierte es una propiedad del historial de discovery, no del diseño. Y **Nostal Vlad la rompe**: sus 3 vídeos más nuevos tienen `channel_seq` **1, 2, 3** — el fondo del canal — porque un lote posterior trajo vídeos más viejos, justo la premisa que el seed asume falsa. Se muestra bien **solo porque esos 3 tienen fecha real** y la fecha manda. Sin fecha se hundirían. Un discovery lo re-observa.
|
||||
|
||||
#### Tres bugs del orden que la prueba de campo no cubría
|
||||
|
||||
Los tres se encontraron auditando la base viva, no razonando:
|
||||
|
||||
- **`sort=oldest` abría con lo que no tiene fecha.** `NO_DATE_SENTINEL` (`"00000000"`) es el valor más bajo, así que lo que lo mantiene fuera de la portada bajo el orden por defecto es exactamente lo que lo ponía **primero** en ascendente: 212 vídeos de un canal sin extraer por delante de una subida real de 2017. `_OLDEST_FIRST` lleva ahora `(sort_date = '00000000')` como primer término. "No lo sabemos" no es "el principio de los tiempos".
|
||||
- **El desempate global rankeaba por tamaño de catálogo.** `channel_seq` cuenta hasta el número de vídeos del canal, así que compararlo **entre** canales ordena por quién tiene más. Con el **92 %** de los pares adyacentes empatados en `sort_date`, eso decidía casi toda la lista: un canal entero delante de otro solo porque 575 > 476. El desempate es ahora `channel_id` **y luego** `channel_seq`, de modo que cada bloque empatado queda contiguo por canal y el orden interno —lo que tiene que cuadrar con YouTube— no se toca. Verificado: 373 bloques de empate, **0** con un canal partido.
|
||||
- **Una fila sin rank no queda desordenada, queda mal colocada.** `COALESCE(channel_seq, -1)` hace que la regla 2 herede el vídeo datado **más viejo** del canal, así que un vídeo que discovery acaba de encontrar —de los más nuevos— se muestra el último. Medido: dos subidas nuevas de Hormozi con `sort_date=20180720`, penúltima y última de 513. Pasa cuando un server de larga vida sigue con el código previo a la columna después de migrar. `_rank_unranked()` corre **siempre**, no solo al añadir la columna, y es no-op si no hay NULLs.
|
||||
|
||||
**Ojo al importar `yt_scraper.webapp.app`:** tiene `app = create_app()` a nivel de módulo ([app.py:115](src/yt_scraper/webapp/app.py#L115)), así que **el simple `import` abre la base real del proyecto** vía `config.yaml`, corre la migración, `auto_import_dir` y lanza el reconcile en un hilo de fondo — aunque después le pases un `Config` distinto a `create_app()`. Para tocar solo una copia, importa `yt_scraper.store` / `yt_scraper.discover` directamente y nunca `webapp.app`.
|
||||
|
||||
### El botón de `.md` descarga `.md`
|
||||
|
||||
`_run_batch` ya no cachea miniaturas. Las traen los caminos de canal (alta de canal, herramienta *Download thumbnails*) y `/api/thumbnails/{id}` redirige al CDN lo que no esté en disco, así que colgarlas del batch gastaba una petición por vídeo por una imagen que la UI ya podía mostrar. Un vídeo cuyo `.md` ya existe en disco ahora se salta entero.
|
||||
|
||||
### Estados terminales y frescura (la UI tiene que reflejar DB + disco)
|
||||
|
||||
`no_subtitles` **no** significa "este vídeo no tiene subtítulos". Se escribe siempre que `data.segments` viene vacío ([pipeline.py:146](src/yt_scraper/pipeline.py#L146)), lo que mezcla tres causas muy distintas: el vídeo no tiene pistas, la política de idiomas rechazó las que sí tiene, o la descarga vino vacía por throttling. `extract.describe_missing_subtitle()` distingue los casos y el motivo se guarda en `videos.error_msg` vía `mark_status(vid, status, reason)`.
|
||||
|
||||
Incidente que motivó esto: 511 vídeos de un canal quedaron en `no_subtitles` porque `config.example.yaml` ponía los idiomas en modo `manual` y el canal solo publica subtítulos automáticos. `_sources_for("manual")` no hace fallback. El default es ahora `any` (manual primero, auto después) — `manual` es opt-in explícito.
|
||||
|
||||
- **Ambos estados son reintentables.** `Store.RETRYABLE_STATUSES` = `("error", "no_subtitles")`. `reset_videos()` los devuelve a `pending`; `reset_errors()` conserva su significado estrecho. `done` nunca se toca. Expuesto en `POST /api/videos/reset`, `GET /api/videos-retryable` y `yt-scraper reset`.
|
||||
- **Contenido bloqueado se detecta en el discovery, no al fallar.** Los entries planos de yt-dlp traen `availability`; `subscriber_only` = vídeo de membresía. Medido: 120 entradas en una petición, 4 marcadas, exactamente las mismas que la DB había aprendido a base de extracciones fallidas. Se guarda en `videos.availability` y `VideoRow.block_reason` lo traduce, con fallback al `error_msg` para las filas antiguas. `get_pending()` las excluye (los `_run_batch` por ID explícito sí las intentan, que es lo que hace recuperable el caso "compré la membresía"). Filtro `blocked` en `query_videos` / `GET /api/videos?blocked=true`, badge `.st-locked` en la UI.
|
||||
- El enum completo es `private | premium_only | subscriber_only | needs_auth | unlisted | public` ([yt_dlp/extractor/common.py:414](https://github.com/yt-dlp/yt-dlp)). Solo los cuatro primeros bloquean: **`unlisted` se descarga sin problema** y marcarlo escondería vídeos que sí se pueden tener. En el SQL del filtro, `COALESCE` es imprescindible: con `availability` NULL, `NOT (NULL OR ...)` es NULL y la rama negada devolvería cero filas.
|
||||
- **Salvo los permanentes.** `Store.PERMANENT_ERROR_PATTERNS` (members-only, private, removed) se excluyen del reset: reintentarlos no los va a desbloquear y gasta peticiones que necesitan los que sí pueden salir. `retryable_counts()` los devuelve aparte en la clave `permanent`. Mantén la lista estrecha: el mensaje de throttling de YouTube ("rate-limited … try again later") **sí** es reintentable y no debe caer ahí.
|
||||
- **`backfill_from_markdown` no arregla estados**: puebla segmentos y metadatos, pero nunca `status` ni `markdown_path`. Para eso está `segments.reconcile_markdown()`, que corre al arrancar la webapp, en `POST /api/tools/reconcile` y en `yt-scraper reconcile`.
|
||||
- `reconcile` es aditivo por defecto. `prune=True` es opt-in porque es destructivo: borra duplicados obsoletos y degrada `done` cuyo `.md` desapareció. Se niega a degradar nada si el árbol de markdown está vacío (root mal configurado).
|
||||
- **La identidad de una nota es el `video_id`, nunca el nombre del fichero.** El nombre sale de `{upload_date}_{slug}` y las dos partes son inestables: YouTube sirve los títulos localizados (el mismo vídeo volvió como "La controversia de Claude Fable 5" en una pasada y "The Claude Fable controversy 5" en la siguiente) y los creadores renombran. Cada cambio de stem escribía un fichero nuevo y dejaba el anterior huérfano.
|
||||
- Las dos rutas que renderizan pasan ahora por `pipeline._render_and_retire()`, que borra el fichero al que apuntaba `markdown_path` si el stem cambió. Existe **porque ya divergieron una vez**: `re_render_videos` usaba la fecha cruda (`20240519_`) y `process_video` la normalizada (`2024-05-19_`) → 94 `.md` para 61 filas. Medido tras arreglarlo: 33 duplicados reales en disco, todos con canónico existente, mezclando las dos causas (`20240519_why-did-the-2000s-look-like-this` contra `2024-05-19_por-que-los-2000-se-veian-asi`). Cero notas únicas perdidas.
|
||||
- `build_filename_stem` acepta `{video_id}` en `filename_template` para quien quiera que el fichero se identifique solo, sin la DB. Es opt-in: el default no cambia, para no renombrar bibliotecas existentes.
|
||||
- `mark_done` sigue siendo el **único** escritor de `markdown_path`.
|
||||
- Al leer `.md`, captura `UnicodeDecodeError` además de `OSError`: es un `ValueError`, y dejarlo escapar abortaba el escaneo entero saltándose todos los ficheros posteriores.
|
||||
|
||||
En el frontend, `refreshLiveState()` es el único punto de invalidación: lo llaman el handler `done` del SSE, un poll de 5 s mientras hay job activo, y `visibilitychange`/`focus`. `setView` recarga **siempre**, no solo cuando la lista está vacía.
|
||||
|
||||
### Rate limiting: `ratelimit.py` es la política, y no es opcional
|
||||
|
||||
Tres piezas distintas, no intercambiables:
|
||||
|
||||
- **`Pacer` (`GLOBAL_PACER`)** — separación mínima entre peticiones, **global al proceso**. Existe porque el `JobManager` serializa *jobs* pero los `/api/tools/*` corren fuera de él, en el threadpool: sin un pacer compartido, dos consumidores machacan YouTube creyendo cada uno que va educado. `wait(cost=N)` reserva N ranuras — `extract_info` de un vídeo son **2** peticiones (watch + player) y cobrarlo como 1 hacía que el pacer contase la mitad.
|
||||
- **`backoff_delay`** — `min(base * 2**n + jitter, cap)`, el algoritmo que Google documenta para sus propias APIs, con el jitter re-sorteado en cada intento.
|
||||
- **`ThrottleGuard`** — el circuit breaker. Cuenta rate-limits **consecutivos**; al llegar a `throttle_threshold` (3) el job para. Lo que no tocó sigue en `pending`, que es el estado recuperable.
|
||||
|
||||
**El incidente que lo motiva:** 343 de los 350 `error` de la DB decían literalmente *"The current session has been rate-limited by YouTube for up to an hour"*. No es que los vídeos fallaran: el scraper se ganaba un ban de una hora y luego **quemaba el resto de la cola en cascada** marcándolos como fallidos. Sin breaker, un throttling se convierte en cientos de filas envenenadas.
|
||||
|
||||
- Un fallo **no** de throttling (vídeo privado, sin subtítulos) **resetea** el contador. Si no, tres vídeos de membresía seguidos abortarían un scrape sano.
|
||||
- `is_rate_limited` e `is_quota_exhausted` son cosas distintas: Google documenta la cuota como diaria (AIP-194), así que reintentar no la arregla y el guard corta en seco sin backoff.
|
||||
- Los mensajes reales de la DB están fijados verbatim en [tests/test_ratelimit.py](tests/test_ratelimit.py). Si el detector deja de reconocerlos, el breaker es decorativo.
|
||||
- **No amplíes `Store.PERMANENT_ERROR_PATTERNS`** con el mensaje de throttling: tiene que seguir siendo reintentable.
|
||||
|
||||
#### Los nombres de las opciones de yt-dlp se validan, no se revisan
|
||||
|
||||
Durante toda la historia del proyecto, los cuatro puntos de red pasaron `sleep_subrequests` a yt-dlp. **Esa opción no existe.** yt-dlp ignora en silencio las claves que no conoce, así que nunca hubo ni un milisegundo de pausa entre las sub-peticiones de una extracción, mientras `config.yaml` aparentaba tenerlo configurado. Lo mismo con `extract_flat_args`.
|
||||
|
||||
El nombre real es **`sleep_interval_requests`**, y es la **única** palanca que afecta a la extracción: `sleep_interval` / `max_sleep_interval` se disparan en el downloader de ficheros y **nunca** actúan bajo `skip_download`, que es todo el camino de metadatos.
|
||||
|
||||
Todo `ydl_opts` de politeness sale ahora de `ratelimit.ydl_throttle_opts()`, y [tests/test_request_economy.py](tests/test_request_economy.py) valida cada clave contra `yt_dlp.parse_options([]).ydl_opts` (172 nombres válidos). Si inventas una opción, el test falla.
|
||||
|
||||
**yt-dlp no reintenta 403/429 en YouTube**: su extractor los excluye explícitamente del `RetryManager`. `retries=10` no te protege de nada; el backoff ante throttling es responsabilidad nuestra.
|
||||
|
||||
#### Economía de peticiones: lo medido, para no re-litigarlo
|
||||
|
||||
| Operación | Peticiones | Nota |
|
||||
|---|---|---|
|
||||
| Extracción de un vídeo | **2** | watch + `youtubei/v1/player`. Es el suelo. |
|
||||
| Transcripción (`timedtext`) | 1 | vía `yt_get` |
|
||||
| Miniatura | 1 | `i.ytimg.com`, CDN estático, no cuenta contra el throttle de la API |
|
||||
| `discover_channel(limit=30)` | **1** | |
|
||||
| Discovery completo de canal de 2564 vídeos | ~86 | ~30 entradas por petición |
|
||||
| `deep_channel_avatar` | **3-4** | antes **735 y subiendo** cuando la medición lo abortó |
|
||||
|
||||
- **`deep_channel_avatar` extraía el canal entero para leer una URL de imagen.** yt-dlp redirige una URL de canal pelada a `/videos`, recorre además `/streams` y `/shorts`, y `download=False` solo evita bajar el media, no la extracción. `extract_flat` + `playlistend: 1` son **carga estructural, no tuning**.
|
||||
- **Restringir `player_client` no ahorra nada y rompe cosas.** Medido: `web_safari` solo → 1 petición pero **0 pistas de subtítulos**. `player_skip=webpage` → pierde subtítulos *y* `duration`/`channel_id`/`tags`. El default de yt-dlp (`android_vr` + `web_safari`, 2 peticiones) ya es el óptimo.
|
||||
- **`playlistend` se ignora en silencio con `process=False`.** Medido sobre 2564 vídeos: procesado + `playlistend=60` = **2** peticiones; sin procesar, el generador perezoso ignora el límite y recorrerlo cuesta **86**. Hay un test que fija `process is not False`.
|
||||
- `check_formats` cuesta **una petición HTTP por formato**. Nunca lo actives.
|
||||
- No desactives `cachedir`: yt-dlp cachea ahí el player JS resuelto y quitarlo añade peticiones.
|
||||
- `POST /api/tools/thumbnails` con solo `channel_id` acota a `THUMBNAIL_AUTO_LIMIT` (60). Antes, añadir un canal disparaba una petición al CDN **por cada vídeo del catálogo** — 2564 de golpe en Platzi.
|
||||
|
||||
#### La transcripción tiene que ser el idioma que se habla, no una traducción
|
||||
|
||||
`languages` es una preferencia **entre idiomas que sabes leer**, no una orden de aceptar una traducción automática cuando el transcript real está ahí. `pick_subtitle` pone delante el idioma hablado si está entre los configurados; solo si no lo está manda el orden del dict.
|
||||
|
||||
Cómo se detecta el original: YouTube publica el ASR como **`<lang>-orig`** y luego una cola larga de traducciones con el código pelado — incluida **una traducción al propio idioma del vídeo**. En un vídeo español existen `es-orig` *y* `es`, y solo el primero es el transcript real. `_is_original_track()` mira ese sufijo, y `original_language()` lo prefiere sobre `info["language"]` porque el sufijo es evidencia de la lista de pistas mientras que `language` es metadato que YouTube localiza.
|
||||
|
||||
**El fallo:** `_normalize_lang("es-orig") == "es"`, así que el sufijo — lo único que distingue el ASR real de una traducción — se tiraba antes de comparar. Con `{es, es-419, en}` sobre el canal de Alex Hormozi (inglés), el picker enganchaba `es` y guardaba una traducción máquina del inglés hablado. Prueba: las URLs de `timedtext` llevaban `lang=en&kind=asr&variant=gemini&tlang=es` — **`tlang=` es la marca de traducción**. Los canales en español acertaban **por casualidad**: yt-dlp lista `es-orig` antes que `es`. 3 filas de 119 afectadas; las otras 116 son `es-orig` legítimo.
|
||||
|
||||
- `extractor_args: {youtube: {skip: ["translated_subs"]}}` **no arregla esto**. El gate está anidado dentro de `if is_manual_subs` (`_video.py:4284-4285`) y `is_manual_subs` es False para pistas ASR. Medido: 158 claves con el flag y sin él, idénticas.
|
||||
- `subtitleslangs` **no influye** en la selección: `pick_subtitle` lee `info["automatic_captions"]` crudo, mientras yt-dlp confina su propia selección a `requested_subtitles`, clave que este repo nunca lee.
|
||||
- Recuperar una fila así **no se puede por las vías normales**: `reset_videos` nunca toca `done`, y `re-render` regeneraría el idioma equivocado desde `segments_json`. Hay que limpiar `transcript_segments`, `transcript_fts`, `segments_json` **y borrar el `.md`** — si lo dejas, `reconcile_markdown()` lo re-marca `done` al arrancar.
|
||||
|
||||
**Pendiente, sin resolver:** los títulos también vienen localizados. El canal Platzi tiene títulos en inglés en la DB ("The Claude Fable controversy 5" por "La controversia de Claude Fable 5") mientras sus `.md` se llamaron en español. No está determinado cuál de los dos lados —el listado flat del tab o la extracción— es el localizado; hace falta una medición contra la red para saberlo.
|
||||
|
||||
#### El detector de throttling no puede depender de la prosa
|
||||
|
||||
Hay dos ortografías y **son cadenas distintas**: yt-dlp lanza `HTTP Error 429: ...` y `requests` lanza `429 Client Error: ... for url: ...`. `_HTTP_STATUS` solo cubría la primera, así que la segunda se detectaba únicamente por el substring `"too many requests"` — y un 429 servido **sin reason phrase**, que es lo normal en HTTP/2, no lleva esa prosa y atravesaba el breaker sin contarse. Igual con 408, que Google documenta como reintentable.
|
||||
|
||||
Ahora `_download_subtitle` normaliza el mensaje anteponiendo `HTTP Error <status>:` desde `exc.response.status_code`, y el detector reconoce ambas formas. Está cubierto por tests parametrizados con las dos ortografías y con los casos que **no** deben disparar.
|
||||
|
||||
#### Los eventos de error usan `message`; los de log usan `msg`
|
||||
|
||||
`jobs.py` emite `{"message": ...}` en los cinco sitios donde manda un evento `error`. El frontend leía `d.msg`, que no existe en esos eventos, así que **todo** fallo de job salía como `"Job failed: connection error"` — incluida la explicación detallada del breaker. Corregido en `app.js` leyendo `d.message || d.msg`, y el flag `throttled` cambia el texto: parar por rate limiting no es un crash, es una parada deliberada con todo lo no alcanzado aún en `pending`.
|
||||
|
||||
#### Los metadatos se persisten antes de la salida por "sin transcripción"
|
||||
|
||||
`process_video` guarda `view_count`/`description`/`thumbnail`/`tags`/`upload_date` **encima** del `return "no_subtitles"`. La extracción ya pagó sus dos peticiones y el `info` está en memoria; tirarlo porque falló la descarga *separada* de subtítulos hace que el reintento las vuelva a gastar para obtener datos que ya teníamos. Medido tras el incidente: cinco filas quedaron con `upload_date`, `view_count`, `description` y `thumbnail` todos NULL.
|
||||
|
||||
Relacionado: `mark_status` trunca el motivo a **2000** caracteres, no 500. A 500 el corte caía veinte caracteres antes del `tlang=` que probaba el diagnóstico.
|
||||
|
||||
#### Los handlers que hacen I/O de red van en `def`, no en `async def`
|
||||
|
||||
FastAPI ejecuta los `async def` **en el event loop**; los `def` van al threadpool. `add_channel`, `sync_channels`, `download_thumbnails`, `download_avatars` y `reconcile` bloquean con yt-dlp o `requests`, así que siendo `async` congelaban el servidor entero — incluido el SSE del job en curso — durante toda su duración. Verificado en producción: añadir Platzi bloquea 295 s, y con el cambio `/api/dashboard` respondió 146/146 sondas con mediana de 0,000 s.
|
||||
|
||||
`stream_job` **sí** debe seguir siendo `async`: devuelve el `EventSourceResponse`.
|
||||
|
||||
#### Calibración
|
||||
|
||||
El único techo publicado es el de la wiki de yt-dlp: **~300 vídeos/hora (~1000 peticiones/hora)** en sesión sin cuenta. Google no documenta límites para acceso sin API, y los umbrales del bot-check no están documentados en ninguna parte: cualquier cifra concreta es prudencia, no norma.
|
||||
|
||||
`config.yaml` apunta por debajo de eso. Medido en producción sobre Platzi: **13,25 s/vídeo → ~272 vídeos/hora, ~815 peticiones/hora**. Si cambias los tiempos, vuelve a medir: 3 peticiones a `youtube.com` por vídeo es la constante de la que sale todo lo demás.
|
||||
|
||||
### Cookies
|
||||
|
||||
Vault híbrido: metadatos en `cookies_meta`, archivos Netscape en `cookies/<uuid>.txt` (gitignored). **Exactamente una** cookie activa a la vez. Cuando no se pasa `--cookies`, CLI, webapp y monitor caen en `cookies.resolve_active_path(store)`. `auto_import_dir()` adopta `.txt` sueltos al arrancar CLI y webapp.
|
||||
|
||||
### Job runner de la webapp
|
||||
|
||||
`JobManager` = un único thread daemon + `deque`, deliberadamente secuencial para no martillear a YouTube. Los eventos van a una lista en memoria por job (`_events`) que el endpoint SSE **poletea** con un cursor cada 250 ms; no hay pub/sub real, y los eventos se pierden al reiniciar (el estado durable está en `scrape_jobs`).
|
||||
|
||||
Despacho por `opts` en `_run_job()`: `mode == "audio"` → `_run_audio`; hay `video_ids` → `_run_batch` (por vídeo: `.md` + thumbnail, saltando los que ya tienen `.md` en disco); si no → `_run_channel` (discovery + pendientes + `polite_sleep` entre vídeos). La cancelación es cooperativa: `_cancel` es un set que los loops consultan.
|
||||
|
||||
### Frontend
|
||||
|
||||
Sin build step, por decisión fija ([.opencode/agent/webapp-builder.md](.opencode/agent/webapp-builder.md)): Tailwind, Alpine 3 y Chart.js por CDN. Todo el estado vive en **un solo** componente Alpine, `window.platform()` en [static/app.js](src/yt_scraper/webapp/static/app.js), y **debe** quedar definido antes de que Alpine inicialice — el orden de los `<script defer>` en `index.html` (app.js antes de alpinejs) es carga funcional, no estilo. Las vistas se sincronizan con la URL (`hydrateURL`/`syncURL`), no con hash routes. Identidad visual: dark "command center", `#0a0a0f` + acento rose `#f43f5e`.
|
||||
|
||||
### Rutas de datos y convenciones
|
||||
|
||||
El data root se deriva siempre como `Path(cfg.output_dir_resolved).parent` → `data/{markdown,audio,thumbnails,avatars,exports,analysis}` y `data/state.db`. `videos.markdown_path` se guarda **relativo a ese root**.
|
||||
|
||||
Gotcha real: el job de audio de la webapp guarda `data/audio/<video_id>.mp3` (`outtmpl` con `%(id)s`) y `GET /api/videos/{id}/audio` solo encuentra ese nombre, mientras que `yt-scraper audio` (CLI) escribe por título. Los MP3 bajados por CLI no los sirve la API.
|
||||
|
||||
### Cualquier petición HTTP a CDNs de YouTube va por `_yt_http.yt_get()`
|
||||
|
||||
Headers compartidos (UA de Chrome + `Referer: https://www.youtube.com/`). Sin eso, `yt3.ggpht.com` (avatares) y `timedtext` (subtítulos) devuelven 403. `yt-dlp` es la **única** interfaz con YouTube: no añadas llamadas directas a InnerTube.
|
||||
|
||||
### Idiomas: dict por-idioma con compatibilidad legacy
|
||||
|
||||
`languages` puede ser lista (`["es","en"]`, todos usan `prefer_manual`) o dict `{lang: "manual"|"auto"|"any"}`. La normalización está duplicada a propósito: `config.parse_languages()` y `extract._coerce_languages()` (para que `extract` no dependa de `config`). Si cambias las reglas, cambia las dos.
|
||||
|
||||
Estados de vídeo: `pending` / `done` / `no_subtitles` / `error`. Estados terminales de job: `Store.TERMINAL_STATUSES`.
|
||||
|
||||
## Tests
|
||||
|
||||
Solo `tmp_path` + `monkeypatch`, cero red: `yt_dlp.YoutubeDL` se sustituye por un fake (ver [tests/test_webapp_jobs.py](tests/test_webapp_jobs.py)) y `pick_subtitle` se testea con dicts `info` sintéticos. Los tests del store construyen un `Store` sobre `tmp_path`, así que también cubren la migración idempotente.
|
||||
|
||||
## Documentos de referencia
|
||||
|
||||
- [AGENTS.md](AGENTS.md) — guía de entrada para cualquier agente de código (pitch, mapa de módulos, reglas de oro, punteros a docs).
|
||||
- [docs/superpowers/specs/2026-07-26-platform-design.md](docs/superpowers/specs/2026-07-26-platform-design.md) — diseño de referencia (esquema, endpoints, alcance, exclusiones). Es la fuente de verdad de las decisiones; donde el código difiere, gana el código.
|
||||
- [docs/backlog.md](docs/backlog.md) — backlog destilado: lo pendiente y qué ya está shipped (sustituye al histórico OPPORTUNITIES.md).
|
||||
- [docs/audits/](docs/audits/) — auditorías: ingeniería inversa del mecanismo InnerTube del Obsidian Web Clipper (`YOUTUBE-TRANSCRIPT-AUDIT.md`, contexto de *por qué* `yt-dlp` es la ruta elegida) y comparativa extensión vs scraper (`CLIPPER-COMPARISON-AUDIT.md`).
|
||||
- `docs/GETTING-STARTED.md`, `docs/CONFIG.md`, `docs/COOKIES.md` — docs de usuario (máquina nueva, referencia de config, guía de cookies).
|
||||
- `data/`, `cookies/`, `.run/` están gitignored y contienen datos reales (sesión de YouTube). No los commitees ni pegues su contenido en respuestas. `config.yaml` ya no está versionado (config privada del usuario; la plantilla versionada es `config.example.yaml`).
|
||||
|
||||
## Actualización de docs
|
||||
|
||||
Checklist rápida al aterrizar un cambio — las docs desactualizadas mienten peor que no existir:
|
||||
|
||||
- **Flags/subcomandos del CLI** nuevos o cambiados → `README.md` (sección CLI) y ejemplos de `docs/GETTING-STARTED.md`.
|
||||
- **Plantilla `templates/video.md.j2`** → este archivo (el contrato bidireccional con `segments.py`) y el ejemplo de nota del `README.md`.
|
||||
- **Endpoints de la webapp o UI** → `README.md` (sección webapp) y las secciones *Job runner* / *Frontend* de aquí.
|
||||
- **Clave de config nueva o default distinto** → `config.example.yaml`, `docs/CONFIG.md` y el resumen de bloques del `README.md`.
|
||||
- **Números que caducan** (cantidad de tests, refs `archivo.py:NN`) → usa órdenes de magnitud (`~240 tests`) y re-verifica las refs de línea antes de citarlas.
|
||||
- **Feature que estaba en el backlog** → márcala como shipped en `docs/backlog.md`; si añade un módulo, actualiza el mapa de `AGENTS.md`.
|
||||
@@ -0,0 +1,44 @@
|
||||
# yt-scraper — atajos para macOS/Linux (y test/clean multiplataforma).
|
||||
#
|
||||
# macOS trae `make` de fabrica, asi que un Mac nuevo queda sirviendo con:
|
||||
# make setup && make serve
|
||||
# En Windows estos targets solo orientan: usa los .bat / .ps1 del proyecto
|
||||
# (start-server.bat, stop-server.bat, scripts/doctor.ps1).
|
||||
|
||||
# En Windows (GNU make define OS=Windows_NT) los scripts .sh no aplican.
|
||||
ifeq ($(OS),Windows_NT)
|
||||
|
||||
setup:
|
||||
@echo [Windows] Los scripts .sh son para macOS/Linux. Usa scripts\doctor.ps1 o instala a mano: pip install -e ".[dev,web,analysis]"
|
||||
|
||||
serve:
|
||||
@echo [Windows] Usa start-server.bat
|
||||
|
||||
stop:
|
||||
@echo [Windows] Usa stop-server.bat
|
||||
|
||||
test:
|
||||
uv run python -m pytest -q || python -m pytest -q
|
||||
|
||||
else
|
||||
|
||||
setup:
|
||||
@bash scripts/bootstrap.sh
|
||||
|
||||
serve:
|
||||
@bash scripts/start-server.sh
|
||||
|
||||
stop:
|
||||
@bash scripts/stop-server.sh
|
||||
|
||||
test:
|
||||
@if command -v uv >/dev/null 2>&1; then uv run python -m pytest -q; else python -m pytest -q; fi
|
||||
|
||||
endif
|
||||
|
||||
# Borra solo caches de Python; no toca datos, venv ni .run/.
|
||||
clean:
|
||||
@find . -path ./.venv -prune -o -path ./.venv-312 -prune -o -type d \( -name __pycache__ -o -name .pytest_cache \) -prune -exec rm -rf {} +
|
||||
@echo caches eliminadas.
|
||||
|
||||
.PHONY: setup serve stop test clean
|
||||
@@ -1,298 +0,0 @@
|
||||
# AUDIT · Oportunidades de extensión
|
||||
|
||||
> Análisis de features, comandos y opciones que se pueden construir sobre la infraestructura actual.
|
||||
> Cada item incluye: valor, esfuerzo estimado, dependencias y comando propuesto.
|
||||
|
||||
---
|
||||
|
||||
## Infraestructura disponible (lo que ya tenemos)
|
||||
|
||||
| Recurso | Estado | Reutilizable para |
|
||||
|---|---|---|
|
||||
| `yt-dlp` instalado | ✅ v2026.7.4 | Audio/video download, thumbnails, metadata enriquecida |
|
||||
| SQLite 3.49 con FTS5 | ✅ verificado | Búsqueda full-text sobre transcripciones |
|
||||
| SQLite JSON1 | ✅ verificado | Queries estructuradas sobre metadatos |
|
||||
| 33 transcripciones parseadas | ✅ en `data/state.db` | Análisis, estadísticas, búsqueda |
|
||||
| Markdown con frontmatter YAML | ✅ 842 KB en `data/markdown/` | Obsidian, Pandoc, static site generators |
|
||||
| Pipeline modular | ✅ 8 módulos | Insertar nuevos pasos sin romper nada |
|
||||
|
||||
---
|
||||
|
||||
## TIER 1 · Alto valor, bajo esfuerzo (1-3 h c/u)
|
||||
|
||||
### 1. `search` — Búsqueda full-text en transcripciones
|
||||
|
||||
```bash
|
||||
yt-scraper search "vaporwave"
|
||||
yt-scraper search "dragon ball" --channel UCmhcYyPg7fsxMzQsY0RJBjw
|
||||
```
|
||||
|
||||
**Qué hace:** busca texto dentro de las transcripciones ya scrapeadas y devuelve vídeo + timestamp exacto.
|
||||
|
||||
**Implementación:**
|
||||
- Nueva tabla `transcript_segments(video_id, start, text)` poblada durante el scrape
|
||||
- Índice `USING fts5(text)` sobre esa tabla
|
||||
- Comando `search` que hace `SELECT ... WHERE transcript_segments MATCH ?`
|
||||
- **Esfuerzo:** 1-2 h. Sin dependencias nuevas.
|
||||
|
||||
### 2. `stats` — Estadísticas del canal
|
||||
|
||||
```bash
|
||||
yt-scraper stats
|
||||
yt-scraper stats --channel UCmhcYyPg7fsxMzQsY0RJBjw
|
||||
```
|
||||
|
||||
**Qué muestra:**
|
||||
```
|
||||
Nostal Vlad (@Nostal-Vlad)
|
||||
Videos: 33 scrapeados / 37 totales
|
||||
Duración total: 11.2 horas
|
||||
Palabras totales: 89,432
|
||||
Fecha rango: 2019-04-28 → 2026-07-26
|
||||
Views promedio: 45,231
|
||||
Likes promedio: 3,201
|
||||
Top tags: videojuegos (12), 2000s (10), nostalgia (8)
|
||||
```
|
||||
|
||||
**Implementación:** queries SQL agregadas + presentación con Rich tables.
|
||||
- **Esfuerzo:** 1 h.
|
||||
|
||||
### 3. `export` — Export multi-formato
|
||||
|
||||
```bash
|
||||
yt-scraper export --format json # un JSON con todo
|
||||
yt-scraper export --format csv # CSV de metadatos
|
||||
yt-scraper export --format srt # subtítulos SRT estándar
|
||||
yt-scraper export --format html # galería navegable
|
||||
```
|
||||
|
||||
**Implementación:**
|
||||
- JSON: serializar `Store.get_all()` + leer transcripciones de los `.md`
|
||||
- CSV: `csv.writer` sobre metadatos
|
||||
- SRT: ya tenemos `{start, end, text}` en `Segment` — conversión trivial
|
||||
- HTML: plantilla Jinja2 con tarjetas por vídeo (thumbnail + título + link)
|
||||
- **Esfuerzo:** 2-3 h los 4 formatos.
|
||||
|
||||
### 4. `clip` — Extracción de segmento por timestamp
|
||||
|
||||
```bash
|
||||
yt-scraper clip gOUyxFwWQqA --from 01:38 --to 03:20
|
||||
```
|
||||
|
||||
**Qué hace:** devuelve el texto de la transcripción entre dos timestamps, listo para citar.
|
||||
|
||||
**Implementación:** cargar el `.md`, parsear segmentos, filtrar por rango.
|
||||
- **Esfuerzo:** 30 min.
|
||||
|
||||
### 5. `audio` — Descarga de audio (podcast)
|
||||
|
||||
```bash
|
||||
yt-scraper --channel "https://www.youtube.com/@Nostal-Vlad/videos" --download-audio
|
||||
```
|
||||
|
||||
**Qué hace:** descarga el audio MP3 de cada vídeo junto a la transcripción.
|
||||
|
||||
**Implementación:** añadir a `extract.py`:
|
||||
```python
|
||||
ydl_opts["format"] = "bestaudio/best"
|
||||
ydl_opts["postprocessors"] = [{"key": "FFmpegExtractAudio", "preferredcodec": "mp3", "preferredquality": "128"}]
|
||||
ydl_opts["outtmpl"] = str(output_dir / "%(title)s.%(ext)s")
|
||||
```
|
||||
Requiere `ffmpeg` instalado.
|
||||
- **Esfuerzo:** 1 h.
|
||||
|
||||
### 6. `re-render` — Regenerar Markdown tras cambiar plantilla
|
||||
|
||||
```bash
|
||||
yt-scraper re-render
|
||||
```
|
||||
|
||||
**Qué hace:** re-procesa todos los `.md` usando la plantilla actual sin re-descargar nada de YouTube. Útil cuando cambias `video.md.j2`.
|
||||
|
||||
**Implementación:** guardar `segments` y `chapters` en SQLite (columnas JSON), o re-parsear los `.md` existentes.
|
||||
- **Esfuerzo:** 1-2 h.
|
||||
|
||||
---
|
||||
|
||||
## TIER 2 · Alto valor, esfuerzo medio (3-8 h c/u)
|
||||
|
||||
### 7. Multi-canal — Gestión de varios canales
|
||||
|
||||
```bash
|
||||
yt-scraper channels add "https://www.youtube.com/@otro-canal"
|
||||
yt-scraper channels list
|
||||
yt-scraper --channel @nostal-vlad search "vaporwave"
|
||||
yt-scraper --all-channels stats
|
||||
```
|
||||
|
||||
**Implementación:** tabla `channels` ya existe; ampliar CLI con subcomandos `channels add/list/remove`. El `Store` ya soporta multi-canal (PK por `channel_id`).
|
||||
- **Esfuerzo:** 3-4 h.
|
||||
|
||||
### 8. `watch` — Monitoreo de nuevos vídeos
|
||||
|
||||
```bash
|
||||
yt-scraper watch --interval 6h
|
||||
```
|
||||
|
||||
**Qué hace:** ejecuta discovery periódicamente, y si hay vídeos nuevos los procesa automáticamente.
|
||||
|
||||
**Implementación:** loop con `time.sleep(interval)` + `discover_channel()`. En Windows se puede dejar corriendo o usar Task Scheduler.
|
||||
- **Esfuerzo:** 2-3 h.
|
||||
|
||||
### 9. Análisis de contenido
|
||||
|
||||
```bash
|
||||
yt-scraper analyze --wordcloud
|
||||
yt-scraper analyze --top-words 50
|
||||
yt-scraper analyze --timeline "vaporwave"
|
||||
```
|
||||
|
||||
**Qué hace:**
|
||||
- Nube de palabras (matplotlib + wordcloud)
|
||||
- Frecuencia de términos con gráfico
|
||||
- Timeline: cuándo se mencionó un tema a lo largo del tiempo
|
||||
|
||||
**Dependencias nuevas:** `matplotlib`, `wordcloud`, `nltk` (o regex simple para español).
|
||||
- **Esfuerzo:** 4-5 h.
|
||||
|
||||
### 10. `thumbnails` — Descarga masiva de miniaturas
|
||||
|
||||
```bash
|
||||
yt-scraper thumbnails --size maxres
|
||||
```
|
||||
|
||||
**Implementación:** ya tenemos `thumbnail` URL en el frontmatter. Un `requests.get` por vídeo + guardar en `data/thumbnails/`.
|
||||
- **Esfuerzo:** 30 min.
|
||||
|
||||
### 11. Traducción de transcripciones
|
||||
|
||||
```bash
|
||||
yt-scraper --channel @nostal-vlad --translate-to en
|
||||
```
|
||||
|
||||
**Qué hace:** traduce cada transcripción al idioma especificado antes de renderizar el Markdown.
|
||||
|
||||
**Opciones:**
|
||||
- Google Translate (gratis, vía `deep-translator`): rápido, baja calidad
|
||||
- LLM (OpenAI/Anthropic): calidad alta, requiere API key
|
||||
- **Esfuerzo:** 3-4 h (cualquier ruta).
|
||||
|
||||
---
|
||||
|
||||
## TIER 3 · Features avanzadas (8+ h)
|
||||
|
||||
### 12. LLM Summaries — Resúmenes por vídeo
|
||||
|
||||
```bash
|
||||
yt-scraper --channel @nostal-vlad --summarize
|
||||
```
|
||||
|
||||
**Qué hace:** genera un resumen de 3-5 párrafos por vídeo usando un LLM. Se añade al Markdown como campo `summary` en el frontmatter.
|
||||
|
||||
**Implementación:**
|
||||
- Nuevo módulo `summarize.py`
|
||||
- Usa la transcripción completa como contexto
|
||||
- Modelos: OpenAI `gpt-4o-mini` ($0.015 por vídeo de 20 min), Anthropic Claude Haiku, o modelo local con `ollama`
|
||||
- Guardar resumen en SQLite para no re-ejecutar
|
||||
- **Esfuerzo:** 4-6 h.
|
||||
|
||||
**Dependencias:** `openai` o `anthropic` o `ollama` (local, gratis).
|
||||
|
||||
### 13. Q&A / RAG — Preguntas sobre el canal
|
||||
|
||||
```bash
|
||||
yt-scraper ask "¿Qué dijo Vlad sobre los juegos flash?"
|
||||
```
|
||||
|
||||
**Qué hace:** busca en todas las transcripciones y responde con citas + timestamps exactos.
|
||||
|
||||
**Implementación:**
|
||||
- **RAG simple:** FTS5 search → mandar top-K segmentos al LLM como contexto
|
||||
- **RAG con embeddings:** `sentence-transformers` → ChromaDB/FAISS → recuperación semántica
|
||||
- **Esfuerzo:** RAG simple 4-5 h, con embeddings 8-10 h.
|
||||
|
||||
### 14. Extracción de entidades (NER)
|
||||
|
||||
```bash
|
||||
yt-scraper analyze --entities
|
||||
```
|
||||
|
||||
**Qué hace:** detecta personas, marcas, lugares, videojuegos mencionados a lo largo de todos los vídeos.
|
||||
|
||||
**Output:** tabla `entities(video_id, type, name, count)` → "Pokémon mencionado en 8 vídeos", "Sonic en 5 vídeos".
|
||||
|
||||
**Implementación:** `spaCy` (modelo `es_core_news_sm`) o LLM con prompt estructurado.
|
||||
- **Esfuerzo:** 4-6 h.
|
||||
|
||||
### 15. Speaker diarization — Separación de hablantes
|
||||
|
||||
**Qué hace:** distingue "narrador" de "entrevistado" o "invitado" en la transcripción.
|
||||
|
||||
**Implementación:** el audit del Web Clipper ya documenta esto (`groupBySpeaker` en `YoutubeExtractor`). Se puede portar a Python o usar `pyannote-audio`.
|
||||
- **Esfuerzo:** port del algoritmo 2-3 h; con modelo de audio 8+ h.
|
||||
|
||||
---
|
||||
|
||||
## Matriz de prioridades
|
||||
|
||||
| # | Feature | Valor | Esfuerzo | ROI |
|
||||
|---|---|---|---|---|
|
||||
| 1 | `search` (FTS5) | 🔥🔥🔥 | 1-2 h | ⭐⭐⭐ |
|
||||
| 2 | `stats` | 🔥🔥 | 1 h | ⭐⭐⭐ |
|
||||
| 3 | `export json/srt/html` | 🔥🔥 | 2-3 h | ⭐⭐⭐ |
|
||||
| 4 | `clip` | 🔥 | 30 min | ⭐⭐ |
|
||||
| 5 | `audio` download | 🔥🔥 | 1 h | ⭐⭐⭐ |
|
||||
| 6 | `re-render` | 🔥 | 1-2 h | ⭐⭐ |
|
||||
| 7 | Multi-canal | 🔥🔥 | 3-4 h | ⭐⭐ |
|
||||
| 8 | `watch` mode | 🔥🔥 | 2-3 h | ⭐⭐ |
|
||||
| 9 | Análisis (wordcloud) | 🔥 | 4-5 h | ⭐ |
|
||||
| 10 | `thumbnails` | 🔥 | 30 min | ⭐⭐⭐ |
|
||||
| 11 | Traducción | 🔥🔥 | 3-4 h | ⭐⭐ |
|
||||
| 12 | LLM summaries | 🔥🔥🔥 | 4-6 h | ⭐⭐⭐ |
|
||||
| 13 | Q&A / RAG | 🔥🔥🔥 | 4-10 h | ⭐⭐ |
|
||||
| 14 | NER (entidades) | 🔥 | 4-6 h | ⭐ |
|
||||
| 15 | Speaker diarization | 🔥 | 2-8 h | ⭐ |
|
||||
|
||||
---
|
||||
|
||||
## Dependencias Python adicionales por feature
|
||||
|
||||
```bash
|
||||
# Tier 1 — sin dependencias nuevas (solo stdlib)
|
||||
# search, stats, export, clip, re-render usan SQLite + Jinja2 ya instalados
|
||||
|
||||
# Tier 2
|
||||
pip install matplotlib wordcloud # análisis / wordcloud
|
||||
pip install deep-translator # traducción gratuita
|
||||
|
||||
# Tier 3
|
||||
pip install openai # LLM summaries / Q&A
|
||||
pip install sentence-transformers # embeddings para RAG semántico
|
||||
pip install spacy && python -m spacy download es_core_news_sm # NER
|
||||
pip install pyannote.audio # speaker diarization (requiere GPU idealmente)
|
||||
|
||||
# Audio download
|
||||
# Requiere ffmpeg instalado en el sistema:
|
||||
# winget install ffmpeg
|
||||
# o: choco install ffmpeg
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Comandos compuestos propuestos (pipelines)
|
||||
|
||||
```bash
|
||||
# Pipeline completo: scrapear + buscar + extraer clip
|
||||
yt-scraper --channel @nostal-vlad scrape
|
||||
yt-scraper search "frutiger aero"
|
||||
yt-scraper clip 8tMQMtIBfTE --from 02:15 --to 04:30 > clip.md
|
||||
|
||||
# Pipeline de análisis
|
||||
yt-scraper stats > stats.txt
|
||||
yt-scraper export --format html > index.html
|
||||
yt-scraper analyze --top-words 30 > palabras.csv
|
||||
|
||||
# Pipeline LLM
|
||||
yt-scraper --channel @nostal-vlad --summarize
|
||||
yt-scraper ask "¿Qué opina Vlad sobre la cultura otaku?"
|
||||
```
|
||||
@@ -1,163 +1,212 @@
|
||||
# yt-channel-scraper
|
||||
|
||||
Scraper de canales de YouTube que extrae transcripciones, capítulos y metadatos completos, generando notas Markdown listas para Obsidian.
|
||||
Scraper local de canales de YouTube: descubre los vídeos de cada canal con `yt-dlp`, extrae metadatos y transcripciones, guarda todo en SQLite (con búsqueda full-text) y genera una nota Markdown por vídeo, lista para un vault de Obsidian. Se maneja desde el CLI `yt-scraper` o desde una webapp local (FastAPI + Alpine) con gestión de canales, jobs en vivo, lector de transcripciones y vault de cookies.
|
||||
|
||||
Basado en la ingeniería inversa de [Obsidian Web Clipper](https://obsidian.md/) — usa `yt-dlp` internamente, que implementa el mismo mecanismo InnerTube (`youtubei/v1/player` con clientes ANDROID/IOS/WEB) que la extensión audita.
|
||||
|
||||
## Instalación
|
||||
|
||||
```bash
|
||||
cd yt-channel-scraper
|
||||
pip install -e ".[dev]"
|
||||
```
|
||||
|
||||
Requiere Python ≥ 3.10.
|
||||
|
||||
## Uso
|
||||
|
||||
### Scrape completo de un canal
|
||||
|
||||
```bash
|
||||
yt-scraper --channel "https://www.youtube.com/@Nostal-Vlad/videos"
|
||||
```
|
||||
|
||||
### Dry run (ver qué descubriría sin descargar)
|
||||
|
||||
```bash
|
||||
yt-scraper --channel "https://www.youtube.com/@Nostal-Vlad/videos" --dry-run --limit 10
|
||||
```
|
||||
|
||||
### Solo vídeos recientes
|
||||
|
||||
```bash
|
||||
yt-scraper --channel "https://www.youtube.com/@Nostal-Vlad/videos" --since 2025-01-01
|
||||
```
|
||||
|
||||
### Reanudar tras interrupción
|
||||
|
||||
```bash
|
||||
yt-scraper --channel "https://www.youtube.com/@Nostal-Vlad/videos"
|
||||
```
|
||||
|
||||
El estado se guarda en `data/state.db` (SQLite). Los vídeos ya procesados se saltan automáticamente.
|
||||
|
||||
### Reintentar vídeos con error
|
||||
|
||||
```bash
|
||||
yt-scraper --channel "https://www.youtube.com/@Nostal-Vlad/videos" --reset-errors
|
||||
```
|
||||
|
||||
## Opciones
|
||||
|
||||
| Flag | Descripción | Default |
|
||||
|---|---|---|
|
||||
| `--channel, -c` | URL del canal | de config.yaml |
|
||||
| `--config` | Ruta al YAML de configuración | `config.yaml` |
|
||||
| `--limit N` | Procesar solo N vídeos | sin límite |
|
||||
| `--since DATE` | Solo vídeos desde YYYY-MM-DD | sin filtro |
|
||||
| `--languages, -l` | Idiomas preferidos (coma-sep) | `es,en` |
|
||||
| `--no-auto` | Ignorar subtítulos auto-generados | false |
|
||||
| `--no-shorts` | Excluir Shorts | de config |
|
||||
| `--include-shorts` | Incluir Shorts | de config |
|
||||
| `--resume/--no-resume` | Saltar procesados | `--resume` |
|
||||
| `--dry-run` | Solo discovery, no descargar | false |
|
||||
| `--reset-errors` | Reintentar vídeos con error | false |
|
||||
| `--verbose, -v` | Logging DEBUG | false |
|
||||
|
||||
## Configuración
|
||||
|
||||
Copia `config.example.yaml` a `config.yaml` y edita:
|
||||
|
||||
```yaml
|
||||
channel_url: "https://www.youtube.com/@Nostal-Vlad/videos"
|
||||
languages: ["es", "es-419", "en"]
|
||||
prefer_manual: true
|
||||
include_shorts: false
|
||||
min_duration_sec: 30
|
||||
|
||||
delay:
|
||||
min_seconds: 1.5
|
||||
max_seconds: 3.5
|
||||
```
|
||||
|
||||
## Output
|
||||
|
||||
Cada vídeo genera un archivo Markdown en `data/markdown/@canal/`:
|
||||
|
||||
```
|
||||
data/markdown/Nostal Vlad/
|
||||
├── 2026-07-26_asi-era-ser-una-adolescente-edgy-en-los-2000.md
|
||||
├── 2026-07-12_la-estetica-que-romantiza-ser-un-perdedor-losercore.md
|
||||
└── ...
|
||||
```
|
||||
|
||||
Formato de cada nota:
|
||||
|
||||
```markdown
|
||||
---
|
||||
video_id: gOUyxFwWQqA
|
||||
title: "La Estética Que ROMANTIZA ser un \"PERDEDOR\" | Losercore"
|
||||
channel: Nostal Vlad
|
||||
upload_date: 2026-07-12
|
||||
duration: 1069
|
||||
url: https://www.youtube.com/watch?v=gOUyxFwWQqA
|
||||
transcript_lang: es-orig
|
||||
transcript_src: auto
|
||||
views: 100523
|
||||
likes: 8196
|
||||
---
|
||||
|
||||
# La Estética Que ROMANTIZA ser un "PERDEDOR" | Losercore
|
||||
|
||||
> [Ver en YouTube](https://www.youtube.com/watch?v=gOUyxFwWQqA)
|
||||
|
||||
## Transcripcion
|
||||
|
||||
### Intro (00:15)
|
||||
|
||||
**00:15** · Una de las estéticas que ha cobrado más relevancia últimamente...
|
||||
**00:28** · que hacer esto. Y es que el loser core como tal es muy difuso...
|
||||
|
||||
### Losercore (01:38)
|
||||
|
||||
**01:38** · ...
|
||||
```
|
||||
|
||||
## Estados en SQLite
|
||||
|
||||
| Status | Significado |
|
||||
|---|---|
|
||||
| `pending` | Descubierto, sin procesar |
|
||||
| `done` | Transcripción extraída y Markdown generado |
|
||||
| `no_subtitles` | El vídeo no tiene subtítulos (ni manuales ni auto) |
|
||||
| `error` | Error al procesar (miembros-only, privado, bloqueo, etc.) |
|
||||
100% local: sin API keys, sin deploy remoto, sin telemetría.
|
||||
|
||||
## Arquitectura
|
||||
|
||||
```
|
||||
cli.py Orquestador (Click + Rich progress)
|
||||
├── discover.py yt-dlp --flat-playlist → lista de videoIds
|
||||
├── store.py SQLite: estado, resume, dedup
|
||||
├── extract.py yt-dlp.extract_info → metadata + subtítulos
|
||||
├── parse.py JSON3/VTT → segmentos {start, end, text}
|
||||
├── chapters.py align_chapters: capítulos ↔ segmentos
|
||||
├── render.py Jinja2 → Markdown con frontmatter YAML
|
||||
└── ratelimit.py delays aleatorios + backoff exponencial
|
||||
┌───────────────────────────────────────────────┐
|
||||
│ webapp (FastAPI) │
|
||||
│ canales · jobs SSE · lector · cookies · UI │
|
||||
└───────────────────────┬───────────────────────┘
|
||||
│
|
||||
CLI (Click) ▼
|
||||
yt-scraper ──► pipeline: discover ─► extract ─► parse ─► store ─► render
|
||||
(yt-dlp (player + (JSON3/ (SQLite (Jinja2
|
||||
flat) subtítulos) VTT) + FTS5) → .md)
|
||||
│
|
||||
data/state.db ◄───────────────────────┘
|
||||
data/markdown/<Canal>/<fecha>_<slug>.md
|
||||
```
|
||||
|
||||
- `discover` lista el canal (1 petición por página); `extract` baja metadatos + subtítulos (~2 peticiones por vídeo); `parse` convierte JSON3/VTT en segmentos; `store` persiste en SQLite con FTS5; `render` escribe el `.md` con frontmatter YAML.
|
||||
- CLI, webapp y `watch` son fachadas sobre el mismo pipeline: una feature se implementa una vez y se expone en ambas.
|
||||
|
||||
## Requisitos
|
||||
|
||||
- **Python >= 3.10** (3.12 recomendado).
|
||||
- **ffmpeg** (opcional): solo para `yt-scraper audio` y el job de audio de la webapp.
|
||||
- **Un runtime JS** (node, deno o bun; opcional pero recomendado): yt-dlp lo usa para resolver los challenges de YouTube.
|
||||
- Plataforma: Windows, macOS (incl. Apple Silicon; todas las dependencias publican wheels arm64) o Linux.
|
||||
|
||||
| Herramienta | Windows (winget) | macOS (brew) | Debian/Ubuntu (apt) |
|
||||
|---|---|---|---|
|
||||
| Python 3.12 | `winget install Python.Python.3.12` | `brew install [email protected]` | `sudo apt install python3 python3-venv` |
|
||||
| ffmpeg | `winget install Gyan.FFmpeg` | `brew install ffmpeg` | `sudo apt install ffmpeg` |
|
||||
| node | `winget install OpenJS.NodeJS.LTS` | `brew install node` | `sudo apt install nodejs` |
|
||||
|
||||
## Instalación rápida
|
||||
|
||||
Ejecuta siempre los comandos **desde la raíz del repo**: las rutas de `config.yaml` (`data/`, `templates/`, la BD) se resuelven relativas al directorio de trabajo.
|
||||
|
||||
```bash
|
||||
# Opción A: uv (recomendado, respeta uv.lock)
|
||||
uv sync --all-extras
|
||||
|
||||
# Opción B: pip
|
||||
python -m venv .venv
|
||||
source .venv/bin/activate # Windows: .venv\Scripts\activate
|
||||
pip install -e ".[dev,web,analysis]"
|
||||
```
|
||||
|
||||
Extras: `dev` (pytest, pytest-cov, httpx), `web` (fastapi, uvicorn, sse-starlette, python-multipart), `analysis` (matplotlib, wordcloud).
|
||||
|
||||
Máquina nueva desde cero: [docs/GETTING-STARTED.md](docs/GETTING-STARTED.md) (paso a paso por OS, con troubleshooting).
|
||||
|
||||
## Arranque
|
||||
|
||||
### Webapp
|
||||
|
||||
**Windows** — doble clic en `start-server.bat`:
|
||||
|
||||
- La primera vez pasa por `scripts/doctor.ps1`, que prepara el entorno si falta (Python 3.12 vía winget + `pip install -e ".[web]"`).
|
||||
- `scripts/start-server.ps1` busca un puerto libre (8000–8100), arranca uvicorn, registra el proceso en `.run/server.info` (reutilizable), comprueba `/healthz` y abre el navegador.
|
||||
- Para parar: `stop-server.bat` (mata el proceso registrado y libera el puerto).
|
||||
|
||||
**macOS / Linux**:
|
||||
|
||||
```bash
|
||||
make setup # primera vez: scripts/bootstrap.sh (python/ffmpeg/node vía brew si faltan + deps)
|
||||
make serve # scripts/start-server.sh
|
||||
make stop # scripts/stop-server.sh
|
||||
# O directamente, sin Makefile:
|
||||
scripts/bootstrap.sh && scripts/start-server.sh
|
||||
```
|
||||
|
||||
**Manual (debug)**:
|
||||
|
||||
```bash
|
||||
python -m uvicorn yt_scraper.webapp.app:app --host 127.0.0.1 --port 8000
|
||||
```
|
||||
|
||||
### CLI
|
||||
|
||||
El entry point es `yt-scraper`. OJO: sin subcomando ejecuta un scrape completo del canal de `config.yaml`. Los flags de scrape van en el subcomando:
|
||||
|
||||
```bash
|
||||
yt-scraper scrape # scrape del canal de config.yaml (incremental)
|
||||
yt-scraper -c "https://www.youtube.com/@Nostal-Vlad/videos" scrape # otro canal sin tocar config
|
||||
|
||||
# Discovery sin descargar nada (ver qué encontraría)
|
||||
yt-scraper scrape --dry-run --limit 10
|
||||
|
||||
# Solo vídeos desde una fecha / limitar la tirada
|
||||
yt-scraper scrape --since 2025-01-01 --limit 5
|
||||
|
||||
# Reintentar vídeos en error/no_subtitles y continuar
|
||||
yt-scraper scrape --reset-errors
|
||||
|
||||
# Recorrer el canal entero (ignora la ventana incremental)
|
||||
yt-scraper scrape --full
|
||||
```
|
||||
|
||||
Flags del grupo raíz (aplican a todo): `--config`, `--channel/-c`, `--all-channels`, `--cookies`, `--cookies-from-browser`, `--verbose/-v`.
|
||||
|
||||
**Canales** (multi-canal):
|
||||
|
||||
```bash
|
||||
yt-scraper channels add "https://www.youtube.com/@Fazt"
|
||||
yt-scraper channels list
|
||||
yt-scraper channels remove "@Fazt"
|
||||
yt-scraper --all-channels scrape # aplica el subcomando a todos los canales
|
||||
```
|
||||
|
||||
**Búsqueda y export** (sobre lo ya scrapeado, sin tocar la red):
|
||||
|
||||
```bash
|
||||
yt-scraper search "vaporwave" # FTS5 en transcripciones, con timestamp
|
||||
yt-scraper -c "@Fazt" search "typescript" -l 50
|
||||
yt-scraper export --format json # json | csv | srt | html → data/exports/
|
||||
yt-scraper --all-channels export --format srt
|
||||
```
|
||||
|
||||
**Recuperación** (DB ↔ disco):
|
||||
|
||||
```bash
|
||||
yt-scraper reset # error/no_subtitles → pending (pregunta antes)
|
||||
yt-scraper reset --status error --yes
|
||||
yt-scraper reconcile # re-escanea data/markdown y cuadra la DB con el disco
|
||||
yt-scraper reconcile --prune # además devuelve a pending los done sin .md
|
||||
yt-scraper re-render # regenera .md desde segmentos guardados (0 descargas)
|
||||
yt-scraper re-render --backfill # antes importa segmentos desde los .md existentes
|
||||
```
|
||||
|
||||
**Audio, monitoreo y análisis**:
|
||||
|
||||
```bash
|
||||
yt-scraper audio --limit 5 # MP3 de los done (requiere ffmpeg) → data/audio/
|
||||
yt-scraper watch --interval 6h # discovery periódico y procesa lo nuevo
|
||||
yt-scraper watch --once # una sola pasada
|
||||
yt-scraper analyze --top-words 50 # frecuencia → data/analysis/top_words.{csv,png}
|
||||
yt-scraper analyze --wordcloud # + data/analysis/wordcloud.png
|
||||
yt-scraper analyze --timeline "ia" # menciones de un término por mes
|
||||
```
|
||||
|
||||
## Configuración
|
||||
|
||||
Copia la plantilla y edita:
|
||||
|
||||
```bash
|
||||
cp config.example.yaml config.yaml # Windows: copy config.example.yaml config.yaml
|
||||
```
|
||||
|
||||
`config.yaml` es privado y está gitignored; la plantilla versionada (`config.example.yaml`) documenta todas las claves. Bloques principales:
|
||||
|
||||
| Bloque | Qué controla |
|
||||
|---|---|
|
||||
| raíz (`channel_url`, `languages`, `include_shorts`, ...) | canal por defecto, idiomas de subtítulos (modo `manual`/`auto`/`any` por idioma), filtros de shorts/live/duración |
|
||||
| `delay` | pacing: pausas entre vídeos, backoff ante rate-limit, circuit breaker, intervalo mínimo global entre peticiones |
|
||||
| `yt_dlp` | reintentos, pausa entre sub-peticiones, timeouts de yt-dlp |
|
||||
| `sync` | discovery incremental: ventana inicial, máximo y solapamiento que da la sincronización por alcanzada |
|
||||
| `database_path`, `output_dir`, `template_path`, `filename_template` | rutas (relativas al CWD) y patrón del nombre de las notas |
|
||||
|
||||
Referencia completa de cada clave con tipos y defaults: [docs/CONFIG.md](docs/CONFIG.md).
|
||||
|
||||
## Cookies
|
||||
|
||||
Las cookies de sesión de YouTube (formato Netscape) sirven para acceder a vídeos de membresía y reducir los bot-checks. El vault (`cookies/`) admite varias, con **exactamente una activa**; se pueden importar:
|
||||
|
||||
- **Webapp**: arrastrar y soltar un `cookies.txt`, o importar directamente desde el navegador local (Brave).
|
||||
- **CLI**: `yt-scraper --cookies ruta/cookies.txt scrape` o `--cookies-from-browser brave` (chrome|firefox|edge|brave).
|
||||
|
||||
Cómo exportarlas desde tu navegador y cómo rotarlas: [docs/COOKIES.md](docs/COOKIES.md). **Son equivalentes a tu contraseña: nunca las commitees** (están en `.gitignore`).
|
||||
|
||||
## Datos generados
|
||||
|
||||
```
|
||||
data/
|
||||
├── state.db # SQLite (WAL): vídeos, canales, segmentos, FTS5, jobs, cookies_meta
|
||||
├── markdown/<Canal>/ # una nota .md por vídeo: <fecha>_<slug>.md con frontmatter YAML
|
||||
├── thumbnails/<id>.jpg # miniaturas
|
||||
├── avatars/ # avatares de canal
|
||||
├── audio/ # MP3 descargados (webapp: <video_id>.mp3; CLI: por título)
|
||||
├── videos/<id>/ # descargas de vídeo completas (.webm)
|
||||
├── exports/ # export json/csv/srt/html
|
||||
└── analysis/ # wordcloud, top_words, timeline
|
||||
cookies/ # vault de cookies Netscape (gitignored, SENSIBLE)
|
||||
.run/ # estado del servidor (puerto, PID)
|
||||
```
|
||||
|
||||
Estados de un vídeo en la BD: `pending` (descubierto, sin procesar), `done` (transcripción + `.md`), `no_subtitles` (sin pista válida según la política de idiomas), `error` (fallo: privado, bloqueo, ...).
|
||||
|
||||
`data/`, `cookies/`, `config.yaml` y `.run/` están gitignored: contienen tu biblioteca y tu sesión. No los commitees.
|
||||
|
||||
## Tests
|
||||
|
||||
```bash
|
||||
python -m pytest tests/ -v
|
||||
uv run python -m pytest -q # o: python -m pytest -q (dentro del venv)
|
||||
```
|
||||
|
||||
## Mantenimiento
|
||||
Suite hermética (~240 tests en 24 archivos, sin red): todo con `tmp_path` + `monkeypatch`; `yt_dlp.YoutubeDL` se sustituye por un fake.
|
||||
|
||||
Si la extracción falla tras una actualización de YouTube:
|
||||
## Migrar el repo a otra máquina
|
||||
|
||||
```bash
|
||||
yt-dlp -U # actualizar yt-dlp
|
||||
pip install -U yt-dlp
|
||||
```
|
||||
- Las rutas se resuelven **relativas al CWD**: sigue ejecutando todo desde la raíz del repo.
|
||||
- Si mueves o clonas el repo, la instalación editable apunta a la ruta vieja: re-ejecuta `uv sync --all-extras` (o `pip install -e ".[dev,web,analysis]"`) y recrea el venv si hace falta.
|
||||
- `config.yaml` y `cookies/` no viajan en el clon: recréalos desde `config.example.yaml` y re-exporta las cookies.
|
||||
|
||||
Las versiones de cliente InnerTube (ANDROID `20.10.x`, IOS `20.10.x`, WEB `2.2024xxxx`) rotan mensualmente. `yt-dlp` las mantiene actualizadas.
|
||||
## Licencia
|
||||
|
||||
TBD.
|
||||
|
||||
+51
-5
@@ -1,20 +1,66 @@
|
||||
channel_url: "https://www.youtube.com/@Nostal-Vlad/videos"
|
||||
|
||||
languages: ["es", "es-419", "en"]
|
||||
prefer_manual: true
|
||||
# Per-language subtitle preference. Either a list (legacy) or a dict.
|
||||
# Dict form (preferred): per-lang mode. Modes are
|
||||
# "manual" - ONLY human-uploaded captions; a video with just auto captions is
|
||||
# skipped and stored as `no_subtitles`
|
||||
# "auto" - ONLY YouTube-auto-generated captions
|
||||
# "any" - try manual first (per `prefer_manual`), fall back to auto
|
||||
# List form (legacy): `languages: ["es", "en"]` maps EVERY entry to the hard
|
||||
# "manual"/"auto" mode picked by `prefer_manual` - note this means manual-ONLY
|
||||
# when prefer_manual is true, which yields nothing on channels that publish
|
||||
# only auto-generated captions. Prefer the dict form below.
|
||||
languages:
|
||||
es: any
|
||||
es-419: any
|
||||
en: any
|
||||
prefer_manual: true # with mode "any": try manual first, then auto
|
||||
include_shorts: false
|
||||
include_live: true
|
||||
min_duration_sec: 30
|
||||
|
||||
# Pacing and what to do when YouTube starts refusing.
|
||||
#
|
||||
# YouTube publishes no rate limits, so none of these numbers are official. The
|
||||
# backoff shape is: yt-dlp's own docs put a guest session at roughly 300 videos
|
||||
# an hour (~1000 requests) before the bot wall, and Google documents truncated
|
||||
# exponential backoff with jitter for its own APIs, which is what backoff_base
|
||||
# and backoff_cap feed.
|
||||
delay:
|
||||
min_seconds: 1.5
|
||||
min_seconds: 1.5 # randomised gap between videos / channels
|
||||
max_seconds: 3.5
|
||||
backoff_base: 2.0
|
||||
backoff_base: 2.0 # after a rate-limit: min(base * 2**n + jitter, cap)
|
||||
backoff_cap: 60.0
|
||||
throttle_threshold: 3 # consecutive rate-limit responses that abort the run
|
||||
# Floor on the gap between ANY two requests to YouTube from this process,
|
||||
# including the /api/tools/* endpoints that run outside the job runner.
|
||||
# 0 = off. Raise it if you still hit the bot wall; it is the one setting that
|
||||
# applies everywhere at once.
|
||||
min_request_interval: 0.0
|
||||
audio_rate_limit: 0 # bytes/sec ceiling for audio downloads, 0 = unlimited
|
||||
|
||||
yt_dlp:
|
||||
retries: 10
|
||||
retries: 10 # download retries (only bites on the audio path)
|
||||
# Seconds between the individual HTTP calls inside one extraction — the watch
|
||||
# page, the InnerTube player call, each continuation page of a channel tab.
|
||||
# Forwarded to yt-dlp as `sleep_interval_requests`.
|
||||
sleep_subrequests: 2
|
||||
extractor_retries: 3 # retries during extraction (5xx and network only:
|
||||
# yt-dlp's YouTube extractor never retries 403/429)
|
||||
socket_timeout: 30.0 # without this a hung connection blocks the worker
|
||||
|
||||
# Incremental channel sync. Re-scanning a tracked channel reads the /videos tab
|
||||
# newest-first and stops once it sees `overlap` videos already in the DB, so a
|
||||
# routine sync costs one page instead of paginating the whole channel. If every
|
||||
# video in the window turns out to be new, the window doubles (up to
|
||||
# max_window) rather than silently missing uploads.
|
||||
# Set incremental: false — or pass --full / tick "Full channel rescan" — to walk
|
||||
# every page, which is only needed when the local catalog is incomplete.
|
||||
sync:
|
||||
incremental: true
|
||||
window: 30 # entries read on the first pass
|
||||
max_window: 300 # ceiling before reporting a truncated scan
|
||||
overlap: 3 # consecutive known videos that prove we caught up
|
||||
|
||||
database_path: "data/state.db"
|
||||
output_dir: "data/markdown"
|
||||
|
||||
@@ -0,0 +1,2 @@
|
||||
# data/ contiene la BD SQLite, notas markdown, miniaturas, audio y video de los canales.
|
||||
# Todo su contenido es privado y esta ignorado por git (ver .gitignore).
|
||||
+110
@@ -0,0 +1,110 @@
|
||||
# Referencia de configuración (`config.yaml`)
|
||||
|
||||
`config.yaml` es tu configuración privada (gitignored). La plantilla versionada `config.example.yaml` documenta todas las claves y es el punto de partida (`cp config.example.yaml config.yaml`). Todas las rutas del archivo se resuelven **relativas al CWD** — ejecuta siempre desde la raíz del repo.
|
||||
|
||||
Los defaults que siguen son los de `config.example.yaml` (entre paréntesis, cuando difiere, el default del código en `src/yt_scraper/config.py`).
|
||||
|
||||
## Bloque raíz — canal, idiomas y filtros
|
||||
|
||||
| Clave | Tipo | Default | Qué afecta |
|
||||
|---|---|---|---|
|
||||
| `channel_url` | str | `"https://www.youtube.com/@Nostal-Vlad/videos"` | Canal por defecto de `yt-scraper scrape` / webapp. Se ignora si pasas `--channel`/`-c`. Acepta URL completa, `/channel/UC...` o handle |
|
||||
| `languages` | dict \| list | `{es: any, es-419: any, en: any}` | Idiomas de subtítulos aceptables y su modo (ver abajo). Forma legacy: lista `["es","en"]` — todos con el modo de `prefer_manual`; evítala, un canal solo-auto quedaría en `no_subtitles` |
|
||||
| `prefer_manual` | bool | `true` | En modo `any`: prueba subtítulos manuales antes que los automáticos |
|
||||
| `include_shorts` | bool | `false` | Descubrir/procesar Shorts. En CLI: `--include-shorts` / `--no-shorts` |
|
||||
| `include_live` | bool | `true` | Procesar directos. En CLI: `--no-live` |
|
||||
| `min_duration_sec` | int | `30` (código: `0`) | Duración mínima del vídeo en segundos; filtra en discovery |
|
||||
|
||||
Modos por idioma de `languages`:
|
||||
|
||||
| Valor | Significado |
|
||||
|---|---|
|
||||
| `manual` | **Solo** subtítulos subidos por el creador; si solo hay automáticos → `no_subtitles` |
|
||||
| `auto` | **Solo** subtítulos automáticos (ASR) |
|
||||
| `any` | Manuales primero (según `prefer_manual`), automáticos como fallback. **Recomendado** |
|
||||
|
||||
Nota: `languages` es una preferencia entre idiomas que sabes leer, no una orden de traducir. El idioma hablado del vídeo siempre gana si está en la lista; nunca se guarda una traducción automática si existe el transcript original.
|
||||
|
||||
## Bloque `delay` — pacing y rate-limit
|
||||
|
||||
YouTube no publica límites; el techo práctico conocido (wiki de yt-dlp) es ~300 vídeos/hora en sesión sin cuenta. Estos valores apuntan por debajo. Con los defaults conservadores, una tirada real cuesta **del orden de 12–13 s por vídeo** (~270 vídeos/hora).
|
||||
|
||||
| Clave | Tipo | Default | Qué afecta |
|
||||
|---|---|---|---|
|
||||
| `min_seconds` | float | `1.5` | Pausa aleatoria mínima entre vídeos (el gap real se sortea entre `min` y `max`) |
|
||||
| `max_seconds` | float | `3.5` | Pausa aleatoria máxima entre vídeos |
|
||||
| `backoff_base` | float | `2.0` | Backoff tras un rate-limit: `min(base * 2**n + jitter, cap)`, con n = fallos consecutivos |
|
||||
| `backoff_cap` | float | `60.0` | Techo del backoff en segundos |
|
||||
| `throttle_threshold` | int | `3` | Rate-limits **consecutivos** que hacen parar la tirada (circuit breaker). Lo no procesado queda `pending` y es recuperable |
|
||||
| `min_request_interval` | float | `0.0` (example) | Separación mínima entre **cualquier** dos peticiones a YouTube del proceso (Pacer global, incluye los `/api/tools/*` fuera del job runner). `0` = desactivado. Es la única palanca que actúa en todas partes a la vez |
|
||||
| `audio_rate_limit` | int | `0` | Techo de bytes/segundo para descargas de audio (ratelimit de yt-dlp). `0` = ilimitado |
|
||||
|
||||
Referencia práctica: la config real del autor usa `min_request_interval: 2.5` para ser aún más conservador cuando la webapp y el CLI conviven.
|
||||
|
||||
## Bloque `yt_dlp` — reintentos y timeouts
|
||||
|
||||
| Clave | Tipo | Default | Qué afecta |
|
||||
|---|---|---|---|
|
||||
| `retries` | int | `10` | Reintentos de descarga; solo muerde en el camino de audio (el extractor de YouTube de yt-dlp no reintenta 403/429) |
|
||||
| `sleep_subrequests` | float | `2` | Segundos entre las peticiones HTTP **dentro** de una extracción (watch page, llamada player, continuations). Se reenvía a yt-dlp como `sleep_interval_requests` (nombre interno del proyecto; el yt-dlp no tiene ninguna opción llamada `sleep_subrequests`) |
|
||||
| `extractor_retries` | int | `3` | Reintentos durante la extracción (solo 5xx y red) |
|
||||
| `socket_timeout` | float | `30.0` | Timeout de socket; sin él una conexión colgada bloquea el worker para siempre |
|
||||
|
||||
## Bloque `sync` — discovery incremental
|
||||
|
||||
Re-escanear un canal trackeado lee la pestaña `/videos` (cronológica inversa) y corta al ver `overlap` vídeos ya conocidos: un sync rutinario cuesta 1 página, no el canal entero.
|
||||
|
||||
| Clave | Tipo | Default | Qué afecta |
|
||||
|---|---|---|---|
|
||||
| `incremental` | bool | `true` | Activar el modo ventana. `false` (o `--full` / "Full rescan" en la webapp) recorre todo el canal |
|
||||
| `window` | int | `30` | Entradas leídas en la primera pasada |
|
||||
| `max_window` | int | `300` | Techo: si TODA la ventana resulta nueva, se duplica hasta aquí antes de declarar el scan truncado (aviso "usa `--full`") |
|
||||
| `overlap` | int | `3` | Vídeos conocidos consecutivos que dan la sincronización por alcanzada |
|
||||
|
||||
## Rutas y plantillas
|
||||
|
||||
| Clave | Tipo | Default | Qué afecta |
|
||||
|---|---|---|---|
|
||||
| `database_path` | str | `"data/state.db"` | Ruta de la BD SQLite (WAL). Relativa al CWD |
|
||||
| `output_dir` | str | `"data/markdown"` | Directorio de notas; el resto de datos (`audio/`, `thumbnails/`, ...) se deriva de su padre |
|
||||
| `template_path` | str | `"templates/video.md.j2"` | Plantilla Jinja2 de la nota. OJO: el formato del `.md` es un contrato bidireccional (los regex de `segments.py` lo re-lean) |
|
||||
| `filename_template` | str | `"{upload_date}_{slug}"` | Patrón del nombre de las notas. Placeholders: `{upload_date}` (fecha de subida; `unknown-date` si se desconoce), `{slug}` (slug del título, máx 60), `{title}` (título crudo), `{video_id}` (id estable de YouTube — útil porque título y fecha son inestables) |
|
||||
|
||||
## Ejemplos
|
||||
|
||||
**Mínima** (defaults razonables, primer contacto):
|
||||
|
||||
```yaml
|
||||
channel_url: "https://www.youtube.com/@MiCanal/videos"
|
||||
languages:
|
||||
es: any
|
||||
```
|
||||
|
||||
**Conservadora** (muchos vídeos, webapp y CLI a la vez, o historial de rate-limits):
|
||||
|
||||
```yaml
|
||||
channel_url: "https://www.youtube.com/@MiCanal/videos"
|
||||
languages:
|
||||
es: any
|
||||
es-419: any
|
||||
en: any
|
||||
prefer_manual: true
|
||||
include_shorts: false
|
||||
min_duration_sec: 30
|
||||
|
||||
delay:
|
||||
min_seconds: 1.5
|
||||
max_seconds: 3.5
|
||||
backoff_base: 2.0
|
||||
backoff_cap: 60.0
|
||||
throttle_threshold: 3
|
||||
min_request_interval: 2.5 # pacer global: la palanca más eficaz contra el bot-wall
|
||||
|
||||
sync:
|
||||
incremental: true
|
||||
window: 30
|
||||
max_window: 300
|
||||
overlap: 3
|
||||
```
|
||||
|
||||
Si cambias los tiempos, vuelve a medir sobre tu canal: la constante es ~3 peticiones a `youtube.com` por vídeo y de ahí sale todo lo demás.
|
||||
@@ -0,0 +1,68 @@
|
||||
# Guía de cookies
|
||||
|
||||
Las cookies son **tu sesión de YouTube**. Este scraper las usa (vía yt-dlp) para presentarse con esa sesión en cada petición.
|
||||
|
||||
## Por qué hacen falta
|
||||
|
||||
- **Vídeos de membresía** (`subscriber_only`): sin sesión es imposible descargarlos; con ella, si estás suscrito al canal, entran como cualquier otro.
|
||||
- **Menos bot-checks / rate-limit**: una sesión autenticada aguanta bastante más tráfico que una anónima antes de toparse con el muro (~300 vídeos/hora en sesión de invitado según la wiki de yt-dlp).
|
||||
- **Challenges y PO tokens**: parte de los retos anti-bot que YouTube sirve se suavizan cuando la petición viaja con cookies de una sesión real.
|
||||
|
||||
Sin cookies el scraper funciona igual; simplemente verás más `error` por throttling y los vídeos de membresía quedan bloqueados.
|
||||
|
||||
## Qué es un `cookies.txt` (formato Netscape)
|
||||
|
||||
Un archivo de texto plano con una cookie por línea, tabulador-separado:
|
||||
|
||||
```
|
||||
.youtube.com\tTRUE\t/\tTRUE\t1798761600\tSID\t"value"
|
||||
.youtube.com\tTRUE\t/\tTRUE\t0\t__Secure-3PSID\t"value"
|
||||
```
|
||||
|
||||
Detalles que importan aquí:
|
||||
|
||||
- Las líneas `#HttpOnly_...` **son datos**, no comentarios: las cookies de login (SID, HSID, ...) son HttpOnly y un export que las omita no sirve.
|
||||
- **Sesión válida** (criterio que aplica el vault al importar): el archivo contiene las tres `SID` + `HSID` + `SSID`, **o** al menos `LOGIN_INFO`. Las que debe traer un export correcto: `SID`, `HSID`, `SSID`, `SAPISID`, `LOGIN_INFO` (y normalmente también `APISID`, `__Secure-3PSID`, ...).
|
||||
- Un archivo con solo `__Secure-3PSID` (export parcial de algunas extensiones) **no es una sesión**: YouTube lo trata como anónimo y la membresía sigue bloqueada.
|
||||
|
||||
## Cómo exportarlas (paso a paso)
|
||||
|
||||
1. Inicia sesión en [youtube.com](https://www.youtube.com) en tu navegador (Chrome, Brave o Firefox).
|
||||
2. Instala la extensión **"Get cookies.txt LOCALLY"** (Chrome Web Store / Firefox Add-ons). La palabra LOCALLY importa: exporta en tu máquina sin mandar nada a un servidor.
|
||||
3. Con youtube.com abierto, abre la extensión y exporta las cookies de **youtube.com** (formato Netscape por defecto).
|
||||
4. Guarda el archivo (`cookies.txt`). Verifica que aparecen `SID`, `HSID`, `SSID`, `SAPISID` y `LOGIN_INFO`.
|
||||
|
||||
## Cómo importarlas
|
||||
|
||||
### Webapp (recomendado)
|
||||
|
||||
Sección **Cookies** de la UI:
|
||||
|
||||
- **Arrastrar y soltar** el `cookies.txt` (o selección manual). El vault lo copia a `cookies/<uuid>.txt`, analiza sesión/caducidad y lo registra con etiqueta.
|
||||
- **Importar desde navegador**: lee las cookies directamente del navegador local (Brave por defecto). **El navegador debe estar cerrado por completo** — Chromium bloquea el archivo de cookies si el proceso vive, y el error que verás es "cookie store locked". Al importar así, la activación es automática.
|
||||
- Si no hay ninguna activa, la primera que subas se activa sola.
|
||||
|
||||
### CLI
|
||||
|
||||
```bash
|
||||
yt-scraper --cookies ruta/a/cookies.txt scrape # usar un archivo concreto
|
||||
yt-scraper --cookies-from-browser brave scrape # chrome|firefox|edge|brave
|
||||
```
|
||||
|
||||
Además, al arrancar el CLI o la webapp, cualquier `.txt` suelto en `cookies/` se adopta automáticamente al vault (`auto_import_dir`, idempotente).
|
||||
|
||||
## Activación
|
||||
|
||||
Hay **exactamente una cookie activa** a la vez: CLI, webapp y watch usan esa si no se pasa `--cookies`. En la webapp puedes cambiar la activa (botón *activate*), probarla (*test* comprueba que la sesión sigue viva) y borrar las demás.
|
||||
|
||||
## Rotación y caducidad
|
||||
|
||||
- Las cookies de login **caducan** (meses) o se invalidan si cierras sesión / cambias contraseña en ese navegador. Cuando el scrape vuelva a ver bloqueos de membresía o un chorreo de rate-limits, re-exporta y sube un archivo nuevo.
|
||||
- El vault marca el estado `expired` según la fecha de expiración del propio archivo; el criterio `has_session` (SID+HSID+SSID o LOGIN_INFO) se comprueba en la importación y en el listado.
|
||||
- No pasa nada por tener varias en el vault: solo la activa se usa.
|
||||
|
||||
## Seguridad
|
||||
|
||||
- **Son equivalentes a tu contraseña de Google para YouTube.** Quien tenga el archivo puede usar tu sesión.
|
||||
- `cookies/` está en `.gitignore` — nunca las commitees, ni las pegues en un chat, ni las subas a ningún sitio.
|
||||
- Si sospechas una fuga: cierra la sesión de YouTube en ese navegador (invalida las cookies) y re-exporta.
|
||||
@@ -0,0 +1,159 @@
|
||||
# Getting Started — máquina nueva desde cero
|
||||
|
||||
De un equipo vacío a la webapp corriendo y el primer scrape hecho. Una sección por OS; salta a la tuya. Todo lo demás (arquitectura, CLI completo, config) está en el [README](../README.md).
|
||||
|
||||
Regla transversal: **todos los comandos se ejecutan desde la raíz del repo** (`yt-channel-scraper/`), porque las rutas de `config.yaml` se resuelven relativas al directorio de trabajo.
|
||||
|
||||
## Windows
|
||||
|
||||
### 1. Python
|
||||
|
||||
PowerShell:
|
||||
|
||||
```powershell
|
||||
winget install Python.Python.3.12
|
||||
# cierra y reabre la terminal para que entre en PATH; verifica:
|
||||
python --version
|
||||
```
|
||||
|
||||
### 2. Obtener el código
|
||||
|
||||
```powershell
|
||||
git clone <URL-del-gitea> yt-channel-scraper
|
||||
cd yt-channel-scraper
|
||||
```
|
||||
|
||||
### 3. Entorno + dependencias
|
||||
|
||||
```powershell
|
||||
# Opción A: uv (recomendado; respeta uv.lock)
|
||||
winget install astral-sh.uv
|
||||
uv sync --all-extras
|
||||
|
||||
# Opción B: venv + pip
|
||||
python -m venv .venv
|
||||
.venv\Scripts\activate
|
||||
pip install -e ".[dev,web,analysis]"
|
||||
```
|
||||
|
||||
Extras: `dev` (tests), `web` (webapp), `analysis` (gráficos). Puedes instalar solo los que uses.
|
||||
|
||||
### 4. ffmpeg y node (opcionales)
|
||||
|
||||
- **ffmpeg**: solo necesario para `yt-scraper audio` / job de audio → `winget install Gyan.FFmpeg`.
|
||||
- **node**: yt-dlp lo usa para resolver challenges de YouTube → `winget install OpenJS.NodeJS.LTS`.
|
||||
|
||||
Reabre la terminal tras instalar y verifica con `ffmpeg -version` / `node -v`.
|
||||
|
||||
### 5. Configuración
|
||||
|
||||
```powershell
|
||||
copy config.example.yaml config.yaml
|
||||
notepad config.yaml
|
||||
```
|
||||
|
||||
Cambia al menos `channel_url` (URL `/videos` de tu canal). Referencia de todas las claves: [CONFIG.md](CONFIG.md).
|
||||
|
||||
### 6. Cookies (recomendado)
|
||||
|
||||
Sin cookies funciona, pero con sesión de YouTube reduce los bot-checks y desbloquea vídeos de membresía. Resumen: exporta `cookies.txt` desde tu navegador con la extensión "Get cookies.txt LOCALLY" (logueado a youtube.com) y arrástralo a la webapp, o usa `--cookies`. Guía completa: [COOKIES.md](COOKIES.md).
|
||||
|
||||
### 7. Primer scrape (sin descargar nada)
|
||||
|
||||
```powershell
|
||||
.venv\Scripts\activate # si usaste venv; con uv: uv run yt-scraper ...
|
||||
yt-scraper scrape --dry-run --limit 10
|
||||
yt-scraper scrape --limit 3 # primeras 3 notas en data/markdown/<Canal>/
|
||||
```
|
||||
|
||||
### 8. Arrancar la webapp
|
||||
|
||||
Doble clic en `start-server.bat`. La primera vez pasa por `scripts/doctor.ps1`, que instala lo que falte (Python 3.12 vía winget + `pip install -e ".[web]"`) y delega en `scripts/start-server.ps1`: busca puerto libre (8000–8100), arranca uvicorn, registra el PID en `.run/server.info`, comprueba `/healthz` y abre el navegador. Para parar: `stop-server.bat`.
|
||||
|
||||
Manual (debug): `python -m uvicorn yt_scraper.webapp.app:app --host 127.0.0.1 --port 8000`.
|
||||
|
||||
## macOS (Apple Silicon)
|
||||
|
||||
### 1. Homebrew + Python
|
||||
|
||||
```bash
|
||||
/bin/bash -c "$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)"
|
||||
brew install [email protected] ffmpeg node
|
||||
```
|
||||
|
||||
Todas las dependencias Python publican wheels arm64; no hay compilación.
|
||||
|
||||
### 2-3. Código y entorno
|
||||
|
||||
```bash
|
||||
git clone <URL-del-gitea> yt-channel-scraper
|
||||
cd yt-channel-scraper
|
||||
|
||||
# Opción A: uv
|
||||
brew install astral-sh.uv && uv sync --all-extras
|
||||
|
||||
# Opción B: venv + pip
|
||||
python3.12 -m venv .venv && source .venv/bin/activate
|
||||
pip install -e ".[dev,web,analysis]"
|
||||
```
|
||||
|
||||
### 4. ffmpeg / node
|
||||
|
||||
Ya instalados en el paso 1 (si los saltaste: `brew install ffmpeg node`).
|
||||
|
||||
### 5-6. Config y cookies
|
||||
|
||||
```bash
|
||||
cp config.example.yaml config.yaml
|
||||
$EDITOR config.yaml # cambia channel_url
|
||||
```
|
||||
|
||||
Cookies: [COOKIES.md](COOKIES.md).
|
||||
|
||||
### 7. Primer scrape
|
||||
|
||||
```bash
|
||||
yt-scraper scrape --dry-run --limit 10
|
||||
yt-scraper scrape --limit 3
|
||||
```
|
||||
|
||||
### 8. Arrancar la webapp
|
||||
|
||||
```bash
|
||||
make setup # primera vez: scripts/bootstrap.sh prepara todo lo que falte
|
||||
make serve # arranca uvicorn y abre el navegador
|
||||
make stop # para el servidor
|
||||
```
|
||||
|
||||
Equivalente sin Makefile: `scripts/bootstrap.sh && scripts/start-server.sh` / `scripts/stop-server.sh`.
|
||||
|
||||
## Linux (Debian/Ubuntu)
|
||||
|
||||
```bash
|
||||
sudo apt update
|
||||
sudo apt install -y python3 python3-venv python3-pip git ffmpeg nodejs # nodejs si el repo lo empaqueta; si no, usa nodesource
|
||||
```
|
||||
|
||||
El resto es idéntico a macOS: `git clone` → `uv sync --all-extras` (o venv + `pip install -e ".[dev,web,analysis]"`) → `cp config.example.yaml config.yaml` → `yt-scraper scrape --dry-run` → `make serve`. En distros sin `make`: `scripts/bootstrap.sh` + `scripts/start-server.sh`.
|
||||
|
||||
Si tu distro trae Python < 3.10, instala uno moderno (pyenv, deadsnakes PPA en Ubuntu, o `brew` en Linux) — el paquete exige `>=3.10`.
|
||||
|
||||
## Verificación final
|
||||
|
||||
```bash
|
||||
uv run python -m pytest -q # debe pasar la suite completa (~240 tests, sin red)
|
||||
yt-scraper channels list # tras el primer scrape: tu canal con sus vídeos
|
||||
```
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
| Síntoma | Causa | Arreglo |
|
||||
|---|---|---|
|
||||
| `ModuleNotFoundError: yt_scraper` tras mover/renombrar el repo | La instalación editable apunta a la ruta vieja | `pip install -e ".[dev,web,analysis]"` otra vez (o recrea el venv: `uv sync --all-extras`) |
|
||||
| `yt-scraper: command not found` | Venv sin activar, o instalaste sin `-e` | Activa el venv, o usa `uv run yt-scraper ...`, o `python -m yt_scraper.cli` |
|
||||
| Los paths apuntan a sitios raros (`data/` vacío en otro lado) | Comandos ejecutados fuera de la raíz del repo | `cd` a la raíz; las rutas son relativas al CWD |
|
||||
| `file cannot be loaded because running scripts is disabled` al lanzar un `.ps1` | PowerShell ExecutionPolicy restringido | `powershell -ExecutionPolicy Bypass -File scripts\start-server.ps1`, o `Set-ExecutionPolicy -Scope CurrentUser RemoteSigned` |
|
||||
| "puerto ocupado" / el navegador abre otra app | Resto de un servidor anterior en el puerto | Borra `.run/server.info` (estado obsoleto) o usa `stop-server.bat` / `make stop`; el arranque busca puerto libre 8000–8100 |
|
||||
| `uv: command not found` o trampolines rotos tras actualizar uv/Python | Los shims del venv quedaron inconsistentes | Ejecuta vía módulo: `uv run python -m pytest -q`, `python -m yt_scraper.cli` |
|
||||
| Tests de red fallan / scrape con muchos `error` seguidos | Rate-limit de YouTube (sin cookies es más frecuente) | Espera; revisa `delay` en [CONFIG.md](CONFIG.md); añade cookies ([COOKIES.md](COOKIES.md)) |
|
||||
| `ffmpeg no encontrado` en `yt-scraper audio` | ffmpeg no está en PATH | `winget install Gyan.FFmpeg` / `brew install ffmpeg` / `apt install ffmpeg`, y reabre la terminal |
|
||||
@@ -0,0 +1,122 @@
|
||||
# Auditoría comparativa · Obsidian Web Clipper 1.7.1 vs `yt-channel-scraper`
|
||||
|
||||
> **Producto auditado:** *Obsidian Web Clipper* 1.7.1 (extensión MV3) en `D:\Obsidian Web Clipper - Chrome Web Store 1.7.1.0`.
|
||||
> **Sistema actual:** `yt-channel-scraper` (yt-dlp como única interfaz con YouTube; techo práctico ~300 videos/h).
|
||||
> **Alcance:** solo auditoría, comparación y análisis de mejoras. Sin cambios de código.
|
||||
> **Base:** verificación directa del bundle `popup.js` (offsets 380236–393083, clase `YoutubeExtractor`) además del doc previo `YOUTUBE-TRANSCRIPT-AUDIT.md`.
|
||||
> **Fecha:** 2026-09-01.
|
||||
|
||||
---
|
||||
|
||||
## 0 · TL;DR
|
||||
|
||||
La extensión resuelve el mismo problema (metadatos + transcripción de YouTube sin API oficial) con una filosofía de red **opuesta y complementaria** a la del scraper:
|
||||
|
||||
| | Scraper (hoy) | Extensión |
|
||||
|---|---|---|
|
||||
| Filosofía ante el fallo | **Backoff temporal**: esperar y reintentar más tarde (Pacer + ThrottleGuard, abort limpio) | **Rotación de identidad**: cambiar de cliente/recurso y degradar, casi nunca esperar |
|
||||
| Requests por video | 3 (watch + player + timedtext), medidos irreducibles vía yt-dlp (`config.yaml:21-33`) | 1–2 (player InnerTube + timedtext); el 3º (`next`) solo si faltan capítulos |
|
||||
| Identidad | 1 cookie activa, cliente yt-dlp default, sin rotación (`cookies.py:151-163`) | 3 clientes InnerTube en cascada (IOS → ANDROID+UA → WEB), **sin cookies** |
|
||||
| Retry del mismo recurso | Nunca (deliberado, `ratelimit.py:18-22`) | Nunca tampoco: 1 intento por cliente, error silenciado, siguiente |
|
||||
| Timeout | 15–30 s (`_yt_http.py:40`, `config.yaml:55`) | **4 s** por intento (`AbortSignal.timeout(4e3)`) |
|
||||
| Criterio de éxito | HTTP status (429/403 → clasificar y abortar) | **Contenido**: `r.ok && captionTracks.length > 0` — un 200 vacío se trata como fallo y se rota |
|
||||
| Degradación del resultado | `skip_reason` con taxonomía de causas | Siempre entrega nota con metadatos; transcripción es opcional |
|
||||
|
||||
Los insights accionables para el scraper están en §2, priorizados en §3.
|
||||
|
||||
---
|
||||
|
||||
## 1 · Qué hace la extensión (verificado en el bundle)
|
||||
|
||||
Cadena de extracción de `YoutubeExtractor.extractAsync()` (offset ~380236 de `popup.js`):
|
||||
|
||||
1. **DOM existente** (costo 0): segmentos ya renderizados en la página.
|
||||
2. **Ruta de red principal** `fetchTranscript()`:
|
||||
- `fetchChapters(videoId)` se **dispara sin await** (la promesa se resuelve en paralelo).
|
||||
- Track inline desde `ytInitialPlayerResponse` del DOM (costo 0), validando que `videoDetails.videoId` coincida con el de la URL (`getValidatedPlayerResponse`).
|
||||
- Si no hay inline: `fetchPlayerData(videoId)` → **POST a `youtubei/v1/player?prettyPrint=false`** con cascada:
|
||||
1. `{clientName:"IOS", clientVersion:"20.10.3"}` — sin UA especial.
|
||||
2. `{clientName:"ANDROID", clientVersion:"20.10.38"}` + `User-Agent: com.google.android.youtube/20.10.38 (Linux; U; Android 14)`.
|
||||
3. `{clientName:"WEB", clientVersion:"2.20240101.00.00"}`.
|
||||
4. Fallback final: JSON embebido del DOM.
|
||||
- Cada intento: timeout 4 s, `try{}catch{}` silenciado, y **se acepta solo si `captionTracks.length > 0`**.
|
||||
- Descarga del track: `GET track.baseUrl` con guard de host (`new URL(baseUrl).hostname.endsWith(".youtube.com")`), UA `Mozilla/5.0`, `Accept-Language` si hay idioma preferido, timeout 4 s.
|
||||
3. **Apertura programática del panel** de transcripción (click + polling `pollFor` cada 250 ms, máx 20 intentos) como último recurso.
|
||||
|
||||
Capítulos (`fetchChapters`): primero inline desde `ytInitialData` (`playerOverlays…multiMarkersPlayerBarRenderer.markersMap`); si vacío, POST a `youtubei/v1/next` con cliente WEB; segundo fallback `engagementPanels[*].macroMarkersListItemRenderer`.
|
||||
|
||||
Puntos de red relevantes:
|
||||
|
||||
- Los POST a InnerTube **no llevan cookies** (el fetch de la popup corre en contexto de extensión, `credentials` same-origin ⇒ youtube.com no recibe sesión). La ruta IOS/ANDROID funciona **anónima**.
|
||||
- El `Origin: https://www.youtube.com` / `Referer` los fuerza la regla DNR 9002 porque un browser no puede setear `Origin` — en Python sería simplemente otro header.
|
||||
- `BilibiliExtractor` (mismo bundle) sí cachea transcripciones: LRU `Map` con tope 300 entradas. `YoutubeExtractor` no cachea nada.
|
||||
|
||||
---
|
||||
|
||||
## 2 · Insights accionables para el scraper
|
||||
|
||||
### I1 · Ruta InnerTube propia como *modo degradado* (impacto alto, esfuerzo medio)
|
||||
|
||||
**Evidencia:** la extensión obtiene metadatos completos + `captionTracks` con **un solo POST anónimo** a `youtubei/v1/player` con cliente IOS. `videoDetails` da título, autor, channelId, lengthSeconds, viewCount, keywords; `microformat` da publishDate, description, ownerChannelName. Nada de watch page.
|
||||
|
||||
**Aplicación:** hoy, cuando el ThrottleGuard trip (`ratelimit.py:221-283`), el job aborta limpio y todo queda `pending` hasta el siguiente pase del monitor. Una vía de salvage — POST directo a `player` (1 petición/video en vez de 3) para lo estrictamente necesario (transcripción + metadatos básicos) — permitiría **seguir produciendo a ⅓ del costo** durante los periodos en que la ruta completa (watch page incluida) está bloqueada. `_yt_http.py` ya es el lugar natural para ese cliente.
|
||||
|
||||
**Advertencias:**
|
||||
- `config.yaml:28-32` dice "3 requests irreducibles — no re-litigar". Esa medición fue sobre **yt-dlp restringido** (`player_skip=webpage`, single-client), que pierde pistas. La ruta de la extensión es distinta: una llamada InnerTube propia con aceptación por contenido. No la invalida, pero **habría que medirla** antes de tratarla como reemplazo; como modo degradado opcional el riesgo es acotado.
|
||||
- Las versiones de cliente hardcodeadas de la extensión (20.10.3 / 20.10.38 / 2.20240101) tienen más de un año de rotación. Si se implementa, tomar las versiones vigentes de yt-dlp (que ya las mantiene) en vez de hardcodear, o aceptar el mismo mantenimiento que la extensión.
|
||||
- YouTube exige PO tokens en algunos clientes para formats/streaming; para metadatos/captions la ruta IOS/ANDROID ha seguido funcionando sin ellos (es la evidencia de esta extensión), pero es el punto que puede romperse.
|
||||
|
||||
### I2 · Aceptación por contenido, no por status (impacto medio, esfuerzo bajo)
|
||||
|
||||
La extensión trata "HTTP 200 con respuesta inútil" como fallo y rota. El scraper ya clasifica causas (`skip_reason`, `describe_missing_subtitle`), pero la aceptación es binaria por status. Aplicable a: timedtext que devuelve 200 con cuerpo vacío/corrupto (hoy parsearía vacío y se marcaría "parsed empty" en vez de reintentable), y a cualquier futura llamada InnerTube propia.
|
||||
|
||||
### I3 · Timeout corto con fail-fast en timedtext (impacto medio, esfuerzo bajo)
|
||||
|
||||
`extract.py:301-321` baja subtítulos con timeout 15 s; `config.yaml:55` pone 30 s de socket. El punto donde el throttling "más aparece" es precisamente timedtext (`extract.py:302-307`). Un timeout más agresivo (configurable, p. ej. 6–8 s para json3/srv1 — payloads pequeños) convertiría cuelgues de 15–30 s en un fallo clasificable como reintentable casi inmediato, liberando el pacer antes. La extensión usa 4 s para todo.
|
||||
|
||||
### I4 · Rotación de cookie al hacer trip el ThrottleGuard (impacto alto, esfuerzo medio)
|
||||
|
||||
La extensión no rota cookies (viaja sobre la sesión real del usuario), pero su patrón estructural — *ante el bloqueo, cambiar de identidad en vez de solo esperar* — traducido al scraper es: el vault ya persiste múltiples cookies (`cookies/<uuid>.txt` + `cookies_meta`), pero `resolve_active_path` usa exactamente una (`cookies.py:151-163`). Al trip del breaker, cambiar a la siguiente cookie no expirada antes de rendirse al reloj multiplicaría el presupuesto efectivo por sesión sin nueva infraestructura. (Insight inspirado en el patrón de la extensión, no copiado de ella.)
|
||||
|
||||
### I5 · `youtubei/v1/next` para capítulos sin watch page (habilitador de I1)
|
||||
|
||||
Si algún día se activa la ruta de 1 petición (I1), los capítulos —que hoy llegan vía info de yt-dlp desde la watch page (`chapters.py:39-49`)— se recuperan con un POST a `next` (cliente WEB) parseando `playerOverlays…markersMap`, con fallback a `engagementPanels`. El bundle de la extensión contiene la implementación de referencia exacta (offset ~393083). Solo necesario para videos con capítulos; el resto no paga el request.
|
||||
|
||||
### I6 · Guard de host antes de descargar timedtext (hardening, esfuerzo mínimo)
|
||||
|
||||
`fetchCaptionXml` exige `hostname.endsWith(".youtube.com")` antes del GET al `baseUrl` del caption track. El scraper baja `pick.url` con `yt_get` sin validar host (`extract.py:301-321`). La URL viene de yt-dlp (confiable hoy), pero una línea de validación cierra la clase de riesgo "baseUrl corrupto/inyectado ⇒ GET con headers de navegador a un host arbitrario".
|
||||
|
||||
### I7 · Detalles menores de protocolo (gratis si se implementa I1)
|
||||
|
||||
- `?prettyPrint=false` en llamadas InnerTube: menos payload.
|
||||
- `Origin: https://www.youtube.com` como header explícito en POSTs propios.
|
||||
- Lanzar la descarga dependiente como promesa paralela (la extensencia dispara `fetchChapters` sin await): en el scraper el equivalente sería solapar la descarga de timedtext con el siguiente video del pacer **solo si** el presupuesto de requests ya lo contempla — cuidado: hoy el pacer es la política de cortesía, no paralelizar contra él.
|
||||
|
||||
### I8 · Paridades confirmadas (sin acción)
|
||||
|
||||
- Selección de pista por idioma: `pick_subtitle` del scraper (política por idioma, rechazo de `tlang=`, detección `-orig`, prioridad json3) es **más rica** que `pickCaptionTrack`/`findPreferredCaptionTrack` de la extensión.
|
||||
- Degradación graciosa del resultado: taxonomía `skip_reason` ≥ "nota siempre con metadatos" de la extensión.
|
||||
- Caching: SQLite del scraper > LRU 300 del BilibiliExtractor; YoutubeExtractor ni siquiera cachea.
|
||||
- Validación de JSON inline contra videoId (`getValidatedPlayerResponse`): patrón correcto a recordar **si** algún día se cachean player responses o se reutilizan continuations entre sesiones (evita atribuir a un video la respuesta de otro).
|
||||
|
||||
---
|
||||
|
||||
## 3 · Priorización sugerida
|
||||
|
||||
| # | Mejora | Contra qué límite ayuda | Esfuerzo | Riesgo |
|
||||
|---|---|---|---|---|
|
||||
| 1 | I3 timeout fail-fast en timedtext | Throughput bajo throttling | Bajo | Bajo (hacerlo configurable) |
|
||||
| 2 | I6 guard de host | Hardening | Mínimo | Nulo |
|
||||
| 3 | I4 rotación de cookies al trip del breaker | Techo de presupuesto por sesión | Medio | Medio (cuenta de la cookie expuesta al mismo ritmo) |
|
||||
| 4 | I1+I2+I5+I7 ruta InnerTube propia como modo degradado | Bot wall: seguir produciendo a ⅓ de costo | Medio-Alto | Medio (requiere medición; mantenimiento de client versions) |
|
||||
|
||||
## 4 · Qué NO copiar de la extensión
|
||||
|
||||
1. **Scraping del DOM del panel de transcripción** (clicks + polling): requiere un browser real con fingerprint real; el scraper es headless. Solo cobraría sentido con Playwright + perfil real, que es otra conversación.
|
||||
2. **Re-litigar los 3 requests/video vía yt-dlp** (`config.yaml:28-32`): la medición del repo sigue en pie para yt-dlp. La vía InnerTube propia es una **ruta paralela degradada**, no un reemplazo de la ruta completa.
|
||||
3. **Silenciado de errores** (`try{}catch{}` sin logging): la extensión degradea muda; el scraper necesita auditabilidad (`skip_reason` ya la da).
|
||||
4. **Versiones de cliente hardcodeadas**: la extensión las parchea por release; el scraper ya delega eso en yt-dlp. Cualquier ruta propia debe heredar las versiones de yt-dlp, no duplicarlas.
|
||||
|
||||
---
|
||||
|
||||
*Fin del informe.*
|
||||
@@ -0,0 +1,8 @@
|
||||
# Auditorías
|
||||
|
||||
Documentos de investigación que explican decisiones técnicas del proyecto mediante ingeniería inversa. No son specs ni manuales: son material de referencia.
|
||||
|
||||
- **`YOUTUBE-TRANSCRIPT-AUDIT.md`** — ingeniería inversa de la extensión *Obsidian Web Clipper* 1.7.1: cómo obtiene metadatos y transcripciones vía la API privada InnerTube (`youtubei/v1/player`) con clientes ANDROID/IOS/WEB. Es el contexto de *por qué* este scraper delega en `yt-dlp` (que implementa el mismo mecanismo) en vez de llamar a InnerTube directamente.
|
||||
- **`CLIPPER-COMPARISON-AUDIT.md`** — comparativa extensión vs `yt-channel-scraper`: filosofías de red opuestas (una petición por vídeo pinchado vs batch educado de canal) y qué ideas de una aplican a la otra.
|
||||
|
||||
Ambos provienen de la raíz del repo y se mantienen sin edición aquí.
|
||||
@@ -0,0 +1,505 @@
|
||||
# Auditoría · Cómo la extensión obtiene la transcripción / información de vídeos de YouTube
|
||||
|
||||
> **Producto auditado:** *Obsidian Web Clipper* (Chrome / Chromium / Firefox / Safari) — versión `1.7.1` del paquete `D:\Obsidian Web Clipper - Chrome Web Store 1.7.1.0`.
|
||||
> **Tipo de extensión:** MV3 (manifest v3) con service worker (`background.js`).
|
||||
> **Alcance de la auditoría:** mecanismo end‑to‑end por el que la extensión extrae metadatos y la transcripción de un vídeo de YouTube (incluye short `youtu.be`, `youtube.com/watch?v=…` y `youtube.com/shorts/…`).
|
||||
> **Fecha:** 2026‑07‑26.
|
||||
> **Audiencia del documento:** LLMs / agentes de mantenimiento. Estructura deliberadamente declarativa, sin prosa narrativa.
|
||||
|
||||
---
|
||||
|
||||
## 0 · TL;DR (resumen ejecutable)
|
||||
|
||||
1. La extensión **no usa `timedtext`, `youtube-transcript` web, ni scraping de `ytd-transcript-segment-renderer` como ruta principal** cuando la URL es un watch normal: usa la **API privada `youtubei/v1/player`** (InnerTube) con cabeceras que imitan clientes oficiales de YouTube (ANDROID, IOS, WEB).
|
||||
2. Para llegar a esa API sin ser bloqueada por CORS / firma, el `service worker` declara una regla `declarativeNetRequest` (id `9002`, nombre interno `enableYouTubeInnertubeRule`) que **fuerza `Origin: https://www.youtube.com` y `Referer: https://www.youtube.com/`** en toda petición XHR iniciada por la propia extensión hacia `||youtube.com/youtubei/`.
|
||||
3. La capa de extracción es una clase `YoutubeExtractor` (en `popup.js` y replicada en `reader-page.js`, ambos `webpack` bundles de Defuddle) que:
|
||||
- 1️⃣ parsea el JSON embebido `ytInitialPlayerResponse` del DOM para sacar `captionTracks` y `baseUrl` sin red.
|
||||
- 2️⃣ si falla, abre el panel "Mostrar transcripción" del propio YouTube haciendo `click()` y espera con `MutationObserver`‑style polling (DOM scraping fallback).
|
||||
- 3️⃣ si la transcripción automática no está disponible, llama a `youtubei/v1/player` con 3 identidades de cliente en cascada (ANDROID → IOS → WEB) hasta que una devuelve `captions.playerCaptionsTracklistRenderer.captionTracks`.
|
||||
- 4️⃣ descarga la pista (`timedtext`-like `baseUrl` con sufijo `&fmt=…`) y la parsea como XML.
|
||||
4. Los capítulos se extraen con una segunda ruta: `youtubei/v1/next` (también con cabeceras de cliente), o desde `ytInitialData` embebido (`playerOverlays.playerOverlayRenderer.decoratedPlayerBarRenderer.multiMarkersPlayerBarRenderer.markersMap`).
|
||||
5. Toda la red de la popup se hace con `globalThis.fetch` directo (no hay proxy interno), aprovechando la regla DNR 9002. El background solo ofrece un *fallback* `sendNativeMessage` para hosts que devuelven CORS (por ejemplo Bilibili).
|
||||
|
||||
---
|
||||
|
||||
## 1 · Vista general de componentes (mapa de archivos)
|
||||
|
||||
| Archivo | Rol respecto a YouTube | Tamaño aprox. | Notas |
|
||||
|---|---|---|---|
|
||||
| `manifest.json` | Declara `host_permissions: ["<all_urls>","http://*/*","https://*/*"]` y `declarativeNetRequest`. | 89 líneas | Sin URL allow‑list específica de YouTube. |
|
||||
| `background.js` (service worker) | Define la **regla DNR 9002** `enableYouTubeInnertubeRule` (set Origin/Referer para `||youtube.com/youtubei/`). Contiene además `enableYouTubeEmbedRule` (9001) que pone `Referer: https://obsidian.md/` en iframes `||youtube.com/embed/`. | ~1 archivo compilado | Comentario interno: `initiatorDomains: [chrome.runtime.id]` ⇒ sólo afecta peticiones de la propia extensión. |
|
||||
| `content.js` (content script) | Sólo contiene el glue de highlights (`getClosestTextBlock` ignora elementos con clase `transcript-segment` para no romper la selección). **No extrae la transcripción.** | 1 bundle webpack | No realiza llamadas a YouTube. |
|
||||
| `popup.js` | Contiene la clase `YoutubeExtractor` real (minificada) y todo el código de extracción. Se carga como `popup.html` y como `side-panel.html` (ver `<script type="module" src="popup.js">`). | ~2.5 MB minificado | Aquí vive toda la lógica de transcripción. |
|
||||
| `reader-page.js` | Réplica exacta de los extractores (incluye otra copia de `YoutubeExtractor`). Se usa cuando se abre la URL en modo *Reader* (`reader.html?url=…`). | ~2.5 MB minificado | Mismo binario que popup. |
|
||||
| `highlighter.js` | Sólo lógica de resaltado (no relevante para transcripción). | — | — |
|
||||
| `reader-script.js` | Inyectado por background con `scripting.executeScript` para modo Reader. | — | — |
|
||||
| `_locales/*/messages.json` | i18n; incluye claves `readerTranscripts`, `readerPinPlayer`, `readerHighlightActiveLine` (configuración visual de la transcripción, no de extracción). | — | — |
|
||||
| `web_accessible_resources` | Lista `reader.css`, `reader-script.js`, `browser-polyfill.min.js`, `style.css`, `side-panel.html`, `flatten-shadow-dom.js`, `highlighter.css`. | — | `popup.js` **no** está en `web_accessible_resources`; por tanto la extracción no se hace desde un script inyectado en la página, sino desde la propia página de extensión. |
|
||||
|
||||
### 1.1 Flujo de control (quién llama a quién)
|
||||
|
||||
```
|
||||
[user clicks action / shortcut / context menu]
|
||||
└─ background.js (service worker)
|
||||
└─ browser.action.openPopup() OR tabs.sendMessage("openPopup")
|
||||
└─ popup.html (extension page, chrome-extension://<id>/popup.html)
|
||||
└─ popup.js (module)
|
||||
├─ Defuddle.parse(doc) ← extractor genérico
|
||||
│ └─ para URL que matchea "youtube.com" o "youtu.be"
|
||||
│ └─ new YoutubeExtractor(document, url, schemaOrg, options)
|
||||
│ └─ extractAsync() ⇒ runExtractor()
|
||||
│ ├─ extractTranscriptFromExistingDom() (1ª opción)
|
||||
│ ├─ fetchTranscript() (2ª opción: red)
|
||||
│ └─ extractTranscriptFromOpenedDom() (3ª opción: click en panel)
|
||||
└─ resultado ⇒ variables { transcript, language } ⇒ se inyecta en la nota Markdown
|
||||
|
||||
(En paralelo, la regla DNR 9002 reescribe Origin/Referer de las XHR
|
||||
lanzadas por la propia extensión hacia youtube.com/youtubei/v1/…)
|
||||
```
|
||||
|
||||
> **Punto importante para LLMs:** la extracción de YouTube **no se ejecuta dentro de la página de YouTube** ni desde el content script. Se ejecuta en el contexto privilegiado de la extensión (`chrome-extension://`). La página de YouTube solo aporta el `document` con el HTML actual y, opcionalmente, el JSON embebido en los `<script>`.
|
||||
|
||||
---
|
||||
|
||||
## 2 · Punto de entrada: ¿cuándo se invoca la extracción?
|
||||
|
||||
`background.js` ofrece cuatro formas de abrir la popup, todas convergen al mismo punto:
|
||||
|
||||
| Acción del usuario | Mensaje / llamada en `background.js` | Destino final |
|
||||
|---|---|---|
|
||||
| Click en icono de la extensión | `action.onClicked` → `openPopup()` | `popup.html` |
|
||||
| Atajo `Ctrl+Shift+O` (`_execute_action`) | `commands.onCommand` → `openPopup()` | `popup.html` |
|
||||
| Atajo `Alt+Shift+O` (`quick_clip`) | `commands.onCommand` → `openPopup()` + 500 ms `triggerQuickClip` | `popup.html` |
|
||||
| Menú contextual "Save this page" / "Add to highlights" | `contextMenus.onClicked` → `openPopup()` | `popup.html` |
|
||||
| Behavior `embedded` | `tabs.sendMessage("toggle-iframe")` → `side-panel.html?context=iframe` | `side-panel.html` |
|
||||
|
||||
> `side-panel.html` y `popup.html` cargan **el mismo `popup.js`** (`<script type="module" src="popup.js">`). Por tanto, la lógica de YouTube es única y se invoca desde dos contenedores distintos.
|
||||
|
||||
Cuando la popup se carga, en `popup.js` se hace algo equivalente a:
|
||||
|
||||
```js
|
||||
const defuddle = new Defuddle(document, { url: location.href });
|
||||
const result = defuddle.parse(); // extracción síncrona
|
||||
const asyncVars = await defuddle.fetchAsyncVariables({ language, fetch }); // asíncrono
|
||||
```
|
||||
|
||||
`Defuddle` consulta su `ExtractorRegistry` (poblado en `ExtractorRegistry.initialize()`); para YouTube registra:
|
||||
|
||||
```js
|
||||
this.register({ patterns: ["youtube.com","youtu.be"], extractor: YoutubeExtractor });
|
||||
this.register({ patterns: ["m.youtube.com"], extractor: YoutubeExtractor }); // implícito
|
||||
this.register({ patterns: [/youtube\.com\/shorts\//], extractor: YoutubeExtractor });
|
||||
```
|
||||
|
||||
El extractor se instancia con `new YoutubeExtractor(document, url, schemaOrgData, options)`. Las `options` que recibe la popup le inyectan `language` (preferida por el usuario) y un `fetch` opcional; si no se inyecta, usa `globalThis.fetch`.
|
||||
|
||||
---
|
||||
|
||||
## 3 · `YoutubeExtractor` (la clase clave) — Anatomía
|
||||
|
||||
> **Ubicación física del código (minificado):**
|
||||
> - En `popup.js`, la clase aparece aproximadamente entre los offsets `382 000`–`405 000` del bundle (texto buscado: `class YoutubeExtractor` o el alias `class A extends o.BaseExtractor`).
|
||||
> - En `reader-page.js` es la **misma clase** con nombre `A` (mismo fingerprint de strings `transcript-segment-view-model`, `ytwTranscriptSegmentViewModelTimestamp`, `ytInitialPlayerResponse`).
|
||||
> - En origen viene del paquete npm `defuddle` (≥ v0.x) — la extensión lo reempaqueta con webpack.
|
||||
|
||||
### 3.1 Identidad de cliente (constantes globales)
|
||||
|
||||
```js
|
||||
// constantes a nivel de módulo, dentro del bundle de popup.js / reader-page.js
|
||||
const TIMEOUT_MS = 4000; // f = 4e3
|
||||
const PLAYER_URL = "https://www.youtube.com/youtubei/v1/player?prettyPrint=false"; // g
|
||||
const NEXT_URL = "https://www.youtube.com/youtubei/v1/next?prettyPrint=false"; // usado en fetchChapters
|
||||
const ANDROID_UA = "com.google.android.youtube/20.10.38 (Linux; U; Android 14)"; // v, usada en b/x
|
||||
|
||||
// Contextos de cliente (probados en cascada)
|
||||
const ANDROID_CLIENT = { client: { clientName: "ANDROID", clientVersion: "20.10.38" } }; // b / x
|
||||
const IOS_CLIENT = { client: { clientName: "IOS", clientVersion: "20.10.3" } }; // y
|
||||
const WEB_CLIENT = { client: { clientName: "WEB", clientVersion: "2.20240101.00.00" } }; // w
|
||||
```
|
||||
|
||||
### 3.2 Selectores DOM (definidos como objetos)
|
||||
|
||||
```js
|
||||
const DESKTOP_SELECTORS = {
|
||||
segments: "ytd-transcript-segment-renderer",
|
||||
timestamp: ".segment-timestamp",
|
||||
text: ".segment-text",
|
||||
};
|
||||
const MOBILE_SELECTORS = {
|
||||
segments: "transcript-segment-view-model",
|
||||
timestamp: ".ytwTranscriptSegmentViewModelTimestamp",
|
||||
text: "span.yt-core-attributed-string",
|
||||
chapters: "timeline-chapter-view-model h3",
|
||||
};
|
||||
```
|
||||
|
||||
`getTranscriptSelectors(container)` elige uno u otro mirando qué nodos existen. Si no existe ninguno devuelve `undefined` (⇒ no hay transcripción en el DOM todavía).
|
||||
|
||||
### 3.3 Métodos principales (firmas y propósito)
|
||||
|
||||
| Método | Tipo | Propósito |
|
||||
|---|---|---|
|
||||
| `getVideoId()` | síncrono | Devuelve el id de 11 chars. Soporta `youtube.com/watch?v=…`, `youtu.be/…`, `youtube.com/shorts/…`. Cachea en `this._videoId`. |
|
||||
| `canExtractAsync()` | síncrono | Devuelve `true` si la URL es de YouTube. |
|
||||
| `extractAsync()` | async | Punto de entrada. Cadena: `extractTranscriptFromExistingDom()` → si vacío `fetchTranscript()` → si vacío `extractTranscriptFromOpenedDom()`. Devuelve `{ html, text, languageCode, … }` o `null`. |
|
||||
| `extractTranscriptFromExistingDom()` | try/catch | Lee los segmentos si YouTube ya renderizó el panel de transcripción (panel abierto por el usuario o cargado por interacción previa). |
|
||||
| `getTranscriptContainer()` | síncrono | Selector: `'ytd-engagement-panel-section-list-renderer[target-id="engagement-panel-searchable-transcript"] #segments-container'`; en `m.youtube.com`: `ytm-macro-markers-list-renderer .ytm-macro-markers-list-container`. |
|
||||
| `buildTranscriptFromContainer(container, chapters)` | síncrono | Itera cada segmento, parsea timestamp (`parseTimestamp` acepta `h:mm:ss` o `mm:ss`), agrupa por hablante si detecta patrón (`groupTranscriptSegments` → `groupBySpeaker` o `groupBySentence`), y emite HTML `<p class="transcript-segment">` y texto plano `**HH:MM:SS** · texto`. Llama a `buildTranscript("youtube", groups, chapters)`. |
|
||||
| `extractTranscriptFromOpenedDom()` | async | Si el panel no estaba abierto pero el DOM lo permite (`canOpenTranscriptPanel()`: `typeof MutationObserver === "function"`): hace **click programático** en `ytd-video-description-transcript-section-renderer button` (o equivalente mobile), espera el contenedor y re-usa `buildTranscriptFromContainer`. |
|
||||
| `openMobileTranscriptPanel()` | async | Variante `m.youtube.com`: clicks en `button[aria-label="Show more"]` → espera `button[aria-label="View all"]` → click → espera segmentos. |
|
||||
| `fetchTranscript()` | async | **Ruta de red principal**. Ver §3.4. |
|
||||
| `fetchPlayerData(videoId)` | async | POST a `youtubei/v1/player` probando 3 clientes en cascada. Ver §3.4. |
|
||||
| `fetchChapters(videoId)` | async | POST a `youtubei/v1/next` (cliente WEB) → extrae capítulos de `playerOverlays…multiMarkersPlayerBarRenderer.markersMap[*].value.chapters[*].chapterRenderer`; fallback a `engagementPanels[*].engagementPanelSectionListRenderer.content.macroMarkersListRenderer.contents[*].macroMarkersListItemRenderer`. Si ya están embebidos en `ytInitialData` no se hace red. |
|
||||
| `getValidatedPlayerResponse()` | síncrono | Devuelve el JSON parseado de `ytInitialPlayerResponse` (parseado inline desde el `<script>`) **solo si** su `videoDetails.videoId` o `microformat.playerMicroformatRenderer.externalVideoId` coincide con `getVideoId()`. |
|
||||
| `parseInlineJson(varName)` | síncrono | Itera todos los `<script>` del documento, encuentra el primero cuyo `textContent` contiene la variable global, **balancea llaves** manualmente y hace `JSON.parse`. Cachea el resultado en `this.inlineJsonCache` (Map). |
|
||||
| `getCaptionTracks(playerResponse)` | síncrono | `playerResponse?.captions?.playerCaptionsTracklistRenderer?.captionTracks` (devuelve `[]` si no es array). |
|
||||
| `pickCaptionTrack(tracks)` | síncrono | Si hay `options.language`: prefiere la pista exacta (`code === lang`) ⇒ mismo idioma base (`code.split("-")[0] === lang.split("-")[0]`) ⇒ mismo prefijo de idioma. Filtra `kind === "asr"` (auto‑generadas) si hay manuales. Si no, devuelve la primera no‑`asr`, o una con `languageCode === "en"`, o la primera. |
|
||||
| `findPreferredCaptionTrack(tracks, lang)` | síncrono | Igual que el anterior pero con scoring explícito. |
|
||||
| `getInlineCaptionTrack()` | síncrono | `getValidatedPlayerResponse()` → `getCaptionTracks()` → `pickCaptionTrack()`. Si hay `baseUrl`, devuelve la pista. |
|
||||
| `fetchCaptionXml(track, chaptersPromise)` | async | `fetch(track.baseUrl, { headers: { "User-Agent":"Mozilla/5.0", "Accept-Language": lang } })` con `AbortSignal.timeout(4000)`. Devuelve solo si la URL acaba en `.youtube.com`. Pasa el texto a `parseTranscriptXml`. |
|
||||
| `parseTranscriptXml(xml, lang, chapters)` | síncrono | Dos regex: `<p t="N">…<s>…</s>…</p>` (formato nuevo) y `<text start="N">…</text>` (formato legacy). Decodifica entidades (`decodeEntities`). Llama a `groupTranscriptSegments` y `buildTranscript`. |
|
||||
| `decodeEntities(str)` | síncrono | Reemplazos para `& < > " ' ' &#xHH; &#NN;`. |
|
||||
| `groupTranscriptSegments(segs)` | síncrono | Decide por regex CJK: si hay mezcla CJK/Latín → `groupBySpeaker` (split por `:` al inicio de línea); si no → `groupBySentence`. |
|
||||
| `getVideoData()` | síncrono | Lee `<script type="application/ld+json">` buscando un `VideoObject` cuyo `embedUrl`/`url`/`@id` contenga el videoId. Fallback a `meta[property="og:title|og:description|og:image|og:url"]`. |
|
||||
| `getChannelNameFromDom()` / `getChannelNameFromPlayerResponse()` | síncrono | `[itemprop="name"]` o `videoDetails.author`/`ownerChannelName`/`microformat.playerMicroformatRenderer.ownerChannelName`. |
|
||||
| `getTranscriptLanguageCodeFromDom()` | síncrono | Lee el botón de `yt-sort-filter-sub-menu-renderer` en el footer del panel y compara con `name.simpleText`/`name.runs[].text` de cada caption track. |
|
||||
| `getInlineChapters()` | síncrono | Desde `ytInitialData`; valida que el videoId en el JSON coincida con el de la URL; cae a `extractChaptersFromEngagementPanels`. |
|
||||
| `buildResult(transcript)` | síncrono | Empaqueta `{ title, author, site:"YouTube", image, published, description }`, añade `transcript` y `language` al `variables`, prepende un `<iframe src="https://www.youtube.com/embed/{id}" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen>`. |
|
||||
| `formatDescription(text)` | síncrono | `<p>…</p>` con `escapeHtml` y `<br>` por `\n`. |
|
||||
|
||||
### 3.4 `fetchTranscript()` — núcleo de la ruta de red
|
||||
|
||||
```js
|
||||
async fetchTranscript() {
|
||||
const videoId = this.getVideoId();
|
||||
const chapters = this.fetchChapters(videoId); // (a)
|
||||
const inlineTrack = this.getInlineCaptionTrack(); // (b) caption del JSON embebido
|
||||
const inlineFetch = inlineTrack ? this.fetchCaptionXml(inlineTrack, chapters) : undefined; // (c)
|
||||
const playerData = await this.fetchPlayerData(videoId); // (d) red: youtubei/v1/player
|
||||
const picked = playerData ? this.pickCaptionTrack(this.getCaptionTracks(playerData)) : undefined;
|
||||
const remoteFetch = (picked?.baseUrl && picked.baseUrl !== inlineTrack?.baseUrl)
|
||||
? this.fetchCaptionXml(picked, chapters) // (e)
|
||||
: undefined;
|
||||
return (await remoteFetch) || (await inlineFetch);
|
||||
}
|
||||
```
|
||||
|
||||
Y `fetchPlayerData()`:
|
||||
|
||||
```js
|
||||
async fetchPlayerData(videoId) {
|
||||
// 1º intento: cliente IOS
|
||||
try {
|
||||
const r = await this.fetch(PLAYER_URL, {
|
||||
method: "POST",
|
||||
headers: { "Content-Type":"application/json", ...(lang && {"Accept-Language": lang}) },
|
||||
signal: AbortSignal.timeout(TIMEOUT_MS),
|
||||
body: JSON.stringify({ context: IOS_CLIENT, videoId })
|
||||
});
|
||||
if (r.ok) { const t = await r.json(); if (this.getCaptionTracks(t).length) return t; }
|
||||
} catch {}
|
||||
|
||||
// 2º intento: cliente ANDROID con UA
|
||||
try {
|
||||
const r = await this.fetch(PLAYER_URL, {
|
||||
method: "POST",
|
||||
headers: { "Content-Type":"application/json", "User-Agent": ANDROID_UA, ...(lang && {"Accept-Language": lang}) },
|
||||
signal: AbortSignal.timeout(TIMEOUT_MS),
|
||||
body: JSON.stringify({ context: ANDROID_CLIENT, videoId })
|
||||
});
|
||||
if (r.ok) { const t = await r.json(); if (this.getCaptionTracks(t).length) return t; }
|
||||
} catch {}
|
||||
|
||||
// 3º intento: cliente WEB clásico
|
||||
try {
|
||||
const r = await this.fetch(PLAYER_URL, {
|
||||
method: "POST",
|
||||
headers: { "Content-Type":"application/json" },
|
||||
signal: AbortSignal.timeout(TIMEOUT_MS),
|
||||
body: JSON.stringify({ context: WEB_CLIENT, videoId })
|
||||
});
|
||||
if (r.ok) { const t = await r.json(); if (this.getCaptionTracks(t).length) return t; }
|
||||
} catch {}
|
||||
|
||||
// 4º fallback: parsear el JSON embebido en el HTML (sin red)
|
||||
const inline = this.parseInlineJson("ytInitialPlayerResponse");
|
||||
if (this.getCaptionTracks(inline).length) return inline;
|
||||
}
|
||||
```
|
||||
|
||||
> **Para LLMs:** los 3 contextos de cliente y el `User-Agent` ANDROID son los mismos que usa `youtube-dl`, `yt-dlp` y la mayoría de librerías no oficiales. Si YouTube empieza a rechazar la firma, el orden de los 3 intentos es lo que hay que tocar; también se puede añadir `TVHTML5_SIMPLY_EMBEDDED_PLAYER` u otros.
|
||||
|
||||
---
|
||||
|
||||
## 4 · Mecanismo de bypass de CORS / firma de YouTube
|
||||
|
||||
YouTube no expone `youtubei/v1/player` con CORS abierto, así que la extensión necesita que las peticiones se *vean* como originadas desde la propia web de YouTube. Hay **dos piezas** que lo permiten:
|
||||
|
||||
### 4.1 `enableYouTubeInnertubeRule` (DNR id `9002`)
|
||||
|
||||
En `background.js` (en la función `initialize()`):
|
||||
|
||||
```js
|
||||
await dnr.updateSessionRules({
|
||||
removeRuleIds: [9002],
|
||||
addRules: [{
|
||||
id: 9002,
|
||||
priority: 1,
|
||||
action: {
|
||||
type: "modifyHeaders",
|
||||
requestHeaders: [
|
||||
{ header: "Origin", operation: "set", value: "https://www.youtube.com" },
|
||||
{ header: "Referer", operation: "set", value: "https://www.youtube.com/" }
|
||||
]
|
||||
},
|
||||
condition: {
|
||||
urlFilter: "||youtube.com/youtubei/",
|
||||
resourceTypes: ["xmlhttprequest"],
|
||||
initiatorDomains: [ chrome.runtime.id ].filter(Boolean)
|
||||
}
|
||||
}]
|
||||
});
|
||||
```
|
||||
|
||||
> Como el popup (`chrome-extension://<id>/popup.html`) lanza `fetch` desde el contexto de la extensión, el `initiator` de la petición es el `extension_id` ⇒ la regla **solo** se aplica a las peticiones que dispara la propia extensión. Las peticiones que haga el content script de la página de YouTube no se tocan aquí.
|
||||
|
||||
### 4.2 `webRequest.onBeforeSendHeaders` (sólo Firefox / WebExtensions)
|
||||
|
||||
En `background.js` también hay (entre `try { … } catch {}` para tolerancia a Safari):
|
||||
|
||||
```js
|
||||
browser_polyfill.webRequest.onBeforeSendHeaders.addListener(details => {
|
||||
if (details.tabId > 0) {
|
||||
const refHeader = details.requestHeaders.find(h => h.name.toLowerCase() === "referer");
|
||||
const refValue = refHeader?.value || "";
|
||||
const originHdr = details.requestHeaders.find(h => h.name.toLowerCase() === "origin");
|
||||
const originValue = originHdr?.value || "";
|
||||
if (!(refValue.startsWith("moz-extension://") || refValue.startsWith("safari-web-extension://"))) {
|
||||
return { requestHeaders: details.requestHeaders }; // no tocar: es la propia web
|
||||
}
|
||||
}
|
||||
const headers = details.requestHeaders || [];
|
||||
const setHeader = (name, value) => {
|
||||
const existing = headers.find(h => h.name.toLowerCase() === name.toLowerCase());
|
||||
existing ? existing.value = value : headers.push({ name, value });
|
||||
};
|
||||
setHeader("Origin", "https://www.youtube.com");
|
||||
setHeader("Referer", "https://www.youtube.com/");
|
||||
return { requestHeaders: headers };
|
||||
}, { urls: ["*://www.youtube.com/*"] }, ["blocking","requestHeaders"]);
|
||||
```
|
||||
|
||||
> **Para LLMs:** este listener sólo aplica a Firefox MV2 / WebExtensions, donde `declarativeNetRequest` no soporta `modifyHeaders` de la misma forma. **No se ejecuta en Chrome** (donde ya tenemos DNR 9002). En Safari se ignora silenciosamente (`catch` lo traga).
|
||||
|
||||
### 4.3 `enableYouTubeEmbedRule` (DNR id `9001`) — caso especial, no transcripción
|
||||
|
||||
No participa en la extracción de transcripción, pero la documentamos para que el lector no se confunda: cuando la popup muestra el `<iframe src="https://www.youtube.com/embed/…">` en modo *Reader*, background fuerza `Referer: https://obsidian.md/` para que el embed no se rompa en vídeos con restricción por referer.
|
||||
|
||||
---
|
||||
|
||||
## 5 · `fetchProxy` y `nativeFetch` — por qué existen (y por qué YouTube NO los usa)
|
||||
|
||||
En `background.js`:
|
||||
|
||||
```js
|
||||
browser_polyfill.runtime.onMessage.addListener(request => {
|
||||
if (request.action !== "fetchProxy") return;
|
||||
return fetch(request.url, request.options)
|
||||
.then(async resp => {
|
||||
const text = await resp.text();
|
||||
const looksLikeHTML = !resp.ok && (text.includes("Sorry") || text.includes("<html"));
|
||||
if (!looksLikeHTML) return { ok: resp.ok, status: resp.status, text, finalUrl: resp.url };
|
||||
return browser_polyfill.runtime.sendNativeMessage ? nativeFetch(request.url, request.options) : { ok:false, status:0, error:"CORS_PERMISSION_NEEDED" };
|
||||
})
|
||||
.catch(() => browser_polyfill.runtime.sendNativeMessage ? nativeFetch(request.url, request.options) : { ok:false, error:"CORS_PERMISSION_NEEDED" });
|
||||
});
|
||||
```
|
||||
|
||||
`nativeFetch` usa `runtime.sendNativeMessage("application.id", { type:"fetchRequest", url, method, headers, body })` que solo está disponible en Safari (App‑bound messaging) — sirve para que la app nativa de Mac de Safari haga la petición y devuelva el cuerpo.
|
||||
|
||||
> **Implicación para YouTube:** la popup **nunca** enruta sus llamadas a YouTube por `fetchProxy` ni por `nativeFetch`. Las llamadas van por `globalThis.fetch` directo, protegidas por la regla DNR 9002.
|
||||
|
||||
`fetchProxy` se usa en el extractor de **Bilibili** (`BilibiliExtractor`):
|
||||
- Llama a `https://api.bilibili.com/x/player/wbi/v2?bvid=…&cid=…` desde la popup.
|
||||
- Bilibili suele devolver CORS abierto, así que normalmente no hace falta el proxy. El proxy queda como fallback de seguridad.
|
||||
|
||||
---
|
||||
|
||||
## 6 · Estructura del output (qué se mete en la nota Markdown)
|
||||
|
||||
`YoutubeExtractor.buildResult(transcript)` produce:
|
||||
|
||||
```js
|
||||
{
|
||||
content: ` <iframe … src="https://www.youtube.com/embed/{id}" …></iframe><p>{descripción}</p>{transcript.html}`,
|
||||
contentHtml: idem,
|
||||
extractedContent: { videoId, author },
|
||||
variables: {
|
||||
title: e.name || "", // videoDetails.title
|
||||
author: r, // canal
|
||||
site: "YouTube",
|
||||
image: Array.isArray(e.thumbnailUrl) ? e.thumbnailUrl[0] : "",
|
||||
published: e.uploadDate, // microformat.publishDate o uploadDate
|
||||
description: n.slice(0, 200).trim(),
|
||||
transcript: transcript.text, // sólo si hay
|
||||
language: transcript.languageCode // "en", "es", "es-419", ...
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
`transcript.text` es texto plano con formato Markdown:
|
||||
|
||||
```
|
||||
**00:00** · Primera frase hablada
|
||||
|
||||
**00:04** · Segunda frase hablada
|
||||
```
|
||||
|
||||
`transcript.html` es:
|
||||
|
||||
```html
|
||||
<div class="youtube transcript">
|
||||
<h2>Transcript</h2>
|
||||
<p class="transcript-segment"><strong><span class="timestamp" data-timestamp="0">00:00</span></strong> · Primera frase hablada</p>
|
||||
…
|
||||
</div>
|
||||
```
|
||||
|
||||
Los `transcript-segment` con `data-timestamp` permiten que el modo *Reader* haga **highlight de la línea activa mientras reproduce** (`readerHighlightActiveLine`, `readerPinPlayer`, `readerAutoScroll` — ver `_locales/*/messages.json`).
|
||||
|
||||
---
|
||||
|
||||
## 7 · Manejo de errores y degradación
|
||||
|
||||
| Escenario | Comportamiento observado |
|
||||
|---|---|
|
||||
| Vídeo sin subtítulos manuales ni auto‑generados | `getCaptionTracks([])` ⇒ `pickCaptionTrack(undefined)` ⇒ `null` ⇒ `buildResult` se llama con `transcript = undefined` ⇒ `variables.transcript` queda ausente. El iframe y los metadatos sí se guardan. |
|
||||
| Vídeo con subtítulos sólo auto (`kind:"asr"`) | `pickCaptionTrack` la prefiere si no hay otra o si la opción `language` coincide; en otro caso la salta. |
|
||||
| `fetchPlayerData` falla en los 3 clientes y no hay JSON embebido | `transcript` queda `undefined`. La popup sigue mostrando la nota con metadatos. |
|
||||
| Red lenta / timeout 4 s | `AbortSignal.timeout(4000)` corta la petición. Se prueba el siguiente cliente. |
|
||||
| CORS aún bloqueado pese a DNR 9002 | En Chrome MV3 no hay fallback automático desde background (no usa `fetchProxy` para YouTube). El error queda silenciado dentro del `try { … } catch {}` y la nota se guarda sin transcripción. |
|
||||
| Página no es `youtube.com` (p.ej. embed de otro dominio) | `canExtractAsync()` devuelve `false`; Defuddle cae a su extractor genérico (`BbcodeDataExtractor` según `ExtractorRegistry.mappings[last]`). |
|
||||
|
||||
---
|
||||
|
||||
## 8 · Datos estáticos y constantes que un LLM debe conocer
|
||||
|
||||
### 8.1 Versiones de cliente (claves de la firma)
|
||||
|
||||
| Variable | Valor | Notas |
|
||||
|---|---|---|
|
||||
| `ANDROID_CLIENT_VERSION` | `20.10.38` | Cliente ANDROID. |
|
||||
| `IOS_CLIENT_VERSION` | `20.10.3` | Cliente IOS. |
|
||||
| `WEB_CLIENT_VERSION` | `2.20240101.00.00` | Cliente WEB clásico. |
|
||||
| `ANDROID_USER_AGENT` | `com.google.android.youtube/20.10.38 (Linux; U; Android 14)` | Solo se envía en el 2º intento. |
|
||||
| `PLAYER_ENDPOINT` | `https://www.youtube.com/youtubei/v1/player?prettyPrint=false` | — |
|
||||
| `NEXT_ENDPOINT` | `https://www.youtube.com/youtubei/v1/next?prettyPrint=false` | Solo `fetchChapters`. |
|
||||
| `TIMEOUT_MS` | `4000` | `AbortSignal.timeout`. |
|
||||
| `LANG_HEADER` | `Accept-Language` (opcional) | Se añade sólo si `options.language` está definido. |
|
||||
|
||||
### 8.2 Reglas DNR declaradas por background
|
||||
|
||||
| id | Nombre interno | Trigger | Acción | Uso |
|
||||
|---|---|---|---|---|
|
||||
| `9001` | `enableYouTubeEmbedRule` | `urlFilter:"||youtube.com/embed/"`, `resourceTypes:["sub_frame"]`, `tabIds:[<sender tab>]` | `set Referer: https://obsidian.md/` | Solo embeds (no transcripción). |
|
||||
| `9002` | `enableYouTubeInnertubeRule` | `urlFilter:"||youtube.com/youtubei/"`, `resourceTypes:["xmlhttprequest"]`, `initiatorDomains:[chrome.runtime.id]` | `set Origin: https://www.youtube.com`, `set Referer: https://www.youtube.com/` | **Clave para que funcione la extracción.** |
|
||||
|
||||
### 8.3 Selectores DOM relevantes
|
||||
|
||||
| Uso | Selector |
|
||||
|---|---|
|
||||
| Panel de transcripción (desktop) | `ytd-engagement-panel-section-list-renderer[target-id="engagement-panel-searchable-transcript"] #segments-container` |
|
||||
| Botón "Mostrar transcripción" | `ytd-video-description-transcript-section-renderer button` |
|
||||
| Idioma seleccionado en panel | `ytd-engagement-panel-section-list-renderer[target-id="engagement-panel-searchable-transcript"] #footer yt-sort-filter-sub-menu-renderer yt-dropdown-menu button` |
|
||||
| Panel mobile | `ytm-macro-markers-list-renderer .ytm-macro-markers-list-container` |
|
||||
| Botón "Show more" mobile | `button[aria-label="Show more"]` |
|
||||
| Botón "View all" mobile | `button[aria-label="View all"]` |
|
||||
| Segmento desktop | `ytd-transcript-segment-renderer` (timestamp `.segment-timestamp`, texto `.segment-text`) |
|
||||
| Segmento mobile | `transcript-segment-view-model` (timestamp `.ytwTranscriptSegmentViewModelTimestamp`, texto `span.yt-core-attributed-string`) |
|
||||
| Metadatos (LD+JSON) | `script[type="application/ld+json"]` buscando `VideoObject` |
|
||||
| OpenGraph | `meta[property="og:title|og:description|og:image|og:url"]` |
|
||||
| Nombre de canal (DOM) | `[itemprop="name"]` / `link[itemprop="name"]` / `a, span` |
|
||||
| JSON embebido | `<script>` con `ytInitialPlayerResponse` o `ytInitialData` |
|
||||
|
||||
### 8.4 Mensajes i18n ligados a la transcripción
|
||||
|
||||
`_locales/*/messages.json` contiene (en cada idioma) entradas como:
|
||||
|
||||
- `readerTranscripts` → "Transcripts" / "Transcripciones" / "Transcriptions" / …
|
||||
- `readerHighlightActiveLine` / `…Description` → toggle de la línea activa
|
||||
- `readerPinPlayer` / `…Description` → fija el reproductor
|
||||
- `readerAutoScroll` / `…Description` → auto‑scroll durante reproducción
|
||||
- `readerThemeSection` → tema de la transcripción en el Reader
|
||||
|
||||
(Solo configuran el render, no el proceso de extracción.)
|
||||
|
||||
---
|
||||
|
||||
## 9 · Diagrama textual del flujo de red
|
||||
|
||||
```
|
||||
┌──────────────────────────────────────────────────────┐
|
||||
│ popup.html / side-panel.html (chrome-extension://) │
|
||||
│ ┌────────────────┐ │
|
||||
│ │ popup.js │ │
|
||||
│ │ Defuddle + │ │
|
||||
│ │ YoutubeExtractor │
|
||||
│ └───┬────────────┘ │
|
||||
│ │ globalThis.fetch (XHR) │
|
||||
│ ▼ │
|
||||
│ https://www.youtube.com/youtubei/v1/player?… │
|
||||
│ Body: { context: <IOS|ANDROID|WEB>, videoId } │
|
||||
└──────────────────────────────────────────────────────┘
|
||||
│ (DNR rule 9002)
|
||||
│ set Origin: https://www.youtube.com
|
||||
│ set Referer: https://www.youtube.com/
|
||||
▼
|
||||
YouTube InnerTube API
|
||||
│ JSON con captionTracks[*].baseUrl
|
||||
▼
|
||||
https://www.youtube.com/api/timedtext?…&fmt=…&v=…&lang=… (track.baseUrl)
|
||||
│ Headers: User-Agent: Mozilla/5.0
|
||||
│ Accept-Language: <opción del usuario>
|
||||
▼
|
||||
XML de subtítulos
|
||||
│ parseTranscriptXml()
|
||||
▼
|
||||
{ html, text, languageCode }
|
||||
│
|
||||
▼
|
||||
variables.transcript + variables.language
|
||||
⇒ se inyecta en la nota Markdown
|
||||
```
|
||||
|
||||
Paralelo: `fetchChapters` → `https://www.youtube.com/youtubei/v1/next?…` con `client: WEB` ⇒ `playerOverlays…markersMap` ⇒ capítulos embebidos como `## H2` dentro del HTML de la transcripción.
|
||||
|
||||
---
|
||||
|
||||
## 10 · Riesgos y consideraciones para mantenimiento
|
||||
|
||||
1. **Versiones hard‑coded de cliente (`20.10.38`, `20.10.3`, `2.20240101.00.00`)**: YouTube rota las firmas mensualmente. Si la transcripción deja de funcionar, lo primero a actualizar son estas tres constantes en `popup.js` y `reader-page.js` (buscar `clientVersion` y `20.10.38`).
|
||||
2. **DNR `9002` solo en `xmlhttprequest`**: si YouTube migra a `fetch` puro o cambia el `resourceType`, la regla deja de aplicar. Verificar `chrome.declarativeNetRequest.getEnabledRulesets()` en la consola de la extensión.
|
||||
3. **`fetchProxy` NO se usa para YouTube**: no intentes enrutar por ahí; el flujo correcto es `globalThis.fetch` + DNR.
|
||||
4. **El content script no extrae transcripción**: si la popup falla, no hay fallback desde `content.js`. Mejorar la extracción implica tocar `popup.js` y `reader-page.js` (que son el mismo bundle).
|
||||
5. **Cache `Map<key, transcript>`**: existe en `BilibiliExtractor.transcriptCache` (LRU con cap 300). `YoutubeExtractor` **no** cachea transcripciones; cada `extract` rehace la red.
|
||||
6. **Permisos**: el manifest pide `<all_urls>` y `declarativeNetRequest`. Si se reduce a un `optional_host_permissions` específico de YouTube, la popup seguirá funcionando porque `youtubei` está cubierto por `host_permissions`, pero el `webRequest` listener de Firefox puede dejar de aplicar.
|
||||
7. **Time limit de service worker (MV3)**: el SW se duerme tras 30 s. `enableYouTubeInnertubeRule` se registra en `initialize()` al arrancar y se elimina solo si se pide `disableYouTubeInnertubeRule` (que no existe en el código actual). En la práctica, la regla es *session‑scoped* y dura lo que dure la sesión de Chrome.
|
||||
8. **`ytInitialPlayerResponse` puede no estar presente** en páginas con cookie consent previo o si el usuario está en `consent.youtube.com`. La cascada ANDROID/IOS/WEB lo cubre.
|
||||
|
||||
---
|
||||
|
||||
## 11 · Resumen para indexar/embeddings
|
||||
|
||||
> Texto generado para ser embeddings‑friendly. 7 frases autocontenidas.
|
||||
|
||||
1. La extensión **Obsidian Web Clipper 1.7.1** extrae la transcripción de YouTube desde la **popup** (contexto de la extensión, `chrome-extension://`), nunca desde el content script.
|
||||
2. La clase **`YoutubeExtractor`** (minificada en `popup.js` y duplicada en `reader-page.js`, originalmente de Defuddle) ofrece una cascada de tres rutas: (a) parseo del JSON embebido `ytInitialPlayerResponse`, (b) scraping del panel de transcripción DOM con selectores `ytd-transcript-segment-renderer` / `transcript-segment-view-model`, (c) peticiones a la API privada **`youtubei/v1/player`** con contextos de cliente ANDROID / IOS / WEB.
|
||||
3. El bypass de CORS / firma se hace con una regla **`declarativeNetRequest` id 9002** que fuerza `Origin: https://www.youtube.com` y `Referer: https://www.youtube.com/` para todas las XHR a `||youtube.com/youtubei/` iniciadas por la propia extensión; Firefox usa `webRequest.onBeforeSendHeaders` en su lugar.
|
||||
4. La pista de subtítulos descargada (`track.baseUrl` con sufijo `fmt=`) es XML y se parsea con dos regex: `<p t="N">…<s>…</s>…</p>` (formato moderno) y `<text start="N">…</text>` (formato legacy), produciendo HTML con clase `transcript-segment` y texto plano con timestamps `HH:MM:SS`.
|
||||
5. Los **capítulos** se extraen de `ytInitialData` o, en su defecto, de `youtubei/v1/next` con cliente WEB, parseando `playerOverlays.playerOverlayRenderer.decoratedPlayerBarRenderer.multiMarkersPlayerBarRenderer.markersMap`.
|
||||
6. La popup usa `globalThis.fetch` directo (sin proxy) para YouTube; el `fetchProxy`/`nativeFetch` de `background.js` solo se utiliza como fallback CORS para Bilibili y otros hosts, no para YouTube.
|
||||
7. El resultado (`{transcript.text, language}`) se inyecta en la nota Markdown final bajo la variable `transcript` / `language`; la nota incluye también un `<iframe src="https://www.youtube.com/embed/{id}">` y metadatos (`title`, `author`, `image`, `published`, `description`).
|
||||
|
||||
---
|
||||
|
||||
*Fin del informe.*
|
||||
@@ -0,0 +1,40 @@
|
||||
# Backlog
|
||||
|
||||
Destilado del histórico `OPPORTUNITIES.md`: qué sigue siendo candidato a construir y qué ya existe. Criterio de inclusión en "pendiente": no implementado hoy en el repo.
|
||||
|
||||
## Ya está hecho (no re-abrir)
|
||||
|
||||
| Feature (ítem original) | Dónde vive |
|
||||
|---|---|
|
||||
| `search` — búsqueda FTS5 en transcripciones | `yt-scraper search "..."` + webapp |
|
||||
| `export` json/csv/srt/html | `yt-scraper export --format ...` → `data/exports/` |
|
||||
| `audio` — descarga MP3 | `yt-scraper audio` + job de audio de la webapp (ffmpeg) |
|
||||
| `re-render` + `--backfill` | `yt-scraper re-render` |
|
||||
| Multi-canal | `yt-scraper channels add/list/remove`, `--all-channels` |
|
||||
| `watch` — monitoreo periódico | `yt-scraper watch --interval 6h` |
|
||||
| Análisis (top-words, wordcloud, timeline) | `yt-scraper analyze` → `data/analysis/` |
|
||||
| Miniaturas | `data/thumbnails/` + herramienta *Download thumbnails* de la webapp |
|
||||
| Vista grid de vídeos | Webapp (grid de tarjetas con miniatura) |
|
||||
| Cookies (vault, import, activación) | Webapp + flags `--cookies` / `--cookies-from-browser` — ver [COOKIES.md](COOKIES.md) |
|
||||
| Sync incremental de canales | `sync:` en config; discovery de ventana con overlap (no estaba en el original) |
|
||||
|
||||
## Pendiente
|
||||
|
||||
| Feature | Por qué (1 línea) | Esfuerzo |
|
||||
|---|---|---|
|
||||
| `clip` — texto entre dos timestamps (`yt-scraper clip <id> --from 01:38 --to 03:20`) | Citar fragmentos sin abrir el `.md`; los segmentos ya están en la BD | ~30 min |
|
||||
| `stats` como subcomando | La webapp tiene dashboard y el CLI imprime un resumen tras scrape, pero no existe `yt-scraper stats` con rango/wordcount/tags top | ~1 h |
|
||||
| Traducción de transcripciones (es→en o inversa) | Bibliotecas bilingües y notas para no hispanohablantes; hoy el picker rechaza traducciones por diseño (guarda el transcript real) | 3-4 h vía `deep-translator`; más si se quiere calidad |
|
||||
| Alertas al terminar un job (notificación de escritorio) | Jobs largos sin vigilancia; hoy solo SSE en la pestaña abierta | 2-3 h |
|
||||
| Export a formatos de note-taking (Obsidian Canvas, Anki) | El contenido ya está segmentado por capítulos; solo cambia la salida | 3-4 h |
|
||||
|
||||
## Fuera de alcance (por decisión de plataforma)
|
||||
|
||||
El spec (`docs/superpowers/specs/2026-07-26-platform-design.md` §14) excluye explícitamente features de IA. Reabrirlas exige cambiar el spec primero, no solo escribir código:
|
||||
|
||||
| Feature (ítem original) | Motivo de la exclusión |
|
||||
|---|---|
|
||||
| LLM summaries (`--summarize`) | Requiere API key / modelo local; rompe el "100% local, sin auth" |
|
||||
| Q&A / RAG (`yt-scraper ask`) | Ídem; FTS5 cubre la parte recuperativa sin LLM |
|
||||
| NER / extracción de entidades | Dependencia pesada (spaCy) y caso de uso cubierto por search + top-words |
|
||||
| Speaker diarization | Complejidad alta (audio models) para ganancia marginal en canales monohablante |
|
||||
@@ -0,0 +1,308 @@
|
||||
# Discovery-Only Scrape Implementation Plan
|
||||
|
||||
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
||||
|
||||
**Goal:** Add one-channel and all-channel discovery jobs that register new YouTube videos as `pending` without downloading or processing their content.
|
||||
|
||||
**Architecture:** Extend the existing sequential `JobManager` with `opts.mode == "discover"`. Reuse `POST /api/scrape`, its SSE stream, and its job widget; add UI actions in Channels and Videos that pass either a channel ID or no channel ID. Keep the existing full scrape and processing endpoints unchanged.
|
||||
|
||||
**Tech Stack:** Python 3.10+, FastAPI, SQLite via the existing `Store`, yt-dlp flat discovery, Alpine.js CDN SPA, pytest.
|
||||
|
||||
## Global Constraints
|
||||
|
||||
- Discovery-only jobs must not call `process_video()`.
|
||||
- Discovery-only jobs must not download transcripts, Markdown, audio, thumbnails, or avatars.
|
||||
- All HTTP requests must use the existing job queue and SSE event stream.
|
||||
- All database access must remain inside `Store` methods.
|
||||
- Preserve existing user changes in the dirty worktree and edit only the files listed below.
|
||||
|
||||
---
|
||||
|
||||
### Task 1: Make catalog insertion counts accurate
|
||||
|
||||
**Files:**
|
||||
- Modify: `src/yt_scraper/store.py:266-283`
|
||||
- Test: `tests/test_store_platform.py`
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: existing `Store.upsert_videos(refs: list[VideoRef])` callers.
|
||||
- Produces: the same integer return type, now equal to the number of distinct references that were not already present.
|
||||
|
||||
- [ ] **Step 1: Write the failing test**
|
||||
|
||||
Add this test after the existing store filtering tests:
|
||||
|
||||
```python
|
||||
def test_upsert_videos_reports_only_new_rows(seeded_store):
|
||||
inserted = seeded_store.upsert_videos([
|
||||
VideoRef("v1", "UC1", "Updated title", "https://y/watch?v=v1", "20240101", 120),
|
||||
VideoRef("v4", "UC1", "New video", "https://y/watch?v=v4", "20240501", 180),
|
||||
])
|
||||
|
||||
assert inserted == 1
|
||||
assert seeded_store.get_video("v1").title == "Updated title"
|
||||
assert seeded_store.get_video("v4").status == "pending"
|
||||
```
|
||||
|
||||
- [ ] **Step 2: Run the focused test to verify it fails**
|
||||
|
||||
Run: `python -m pytest tests/test_store_platform.py::test_upsert_videos_reports_only_new_rows -q`
|
||||
|
||||
Expected: FAIL because SQLite's `ON CONFLICT DO UPDATE` currently makes `rowcount` nonzero for an existing row.
|
||||
|
||||
- [ ] **Step 3: Implement the minimal count fix**
|
||||
|
||||
Inside `Store.upsert_videos`, load the existing IDs for the incoming references before the loop, deduplicate the incoming IDs with a set, and increment `inserted` only when an incoming distinct ID was absent. Keep the current upsert SQL so existing rows still refresh title, upload date, and duration.
|
||||
|
||||
The essential implementation shape is:
|
||||
|
||||
```python
|
||||
incoming = {r.video_id for r in refs}
|
||||
with self._cursor() as cur:
|
||||
existing = set()
|
||||
if incoming:
|
||||
placeholders = ",".join("?" for _ in incoming)
|
||||
cur.execute(
|
||||
f"SELECT video_id FROM videos WHERE video_id IN ({placeholders})",
|
||||
list(incoming),
|
||||
)
|
||||
existing = {row["video_id"] for row in cur.fetchall()}
|
||||
for r in refs:
|
||||
# existing SQL remains here
|
||||
if r.video_id not in existing:
|
||||
inserted += 1
|
||||
existing.add(r.video_id)
|
||||
```
|
||||
|
||||
- [ ] **Step 4: Run the focused and regression tests**
|
||||
|
||||
Run: `python -m pytest tests/test_store_platform.py -q`
|
||||
|
||||
Expected: PASS for all store tests.
|
||||
|
||||
- [ ] **Step 5: Commit**
|
||||
|
||||
Do not commit automatically in this workspace unless the user explicitly requests a commit. Leave the focused diff ready for review.
|
||||
|
||||
### Task 2: Add discovery-only job execution
|
||||
|
||||
**Files:**
|
||||
- Modify: `src/yt_scraper/webapp/jobs.py:88-269`
|
||||
- Test: `tests/test_webapp_jobs.py`
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: `JobManager.enqueue(channel_id, opts)` with `opts={"mode": "discover"}` and optional `channel_id`.
|
||||
- Produces: queued jobs whose SSE events contain per-channel discovery counts and whose terminal event contains a summary.
|
||||
|
||||
- [ ] **Step 1: Write failing job tests**
|
||||
|
||||
Append tests using a fake `discover_channel` and a fake `process_video`:
|
||||
|
||||
```python
|
||||
def test_discovery_job_registers_new_videos_without_processing(tmp_path, monkeypatch):
|
||||
store = Store(tmp_path / "state.db")
|
||||
store.upsert_channel("UC1", "@alpha", "Alpha", 1)
|
||||
store.upsert_videos([
|
||||
VideoRef("old", "UC1", "Old", "https://y/watch?v=old", "20240101", 60)
|
||||
])
|
||||
store.mark_status("old", "done")
|
||||
store.create_job("discover-job", "UC1", {"mode": "discover"})
|
||||
calls = []
|
||||
|
||||
def fake_discover(url, sleep_subrequests=2.0):
|
||||
calls.append(url)
|
||||
return ("UC1", "Alpha", None, [
|
||||
VideoRef("new", "UC1", "New", "https://y/watch?v=new", "20240501", 90),
|
||||
VideoRef("old", "UC1", "Old", "https://y/watch?v=old", "20240101", 60),
|
||||
])
|
||||
|
||||
processed = []
|
||||
monkeypatch.setattr("yt_scraper.webapp.jobs.discover_channel", fake_discover)
|
||||
monkeypatch.setattr("yt_scraper.webapp.jobs.process_video", lambda *a, **k: processed.append(a))
|
||||
manager = JobManager(store, Config(database_path=str(tmp_path / "state.db"), output_dir=str(tmp_path / "markdown")))
|
||||
|
||||
manager._run_job("discover-job")
|
||||
|
||||
assert calls == ["https://www.youtube.com/@alpha/videos"]
|
||||
assert store.get_video("new").status == "pending"
|
||||
assert store.get_video("old").status == "done"
|
||||
assert processed == []
|
||||
assert store.get_job("discover-job").status == "done"
|
||||
assert any(event["event"] == "done" and event["data"]["new_videos"] == 1
|
||||
for event in manager.events_since("discover-job", 0))
|
||||
|
||||
|
||||
def test_all_channel_discovery_continues_after_one_error(tmp_path, monkeypatch):
|
||||
store = Store(tmp_path / "state.db")
|
||||
store.upsert_channel("UC1", "@one", "One", 0)
|
||||
store.upsert_channel("UC2", "@two", "Two", 0)
|
||||
store.create_job("discover-all", None, {"mode": "discover"})
|
||||
|
||||
def fake_discover(url, sleep_subrequests=2.0):
|
||||
if "@one" in url:
|
||||
raise RuntimeError("temporary failure")
|
||||
return ("UC2", "Two", None, [VideoRef("new2", "UC2", "New", "https://y/watch?v=new2")])
|
||||
|
||||
monkeypatch.setattr("yt_scraper.webapp.jobs.discover_channel", fake_discover)
|
||||
manager = JobManager(store, Config(database_path=str(tmp_path / "state.db"), output_dir=str(tmp_path / "markdown")))
|
||||
|
||||
manager._run_job("discover-all")
|
||||
|
||||
assert store.get_video("new2").status == "pending"
|
||||
assert store.get_job("discover-all").status == "done"
|
||||
done = [e for e in manager.events_since("discover-all", 0) if e["event"] == "done"][-1]
|
||||
assert done["data"]["errors"] == 1
|
||||
```
|
||||
|
||||
- [ ] **Step 2: Run the focused tests to verify they fail**
|
||||
|
||||
Run: `python -m pytest tests/test_webapp_jobs.py::test_discovery_job_registers_new_videos_without_processing tests/test_webapp_jobs.py::test_all_channel_discovery_continues_after_one_error -q`
|
||||
|
||||
Expected: FAIL because `_run_job` currently routes both cases to `_run_channel`, which tries to process pending videos and cannot interpret an all-channel discovery job.
|
||||
|
||||
- [ ] **Step 3: Implement discovery dispatch and execution**
|
||||
|
||||
In `JobManager._run_job`, route `opts.get("mode") == "discover"` to a new `_run_discovery(job_id)` before the existing audio/video/channel branches.
|
||||
|
||||
Implement `_run_discovery` with this behavior:
|
||||
|
||||
```python
|
||||
def _run_discovery(self, job_id: str) -> None:
|
||||
job = self.store.get_job(job_id)
|
||||
if not job:
|
||||
return
|
||||
channels = ([self.store.get_channel(job.channel_id)] if job.channel_id
|
||||
else self.store.list_channels())
|
||||
channels = [c for c in channels if c]
|
||||
self.store.update_job(job_id, status="running", total=len(channels), completed=0)
|
||||
completed = 0
|
||||
totals = {"new_videos": 0, "known_videos": 0, "errors": 0}
|
||||
for channel in channels:
|
||||
if job_id in self._cancel:
|
||||
self.store.update_job(job_id, status="cancelled", completed=completed, finished=True)
|
||||
self._emit(job_id, "cancelled", {"completed": completed, "total": len(channels)})
|
||||
return
|
||||
try:
|
||||
# Build the stored handle/channel URL, call discover_channel, filter
|
||||
# using cfg.include_shorts/cfg.include_live, upsert the channel and refs.
|
||||
# Do not call deep_channel_avatar, cache helpers, or process_video.
|
||||
new_count = self.store.upsert_videos(refs)
|
||||
known_count = len({r.video_id for r in refs}) - new_count
|
||||
totals["new_videos"] += new_count
|
||||
totals["known_videos"] += known_count
|
||||
self._emit(job_id, "progress", {"channel_id": channel["channel_id"], "new_videos": new_count, "known_videos": known_count, "completed": completed + 1, "total": len(channels)})
|
||||
except Exception as exc:
|
||||
totals["errors"] += 1
|
||||
self._emit(job_id, "log", {"msg": f"discovery failed for {channel.get('name') or channel['channel_id']}: {exc}"})
|
||||
completed += 1
|
||||
self.store.update_job(job_id, completed=completed)
|
||||
self.store.update_job(job_id, status="done", completed=completed, finished=True)
|
||||
self._emit(job_id, "done", {"completed": completed, "total": len(channels), **totals})
|
||||
```
|
||||
|
||||
Use the existing `_resolve_channel_url` for each stored channel. Preserve the existing processing path untouched. For a single-channel exception, emit an error terminal event/status; for all-channel jobs, continue and report the error count as above.
|
||||
|
||||
- [ ] **Step 4: Run focused job tests**
|
||||
|
||||
Run: `python -m pytest tests/test_webapp_jobs.py -q`
|
||||
|
||||
Expected: PASS for both existing audio tests and the new discovery tests.
|
||||
|
||||
- [ ] **Step 5: Run the full Python suite**
|
||||
|
||||
Run: `python -m pytest tests/ -q`
|
||||
|
||||
Expected: all tests pass with no network access.
|
||||
|
||||
### Task 3: Expose the discovery mode through the API
|
||||
|
||||
**Files:**
|
||||
- Modify: `src/yt_scraper/webapp/api.py:111-117`
|
||||
- Test: `tests/test_webapp_jobs.py` (job boundary coverage; no API test fixture exists in the repository).
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: JSON `{"channel_id": "UC...", "opts": {"mode": "discover"}}` or the same body without `channel_id` for all channels.
|
||||
- Produces: the existing `{"job_id": "..."}` response and a queued `JobManager` job.
|
||||
|
||||
- [ ] **Step 1: Add request-shape coverage at the job boundary**
|
||||
|
||||
Extend the discovery job tests to enqueue `manager.enqueue(None, {"mode": "discover"})` and verify that it completes as an all-channel job. This covers the exact payload shape the API forwards; the repository has no existing FastAPI router test fixture.
|
||||
|
||||
- [ ] **Step 2: Implement the minimal API change**
|
||||
|
||||
Keep the endpoint response unchanged. Normalize discovery options and permit a missing channel only for discovery:
|
||||
|
||||
```python
|
||||
opts = (payload or {}).get("opts", {}) or {}
|
||||
if opts.get("mode") == "discover":
|
||||
opts = {"mode": "discover"}
|
||||
job_id = jobs.enqueue(channel_id, opts)
|
||||
return {"job_id": job_id}
|
||||
```
|
||||
|
||||
Do not add a second synchronous discovery endpoint.
|
||||
|
||||
- [ ] **Step 3: Run the API/job regression tests**
|
||||
|
||||
Run: `python -m pytest tests/test_webapp_jobs.py tests/test_store_platform.py -q`
|
||||
|
||||
Expected: PASS.
|
||||
|
||||
### Task 4: Add one-channel and all-channel UI controls
|
||||
|
||||
**Files:**
|
||||
- Modify: `src/yt_scraper/webapp/static/app.js:45-73, 758-815`
|
||||
- Modify: `src/yt_scraper/webapp/static/index.html:167-229, 231-319`
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: `POST /api/scrape` discovery jobs and SSE `progress`/`done` payloads.
|
||||
- Produces: buttons for channel-specific and all-channel discovery, with refresh and result feedback.
|
||||
|
||||
- [ ] **Step 1: Add the failing static contract checks**
|
||||
|
||||
Before editing, use a lightweight repository check to confirm the new labels and method are absent:
|
||||
|
||||
Run: `rg "startDiscovery|Investigar nuevos|Investigar todos" src/yt_scraper/webapp/static`
|
||||
|
||||
Expected: no matches.
|
||||
|
||||
- [ ] **Step 2: Add the Alpine discovery action**
|
||||
|
||||
Add `startDiscovery(channelId)` near `startScrape()`. It posts `{channel_id: channelId || null, opts: {mode: "discover"}}`, subscribes with kind `discovery`, and prevents duplicate active jobs. Extend the SSE `done` handler to parse the payload and show `N video(s) nuevo(s) encontrado(s)` for discovery jobs. Refresh channels, videos, and dashboard after completion.
|
||||
|
||||
- [ ] **Step 3: Add controls to Channels and Videos**
|
||||
|
||||
In Channels, add the global header button and one row button calling `startDiscovery(c.channel_id)`. In Videos, add a header action calling `startDiscovery(filters.channel || null)`. Disable each while `jobActive()` is true and while the channel list is empty. Preserve `.md`, `Process`, and `Audio` actions.
|
||||
|
||||
- [ ] **Step 4: Verify the static contract and inspect the rendered paths**
|
||||
|
||||
Run: `rg "startDiscovery|Investigar nuevos|Investigar todos|mode: \"discover\"" src/yt_scraper/webapp/static`
|
||||
|
||||
Expected: matches in `app.js` and `index.html` for the method, payload, and both scopes. Inspect the changed Alpine expressions to ensure the existing `app.js` script remains before Alpine in `index.html`.
|
||||
|
||||
### Task 5: Final verification and review
|
||||
|
||||
**Files:**
|
||||
- Review: `src/yt_scraper/store.py`
|
||||
- Review: `src/yt_scraper/webapp/jobs.py`
|
||||
- Review: `src/yt_scraper/webapp/api.py`
|
||||
- Review: `src/yt_scraper/webapp/static/app.js`
|
||||
- Review: `src/yt_scraper/webapp/static/index.html`
|
||||
|
||||
- [ ] **Step 1: Run the complete test suite**
|
||||
|
||||
Run: `python -m pytest tests/ -q`
|
||||
|
||||
Expected: all tests pass.
|
||||
|
||||
- [ ] **Step 2: Review the diff for scope and unintended downloads**
|
||||
|
||||
Run: `git diff -- src/yt_scraper/store.py src/yt_scraper/webapp/jobs.py src/yt_scraper/webapp/api.py src/yt_scraper/webapp/static/app.js src/yt_scraper/webapp/static/index.html tests/test_store_platform.py tests/test_webapp_jobs.py`
|
||||
|
||||
Confirm discovery code contains no calls to `process_video`, `cache_thumbnail`, `cache_channel_avatar`, `extract_video`, or audio routes.
|
||||
|
||||
- [ ] **Step 3: Check worktree status**
|
||||
|
||||
Run: `git status --short`
|
||||
|
||||
Confirm only intended feature files and the two planning documents are changed or untracked; do not modify or revert unrelated existing user changes.
|
||||
@@ -0,0 +1,195 @@
|
||||
# Fechas aproximadas en discovery Implementation Plan
|
||||
|
||||
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
||||
|
||||
**Goal:** Llenar `videos.upload_date` con fechas aproximadas (flag `upload_date_approx`) durante discovery/sync a costo cero, con upgrade automático a real al extraer.
|
||||
|
||||
**Architecture:** `youtubetab:approximate_date` en los ydl_opts de discovery → entries flat traen `timestamp` → helper puro convierte a `YYYYMMDD` + flag. `upsert_videos` gana matriz de precedencia; `set_upload_date` hace el upgrade approx→real. UI marca `~`.
|
||||
|
||||
**Tech Stack:** Python 3.12, SQLite sin ORM, pytest con tmp_path.
|
||||
|
||||
**Spec:** `docs/superpowers/specs/2026-08-22-approximate-upload-dates-design.md`
|
||||
|
||||
## Global Constraints
|
||||
|
||||
- 0 peticiones HTTP nuevas por design; no tocar economía de requests.
|
||||
- Migración SOLO vía dict `_VIDEO_COLUMNS` (patrón del repo).
|
||||
- `NO_DATE_SENTINEL` y cascada `SORT_DATE_SQL` intocables.
|
||||
- Tests sin red: helpers puros y Store(tmp_path).
|
||||
|
||||
---
|
||||
|
||||
### Task 1: store.py — columna, bandera y matriz de upsert
|
||||
|
||||
**Files:**
|
||||
- Modify: `src/yt_scraper/store.py` (`_VIDEO_COLUMNS` ~44, `VideoRef` ~251, `VideoRow` ~268, `upsert_videos` ~500, `set_upload_date` ~756, `_row_to_videorow` ~1146)
|
||||
- Test: `tests/test_approx_dates.py` (nuevo)
|
||||
|
||||
**Interfaces:**
|
||||
- Produces: `VideoRef.date_approx: int = 0`; `VideoRow.upload_date_approx: int | None = None`; semántica de upsert según spec.
|
||||
|
||||
- [ ] **Step 1: test que falla** — crear `tests/test_approx_dates.py` con la matriz completa (4 casos + upgrade). Ver patrón de construcción de Store en tests/test_store_platform.py.
|
||||
|
||||
```python
|
||||
from yt_scraper.store import Store, VideoRef
|
||||
|
||||
def _ref(vid, date=None, approx=0, ch="ch1"):
|
||||
return VideoRef(video_id=vid, channel_id=ch, title="t " + vid,
|
||||
url=f"https://youtu.be/{vid}", upload_date=date, date_approx=approx)
|
||||
|
||||
def _store(tmp_path):
|
||||
return Store(tmp_path)
|
||||
|
||||
def test_approx_into_empty(tmp_path):
|
||||
s = _store(tmp_path); s.upsert_videos([_ref("v1", date="20260820", approx=1)])
|
||||
row = s.get_video("v1")
|
||||
assert row.upload_date == "20260820" and row.upload_date_approx == 1
|
||||
|
||||
def test_real_not_downgraded_by_approx(tmp_path):
|
||||
s = _store(tmp_path)
|
||||
s.upsert_videos([_ref("v1", date="20240101", approx=0)])
|
||||
s.upsert_videos([_ref("v1", date="20260820", approx=1)])
|
||||
row = s.get_video("v1")
|
||||
assert row.upload_date == "20240101" and row.upload_date_approx == 0
|
||||
|
||||
def test_approx_refreshed_by_newer_approx(tmp_path):
|
||||
s = _store(tmp_path)
|
||||
s.upsert_videos([_ref("v1", date="20260101", approx=1)])
|
||||
s.upsert_videos([_ref("v1", date="20260820", approx=1)])
|
||||
row = s.get_video("v1")
|
||||
assert row.upload_date == "20260820" and row.upload_date_approx == 1
|
||||
|
||||
def test_extraction_upgrades_approx_to_real(tmp_path):
|
||||
s = _store(tmp_path)
|
||||
s.upsert_videos([_ref("v1", date="20260820", approx=1)])
|
||||
s.set_upload_date("v1", "20260818")
|
||||
row = s.get_video("v1")
|
||||
assert row.upload_date == "20260818" and row.upload_date_approx == 0
|
||||
|
||||
def test_set_upload_date_still_skips_existing_real(tmp_path):
|
||||
s = _store(tmp_path)
|
||||
s.upsert_videos([_ref("v1", date="20240101", approx=0)])
|
||||
s.set_upload_date("v1", "20260818")
|
||||
assert s.get_video("v1").upload_date == "20240101"
|
||||
```
|
||||
|
||||
(Adaptar nombres reales: constructor `Store`, getter `get_video`; verificar en el código antes.)
|
||||
|
||||
- [ ] **Step 2: correr y ver fallo** — `python -m pytest tests/test_approx_dates.py -q` → FAIL (TypeError/AttributeError por falta de campos).
|
||||
|
||||
- [ ] **Step 3: implementar mínimo**
|
||||
- `_VIDEO_COLUMNS["upload_date_approx"] = "INTEGER DEFAULT 0"`
|
||||
- `VideoRef`: `date_approx: int = 0`
|
||||
- `VideoRow`: `upload_date_approx: int | None = None` (antes de sort_date)
|
||||
- `_row_to_videorow`: `upload_date_approx=row["upload_date_approx"] if "upload_date_approx" in keys else None`
|
||||
- `upsert_videos`: INSERT incluye columna+valor; ON CONFLICT:
|
||||
```sql
|
||||
upload_date = CASE
|
||||
WHEN excluded.upload_date IS NULL THEN videos.upload_date
|
||||
WHEN videos.upload_date IS NULL THEN excluded.upload_date
|
||||
WHEN videos.upload_date_approx = 1 THEN excluded.upload_date
|
||||
ELSE videos.upload_date END,
|
||||
upload_date_approx = CASE
|
||||
WHEN videos.upload_date IS NOT NULL AND COALESCE(videos.upload_date_approx, 0) = 0
|
||||
THEN videos.upload_date_approx
|
||||
ELSE COALESCE(excluded.upload_date_approx, videos.upload_date_approx) END,
|
||||
```
|
||||
- `set_upload_date`:
|
||||
```sql
|
||||
UPDATE videos SET upload_date = ?, upload_date_approx = 0
|
||||
WHERE video_id = ? AND (upload_date IS NULL OR COALESCE(upload_date_approx, 0) = 1)
|
||||
```
|
||||
|
||||
- [ ] **Step 4: tests pasan** — mismo comando → PASS. Suite completa verde.
|
||||
|
||||
- [ ] **Step 5: commit** — `feat: bandera upload_date_approx + matriz de precedencia en upsert`
|
||||
|
||||
### Task 2: discover.py — extractor arg y conversión timestamp
|
||||
|
||||
**Files:**
|
||||
- Modify: `src/yt_scraper/discover.py` (~31 opts, ~62 loop)
|
||||
- Test: `tests/test_approx_dates.py` (añadir)
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: `VideoRef.date_approx` (Task 1).
|
||||
- Produces: `_entry_upload_date(entry) -> tuple[str | None, int]`.
|
||||
|
||||
- [ ] **Step 1: test del helper que falla**
|
||||
|
||||
```python
|
||||
from yt_scraper.discover import _entry_upload_date
|
||||
|
||||
def test_entry_upload_date_prefers_exact():
|
||||
assert _entry_upload_date({"upload_date": "20240101"}) == ("20240101", 0)
|
||||
|
||||
def test_entry_upload_date_from_timestamp():
|
||||
# 2024-05-19 12:00 UTC
|
||||
ts = 1716115200
|
||||
assert _entry_upload_date({"timestamp": ts}) == ("20240519", 1)
|
||||
|
||||
def test_entry_upload_date_none():
|
||||
assert _entry_upload_date({}) == (None, 0)
|
||||
```
|
||||
|
||||
- [ ] **Step 2: fallo** — ImportError.
|
||||
|
||||
- [ ] **Step 3: implementar**
|
||||
|
||||
```python
|
||||
def _entry_upload_date(entry: dict[str, Any]) -> tuple[str | None, int]:
|
||||
exact = entry.get("upload_date")
|
||||
if exact:
|
||||
return str(exact), 0
|
||||
ts = entry.get("timestamp")
|
||||
if not ts:
|
||||
return None, 0
|
||||
from datetime import datetime, timezone
|
||||
return datetime.fromtimestamp(int(ts), tz=timezone.utc).strftime("%Y%m%d"), 1
|
||||
```
|
||||
|
||||
En `discover_channel`: añadir `"extractor_args": {"youtubetab": {"approximate_date": ["true"]}}` a ydl_opts; loop usa `(upload_date, approx) = _entry_upload_date(entry)` y pasa `date_approx=approx` a VideoRef. Import de datetime a nivel módulo.
|
||||
|
||||
- [ ] **Step 4: pasa** + suite verde. **[ ] Step 5: commit** — `feat: discovery llena fechas aproximadas (approximate_date)`
|
||||
|
||||
### Task 3: API + UI — marca visual
|
||||
|
||||
**Files:**
|
||||
- Modify: `src/yt_scraper/webapp/api.py` (`_video_dict` ~628)
|
||||
- Modify: `src/yt_scraper/webapp/static/app.js` (`videoDateText` ~1266, tooltip ~1274)
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: `VideoRow.upload_date_approx`.
|
||||
- Produces: JSON `upload_date_approx: bool`, `date_estimated` ampliado.
|
||||
|
||||
- [ ] **Step 1: implementar** (sin test automatizado de UI; verificación manual tras reinicio)
|
||||
|
||||
```python
|
||||
"upload_date_approx": bool(v.upload_date_approx),
|
||||
"date_estimated": bool(
|
||||
(v.sort_date and not v.upload_date and v.sort_date != store_mod.NO_DATE_SENTINEL)
|
||||
or v.upload_date_approx
|
||||
),
|
||||
```
|
||||
|
||||
app.js videoDateText:
|
||||
|
||||
```js
|
||||
if (v.upload_date && !v.upload_date_approx) return this.fmtDate(v.upload_date);
|
||||
if (v.upload_date) return "~ " + this.fmtDate(v.upload_date);
|
||||
if (v.sort_date) return "~ " + this.fmtDate(v.sort_date);
|
||||
```
|
||||
|
||||
Tooltip nuevo ANTES del branch date_estimated existente:
|
||||
|
||||
```js
|
||||
if (v.upload_date_approx) return "Approximate upload date — YouTube shows relative text ('3 weeks ago') in listings. Download the .md to learn the exact one.";
|
||||
```
|
||||
|
||||
- [ ] **Step 2: suite verde** (`python -m pytest tests/ -q`). **[ ] Step 3: commit** — `feat: UI marca fechas aproximadas con ~ y tooltip propio`
|
||||
|
||||
### Task 4: verificación end-to-end
|
||||
|
||||
- [ ] **Step 1:** `python -m pytest tests/ -q` completo verde.
|
||||
- [ ] **Step 2:** reiniciar server (stop-server.bat + doctor) para servir api.py nueva.
|
||||
- [ ] **Step 3:** smoke `/api/videos?limit=5` responde 200 con campo nuevo presente.
|
||||
- [ ] **Step 4:** informe final; sugerir al usuario un Full rescan por canal para llenar backlog (acción manual suya, no automática).
|
||||
@@ -0,0 +1,409 @@
|
||||
# Arranque automático e inteligente (.bat doctor) Implementation Plan
|
||||
|
||||
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
||||
|
||||
**Goal:** Doble clic en `start-server.bat` abre siempre la aplicación: el nuevo `doctor.ps1` detecta y repara el entorno (Python falso/faltante, dependencias, ffmpeg) antes de delegar al launcher existente.
|
||||
|
||||
**Architecture:** Separación entorno vs ciclo de vida. Nuevo `scripts/doctor.ps1` prepara el entorno y llama a `scripts/start-server.ps1` (lógica de puertos/healthz/browser intacta, único cambio: parámetro `-PythonExe`). Los `.bat` ganan regla de ventana: error → ventana persistente con log; éxito → cierre automático.
|
||||
|
||||
**Tech Stack:** Windows PowerShell 5.1 (sin PS7), cmd batch, winget como instalador de Python.
|
||||
|
||||
**Spec:** `docs/superpowers/specs/2026-08-22-bat-doctor-design.md`
|
||||
|
||||
## Global Constraints
|
||||
|
||||
- PS 5.1 estricto: sin operador ternario, sin `??`, sin `-AsHashtable`. Compatibilidad obligatoria.
|
||||
- Mensajes **sin tildes** (convención de los scripts existentes: "automaticamente", "no se encontro") para sobrevivir cambios de codepage.
|
||||
- El doctor NUNCA importa `yt_scraper.webapp.app` (abre la DB real del proyecto y corre migraciones/backfill — gotcha documentado en CLAUDE.md). El import-test usa módulos terceros directamente.
|
||||
- Lógica existente de puertos/healthz/cancelación en `start-server.ps1`: intocable. Únicos cambios permitidos: parámetro `-PythonExe` (default `python`) y copys.
|
||||
- Sin dependencias nuevas; winget es la única vía de instalación de Python.
|
||||
- Exit codes doctor: 0 ok · 10 sin Python instalable · 11 deps no reparadas · resto propagado del launcher.
|
||||
|
||||
---
|
||||
|
||||
### Task 1: scripts/doctor.ps1 — diagnóstico y reparación de entorno
|
||||
|
||||
**Files:**
|
||||
- Create: `scripts/doctor.ps1`
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: `.run/server.info` (línea `PORT PID TIMESTAMP`), ejecutables del sistema (`py`, `python`, `python3`, `winget`, `ffmpeg`).
|
||||
- Produces: invoca `scripts/start-server.ps1 -PythonExe <ruta> [args...]` pasando args posicionales restantes. Exit codes 0/10/11 según Global Constraints. Flag `-CheckOnly`: solo diagnostica, no muta nada, mismo esquema de exit codes.
|
||||
|
||||
- [ ] **Step 1: Escribir doctor.ps1 completo**
|
||||
|
||||
Crear `scripts/doctor.ps1` con este contenido exacto:
|
||||
|
||||
```powershell
|
||||
# doctor.ps1 -- prepara el entorno para arrancar yt-scraper y delega en
|
||||
# start-server.ps1. Pensado para usuarios no tecnicos: doble clic y listo.
|
||||
#
|
||||
# Uso: powershell -File doctor.ps1 [-CheckOnly] [args para start-server]
|
||||
# -CheckOnly solo diagnostica e informa que haria; no instala ni abre nada.
|
||||
#
|
||||
# Exit codes: 0 ok | 10 sin Python instalable | 11 dependencias no reparadas
|
||||
# (cualquier otro codigo viene propagado de start-server.ps1)
|
||||
|
||||
param(
|
||||
[switch]$CheckOnly
|
||||
)
|
||||
|
||||
$ErrorActionPreference = 'Stop'
|
||||
$root = Split-Path -Parent $PSScriptRoot
|
||||
Set-Location -LiteralPath $root
|
||||
try { [Console]::OutputEncoding = [System.Text.Encoding]::UTF8 } catch {}
|
||||
|
||||
function Out-Line($msg, $color = $null) {
|
||||
if ($color) { Write-Host $msg -ForegroundColor $color }
|
||||
else { Write-Host $msg }
|
||||
[Console]::Out.Flush()
|
||||
}
|
||||
|
||||
function Test-PortOpen($port, $timeoutMs = 350) {
|
||||
$client = New-Object System.Net.Sockets.TcpClient
|
||||
try {
|
||||
$iar = $client.BeginConnect('127.0.0.1', [int]$port, $null, $null)
|
||||
if (-not $iar.AsyncWaitHandle.WaitOne($timeoutMs)) { return $false }
|
||||
$client.EndConnect($iar)
|
||||
return $true
|
||||
} catch { return $false }
|
||||
finally { $client.Close() }
|
||||
}
|
||||
|
||||
# Devuelve la ruta del exe si el candidato es un Python >= 3.10 REAL;
|
||||
# $null si no existe, es el stub de la Microsoft Store o es demasiado viejo.
|
||||
function Resolve-PythonCandidate($tokens) {
|
||||
$name = $tokens[0]
|
||||
$cmd = Get-Command $name -ErrorAction SilentlyContinue
|
||||
if (-not $cmd) { return $null }
|
||||
# Stub de la Microsoft Store: vive en WindowsApps y no ejecuta nada.
|
||||
if ($cmd.Source -and $cmd.Source -like '*WindowsApps*') { return $null }
|
||||
try {
|
||||
$out = & $cmd.Source @($tokens | Select-Object -Skip 1) --version 2>&1
|
||||
if ($LASTEXITCODE -ne 0) { return $null }
|
||||
$s = ($out | Out-String).Trim()
|
||||
if ($s -notmatch '^Python\s+(\d+)\.(\d+)') { return $null }
|
||||
if ([int]$Matches[1] -lt 3 -or ([int]$Matches[1] -eq 3 -and [int]$Matches[2] -lt 10)) { return $null }
|
||||
return $cmd.Source
|
||||
} catch { return $null }
|
||||
}
|
||||
|
||||
Out-Line ""
|
||||
Out-Line " [yt-scraper] verificando tu sistema..." Cyan
|
||||
|
||||
# --- (1) Server ya vivo? Abrir navegador y terminar. -----------------------
|
||||
if (Test-Path '.run\server.info') {
|
||||
$parts = ((Get-Content '.run\server.info' -TotalCount 1) -split '\s+')
|
||||
if ($parts.Count -ge 1 -and $parts[0] -match '^\d+$' -and (Test-PortOpen ([int]$parts[0]) 350)) {
|
||||
Out-Line " [ok] El servidor ya estaba corriendo - abriendo..." Green
|
||||
Out-Line " http://127.0.0.1:$($parts[0])"
|
||||
if (-not $CheckOnly) { try { Start-Process "http://127.0.0.1:$($parts[0])" } catch {} }
|
||||
exit 0
|
||||
}
|
||||
}
|
||||
|
||||
# --- (2) Resolver Python ----------------------------------------------------
|
||||
$candidates = @(,@('py','-3')) + @(,@('python')) + @(,@('python3'))
|
||||
$pythonExe = $null
|
||||
foreach ($cand in $candidates) {
|
||||
$pythonExe = Resolve-PythonCandidate $cand
|
||||
if ($pythonExe) { break }
|
||||
}
|
||||
|
||||
if (-not $pythonExe) {
|
||||
Out-Line " [info] Python no esta instalado. Instalandolo automaticamente (~25 MB)..." Yellow
|
||||
if ($CheckOnly) {
|
||||
Out-Line " [check] AQUI: winget install Python.Python.3.12 + busqueda de ruta nueva" DarkGray
|
||||
exit 10
|
||||
}
|
||||
$winget = Get-Command winget -ErrorAction SilentlyContinue
|
||||
if (-not $winget) {
|
||||
Out-Line " [!] No pude instalar Python automaticamente (falta winget)." Red
|
||||
Out-Line " Abriendo la pagina de descarga de Python..." Gray
|
||||
Out-Line " Instalalo (marca 'Add python.exe to PATH') y vuelve a hacer doble clic." Gray
|
||||
try { Start-Process 'https://www.python.org/downloads/' } catch {}
|
||||
exit 10
|
||||
}
|
||||
& winget install --id Python.Python.3.12 --silent --accept-package-agreements --accept-source-agreements | Out-Null
|
||||
if ($LASTEXITCODE -ne 0) {
|
||||
Out-Line " [!] La instalacion de Python fallo (codigo $LASTEXITCODE). Reinstala manualmente:" Red
|
||||
Out-Line " https://www.python.org/downloads/" Gray
|
||||
exit 10
|
||||
}
|
||||
# El PATH de esta sesion no se refresca: sondear rutas conocidas.
|
||||
$fresh = @()
|
||||
$fresh += Get-ChildItem "$env:LOCALAPPDATA\Programs\Python\Python3*\python.exe" -ErrorAction SilentlyContinue
|
||||
$fresh += Get-ChildItem 'C:\Program Files\Python3*\python.exe' -ErrorAction SilentlyContinue
|
||||
foreach ($f in ($fresh | Sort-Object FullName -Descending)) {
|
||||
$pythonExe = Resolve-PythonCandidate @($f.FullName)
|
||||
if ($pythonExe) { break }
|
||||
}
|
||||
if (-not $pythonExe) { $pythonExe = Resolve-PythonCandidate @('py','-3') }
|
||||
if (-not $pythonExe) {
|
||||
Out-Line " [!] Se instalo Python pero no lo encuentro. Cierra esta ventana," Red
|
||||
Out-Line " abre una nueva e intenta de nuevo (el PATH se refresca al reabrir)." Gray
|
||||
exit 10
|
||||
}
|
||||
}
|
||||
Out-Line " [ok] Python encontrado: $pythonExe" DarkGray
|
||||
|
||||
# --- (3) Dependencias --------------------------------------------------------
|
||||
$importTest = 'import fastapi, uvicorn, sse_starlette, jinja2, yaml, yt_dlp, requests, slugify'
|
||||
$depsOk = $false
|
||||
try {
|
||||
& $pythonExe -c $importTest *> $null
|
||||
$depsOk = ($LASTEXITCODE -eq 0)
|
||||
} catch { $depsOk = $false }
|
||||
|
||||
if (-not $depsOk) {
|
||||
if ($CheckOnly) {
|
||||
Out-Line " [check] AQUI: instalar dependencias -> & '$pythonExe' -m pip install -e `".[web]`"" DarkGray
|
||||
exit 11
|
||||
}
|
||||
Out-Line " [info] Instalando las piezas que faltan por primera vez (puede tardar 1-2 min)..." Yellow
|
||||
try {
|
||||
& $pythonExe -m pip --version *> $null
|
||||
if ($LASTEXITCODE -ne 0) { & $pythonExe -m ensurepip --upgrade | Out-Null }
|
||||
} catch {
|
||||
& $pythonExe -m ensurepip --upgrade | Out-Null
|
||||
}
|
||||
& $pythonExe -m pip install -e ".[web]"
|
||||
if ($LASTEXITCODE -ne 0) {
|
||||
Out-Line " [!] No pude instalar las dependencias (revisa tu conexion a internet)" Red
|
||||
Out-Line " y vuelve a hacer doble clic en start-server.bat" Gray
|
||||
exit 11
|
||||
}
|
||||
}
|
||||
if ($CheckOnly) { Out-Line " [check] dependencias: OK" DarkGray }
|
||||
|
||||
# --- (4) ffmpeg (opcional: solo audio) --------------------------------------
|
||||
if (-not (Get-Command ffmpeg -ErrorAction SilentlyContinue)) {
|
||||
Out-Line " [aviso] La descarga de AUDIO no estara disponible (falta ffmpeg)." Yellow
|
||||
Out-Line " Todo lo demas funciona perfecto. Puedes ignorarlo." Gray
|
||||
}
|
||||
|
||||
# --- (5) Delegar al launcher -------------------------------------------------
|
||||
if ($CheckOnly) {
|
||||
Out-Line " [check] AQUI: lanzaria scripts\start-server.ps1 -PythonExe '$pythonExe'" DarkGray
|
||||
Out-Line ""
|
||||
Out-Line " Diagnostico completo. Todo listo para arrancar." Green
|
||||
exit 0
|
||||
}
|
||||
|
||||
Out-Line ""
|
||||
& (Join-Path $PSScriptRoot 'start-server.ps1') -PythonExe $pythonExe @args
|
||||
exit $LASTEXITCODE
|
||||
```
|
||||
|
||||
- [ ] **Step 2: Verificar sintaxis y modo diagnóstico**
|
||||
|
||||
Run: `powershell -NoProfile -ExecutionPolicy Bypass -File scripts\doctor.ps1 -CheckOnly`
|
||||
Expected: líneas `[yt-scraper] verificando tu sistema...`, `[ok] Python encontrado: ...`, `[check] dependencias: OK` o `[check] AQUI: instalar...`, `[check] AQUI: lanzaria ... start-server.ps1 -PythonExe '<ruta>'`, `Diagnostico completo.` — y **ningún efecto colateral** (sin browser, sin installs). Exit 0 si el entorno actual está sano.
|
||||
|
||||
Nota: el server sigue vivo en 8001 de la sesión anterior → si `.run/server.info` apunta ahí, la salida esperada es `[ok] El servidor ya estaba corriendo - abriendo...` **sin abrir navegador** (por -CheckOnly) y exit 0. Ambas salidas son correctas; registrar cuál ocurrió.
|
||||
|
||||
- [ ] **Step 3: Commit**
|
||||
|
||||
```bash
|
||||
git add scripts/doctor.ps1
|
||||
git commit -m "feat: doctor.ps1 prepara entorno (python/deps) antes del arranque"
|
||||
```
|
||||
|
||||
### Task 2: start-server.ps1 — parámetro -PythonExe + pulido de copys
|
||||
|
||||
**Files:**
|
||||
- Modify: `scripts/start-server.ps1` (línea ~1 param block nuevo; línea ~131 uso del parámetro; copys de mensajes visibles)
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: nada nuevo.
|
||||
- Produces: `param([string]$PythonExe = 'python')` — consumido por doctor.ps1 Task 1 vía `-PythonExe $pythonExe`.
|
||||
|
||||
- [ ] **Step 1: Añadir param block**
|
||||
|
||||
Al inicio del archivo (antes de `$ErrorActionPreference`), añadir:
|
||||
|
||||
```powershell
|
||||
param(
|
||||
# Ruta absoluta o nombre del interprete Python que lanza uvicorn.
|
||||
# doctor.ps1 la resuelve porque tras una instalacion fresca el PATH
|
||||
# de esta sesion aun no ve el Python nuevo.
|
||||
[string]$PythonExe = 'python'
|
||||
)
|
||||
```
|
||||
|
||||
- [ ] **Step 2: Usar el parámetro en la línea de uvicorn**
|
||||
|
||||
Reemplazar (línea ~131):
|
||||
|
||||
```powershell
|
||||
$cmdLine = '/c python -m uvicorn yt_scraper.webapp.app:app --host 127.0.0.1 --port ' + $candidate + ' --log-level info > ".run\server.log" 2>&1'
|
||||
```
|
||||
|
||||
por:
|
||||
|
||||
```powershell
|
||||
$cmdLine = '/c "' + $PythonExe + '" -m uvicorn yt_scraper.webapp.app:app --host 127.0.0.1 --port ' + $candidate + ' --log-level info > ".run\server.log" 2>&1'
|
||||
```
|
||||
|
||||
(Comillas alrededor de la ruta: necesarios si contiene espacios.)
|
||||
|
||||
- [ ] **Step 3: Pulir copys (solo texto, cero lógica)**
|
||||
|
||||
Sustituir estos mensajes por versión amigable (mantener colores y estructura):
|
||||
|
||||
| Original | Nuevo |
|
||||
|---|---|
|
||||
| `" [yt-scraper] iniciando servidor local..." Cyan` | igual (ya es claro) |
|
||||
| `" [info] server.info apuntaba a puerto={0} pid={1} pero esta obsoleto (portBusy=False pidAlive={2}); limpiando."` | `" [info] Habia un registro viejo de otra sesion; limpiandolo..." Yellow` |
|
||||
| `" [scan] puertos candidatos: $($candidates -join ', ')" DarkGray` | eliminar línea (ruido técnico) |
|
||||
| `" [try] lanzando uvicorn en puerto $candidate ..." Gray` | `" [...] Encendiendo el servidor (intento $candidate)..." Gray` |
|
||||
| `" cmd PID=$($proc.Id) -> ventana minimizada (.run\server.log)"` | `" El servidor corre en segundo plano. Registro: .run\server.log"` |
|
||||
| `" [warn] puerto $candidate no respondio en ...s; probando siguiente"` | `" [warn] Este intento no respondio; probando el siguiente..." Yellow` |
|
||||
| `" [ok] puerto $port aceptando conexiones (t=$([int]$boundAt.TotalSeconds)s)" Green` | `" [ok] Servidor encendido." Green` |
|
||||
| `" (esperando /healthz para confirmar arranque completo...)" DarkGray` | `" Confirmando que todo cargo bien..." DarkGray` |
|
||||
| `" [warn] puerto abierto pero /healthz no respondio; revisa .run\server.log" Yellow` | `" [warn] El servidor abrio pero tardo en responder; revisa .run\server.log" Yellow` |
|
||||
|
||||
Las líneas `URL:` / `log:` / `stop:` se mantienen (útiles también para no técnicos).
|
||||
|
||||
- [ ] **Step 4: Verificar camino feliz con server vivo**
|
||||
|
||||
Run: `powershell -NoProfile -ExecutionPolicy Bypass -File scripts\start-server.ps1`
|
||||
Expected: `[ok] servidor ya activo en puerto 8001` + apertura de navegador + exit 0. (El server sigue corriendo desde la sesión anterior.)
|
||||
|
||||
- [ ] **Step 5: Commit**
|
||||
|
||||
```bash
|
||||
git add scripts/start-server.ps1
|
||||
git commit -m "feat: start-server.ps1 acepta -PythonExe y copys amigables"
|
||||
```
|
||||
|
||||
### Task 3: stop-server.ps1 — pulido de copys (sin lógica nueva)
|
||||
|
||||
**Files:**
|
||||
- Modify: `scripts/stop-server.ps1` (solo strings de Out-Line)
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes/Produces: nada cambia funcionalmente.
|
||||
|
||||
- [ ] **Step 1: Pulir copys**
|
||||
|
||||
| Original | Nuevo |
|
||||
|---|---|
|
||||
| `" [yt-scraper] deteniendo servidor (PID $pidv, puerto $port)..." Cyan` | `" [yt-scraper] Apagando el servidor..." Cyan` |
|
||||
| `" [kill] puerto $port ocupado por PID $($_.OwningProcess); terminando..." Gray` | `" [kill] Cerrando un proceso que quedaba suelto..." Gray` |
|
||||
| `" [yt-scraper] no hay servidor registrado. Buscando procesos uvicorn sueltos..." Cyan` | `" [yt-scraper] Buscando servidores que quedaron sueltos..." Cyan` |
|
||||
| `" [info] el PID $pidv ya no existe; el servidor estaba muerto." Yellow` | `" [info] El servidor ya estaba apagado." Yellow` |
|
||||
| `" [warn] el puerto $port sigue ocupado; revisa procesos python manualmente." Yellow` | `" [warn] Algo sigue ocupando la conexion. Reinicia la PC si vuelve a pasar." Yellow` |
|
||||
|
||||
- [ ] **Step 2: Verificar stop limpio**
|
||||
|
||||
Run: `powershell -NoProfile -ExecutionPolicy Bypass -File scripts\stop-server.ps1`
|
||||
Expected: `Apagando el servidor...` → `puerto liberado, servidor detenido.` y el proceso uvicorn de 8001 muerto (`Invoke-WebRequest http://127.0.0.1:8001/api/dashboard` falla).
|
||||
|
||||
- [ ] **Step 3: Commit**
|
||||
|
||||
```bash
|
||||
git add scripts/stop-server.ps1
|
||||
git commit -m "chore: copys amigables en stop-server.ps1"
|
||||
```
|
||||
|
||||
### Task 4: Reglas de ventana en los .bat
|
||||
|
||||
**Files:**
|
||||
- Modify: `start-server.bat`
|
||||
- Modify: `stop-server.bat`
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: exit codes de doctor.ps1 (Task 1) y stop-server.ps1.
|
||||
|
||||
- [ ] **Step 1: Reescribir start-server.bat**
|
||||
|
||||
Contenido exacto:
|
||||
|
||||
```bat
|
||||
@echo off
|
||||
chcp 65001 >nul
|
||||
powershell -NoProfile -ExecutionPolicy Bypass -File "%~dp0scripts\doctor.ps1" %*
|
||||
set EC=%errorlevel%
|
||||
if not "%EC%"=="0" (
|
||||
echo.
|
||||
echo ^[!] Algo fallo. Esta ventana queda abierta para que puedas leerlo.
|
||||
echo Si llamas a alguien por ayuda, enviale una captura de pantalla.
|
||||
echo.
|
||||
pause
|
||||
) else (
|
||||
timeout /t 5 /nobreak >nul
|
||||
)
|
||||
exit /b %EC%
|
||||
```
|
||||
|
||||
- [ ] **Step 2: Reescribir stop-server.bat**
|
||||
|
||||
Contenido exacto:
|
||||
|
||||
```bat
|
||||
@echo off
|
||||
chcp 65001 >nul
|
||||
powershell -NoProfile -ExecutionPolicy Bypass -File "%~dp0scripts\stop-server.ps1" %*
|
||||
set EC=%errorlevel%
|
||||
if not "%EC%"=="0" (
|
||||
echo.
|
||||
echo ^[!] Algo fallo al apagar. Esta ventana queda abierta.
|
||||
echo.
|
||||
pause
|
||||
)
|
||||
exit /b %EC%
|
||||
```
|
||||
|
||||
(Éxito en stop → cierra de inmediato: apagar es instantáneo, no hace falta resumen.)
|
||||
|
||||
- [ ] **Step 3: Verificar reglas de ventana simulando error**
|
||||
|
||||
Run (en cmd, PATH vacío para forzar rama sin Python):
|
||||
`cmd /c "set PATH=C:\Windows\System32 && start-server.bat"`
|
||||
Expected: ventana muestra mensaje de instalación de Python/winget y queda abierta en `Presione una tecla...`. (Verificación manual del executor: lanzar con `cmd /c` captura el texto; confirmar que `pause` está presente en el flujo de error.)
|
||||
|
||||
Nota: en este entorno el executor valida la rama de éxito automáticamente (Task 5); la rama de error se valida leyendo el flujo: `EC≠0` → `pause`.
|
||||
|
||||
- [ ] **Step 4: Commit**
|
||||
|
||||
```bash
|
||||
git add start-server.bat stop-server.bat
|
||||
git commit -m "feat: .bat con ventanas inteligentes (persistente en error, autocierre en exito)"
|
||||
```
|
||||
|
||||
### Task 5: Verificación end-to-end completa
|
||||
|
||||
**Files:** ninguno nuevo (verificación).
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: todo lo anterior.
|
||||
|
||||
- [ ] **Step 1: Estado inicial limpio**
|
||||
|
||||
Confirmar sin servidor vivo: `powershell -NoProfile -Command "Test-NetConnection -ComputerName 127.0.0.1 -Port 8001 -InformationLevel Quiet -WarningAction SilentlyContinue"`
|
||||
Expected: `False`. (Si algo vive en 8000-8100, ejecutar `scripts\stop-server.ps1` primero.)
|
||||
|
||||
- [ ] **Step 2: Arranque completo vía doctor**
|
||||
|
||||
Run: `powershell -NoProfile -ExecutionPolicy Bypass -File scripts\doctor.ps1`
|
||||
Expected: secuencia `[ok] Python encontrado` → `[check]/dependencias OK implícito` → salida del launcher `Servidor encendido` en algún puerto 8000-8100 → `Confirmando que todo cargo bien...` + `/healthz OK`. Exit 0.
|
||||
|
||||
- [ ] **Step 3: Health check externo**
|
||||
|
||||
Run: `powershell -NoProfile -Command "(Invoke-WebRequest -UseBasicParsing http://127.0.0.1:<PUERTO>/healthz -TimeoutSec 5).StatusCode"`
|
||||
Expected: `200`.
|
||||
|
||||
- [ ] **Step 4: Segunda invocación = camino rápido**
|
||||
|
||||
Run: `powershell -NoProfile -ExecutionPolicy Bypass -File scripts\doctor.ps1`
|
||||
Expected: `[ok] El servidor ya estaba corriendo - abriendo...` + navegador abierto + exit 0, en < 2 s.
|
||||
|
||||
- [ ] **Step 5: Dejar el servidor corriendo para el usuario + informe final**
|
||||
|
||||
No apagar al terminar: el usuario quedó con la app funcionando. Informar URL final y qué cambió.
|
||||
|
||||
- [ ] **Step 6: Sin commit** (no hay archivos nuevos; verificar `git status` limpio)
|
||||
|
||||
Run: `git status --short`
|
||||
Expected: sin cambios pendientes (todo commiteado en Tasks 1-4).
|
||||
@@ -0,0 +1,250 @@
|
||||
# Vista Grid estilo YouTube (sección Vídeos) — Implementation Plan
|
||||
|
||||
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
||||
|
||||
**Goal:** Añadir a la sección Vídeos un conmutador tabla ⇄ grid donde el grid replica el look de YouTube (tarjetas 16:9 con miniatura, duración, título y metadatos) manteniendo filtros, orden, selección masiva y paginación compartidos.
|
||||
|
||||
**Architecture:** Solo frontend. El estado `videos.view` (`'table'|'grid'`) vive en el componente Alpine único `window.platform()`; la elección se persiste en `localStorage["videos-view"]`. En `index.html` los dos layouts son bloques hermanos alternados con `<template x-if>`; el toggle es una fila fina sobre ellos.
|
||||
|
||||
**Tech Stack:** Alpine 3 + Tailwind CDN (sin build step), estáticos servidos por FastAPI.
|
||||
|
||||
**Spec:** [docs/superpowers/specs/2026-08-22-videos-grid-view-design.md](../specs/2026-08-22-videos-grid-view-design.md)
|
||||
|
||||
## Global Constraints
|
||||
|
||||
- Sin build step: todo cambio es HTML/JS/CSS estático editado a mano ([.opencode/agent/webapp-builder.md](../../.opencode/agent/webapp-builder.md)).
|
||||
- Todo el estado vive en **un solo** componente Alpine, `window.platform()` en `src/yt_scraper/webapp/static/app.js`; no crear componentes nuevos.
|
||||
- El orden de `<script defer>` en `index.html` (app.js antes de alpinejs) es carga funcional: no reordenar scripts.
|
||||
- No tocar backend, endpoints ni `store.py`: `/api/videos` ya devuelve todos los campos que necesita la tarjeta.
|
||||
- Sin tests automatizados de frontend (el repo no tiene infra JS): verificación = `node --check` + smoke manual con `start-server.bat`.
|
||||
- Estilo visual existente: dark "command center", `glass`, `accent-grad btn-primary` como tratamiento activo (mismo patrón que la paginación).
|
||||
- Sin acciones por tarjeta (`.md`/`Clip` quedan en tabla y detalle).
|
||||
|
||||
---
|
||||
|
||||
### Task 1: Estado y persistencia en app.js
|
||||
|
||||
**Files:**
|
||||
- Modify: `src/yt_scraper/webapp/static/app.js` (línea 39 store `videos`; `init()` líneas 92-93; método nuevo junto a los helpers de selección ~línea 463)
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: nada nuevo.
|
||||
- Produces: `this.videos.view` (`'table' | 'grid'`, default `'table'`) y `setVideoView(v)` — el HTML del Task 2/3 los usa vía `videos.view` / `@click="setVideoView('grid')"`.
|
||||
|
||||
- [ ] **Step 1: Añadir `view` al store**
|
||||
|
||||
En `app.js:39`, cambiar:
|
||||
|
||||
```js
|
||||
videos: { items: [], total: 0, page: 1, size: 25, selected: [] },
|
||||
```
|
||||
|
||||
por:
|
||||
|
||||
```js
|
||||
videos: { items: [], total: 0, page: 1, size: 25, selected: [], view: "table" },
|
||||
```
|
||||
|
||||
- [ ] **Step 2: Restaurar la elección en init()**
|
||||
|
||||
En `app.js` dentro de `init()`, justo después del bloque:
|
||||
|
||||
```js
|
||||
const size = Number(localStorage.getItem("videos-size"));
|
||||
if ([10, 25, 50, 100].includes(size)) this.videos.size = size;
|
||||
```
|
||||
|
||||
añadir:
|
||||
|
||||
```js
|
||||
const savedView = localStorage.getItem("videos-view");
|
||||
if (savedView === "table" || savedView === "grid") this.videos.view = savedView;
|
||||
```
|
||||
|
||||
- [ ] **Step 3: Añadir setVideoView()**
|
||||
|
||||
En `app.js`, inmediatamente antes de `toggleSelect(id) {` (la sección de helpers de selección, ~línea 463), añadir:
|
||||
|
||||
```js
|
||||
// View mode of the Videos section ('table' | 'grid'). UI preference like
|
||||
// videos-size: persisted in localStorage, never in the URL.
|
||||
setVideoView(v) {
|
||||
if (v !== "table" && v !== "grid") return;
|
||||
this.videos.view = v;
|
||||
try { localStorage.setItem("videos-view", v); } catch (_) {}
|
||||
},
|
||||
|
||||
```
|
||||
|
||||
- [ ] **Step 4: Verificar sintaxis JS**
|
||||
|
||||
Run: `node --check src/yt_scraper/webapp/static/app.js`
|
||||
Expected: sin salida (exit 0). Si `node` no está disponible, verificar abriendo la webapp en el Task 4.
|
||||
|
||||
- [ ] **Step 5: Commit**
|
||||
|
||||
```bash
|
||||
git add src/yt_scraper/webapp/static/app.js
|
||||
git commit -m "feat(webapp): estado y persistencia de vista tabla/grid en Videos"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 2: Conmutador de vista + alternancia x-if en index.html
|
||||
|
||||
**Files:**
|
||||
- Modify: `src/yt_scraper/webapp/static/index.html` (contenedor de resultados, líneas 301-352)
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: `videos.view` y `setVideoView(v)` del Task 1.
|
||||
- Produces: la estructura de dos bloques `<template x-if>` que el Task 3 completa con el grid (este task deja la tabla funcional tal cual).
|
||||
|
||||
- [ ] **Step 1: Envolver la tabla en template x-if**
|
||||
|
||||
En `index.html`, la tabla vive en:
|
||||
|
||||
```html
|
||||
<div class="glass overflow-hidden">
|
||||
<table class="tbl">
|
||||
```
|
||||
|
||||
(que cierra con `</table>` + `</div>` en las líneas ~351-352, justo antes del comentario `<!-- pagination -->`). Envolver ese `<div class="glass overflow-hidden">…</div>` completo en:
|
||||
|
||||
```html
|
||||
<template x-if="videos.view==='table'">
|
||||
<div class="glass overflow-hidden">
|
||||
… contenido actual de la tabla SIN CAMBIOS …
|
||||
</div>
|
||||
</template>
|
||||
```
|
||||
|
||||
Reindentar una nivel el interior. Ningún atributo ni clase de la tabla cambia.
|
||||
|
||||
- [ ] **Step 2: Insertar la fila del toggle**
|
||||
|
||||
Justo encima del `<template x-if="videos.view==='table'">` recién creado, añadir:
|
||||
|
||||
```html
|
||||
<!-- table ⇄ grid switch -->
|
||||
<div class="flex justify-end">
|
||||
<div class="inline-flex rounded-lg border border-zinc-800 bg-zinc-900/60 p-1 gap-1" role="group" aria-label="View mode">
|
||||
<button type="button" class="btn !py-1 !px-2" :class="videos.view==='table' ? 'accent-grad btn-primary' : 'btn-ghost'" :aria-pressed="(videos.view==='table').toString()" @click="setVideoView('table')" title="Table view">
|
||||
<svg class="w-4 h-4" viewBox="0 0 24 24" fill="currentColor"><path d="M3 5h18v2H3zm0 6h18v2H3zm0 6h18v2H3z"/></svg>
|
||||
</button>
|
||||
<button type="button" class="btn !py-1 !px-2" :class="videos.view==='grid' ? 'accent-grad btn-primary' : 'btn-ghost'" :aria-pressed="(videos.view==='grid').toString()" @click="setVideoView('grid')" title="Grid view">
|
||||
<svg class="w-4 h-4" viewBox="0 0 24 24" fill="currentColor"><path d="M3 3h8v8H3zm10 0h8v8h-8zM3 13h8v8H3zm10 0h8v8h-8z"/></svg>
|
||||
</button>
|
||||
</div>
|
||||
</div>
|
||||
```
|
||||
|
||||
- [ ] **Step 3: Smoke manual**
|
||||
|
||||
Run: `start-server.bat`, abrir Vídeos.
|
||||
Expected: la tabla se ve idéntica a antes; el toggle aparece arriba a la derecha; pulsar el icono de grid NO cambia nada visible todavía pero el botón grid queda activo (accent), se refresca la página y sigue activo (localStorage); volver a tabla también persiste.
|
||||
|
||||
- [ ] **Step 4: Commit**
|
||||
|
||||
```bash
|
||||
git add src/yt_scraper/webapp/static/index.html
|
||||
git commit -m "feat(webapp): conmutador de vista tabla/grid en Videos"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 3: Bloque grid de tarjetas estilo YouTube
|
||||
|
||||
**Files:**
|
||||
- Modify: `src/yt_scraper/webapp/static/index.html` (insertar tras el cierre `</template>` del bloque tabla, antes de `<!-- pagination -->`)
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: `videos.view`, `setVideoView` (Task 1); helpers existentes de `app.js`: `ts(sec)`, `fmtNum(n)`, `chanName(id)`, `videoDate(v)`, `videoDateTitle(v)`, `statusClass(s)`, `isBlocked(v)`, `blockLabel(v)`, `blockTitle(v)`, `isSelected(id)`, `toggleSelect(id)`, `openVideo(id)`, `loading.videos`, `videos.items`.
|
||||
- Produces: la vista grid completa (fin de feature).
|
||||
|
||||
- [ ] **Step 1: Insertar el bloque grid**
|
||||
|
||||
Entre el `</template>` que cierra el bloque tabla y `<!-- pagination -->`, añadir:
|
||||
|
||||
```html
|
||||
<template x-if="videos.view==='grid'">
|
||||
<div>
|
||||
<template x-if="loading.videos"><div class="text-center text-zinc-500 py-10">loading…</div></template>
|
||||
<template x-if="!loading.videos && videos.items.length===0"><div class="text-center text-zinc-500 py-10"><div>No videos match these filters.</div><button class="btn btn-ghost mt-3" @click.stop="filters={ channel:'', status:'', from:'', to:'', min_dur:'', q:'', sort:'upload_date' }; loadVideos(1)">Clear filters</button></div></template>
|
||||
<template x-if="!loading.videos && videos.items.length>0">
|
||||
<div class="grid grid-cols-2 sm:grid-cols-3 lg:grid-cols-4 2xl:grid-cols-5 gap-4">
|
||||
<template x-for="v in videos.items" :key="v.video_id">
|
||||
<div class="glass group relative cursor-pointer hover:border-rose-500/40 transition-colors overflow-hidden"
|
||||
:class="isSelected(v.video_id) ? 'row-selected' : ''"
|
||||
role="button" tabindex="0"
|
||||
@click="openVideo(v.video_id)"
|
||||
@keydown.enter.prevent="openVideo(v.video_id)">
|
||||
<div class="relative aspect-video bg-zinc-900">
|
||||
<img class="absolute inset-0 w-full h-full object-cover" :src="'/api/thumbnails/'+v.video_id" :alt="v.title" onerror="this.style.visibility='hidden'" />
|
||||
<span x-show="v.duration" class="absolute bottom-1.5 right-1.5 font-mono text-[0.7rem] leading-none px-1.5 py-1 rounded bg-black/80 text-white" x-text="ts(v.duration)"></span>
|
||||
<button type="button" @click.stop="toggleSelect(v.video_id)"
|
||||
class="absolute top-1.5 left-1.5 w-6 h-6 rounded-md border flex items-center justify-center transition-opacity"
|
||||
:class="isSelected(v.video_id) ? 'bg-rose-500 border-rose-400 opacity-100' : 'bg-black/70 border-zinc-300/70 opacity-0 group-hover:opacity-100'"
|
||||
:aria-pressed="isSelected(v.video_id).toString()" title="Select video">
|
||||
<svg x-show="isSelected(v.video_id)" class="w-4 h-4 text-white" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="3"><path stroke-linecap="round" stroke-linejoin="round" d="M5 13l4 4L19 7"/></svg>
|
||||
</button>
|
||||
</div>
|
||||
<div class="p-3 space-y-1.5">
|
||||
<div class="text-sm font-medium text-zinc-100 line-clamp-2 leading-snug" x-text="v.title" :title="v.title"></div>
|
||||
<div class="text-xs text-zinc-400 truncate" x-text="chanName(v.channel_id)" :title="chanName(v.channel_id)"></div>
|
||||
<div class="flex items-center gap-2 font-mono text-xs text-zinc-400 flex-wrap">
|
||||
<span x-text="fmtNum(v.view_count)"></span>
|
||||
<span class="text-zinc-600">·</span>
|
||||
<span :class="v.upload_date ? '' : 'italic text-zinc-600'" x-text="videoDate(v)" :title="videoDateTitle(v)"></span>
|
||||
</div>
|
||||
<div class="flex items-center gap-1 flex-wrap pt-0.5">
|
||||
<span class="pill" :class="statusClass(v.status)" x-text="v.status" :title="v.error_msg || v.status"></span>
|
||||
<span x-show="isBlocked(v)" class="pill st-locked" :title="blockTitle(v)">
|
||||
<svg class="w-3 h-3" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" aria-hidden="true"><rect x="4" y="10" width="16" height="10" rx="2"/><path d="M8 10V7a4 4 0 1 1 8 0v3"/></svg>
|
||||
<span x-text="blockLabel(v)"></span>
|
||||
</span>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
</template>
|
||||
</div>
|
||||
</template>
|
||||
</div>
|
||||
</template>
|
||||
```
|
||||
|
||||
Notas de diseño ya decididas:
|
||||
- `.row-selected` existe en styles.css:179 con `!important` → aplica igual sobre la tarjeta div que sobre el `<tr>`.
|
||||
- `line-clamp-2` ya se usa en index.html:553 → Tailwind CDN lo resuelve.
|
||||
- Checkbox estilo YouTube: aparece al hover (`group-hover`) o permanece si está seleccionada; click con `.stop` para no abrir el detalle.
|
||||
- La selección masiva, barra bulk, paginación y filtros NO se tocan: viven fuera de estos bloques.
|
||||
|
||||
- [ ] **Step 2: Verificación manual completa (checklist del spec)**
|
||||
|
||||
Run: `start-server.bat`
|
||||
|
||||
1. Alternar tabla ⇄ grid: ambas muestran los mismos vídeos con el filtro activo.
|
||||
2. En grid: pasar el ratón sobre una tarjeta → checkbox visible; marcar 3 → barra bulk aparece con "3 selected" → "Download .md" encola el job.
|
||||
3. Click en tarjeta (fuera del checkbox) → abre el detalle; Back → vuelve a grid conservando página.
|
||||
4. Paginar en grid: Prev/Next/números funcionan, scroll arriba, sin repeticiones.
|
||||
5. Recargar página: la vista elegida se conserva; borrar `localStorage["videos-view"]` → cae a tabla.
|
||||
6. Tarjeta de vídeo blocked: pill candado con label; badge de duración visible en tarjetas con duración.
|
||||
7. Miniatura rota: `onerror` la oculta sin romper el layout (contenedor aspect-video mantiene proporción).
|
||||
8. La tabla sigue funcionando exactamente igual que antes.
|
||||
|
||||
Run: `python -m pytest tests/ -q`
|
||||
Expected: suite verde (backend intacto; sanity check barato).
|
||||
|
||||
- [ ] **Step 3: Commit**
|
||||
|
||||
```bash
|
||||
git add src/yt_scraper/webapp/static/index.html
|
||||
git commit -m "feat(webapp): grid de tarjetas estilo YouTube en Videos"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Self-review hecho
|
||||
|
||||
- **Cobertura del spec:** estado+persistencia (Task 1), toggle (Task 2), tarjetas/clamps/badges/pills/loading/empty (Task 3), verificación manual ítem a ítem (Task 3 Step 2). Fuera de alcance respetado: cero cambios backend, sin acciones por tarjeta.
|
||||
- **Placeholders:** ninguno; todo step lleva código literal o comando exacto.
|
||||
- **Consistencia de nombres:** `videos.view`, `setVideoView`, `"videos-view"` usados igual en Tasks 1-3; helpers referenciados existen en app.js (verificado: ts:1265, fmtNum:1293, videoDate:1317, videoDateTitle:1324, chanName:1345, blockLabel:1355, blockTitle:1365, isBlocked:1374, statusClass:1376, toggleSelect:464, isSelected:469).
|
||||
@@ -0,0 +1,47 @@
|
||||
# Discovery-Only Scrape Design
|
||||
|
||||
## Goal
|
||||
|
||||
Add manual discovery controls for one channel or all tracked channels. Discovery checks the complete recent video list returned by YouTube, registers unknown videos as `pending`, and does not download transcripts, Markdown, audio, thumbnails, or avatars.
|
||||
|
||||
## Existing Context
|
||||
|
||||
The webapp already has a single-worker `JobManager`, SSE job events, and a `POST /api/scrape` endpoint. The current channel job combines discovery with `pipeline.process_video()`, so every pending video is immediately extracted and rendered. The channels table has per-channel actions, while the videos view can be filtered to one channel.
|
||||
|
||||
## Architecture
|
||||
|
||||
- Extend the existing job dispatch with `opts.mode == "discover"`.
|
||||
- `channel_id` selects one tracked channel. A missing `channel_id` means all tracked channels for discovery jobs only.
|
||||
- Discovery uses the existing `discover_channel()` call without a playlist cap. It compares the returned IDs with the SQLite catalog and upserts only catalog metadata. It never calls `process_video()`.
|
||||
- The existing scrape mode remains unchanged for users who want discovery plus transcript/Markdown processing.
|
||||
- All discovery jobs stay in the existing sequential queue and use the existing SSE stream, cancellation, progress, and history.
|
||||
|
||||
## Persistence and Results
|
||||
|
||||
- New `VideoRef` rows are inserted with the existing default status `pending`.
|
||||
- Existing video rows retain their status and processed data; rediscovery only refreshes title, upload date, and duration through the existing upsert behavior.
|
||||
- The channel record is refreshed with its name, handle, catalog count, and `last_scraped`; no avatar fetch/cache is performed by discovery-only jobs.
|
||||
- A per-channel progress event reports `new_videos`, `known_videos`, and the channel ID.
|
||||
- The terminal event reports total channels, discovered entries, new videos, known videos, and errors.
|
||||
- A one-channel discovery failure marks the job `error`. An all-channel job continues after individual failures and finishes with a summary so one broken channel does not prevent other channels from being scanned.
|
||||
|
||||
## UI
|
||||
|
||||
- Channels view: add `Investigar todos` in the header and `Investigar` in each channel row.
|
||||
- Videos view: add an investigation button in the header. It investigates the selected channel when `filters.channel` is set, otherwise all channels.
|
||||
- Buttons use the existing job widget/SSE stream, are disabled while an active job is being followed, and show the number of new videos on completion.
|
||||
- Completion refreshes channels, videos, and dashboard data. Existing `.md`, `Process`, and `Audio` controls remain separate.
|
||||
|
||||
## Error Handling
|
||||
|
||||
- Empty channel catalogs complete successfully with zero channels scanned.
|
||||
- Discovery errors are emitted in the job log and do not invoke video processing.
|
||||
- A cancellation checks the existing cancellation set between channels and leaves already inserted catalog rows intact.
|
||||
- Duplicate IDs from a discovery response are counted once by the catalog comparison.
|
||||
|
||||
## Testing
|
||||
|
||||
- Store tests verify that upserting an existing and a new reference reports only the new reference.
|
||||
- Job tests fake `discover_channel()`, verify new rows are `pending`, existing rows retain their status, and `process_video()` is never called.
|
||||
- Job tests verify all-channel discovery continues after one channel fails and reports the successful channel.
|
||||
- The full existing Python test suite remains the regression check. The static UI is verified by code inspection and the existing webapp smoke path; no frontend build step exists.
|
||||
@@ -0,0 +1,76 @@
|
||||
# Diseño: fechas aproximadas de subida vía discovery (approximate_date)
|
||||
|
||||
Fecha: 2026-08-22
|
||||
Estado: aprobado en conversación
|
||||
|
||||
## Problema
|
||||
|
||||
La tabla de vídeos no muestra fecha de subida para la mayoría de filas: el
|
||||
discovery flat de yt-dlp no trae `upload_date` (verificado en el proyecto), y
|
||||
la fecha exacta solo se aprende al extraer cada vídeo (2 peticiones/vídeo).
|
||||
Hoy la UI muestra `~fecha-inferida` derivada del rank del canal.
|
||||
|
||||
Hallazgo verificado con yt-dlp 2026.07.04 instalado:
|
||||
`extractor_args: {"youtubetab": {"approximate_date": ["true"]}}` hace que los
|
||||
entries flat traigan `timestamp` parseado del texto relativo que YouTube ya
|
||||
incluye en el listado ("hace 3 semanas"). **Costo: cero peticiones extra**;
|
||||
viaja en las mismas respuestas del discovery. Precisión: día para recientes,
|
||||
más gruesa para antiguos.
|
||||
|
||||
## Alcance elegido
|
||||
|
||||
Solo el flag gratis. Sin RSS, sin extracción masiva (descartadas por costo o
|
||||
valor marginal). El backlog se llena con un **Full rescan** por canal usando
|
||||
los mismos requests de siempre.
|
||||
|
||||
## Decisiones
|
||||
|
||||
### Datos: misma columna + bandera
|
||||
|
||||
Nueva columna `videos.upload_date_approx INTEGER DEFAULT 0` vía `_VIDEO_COLUMNS`.
|
||||
La fecha aproximada vive en `upload_date` normal: `SORT_DATE_SQL`, orden y
|
||||
consumidores existentes la aprovechan sin cambios. La bandera preserva la
|
||||
distinción aprox/real.
|
||||
|
||||
Matriz de `upsert_videos` (reemplaza el `COALESCE` simple en `upload_date`):
|
||||
|
||||
| En DB \ Entra | Aproximada | Real |
|
||||
|---|---|---|
|
||||
| NULL | escribe, flag=1 | escribe, flag=0 |
|
||||
| Aproximada | re-escribe estimación fresca, flag=1 | escribe real, flag=0 |
|
||||
| Real | conserva real, flag=0 | conserva |
|
||||
|
||||
Upgrade approx→real: `set_upload_date()` hoy solo escribe si `upload_date IS
|
||||
NULL`; pasa a escribir también cuando `upload_date_approx = 1` (y apaga la
|
||||
bandera). Ahí entra la fecha real de la extracción.
|
||||
|
||||
### Discovery: timestamp → YYYYMMDD
|
||||
|
||||
Los entries flat traen `timestamp` (epoch), no `upload_date`. Nuevo helper
|
||||
puro `_entry_upload_date(entry) -> tuple[str | None, int]` en `discover.py`:
|
||||
prefiere `upload_date` crudo (flag 0); si no, deriva de `timestamp` UTC a
|
||||
`YYYYMMDD` (flag 1). Los dos constructores de `VideoRef` lo usan.
|
||||
`deep_channel_avatar` no lo necesita (limit=1, solo avatar).
|
||||
|
||||
Extractor arg siempre activo en `discover_channel`; sin knob de config.
|
||||
|
||||
### UI
|
||||
|
||||
- API: `_video_dict` expone `upload_date_approx` y `date_estimated` pasa a ser
|
||||
`inferida-por-rank OR bandera-aprox`.
|
||||
- Tabla: fecha con `~` cuando `upload_date` viene aproximada; tooltip explica
|
||||
el origen ("YouTube lo muestra relativo"); tooltip previo de rank-inference
|
||||
queda para el caso sin fecha alguna.
|
||||
|
||||
### No afectado
|
||||
|
||||
`.md` filenames (se renderizan tras extracción, con fecha real),
|
||||
`channel_seq`/maquinaria de inferencia (queda como red), economía de
|
||||
peticiones (0 extra).
|
||||
|
||||
## Testing
|
||||
|
||||
Tests nuevos sobre `Store(tmp_path)` sintético: matriz de upsert (4 casos) +
|
||||
upgrade approx→real vía `set_upload_date`; unit del helper de conversión con
|
||||
entries fake. Migración idempotente queda cubierta por los tests existentes
|
||||
del Store. Suite completa debe seguir verde.
|
||||
@@ -0,0 +1,131 @@
|
||||
# Diseño: arranque automático e inteligente para no-técnicos (doctor.ps1)
|
||||
|
||||
Fecha: 2026-08-22
|
||||
Estado: aprobado en conversación
|
||||
|
||||
## Problema
|
||||
|
||||
Los `.bat` actuales son wrappers finos de dos `.ps1` robustos, pero asumen un
|
||||
entorno ya preparado. Para una persona no técnica:
|
||||
|
||||
- Si falta Python (o `python` es el stub falso de la Microsoft Store), el
|
||||
error parpadea y la ventana desaparece sin que se entienda nada.
|
||||
- Si faltan dependencias (`pip install -e ".[web]"` nunca se corrió), uvicorn
|
||||
muere con un traceback ilegible.
|
||||
- Los errores terminan con la ventana cerrándose: imposible leer qué pasó.
|
||||
|
||||
## Objetivo
|
||||
|
||||
Doble clic en `start-server.bat` = la aplicación abre en el navegador,
|
||||
siempre. El script detecta y repara el entorno por sí mismo. Un solo icono
|
||||
para el usuario; `stop-server.bat` queda como acompañante.
|
||||
|
||||
## Enfoque elegido (B)
|
||||
|
||||
Separación de responsabilidades: **entorno** vs **ciclo de vida**.
|
||||
|
||||
```
|
||||
start-server.bat ──> scripts/doctor.ps1 (NUEVO)
|
||||
├─ ¿server ya vivo? ── sí --> abrir navegador, fin
|
||||
├─ resolver Python (py -3 > python > python3)
|
||||
│ └─ ninguno válido --> winget install
|
||||
│ └─ sin winget --> abrir python.org + ERROR
|
||||
├─ import-test de deps
|
||||
│ └─ falla --> pip install -e ".[web]"
|
||||
├─ ffmpeg presente? -- no --> aviso suave NO bloqueante
|
||||
└─ ok --> scripts/start-server.ps1 -PythonExe $py
|
||||
└─ puertos, healthz, browser (SIN CAMBIOS de lógica)
|
||||
```
|
||||
|
||||
`start-server.ps1` conserva toda su lógica verificada (sondas TCP rápidas,
|
||||
limpieza de entradas obsoletas, barrido 8000-8100, apertura de navegador).
|
||||
Único cambio permitido: parámetro opcional `-PythonExe` (default `python`)
|
||||
usado en la línea de comando de uvicorn.
|
||||
|
||||
### Por qué `-PythonExe` es necesario
|
||||
|
||||
Tras una instalación fresca vía winget, el PATH de la sesión actual **no**
|
||||
incluye el Python nuevo. El doctor resuelve la ruta concreta
|
||||
(`%LocalAppData%\Programs\Python\Python312\python.exe`) y se la pasa al
|
||||
launcher; sin esto, el primer arranque post-instalación fallaría.
|
||||
|
||||
## Flujo detallado del doctor
|
||||
|
||||
1. **Server vivo primero** (antes de cualquier chequeo de entorno): sondea el
|
||||
puerto registrado en `.run/server.info`; si responde, abre el navegador y
|
||||
termina con exit 0. Evita instalar dependencias solo para decir "ya está
|
||||
corriendo".
|
||||
2. **Resolver Python**, candidatos en orden: `py -3`, `python`, `python3`.
|
||||
Cada candidato es válido solo si:
|
||||
- ejecuta `--version` y la salida matchea `Python 3.` con versión ≥ 3.10;
|
||||
- su ruta **no** contiene `WindowsApps` (stub de la Store que no ejecuta nada).
|
||||
3. **Instalar Python si falta**: `winget install --id Python.Python.3.12
|
||||
--silent --accept-package-agreements --accept-source-agreements`. Después
|
||||
de instalar, sondear rutas conocidas (`%LocalAppData%\Programs\Python\*`,
|
||||
`C:\Program Files\Python312\`) antes de rendirse. Sin winget disponible:
|
||||
abrir `https://www.python.org/downloads/` en el navegador + instrucción en
|
||||
lenguaje simple + exit ≠ 0.
|
||||
4. **Import-test de dependencias**:
|
||||
`& $py -c "import fastapi, uvicorn, jinja2, sse_starlette, yt_dlp"`.
|
||||
**Nunca** importa `yt_scraper.webapp.app` aquí: ese import abre la DB real
|
||||
del proyecto vía `config.yaml`, corre migraciones y `reconcile_markdown`
|
||||
(gotcha documentado en CLAUDE.md). Si falla el test:
|
||||
`-m pip install -e ".[web]"` mostrando progreso ("instalando por primera
|
||||
vez, puede tardar un par de minutos"). Si falta pip: `-m ensurepip`.
|
||||
5. **ffmpeg**: `Get-Command ffmpeg`; ausente → aviso amarillo no bloqueante
|
||||
("la descarga de audio no estará disponible; todo lo demás funciona").
|
||||
6. Delegar: `& start-server.ps1 -PythonExe $py` + passthrough de args.
|
||||
7. Propagar exit code.
|
||||
|
||||
## Reglas de ventana (.bat)
|
||||
|
||||
| Resultado | Ventana |
|
||||
|---|---|
|
||||
| Exit 0 | mensaje verde + cuenta regresiva ~5 s → cierra sola |
|
||||
| Exit ≠ 0 | **queda abierta** (`pause`) con causa legible + últimas líneas de log |
|
||||
|
||||
Hoy los errores parpadean y desaparecen: es la regla más importante del
|
||||
cambio. Aplica a ambos `.bat`.
|
||||
|
||||
## Mensajes
|
||||
|
||||
Español plano, sin jerga dirigida al usuario final ("servidor", "conexión";
|
||||
no "puerto", "PID", "uvicorn", "PATH"):
|
||||
|
||||
- `[ok] Servidor listo en http://localhost:8000 — abriendo tu navegador...`
|
||||
- `[ok] El servidor ya estaba corriendo — abriendo...`
|
||||
- `[info] Instalando Python automáticamente (~25 MB)…`
|
||||
- `[info] Instalando las piezas que faltan por primera vez (1-2 min)…`
|
||||
- Error: `Algo falló al iniciar el servidor. Esto fue lo último que hizo:`
|
||||
|
||||
El pulido de copys aplica también a los mensajes visibles de
|
||||
`start-server.ps1` y `stop-server.ps1` (solo texto; cero cambios de lógica).
|
||||
|
||||
## Errores cubiertos
|
||||
|
||||
| Caso | Detección | Acción |
|
||||
|---|---|---|
|
||||
| Stub falso Microsoft Store | `--version` falla o ruta `*WindowsApps*` | Siguiente candidato |
|
||||
| Sin Python válido | 3 candidatos fallan | winget silencioso; sin winget → python.org + ERROR |
|
||||
| PATH fresco post-instalación | winget OK | Sondeo de rutas conocidas |
|
||||
| pip ausente | `-m pip --version` falla | `ensurepip` |
|
||||
| Deps rotas/faltantes | import-test falla | `pip install -e ".[web]"` |
|
||||
| ffmpeg ausente | `Get-Command` vacío | Aviso amarillo no bloqueante |
|
||||
| Error de arranque | exit ≠ 0 del launcher | Ventana queda abierta con log |
|
||||
|
||||
## Fuera de alcance
|
||||
|
||||
- No se toca lógica de puertos, healthz, browser ni cancelación del launcher.
|
||||
- No hay `setup.bat` separado (el doctor es el setup).
|
||||
- No se instalan `[dev]` ni `[analysis]`: solo core + `[web]`, lo necesario
|
||||
para correr la aplicación.
|
||||
- Sin tests automatizados nuevos de PowerShell (el repo no tiene infra para
|
||||
eso); verificación manual guiada por `-CheckOnly`.
|
||||
|
||||
## Verificación
|
||||
|
||||
- `doctor.ps1 -CheckOnly`: reporta qué haría sin cambiar nada (cada rama de
|
||||
decisión observable sin efectos).
|
||||
- Escenarios manuales: server ya vivo → solo navegador; server apagado →
|
||||
arranque completo; stop limpio tras arranque; ventana persistente ante
|
||||
error simulado (p. ej. PATH sin Python en un subproceso controlado).
|
||||
@@ -0,0 +1,104 @@
|
||||
# Diseño: vista de grid estilo YouTube para la sección Vídeos
|
||||
|
||||
Fecha: 2026-08-22
|
||||
Estado: aprobado en conversación
|
||||
|
||||
## Problema
|
||||
|
||||
La sección Vídeos solo tiene visualización de tabla. Para navegar el catálogo
|
||||
visualmente (miniaturas grandes, explorar por canal/tema) falta una vista tipo
|
||||
YouTube: grid de tarjetas con miniatura 16:9, duración y metadatos compactos.
|
||||
|
||||
## Objetivo
|
||||
|
||||
Añadir un conmutador tabla ⇄ grid en la sección Vídeos. El grid replica el
|
||||
look de YouTube (tarjetas 16:9) manteniendo toda la funcionalidad existente:
|
||||
filtros, orden, selección masiva ("Download .md" / "Audio") y paginación.
|
||||
|
||||
## Decisiones tomadas en conversación
|
||||
|
||||
| Pregunta | Decisión |
|
||||
|---|---|
|
||||
| ¿Selección masiva en grid? | Sí, checkboxes como en la tabla; la barra masiva se comparte |
|
||||
| ¿"Square" literal o estilo YouTube? | Miniatura 16:9 estilo YouTube (no recorte 1:1) |
|
||||
| ¿Vista inicial y persistencia? | Tabla por defecto; la elección se recuerda en `localStorage` |
|
||||
| ¿Estructura en index.html? | Dos bloques hermanos alternados con `x-if`; filtros/barra/paginación compartidos |
|
||||
|
||||
## Cambios
|
||||
|
||||
### 1. Estado y persistencia (`static/app.js`)
|
||||
|
||||
- Nuevo campo en el store de `videos`: `view`, valores `'table' | 'grid'`,
|
||||
default `'table'`.
|
||||
- En `init()`: leer `localStorage["videos-view"]` y aceptarlo solo si es
|
||||
`'table'` o `'grid'` (mismo patrón defensivo que `videos-size`).
|
||||
- Método `setVideoView(v)`: asigna y guarda en `localStorage`. No toca
|
||||
`syncURL`/`hydrateURL`: es preferencia de UI como `videos-size`, no estado
|
||||
navegable.
|
||||
|
||||
### 2. Conmutador de vista (`index.html`)
|
||||
|
||||
- Control segmentado de dos botones icono (lista / grid) alineado a la
|
||||
derecha, en una fila fina inmediatamente encima del contenedor de
|
||||
resultados.
|
||||
- Una sola instancia visible en ambas vistas: fuera del `<thead>` de la
|
||||
tabla a propósito, porque el grid no tiene cabecera donde colgarlo.
|
||||
- Botón activo con `accent-grad btn-primary`, el mismo tratamiento que usa el
|
||||
número de página activo en la paginación; aria-pressed en cada botón.
|
||||
|
||||
### 3. Tarjetas del grid (`index.html`)
|
||||
|
||||
- Contenedor: `grid grid-cols-2 sm:grid-cols-3 lg:grid-cols-4 2xl:grid-cols-5 gap-4`.
|
||||
- Tarjeta: bloque `glass` redondeado, cursor-pointer, hover con borde rose
|
||||
(mismo lenguaje que las tarjetas de canales). Click → `openVideo(id)`;
|
||||
clase `row-selected` si está seleccionada.
|
||||
- Miniatura: `/api/thumbnails/{id}`, `aspect-video w-full object-cover`,
|
||||
esquinas superiores redondeadas, mismo `onerror="this.style.visibility='hidden'"`
|
||||
que la tabla.
|
||||
- Badge de duración abajo-derecha sobre la miniatura: `ts(v.duration)`,
|
||||
font-mono text-xs, fondo negro semitransparente.
|
||||
- Checkbox de selección arriba-izquierda sobre la miniatura: aparece al
|
||||
hover de la tarjeta o permanece visible si está seleccionada; click con
|
||||
`.stop` → `toggleSelect(id)`.
|
||||
- Cuerpo:
|
||||
- Título con clamp a 2 líneas (`line-clamp-2`), text-zinc-100.
|
||||
- Línea de metadatos: `chanName(channel_id)` · `fmtNum(view_count)` vistas ·
|
||||
`videoDate(v)` (con `~` si estimada; tooltip `videoDateTitle(v)`).
|
||||
- Pills de estado y candado de bloqueo idénticos a los de la tabla
|
||||
(`statusClass`, `isBlocked`, `blockLabel`, `blockTitle`).
|
||||
|
||||
### 4. Comportamiento compartido (sin cambios)
|
||||
|
||||
- Filtros, barra masiva, paginación y select de orden: intactos, aplican a
|
||||
ambas vistas.
|
||||
- Estados loading/vacío duplicados dentro del bloque del grid con los mismos
|
||||
mensajes ("loading…", "No videos match these filters.", "Clear filters").
|
||||
- Las acciones por vídeo `.md`/`Process`/`Clip` **no** se duplican en las
|
||||
tarjetas: quedan en la tabla y en la vista de detalle (look limpio tipo
|
||||
YouTube; un click abre el detalle que las tiene todas).
|
||||
|
||||
## Errores cubiertos
|
||||
|
||||
| Caso | Comportamiento |
|
||||
|---|---|
|
||||
| Miniatura ausente | `onerror` oculta el `<img>`; la tarjeta conserva proporción por el contenedor aspect-video |
|
||||
| Valor inválido/corrupto en localStorage | Se ignora y cae a `'table'` |
|
||||
| Selección activa + cambio de vista | `videos.selected` no se toca; la barra masiva sigue funcionando en ambas |
|
||||
|
||||
## Fuera de alcance
|
||||
|
||||
- No se toca backend ni endpoints: `/api/videos` ya devuelve todos los campos
|
||||
que la tarjeta necesita.
|
||||
- Sin acciones por tarjeta, sin hover-preview de vídeo, sin infinite scroll.
|
||||
- Sin tests automatizados nuevos: el repo no tiene infra de tests de frontend
|
||||
(pytest es backend sin red). Verificación manual.
|
||||
|
||||
## Verificación manual
|
||||
|
||||
Con `start-server.bat`:
|
||||
|
||||
1. Alternar tabla ⇄ grid: ambas renderizan los mismos vídeos del filtro activo.
|
||||
2. Marcar 3 tarjetas → barra masiva aparece → "Download .md" encola el job.
|
||||
3. Paginar en grid: sin repeticiones ni saltos (mismo desempate que tabla).
|
||||
4. Recargar la página: la vista elegida se conserva; primera visita → tabla.
|
||||
5. La tabla existente funciona exactamente igual que antes del cambio.
|
||||
@@ -1 +0,0 @@
|
||||
{"0": "Webapp Frontend (JS)", "1": "Transcript Extraction & Parsing", "2": "Config & Discovery Pipeline", "3": "Store & Export Layer", "4": "CLI Command Layer", "5": "Platform Design & Proposals", "6": "Segments Search & Backfill", "7": "Store Platform Tests", "8": "Content Analysis", "9": "README & Config Docs", "10": "Server Start Script", "11": "Package Manifest", "12": "Server Stop Script", "13": "Webapp Package Init"}
|
||||
@@ -1 +0,0 @@
|
||||
C:\Users\Uriel Jareth\AppData\Local\Programs\Python\Python312\python.exe
|
||||
@@ -1 +0,0 @@
|
||||
D:\yt-channel-scraper
|
||||
@@ -1,115 +0,0 @@
|
||||
# Graph Report - . (2026-07-26)
|
||||
|
||||
## Corpus Check
|
||||
- Corpus is ~28,854 words - fits in a single context window. You may not need a graph.
|
||||
|
||||
## Summary
|
||||
- 392 nodes · 951 edges · 14 communities (12 shown, 2 thin omitted)
|
||||
- Extraction: 93% EXTRACTED · 7% INFERRED · 0% AMBIGUOUS · INFERRED: 65 edges (avg confidence: 0.76)
|
||||
- Token cost: 11,800 input · 4,100 output
|
||||
|
||||
## Community Hubs (Navigation)
|
||||
- Webapp Frontend (JS)
|
||||
- Transcript Extraction & Parsing
|
||||
- Config & Discovery Pipeline
|
||||
- Store & Export Layer
|
||||
- CLI Command Layer
|
||||
- Platform Design & Proposals
|
||||
- Segments Search & Backfill
|
||||
- Store Platform Tests
|
||||
- Content Analysis
|
||||
- README & Config Docs
|
||||
- Package Manifest
|
||||
|
||||
## God Nodes (most connected - your core abstractions)
|
||||
1. `Store` - 81 edges
|
||||
2. `Segment` - 32 edges
|
||||
3. `api()` - 31 edges
|
||||
4. `toast()` - 29 edges
|
||||
5. `Config` - 23 edges
|
||||
6. `process_video()` - 19 edges
|
||||
7. `JobManager` - 17 edges
|
||||
8. `resolve_active_path()` - 13 edges
|
||||
9. `discover_channel()` - 13 edges
|
||||
10. `backfill_from_markdown()` - 13 edges
|
||||
|
||||
## Surprising Connections (you probably didn't know these)
|
||||
- `test_dashboard_aggregates()` --calls--> `Segment` [INFERRED]
|
||||
tests/test_store_platform.py → src/yt_scraper/parse.py
|
||||
- `test_store_and_search_segments()` --calls--> `Segment` [INFERRED]
|
||||
tests/test_store_platform.py → src/yt_scraper/parse.py
|
||||
- `test_store_segments_overwrites()` --calls--> `Segment` [INFERRED]
|
||||
tests/test_store_platform.py → src/yt_scraper/parse.py
|
||||
- `store()` --calls--> `Store` [INFERRED]
|
||||
tests/test_store_platform.py → src/yt_scraper/store.py
|
||||
- `test_migration_idempotent()` --calls--> `Store` [INFERRED]
|
||||
tests/test_store_platform.py → src/yt_scraper/store.py
|
||||
|
||||
## Import Cycles
|
||||
- None detected.
|
||||
|
||||
## Hyperedges (group relationships)
|
||||
- **Workstream B webapp stack (FastAPI + Alpine SPA + subagent)** — opencode_agent_webapp_builder, opencode_goals_webapp_build, src_yt_scraper_webapp_static_index, docs_superpowers_specs_2026_07_26_platform_design_dark_command_center, docs_superpowers_specs_2026_07_26_platform_design_sse_job_runner [EXTRACTED 0.95]
|
||||
- **FTS5 transcript search flow (proposal -> spec -> SPA view)** — opportunities_search_fts5, docs_superpowers_specs_2026_07_26_platform_design_transcript_fts5, src_yt_scraper_webapp_static_index [INFERRED 0.85]
|
||||
- **Cookie auth chain (vault -> resolve_active_path -> scrape job)** — docs_superpowers_specs_2026_07_26_platform_design_cookie_vault_ux, docs_superpowers_specs_2026_07_26_platform_design_sse_job_runner, docs_superpowers_specs_2026_07_26_platform_design_cookies_bug_fix, src_yt_scraper_webapp_static_index [INFERRED 0.85]
|
||||
|
||||
## Communities (14 total, 2 thin omitted)
|
||||
|
||||
### Community 0 - "Webapp Frontend (JS)"
|
||||
Cohesion: 0.06
|
||||
Nodes (59): activateCookie(), addChannel(), api(), axisOpts(), barOpts(), cancelScrape(), checkHealth(), closeStream() (+51 more)
|
||||
|
||||
### Community 1 - "Transcript Extraction & Parsing"
|
||||
Cohesion: 0.07
|
||||
Nodes (57): Environment, align_chapters(), Chapter, chapters_from_info(), Section, _parse_tags(), Regenerar Markdown desde segmentos almacenados., re_render_cmd() (+49 more)
|
||||
|
||||
### Community 2 - "Config & Discovery Pipeline"
|
||||
Cohesion: 0.06
|
||||
Nodes (40): APIRouter, FastAPI, _build_config(), Config, DelayConfig, load_config(), Any, Path (+32 more)
|
||||
|
||||
### Community 3 - "Store & Export Layer"
|
||||
Cohesion: 0.07
|
||||
Nodes (28): Connection, Cursor, Row, is_expired(), list_vault(), export_csv(), export_html(), export_json() (+20 more)
|
||||
|
||||
### Community 4 - "CLI Command Layer"
|
||||
Cohesion: 0.08
|
||||
Nodes (41): analyze_cmd(), _apply_filters(), audio_cmd(), _channel_targets(), channels_add(), cli(), export_cmd(), _extract_handle() (+33 more)
|
||||
|
||||
### Community 5 - "Platform Design & Proposals"
|
||||
Cohesion: 0.12
|
||||
Nodes (20): Platform Implementation Plan, Platform Design Spec, Cookie drag-and-drop vault UX, cli.py cookies scope bug fix, Design language: dark data command center, Architecture: Monorepo in-package webapp (Approach A), Scope rule: No AI features, Single-threaded scrape job runner + SSE event bus (+12 more)
|
||||
|
||||
### Community 6 - "Segments Search & Backfill"
|
||||
Cohesion: 0.26
|
||||
Nodes (11): backfill_from_markdown(), parse_markdown(), ParsedMarkdown, Path, Parse all done .md files under md_root, populate segments/FTS/metadata. Idempote, search(), store_segments(), _strip_quotes() (+3 more)
|
||||
|
||||
### Community 7 - "Store Platform Tests"
|
||||
Cohesion: 0.17
|
||||
Nodes (5): store(), test_dashboard_aggregates(), test_migration_idempotent(), test_store_and_search_segments(), test_store_segments_overwrites()
|
||||
|
||||
### Community 8 - "Content Analysis"
|
||||
Cohesion: 0.27
|
||||
Nodes (10): _iter_texts(), Path, Return [(YYYY-MM, count)] of months where `term` appears in transcripts., render_timeline_chart(), render_top_words_chart(), render_wordcloud(), term_timeline(), _tokenize() (+2 more)
|
||||
|
||||
## Knowledge Gaps
|
||||
- **10 isolated node(s):** `yt-channel-scraper`, `OPPORTUNITIES.md - extension audit`, `Design language: dark data command center`, `yt-dlp InnerTube mechanism (ANDROID/IOS/WEB clients)`, `Feature proposal: search (FTS5 transcript search)` (+5 more)
|
||||
These have ≤1 connection - possible missing edges or undocumented components.
|
||||
- **2 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
|
||||
|
||||
## Suggested Questions
|
||||
_Questions this graph is uniquely positioned to answer:_
|
||||
|
||||
- **Why does `Store` connect `Store & Export Layer` to `Transcript Extraction & Parsing`, `Config & Discovery Pipeline`, `CLI Command Layer`, `Segments Search & Backfill`, `Store Platform Tests`, `Content Analysis`?**
|
||||
_High betweenness centrality (0.194) - this node is a cross-community bridge._
|
||||
- **Why does `Segment` connect `Transcript Extraction & Parsing` to `CLI Command Layer`, `Segments Search & Backfill`, `Store Platform Tests`?**
|
||||
_High betweenness centrality (0.057) - this node is a cross-community bridge._
|
||||
- **Why does `Config` connect `Config & Discovery Pipeline` to `Transcript Extraction & Parsing`, `CLI Command Layer`?**
|
||||
_High betweenness centrality (0.024) - this node is a cross-community bridge._
|
||||
- **Are the 4 inferred relationships involving `Store` (e.g. with `ParsedMarkdown` and `store()`) actually correct?**
|
||||
_`Store` has 4 INFERRED edges - model-reasoned connections that need verification._
|
||||
- **Are the 17 inferred relationships involving `Segment` (e.g. with `Chapter` and `Section`) actually correct?**
|
||||
_`Segment` has 17 INFERRED edges - model-reasoned connections that need verification._
|
||||
- **What connects `yt-channel-scraper`, `OPPORTUNITIES.md - extension audit`, `Design language: dark data command center` to the rest of the system?**
|
||||
_10 weakly-connected nodes found - possible documentation gaps or missing edges._
|
||||
- **Should `Webapp Frontend (JS)` be split into smaller, more focused modules?**
|
||||
_Cohesion score 0.062317429406037 - nodes in this community are weakly interconnected._
|
||||
graphify-out/cache/ast/v0.9.22/017d7d8cd5411dff3256656fa40fc036ebae1576f37e719d258f54ea66ded544.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/09a73a233d94b95f31f750ca2d96fe89ac48cf7ac69cff22e9f5112018c0c2ac.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/0c459d37d23ee068ff0a98b53a93f6c65070e09c18153545989b049560f4f594.json
Vendored
-1
@@ -1 +0,0 @@
|
||||
{"nodes": [{"id": "d_yt_channel_scraper_scripts_start_server_ps1", "label": "start-server.ps1", "file_type": "code", "source_file": "scripts/start-server.ps1", "source_location": "L1"}, {"id": "d_yt_channel_scraper_scripts_start_server_open_browser", "label": "Open-Browser()", "file_type": "code", "source_file": "scripts/start-server.ps1", "source_location": "L6"}], "edges": [{"source": "d_yt_channel_scraper_scripts_start_server_ps1", "target": "d_yt_channel_scraper_scripts_start_server_open_browser", "relation": "contains", "confidence": "EXTRACTED", "source_file": "scripts/start-server.ps1", "source_location": "L6", "weight": 1.0}], "raw_calls": [{"caller_nid": "d_yt_channel_scraper_scripts_start_server_open_browser", "callee": "Start-Process", "is_member_call": false, "source_file": "scripts/start-server.ps1", "source_location": "L9"}, {"caller_nid": "d_yt_channel_scraper_scripts_start_server_open_browser", "callee": "Write-Host", "is_member_call": false, "source_file": "scripts/start-server.ps1", "source_location": "L18"}, {"caller_nid": "d_yt_channel_scraper_scripts_start_server_open_browser", "callee": "Write-Host", "is_member_call": false, "source_file": "scripts/start-server.ps1", "source_location": "L19"}]}
|
||||
graphify-out/cache/ast/v0.9.22/0c924be6b3e5c4aefe549a28f8a314ae236b1b76207023e80fefa32691dce799.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/142ed79d04b7ffc16d45c4fd83b668cbc00e32ba622283aeca1f31210664b0f6.json
Vendored
-1
@@ -1 +0,0 @@
|
||||
{"nodes": [{"id": "d_yt_channel_scraper_src_yt_scraper_ratelimit_py", "label": "ratelimit.py", "file_type": "code", "source_file": "src/yt_scraper/ratelimit.py", "source_location": "L1"}, {"id": "d_yt_channel_scraper_src_yt_scraper_ratelimit_polite_sleep", "label": "polite_sleep()", "file_type": "code", "source_file": "src/yt_scraper/ratelimit.py", "source_location": "L7", "_callable": true}, {"id": "d_yt_channel_scraper_src_yt_scraper_ratelimit_backoff_sleep", "label": "backoff_sleep()", "file_type": "code", "source_file": "src/yt_scraper/ratelimit.py", "source_location": "L12", "_callable": true}], "edges": [{"source": "d_yt_channel_scraper_src_yt_scraper_ratelimit_py", "target": "random", "relation": "imports", "context": "import", "confidence": "EXTRACTED", "source_file": "src/yt_scraper/ratelimit.py", "source_location": "L3", "weight": 1.0}, {"source": "d_yt_channel_scraper_src_yt_scraper_ratelimit_py", "target": "time", "relation": "imports", "context": "import", "confidence": "EXTRACTED", "source_file": "src/yt_scraper/ratelimit.py", "source_location": "L4", "weight": 1.0}, {"source": "d_yt_channel_scraper_src_yt_scraper_ratelimit_py", "target": "d_yt_channel_scraper_src_yt_scraper_ratelimit_polite_sleep", "relation": "contains", "confidence": "EXTRACTED", "source_file": "src/yt_scraper/ratelimit.py", "source_location": "L7", "weight": 1.0}, {"source": "d_yt_channel_scraper_src_yt_scraper_ratelimit_py", "target": "d_yt_channel_scraper_src_yt_scraper_ratelimit_backoff_sleep", "relation": "contains", "confidence": "EXTRACTED", "source_file": "src/yt_scraper/ratelimit.py", "source_location": "L12", "weight": 1.0}], "raw_calls": [{"caller_nid": "d_yt_channel_scraper_src_yt_scraper_ratelimit_polite_sleep", "callee": "uniform", "is_member_call": true, "source_file": "src/yt_scraper/ratelimit.py", "source_location": "L8", "receiver": "random"}, {"caller_nid": "d_yt_channel_scraper_src_yt_scraper_ratelimit_polite_sleep", "callee": "sleep", "is_member_call": true, "source_file": "src/yt_scraper/ratelimit.py", "source_location": "L9", "receiver": "time"}, {"caller_nid": "d_yt_channel_scraper_src_yt_scraper_ratelimit_backoff_sleep", "callee": "uniform", "is_member_call": true, "source_file": "src/yt_scraper/ratelimit.py", "source_location": "L13", "receiver": "random"}, {"caller_nid": "d_yt_channel_scraper_src_yt_scraper_ratelimit_backoff_sleep", "callee": "sleep", "is_member_call": true, "source_file": "src/yt_scraper/ratelimit.py", "source_location": "L14", "receiver": "time"}]}
|
||||
graphify-out/cache/ast/v0.9.22/285ff3f97a258a1caa08b81e25d78590eb1aace9284a1a69d921f73bf80610b0.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/2f15912ebacf22f457d91abb145e94909b16faa74732de18a2f07e31afac47a9.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/301d5a1fde90564e824d05b8b427da6644cbd693ad8ac844821e5563389c2ce8.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/33bef431555f5edd76ea2c4a39cfca53964da20dae91d0b46c81b5e21e5bfe81.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/39d6fb49fe1a6399b04c85f41a27c66deef9a756a36a33d61f45cbbadc2306c6.json
Vendored
-1
@@ -1 +0,0 @@
|
||||
{"nodes": [{"id": "d_yt_channel_scraper_src_yt_scraper_webapp_init_py", "label": "__init__.py", "file_type": "code", "source_file": "src/yt_scraper/webapp/__init__.py", "source_location": "L1"}], "edges": [], "raw_calls": []}
|
||||
graphify-out/cache/ast/v0.9.22/415969ae62c7c3b54477a743db17ff1d00b8020b7e4e469d4191a24ecb220299.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/5c0cce89c92ed6eab1cd14158c3f5f020a68cc0469d2129f7c3af424d184fa6d.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/5eb1ea5adee7b3a6de0470fbefdc34a0081f0886289248d1a6985eebd5d048f2.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/5f5648fe20db27632fcc9468b33427f8525e761bd11953a5cb6fe268c27823b8.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/64a9963471c61688046509288debd54d5eb992be7e3bd86fd0e0d23714316585.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/7aa050f1ed896a0e062669f2f287f75ceb7a57b22ddcc5b635bed3d59779b759.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/7b6c9aede202fa24c4070eff2a45987f38369a6d73121e2cc8bbaafdaaed6ec3.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/92ebf7566766c51a5116f810a79d6444db5472466580806cb39777c868b5a5b4.json
Vendored
-1
@@ -1 +0,0 @@
|
||||
{"nodes": [{"id": "d_yt_channel_scraper_src_yt_scraper_init_py", "label": "__init__.py", "file_type": "code", "source_file": "src/yt_scraper/__init__.py", "source_location": "L1"}], "edges": [], "raw_calls": []}
|
||||
graphify-out/cache/ast/v0.9.22/9b4e29bbe5f0ed1517912d40ea1ca2d60555bef6e8416b4c5cb2e55d4227267d.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/a1dcd3fd32ecc99846a876848d9b1f874e34eaf2fc4de10ff245b757a53209da.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/b18e8b9c6f644dc007f70bf7048f2b885991aeade60c75cd18b33df95dec0a3c.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/bc8259cf378616dd8bcb0d0d75ba8b67e3355f918fcc614fe61389ba3795097c.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/bcb2c3a21e73d7c8c00161c1585bc15ed856116eac5c37608af0699fb5cce7ad.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/bd62de6ce97df98a14ebd1098c91707fc0dd9ba568437c71bb3a075c519b7882.json
Vendored
-1
@@ -1 +0,0 @@
|
||||
{"nodes": [{"id": "d_yt_channel_scraper_scripts_stop_server_ps1", "label": "stop-server.ps1", "file_type": "code", "source_file": "scripts/stop-server.ps1", "source_location": "L1"}], "edges": [], "raw_calls": []}
|
||||
graphify-out/cache/ast/v0.9.22/bd9731c51b005a5cef7bd97a9555a9d5609eaa8c4da9f147a8708571c6440ff7.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/bdd04141d70cba0f2a1b2d4375870c2c475971df35c1d55cf38ca92a1a78b6b7.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/cca26b54622e61299b9e7165fb56629045ac8ba8bdc08fb67d0f1b4422696157.json
Vendored
-1
File diff suppressed because one or more lines are too long
graphify-out/cache/ast/v0.9.22/d769492446cb53e6915aca1bd4ca33178a34670ec6b919dd2c150fd236b23c9e.json
Vendored
-1
@@ -1 +0,0 @@
|
||||
{"nodes": [{"id": "pkg_yt_channel_scraper", "label": "yt-channel-scraper", "file_type": "code", "type": "package", "ecosystem": "python", "source_file": "pyproject.toml", "source_location": "L1", "version": "0.1.0"}], "edges": [{"source": "pkg_yt_channel_scraper", "target": "pkg_yt_dlp", "relation": "depends_on", "context": "dependency", "confidence": "EXTRACTED", "confidence_score": 1.0, "source_file": "pyproject.toml", "source_location": "L1", "weight": 1.0}, {"source": "pkg_yt_channel_scraper", "target": "pkg_click", "relation": "depends_on", "context": "dependency", "confidence": "EXTRACTED", "confidence_score": 1.0, "source_file": "pyproject.toml", "source_location": "L1", "weight": 1.0}, {"source": "pkg_yt_channel_scraper", "target": "pkg_rich", "relation": "depends_on", "context": "dependency", "confidence": "EXTRACTED", "confidence_score": 1.0, "source_file": "pyproject.toml", "source_location": "L1", "weight": 1.0}, {"source": "pkg_yt_channel_scraper", "target": "pkg_jinja2", "relation": "depends_on", "context": "dependency", "confidence": "EXTRACTED", "confidence_score": 1.0, "source_file": "pyproject.toml", "source_location": "L1", "weight": 1.0}, {"source": "pkg_yt_channel_scraper", "target": "pkg_pyyaml", "relation": "depends_on", "context": "dependency", "confidence": "EXTRACTED", "confidence_score": 1.0, "source_file": "pyproject.toml", "source_location": "L1", "weight": 1.0}, {"source": "pkg_yt_channel_scraper", "target": "pkg_python_slugify", "relation": "depends_on", "context": "dependency", "confidence": "EXTRACTED", "confidence_score": 1.0, "source_file": "pyproject.toml", "source_location": "L1", "weight": 1.0}, {"source": "pkg_yt_channel_scraper", "target": "pkg_requests", "relation": "depends_on", "context": "dependency", "confidence": "EXTRACTED", "confidence_score": 1.0, "source_file": "pyproject.toml", "source_location": "L1", "weight": 1.0}]}
|
||||
-1
@@ -1 +0,0 @@
|
||||
{"nodes": [{"id": "src_yt_scraper_webapp_static_index", "label": "SPA shell (index.html)", "file_type": "document", "source_file": "src/yt_scraper/webapp/static/index.html", "source_location": null, "source_url": null, "captured_at": null, "author": null, "contributor": null}], "edges": [{"source": "src_yt_scraper_webapp_static_index", "target": "docs_superpowers_specs_2026_07_26_platform_design_dark_command_center", "relation": "implements", "confidence": "INFERRED", "confidence_score": 0.95, "source_file": "src/yt_scraper/webapp/static/index.html", "source_location": null, "weight": 1.0}, {"source": "src_yt_scraper_webapp_static_index", "target": "docs_superpowers_specs_2026_07_26_platform_design_cookie_vault_ux", "relation": "implements", "confidence": "INFERRED", "confidence_score": 0.95, "source_file": "src/yt_scraper/webapp/static/index.html", "source_location": "Cookies view dropzone", "weight": 1.0}], "hyperedges": []}
|
||||
-1
@@ -1 +0,0 @@
|
||||
{"nodes": [{"id": "opencode_agent_webapp_builder", "label": "webapp-builder agent definition", "file_type": "document", "source_file": ".opencode/agent/webapp-builder.md", "source_location": null, "source_url": null, "captured_at": null, "author": null, "contributor": null}, {"id": "opencode_agent_webapp_builder_store_data_layer", "label": "Store API surface (data layer the webapp reuses)", "file_type": "concept", "source_file": ".opencode/agent/webapp-builder.md", "source_location": null, "source_url": null, "captured_at": null, "author": null, "contributor": null, "rationale": "Agent must REUSE not duplicate: Store(db_path), cookies.py, segments.py, analysis.py, export.py, pipeline.process_video, config.Config. Active cookie (cookies.resolve_active_path) MUST be passed as cookies_file to every scrape job."}], "edges": [{"source": "opencode_agent_webapp_builder", "target": "docs_superpowers_specs_2026_07_26_platform_design", "relation": "references", "confidence": "EXTRACTED", "confidence_score": 1.0, "source_file": ".opencode/agent/webapp-builder.md", "source_location": "last line", "weight": 1.0}, {"source": "opencode_agent_webapp_builder", "target": "docs_superpowers_plans_2026_07_26_platform_build", "relation": "references", "confidence": "EXTRACTED", "confidence_score": 1.0, "source_file": ".opencode/agent/webapp-builder.md", "source_location": "last line", "weight": 1.0}, {"source": "opencode_agent_webapp_builder", "target": "opencode_agent_webapp_builder_store_data_layer", "relation": "references", "confidence": "EXTRACTED", "confidence_score": 1.0, "source_file": ".opencode/agent/webapp-builder.md", "source_location": "What already exists section", "weight": 1.0}, {"source": "opencode_agent_webapp_builder", "target": "src_yt_scraper_webapp_static_index", "relation": "references", "confidence": "INFERRED", "confidence_score": 0.85, "source_file": ".opencode/agent/webapp-builder.md", "source_location": null, "weight": 1.0}], "hyperedges": [{"id": "webapp_stack_b", "label": "Workstream B webapp stack (FastAPI + Alpine SPA + subagent)", "nodes": ["opencode_agent_webapp_builder", "opencode_goals_webapp_build", "src_yt_scraper_webapp_static_index", "docs_superpowers_specs_2026_07_26_platform_design_dark_command_center", "docs_superpowers_specs_2026_07_26_platform_design_sse_job_runner"], "relation": "participate_in", "confidence": "EXTRACTED", "confidence_score": 0.95, "source_file": ".opencode/agent/webapp-builder.md"}]}
|
||||
-1
@@ -1 +0,0 @@
|
||||
{"nodes": [{"id": "opencode_goals_webapp_build", "label": "Sub-goal: Build the local webapp (Workstream B)", "file_type": "document", "source_file": ".opencode/goals/webapp-build.md", "source_location": null, "source_url": null, "captured_at": null, "author": null, "contributor": null}], "edges": [{"source": "opencode_goals_webapp_build", "target": "opencode_agent_webapp_builder", "relation": "references", "confidence": "EXTRACTED", "confidence_score": 1.0, "source_file": ".opencode/goals/webapp-build.md", "source_location": "Owner agent line", "weight": 1.0}, {"source": "opencode_goals_webapp_build", "target": "docs_superpowers_specs_2026_07_26_platform_design", "relation": "references", "confidence": "EXTRACTED", "confidence_score": 1.0, "source_file": ".opencode/goals/webapp-build.md", "source_location": "Spec line", "weight": 1.0}], "hyperedges": []}
|
||||
-1
@@ -1 +0,0 @@
|
||||
{"nodes": [{"id": "docs_superpowers_plans_2026_07_26_platform_build", "label": "Platform Implementation Plan", "file_type": "document", "source_file": "docs/superpowers/plans/2026-07-26-platform-build.md", "source_location": null, "source_url": null, "captured_at": null, "author": null, "contributor": null}], "edges": [{"source": "docs_superpowers_plans_2026_07_26_platform_build", "target": "docs_superpowers_specs_2026_07_26_platform_design", "relation": "references", "confidence": "EXTRACTED", "confidence_score": 1.0, "source_file": "docs/superpowers/plans/2026-07-26-platform-build.md", "source_location": "Spec line", "weight": 1.0}], "hyperedges": []}
|
||||
-1
@@ -1 +0,0 @@
|
||||
{"nodes": [{"id": "readme", "label": "yt-channel-scraper README", "file_type": "document", "source_file": "README.md", "source_location": null, "source_url": null, "captured_at": null, "author": null, "contributor": null}, {"id": "readme_yt_dlp_innertube", "label": "yt-dlp InnerTube mechanism (ANDROID/IOS/WEB clients)", "file_type": "concept", "source_file": "README.md", "source_location": null, "source_url": null, "captured_at": null, "author": null, "contributor": null}], "edges": [{"source": "readme", "target": "config_example", "relation": "references", "confidence": "EXTRACTED", "confidence_score": 1.0, "source_file": "README.md", "source_location": "Configuraci\u00f3n section", "weight": 1.0}, {"source": "readme", "target": "readme_yt_dlp_innertube", "relation": "references", "confidence": "EXTRACTED", "confidence_score": 1.0, "source_file": "README.md", "source_location": "header", "weight": 1.0}], "hyperedges": []}
|
||||
-1
@@ -1 +0,0 @@
|
||||
{"nodes": [{"id": "opportunities", "label": "OPPORTUNITIES.md - extension audit", "file_type": "document", "source_file": "OPPORTUNITIES.md", "source_location": null, "source_url": null, "captured_at": null, "author": null, "contributor": null}, {"id": "opportunities_search_fts5", "label": "Feature proposal: search (FTS5 transcript search)", "file_type": "concept", "source_file": "OPPORTUNITIES.md", "source_location": "Tier 1 #1", "source_url": null, "captured_at": null, "author": null, "contributor": null}, {"id": "opportunities_llm_summaries", "label": "Feature proposal: LLM summaries per video", "file_type": "concept", "source_file": "OPPORTUNITIES.md", "source_location": "Tier 3 #12", "source_url": null, "captured_at": null, "author": null, "contributor": null}, {"id": "opportunities_rag_qa", "label": "Feature proposal: Q&A / RAG over channel", "file_type": "concept", "source_file": "OPPORTUNITIES.md", "source_location": "Tier 3 #13", "source_url": null, "captured_at": null, "author": null, "contributor": null}, {"id": "opportunities_export_formats", "label": "Feature proposal: multi-format export (json/csv/srt/html)", "file_type": "concept", "source_file": "OPPORTUNITIES.md", "source_location": "Tier 1 #3", "source_url": null, "captured_at": null, "author": null, "contributor": null}, {"id": "opportunities_analysis_wordcloud", "label": "Feature proposal: content analysis (wordcloud/top-words/timeline)", "file_type": "concept", "source_file": "OPPORTUNITIES.md", "source_location": "Tier 2 #9", "source_url": null, "captured_at": null, "author": null, "contributor": null}, {"id": "opportunities_watch_mode", "label": "Feature proposal: watch mode (periodic discovery)", "file_type": "concept", "source_file": "OPPORTUNITIES.md", "source_location": "Tier 2 #8", "source_url": null, "captured_at": null, "author": null, "contributor": null}], "edges": [{"source": "opportunities_search_fts5", "target": "docs_superpowers_specs_2026_07_26_platform_design_transcript_fts5", "relation": "semantically_similar_to", "confidence": "INFERRED", "confidence_score": 0.95, "source_file": "OPPORTUNITIES.md", "source_location": null, "weight": 1.0}, {"source": "opportunities_export_formats", "target": "opencode_agent_webapp_builder_store_data_layer", "relation": "semantically_similar_to", "confidence": "INFERRED", "confidence_score": 0.65, "source_file": "OPPORTUNITIES.md", "source_location": null, "weight": 1.0}, {"source": "opportunities_watch_mode", "target": "docs_superpowers_specs_2026_07_26_platform_design_sse_job_runner", "relation": "conceptually_related_to", "confidence": "INFERRED", "confidence_score": 0.65, "source_file": "OPPORTUNITIES.md", "source_location": null, "weight": 1.0}, {"source": "opportunities_analysis_wordcloud", "target": "opencode_agent_webapp_builder_store_data_layer", "relation": "conceptually_related_to", "confidence": "INFERRED", "confidence_score": 0.65, "source_file": "OPPORTUNITIES.md", "source_location": null, "weight": 1.0}], "hyperedges": []}
|
||||
-1
@@ -1 +0,0 @@
|
||||
{"nodes": [{"id": "config_example", "label": "config.example.yaml", "file_type": "document", "source_file": "config.example.yaml", "source_location": null, "source_url": null, "captured_at": null, "author": null, "contributor": null}], "edges": [], "hyperedges": []}
|
||||
-1
File diff suppressed because one or more lines are too long
Vendored
-1
File diff suppressed because one or more lines are too long
@@ -1,12 +0,0 @@
|
||||
{
|
||||
"runs": [
|
||||
{
|
||||
"date": "2026-07-27T05:54:56.822038+00:00",
|
||||
"input_tokens": 11800,
|
||||
"output_tokens": 4100,
|
||||
"files": 37
|
||||
}
|
||||
],
|
||||
"total_input_tokens": 11800,
|
||||
"total_output_tokens": 4100
|
||||
}
|
||||
File diff suppressed because one or more lines are too long
-15135
File diff suppressed because it is too large
Load Diff
@@ -1,187 +0,0 @@
|
||||
{
|
||||
"pyproject.toml": {
|
||||
"mtime": 1785126752.9759955,
|
||||
"ast_hash": "f80dc007b3b67c7f7a78d9df01a81ead",
|
||||
"semantic_hash": "f80dc007b3b67c7f7a78d9df01a81ead"
|
||||
},
|
||||
"scripts/start-server.ps1": {
|
||||
"mtime": 1785130613.533591,
|
||||
"ast_hash": "9baf397df6c14b7917ca59fbe6df61cc",
|
||||
"semantic_hash": "9baf397df6c14b7917ca59fbe6df61cc"
|
||||
},
|
||||
"scripts/stop-server.ps1": {
|
||||
"mtime": 1785127403.2329483,
|
||||
"ast_hash": "4398501c202a3eb624fb439f6d03fcf5",
|
||||
"semantic_hash": "4398501c202a3eb624fb439f6d03fcf5"
|
||||
},
|
||||
"src/yt_scraper/__init__.py": {
|
||||
"mtime": 1785120608.3336751,
|
||||
"ast_hash": "4867131295172353fe6c2295851defc7",
|
||||
"semantic_hash": "4867131295172353fe6c2295851defc7"
|
||||
},
|
||||
"src/yt_scraper/analysis.py": {
|
||||
"mtime": 1785126463.1545029,
|
||||
"ast_hash": "c28d08131e6594e1e7c6111ff01b2d37",
|
||||
"semantic_hash": "c28d08131e6594e1e7c6111ff01b2d37"
|
||||
},
|
||||
"src/yt_scraper/chapters.py": {
|
||||
"mtime": 1785120675.3308973,
|
||||
"ast_hash": "1bf51b7c13486b9db9bcc3420fba2593",
|
||||
"semantic_hash": "1bf51b7c13486b9db9bcc3420fba2593"
|
||||
},
|
||||
"src/yt_scraper/cli.py": {
|
||||
"mtime": 1785126572.0575886,
|
||||
"ast_hash": "9734cedf7c1e045e1ab0e4a102860e29",
|
||||
"semantic_hash": "9734cedf7c1e045e1ab0e4a102860e29"
|
||||
},
|
||||
"src/yt_scraper/config.py": {
|
||||
"mtime": 1785120620.335521,
|
||||
"ast_hash": "2f3c8f7768a942e12c0a9e41b54891be",
|
||||
"semantic_hash": "2f3c8f7768a942e12c0a9e41b54891be"
|
||||
},
|
||||
"src/yt_scraper/cookies.py": {
|
||||
"mtime": 1785126417.4218228,
|
||||
"ast_hash": "dd5ec44bb4c6e6375be80f21307fc73d",
|
||||
"semantic_hash": "dd5ec44bb4c6e6375be80f21307fc73d"
|
||||
},
|
||||
"src/yt_scraper/discover.py": {
|
||||
"mtime": 1785120708.331971,
|
||||
"ast_hash": "8db06d0b131575f709b55467eabdc2a9",
|
||||
"semantic_hash": "8db06d0b131575f709b55467eabdc2a9"
|
||||
},
|
||||
"src/yt_scraper/export.py": {
|
||||
"mtime": 1785126439.6829402,
|
||||
"ast_hash": "f1cff31fb1004f08e997655f838e3a58",
|
||||
"semantic_hash": "f1cff31fb1004f08e997655f838e3a58"
|
||||
},
|
||||
"src/yt_scraper/extract.py": {
|
||||
"mtime": 1785122183.7294595,
|
||||
"ast_hash": "3e2179db2f7fd316d2b3acedee536086",
|
||||
"semantic_hash": "3e2179db2f7fd316d2b3acedee536086"
|
||||
},
|
||||
"src/yt_scraper/monitor.py": {
|
||||
"mtime": 1785126496.288271,
|
||||
"ast_hash": "a64188f938dc4f8917335ef41d333ead",
|
||||
"semantic_hash": "a64188f938dc4f8917335ef41d333ead"
|
||||
},
|
||||
"src/yt_scraper/parse.py": {
|
||||
"mtime": 1785120665.3384771,
|
||||
"ast_hash": "25354365ac1ab52f1561c97256573806",
|
||||
"semantic_hash": "25354365ac1ab52f1561c97256573806"
|
||||
},
|
||||
"src/yt_scraper/pipeline.py": {
|
||||
"mtime": 1785131195.7311583,
|
||||
"ast_hash": "ebce1b294140bf40f809848ef280c290",
|
||||
"semantic_hash": "ebce1b294140bf40f809848ef280c290"
|
||||
},
|
||||
"src/yt_scraper/ratelimit.py": {
|
||||
"mtime": 1785120650.3795395,
|
||||
"ast_hash": "6082befe56d342fc43abd5299dd6c6ea",
|
||||
"semantic_hash": "6082befe56d342fc43abd5299dd6c6ea"
|
||||
},
|
||||
"src/yt_scraper/render.py": {
|
||||
"mtime": 1785120688.3247268,
|
||||
"ast_hash": "ff83764f3a8c8874557993e11eab1f89",
|
||||
"semantic_hash": "ff83764f3a8c8874557993e11eab1f89"
|
||||
},
|
||||
"src/yt_scraper/segments.py": {
|
||||
"mtime": 1785128385.6652262,
|
||||
"ast_hash": "09c7da35bc09a3f75fd0a627c0534746",
|
||||
"semantic_hash": "09c7da35bc09a3f75fd0a627c0534746"
|
||||
},
|
||||
"src/yt_scraper/store.py": {
|
||||
"mtime": 1785131583.548138,
|
||||
"ast_hash": "9a3d4286214aa54ae4ac4df2097b643d",
|
||||
"semantic_hash": "9a3d4286214aa54ae4ac4df2097b643d"
|
||||
},
|
||||
"src/yt_scraper/webapp/__init__.py": {
|
||||
"mtime": 1785126827.4957752,
|
||||
"ast_hash": "d41d8cd98f00b204e9800998ecf8427e",
|
||||
"semantic_hash": "d41d8cd98f00b204e9800998ecf8427e"
|
||||
},
|
||||
"src/yt_scraper/webapp/api.py": {
|
||||
"mtime": 1785131595.9914868,
|
||||
"ast_hash": "e64174f04d6eab0747b1d2a7c2a96a01",
|
||||
"semantic_hash": "e64174f04d6eab0747b1d2a7c2a96a01"
|
||||
},
|
||||
"src/yt_scraper/webapp/app.py": {
|
||||
"mtime": 1785126888.8032887,
|
||||
"ast_hash": "e0b1925b4318305dcf746d360373ff50",
|
||||
"semantic_hash": "e0b1925b4318305dcf746d360373ff50"
|
||||
},
|
||||
"src/yt_scraper/webapp/jobs.py": {
|
||||
"mtime": 1785131211.0397103,
|
||||
"ast_hash": "8b6c80001ff8c95eaaa16f7b61d1f207",
|
||||
"semantic_hash": "8b6c80001ff8c95eaaa16f7b61d1f207"
|
||||
},
|
||||
"src/yt_scraper/webapp/static/app.js": {
|
||||
"mtime": 1785131664.8792574,
|
||||
"ast_hash": "f6b081fadc5884d379152abbd1e59700",
|
||||
"semantic_hash": "f6b081fadc5884d379152abbd1e59700"
|
||||
},
|
||||
"tests/test_chapters.py": {
|
||||
"mtime": 1785120789.3300896,
|
||||
"ast_hash": "6d5303fa86de8dc99af6ae2ee17da033",
|
||||
"semantic_hash": "6d5303fa86de8dc99af6ae2ee17da033"
|
||||
},
|
||||
"tests/test_features.py": {
|
||||
"mtime": 1785126709.65728,
|
||||
"ast_hash": "21f9c44ce7cc1d618ece402ab7c2b057",
|
||||
"semantic_hash": "21f9c44ce7cc1d618ece402ab7c2b057"
|
||||
},
|
||||
"tests/test_parse.py": {
|
||||
"mtime": 1785120974.3283188,
|
||||
"ast_hash": "57ce74ee10fe6f9a90126ec7d9fd96b0",
|
||||
"semantic_hash": "57ce74ee10fe6f9a90126ec7d9fd96b0"
|
||||
},
|
||||
"tests/test_render.py": {
|
||||
"mtime": 1785120806.3260238,
|
||||
"ast_hash": "27125960e235813724fd669001faab59",
|
||||
"semantic_hash": "27125960e235813724fd669001faab59"
|
||||
},
|
||||
"tests/test_store_platform.py": {
|
||||
"mtime": 1785126662.4772756,
|
||||
"ast_hash": "bad13dcf37be7a414ffae8018c29e55b",
|
||||
"semantic_hash": "bad13dcf37be7a414ffae8018c29e55b"
|
||||
},
|
||||
".opencode/agent/webapp-builder.md": {
|
||||
"mtime": 1785126783.7152567,
|
||||
"ast_hash": "58e92641724400370d7823dce1e02ca3",
|
||||
"semantic_hash": "58e92641724400370d7823dce1e02ca3"
|
||||
},
|
||||
".opencode/goals/webapp-build.md": {
|
||||
"mtime": 1785126796.8402326,
|
||||
"ast_hash": "478e7a2350ded2fd273cfdb0714d228d",
|
||||
"semantic_hash": "478e7a2350ded2fd273cfdb0714d228d"
|
||||
},
|
||||
"OPPORTUNITIES.md": {
|
||||
"mtime": 1785122706.3466253,
|
||||
"ast_hash": "4af60e3106b394664fbf5b5c5f645804",
|
||||
"semantic_hash": "4af60e3106b394664fbf5b5c5f645804"
|
||||
},
|
||||
"README.md": {
|
||||
"mtime": 1785121499.3282604,
|
||||
"ast_hash": "a0a8e6b62a991fce5d6c58171f80075e",
|
||||
"semantic_hash": "a0a8e6b62a991fce5d6c58171f80075e"
|
||||
},
|
||||
"config.example.yaml": {
|
||||
"mtime": 1785120602.4114335,
|
||||
"ast_hash": "415b0fc8e71ecc305d21e9c3c08d585a",
|
||||
"semantic_hash": "415b0fc8e71ecc305d21e9c3c08d585a"
|
||||
},
|
||||
"docs/superpowers/plans/2026-07-26-platform-build.md": {
|
||||
"mtime": 1785126240.4211702,
|
||||
"ast_hash": "8097056ef23068f22c48151746de305b",
|
||||
"semantic_hash": "8097056ef23068f22c48151746de305b"
|
||||
},
|
||||
"docs/superpowers/specs/2026-07-26-platform-design.md": {
|
||||
"mtime": 1785126024.4879057,
|
||||
"ast_hash": "9b85f02fd45c93cdabfd2e1c6d1e0c9e",
|
||||
"semantic_hash": "9b85f02fd45c93cdabfd2e1c6d1e0c9e"
|
||||
},
|
||||
"src/yt_scraper/webapp/static/index.html": {
|
||||
"mtime": 1785131695.7653384,
|
||||
"ast_hash": "13591b7c8c662d8e7c2540f3032ae26d",
|
||||
"semantic_hash": "13591b7c8c662d8e7c2540f3032ae26d"
|
||||
}
|
||||
}
|
||||
+3
-1
@@ -8,7 +8,7 @@ version = "0.1.0"
|
||||
description = "Scraper de canales de YouTube hacia notas Markdown para Obsidian"
|
||||
requires-python = ">=3.10"
|
||||
dependencies = [
|
||||
"yt-dlp>=2024.10.7",
|
||||
"yt-dlp[default]>=2026.8.19",
|
||||
"click>=8.1",
|
||||
"rich>=13.7",
|
||||
"jinja2>=3.1",
|
||||
@@ -21,11 +21,13 @@ dependencies = [
|
||||
dev = [
|
||||
"pytest>=8.0",
|
||||
"pytest-cov>=4.1",
|
||||
"httpx>=0.27",
|
||||
]
|
||||
web = [
|
||||
"fastapi>=0.110",
|
||||
"uvicorn[standard]>=0.27",
|
||||
"sse-starlette>=2.0",
|
||||
"python-multipart>=0.0.9",
|
||||
]
|
||||
analysis = [
|
||||
"matplotlib>=3.8",
|
||||
|
||||
@@ -0,0 +1,115 @@
|
||||
#!/usr/bin/env bash
|
||||
# bootstrap.sh — prepara una maquina nueva (macOS/Linux) para yt-scraper.
|
||||
# Idempotente: seguro de ejecutar tantas veces como haga falta.
|
||||
#
|
||||
# bash scripts/bootstrap.sh (o `make setup`)
|
||||
#
|
||||
# Que hace:
|
||||
# 1. Localiza Python >= 3.10 (prefiere 3.12; pista de brew si falta).
|
||||
# 2. Crea el venv .venv/ si no existe (o lo reusa).
|
||||
# 3. Instala el paquete con los extras dev, web y analysis.
|
||||
# 4. Revisa herramientas opcionales: ffmpeg y node (con pistas de instalacion).
|
||||
# 5. Copia config.example.yaml -> config.yaml si aun no existe.
|
||||
#
|
||||
# Compatible con el bash 3.2 de macOS y con POSIX sh.
|
||||
|
||||
set -u
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)" || exit 1
|
||||
ROOT="$(cd "$SCRIPT_DIR/.." && pwd)" || exit 1
|
||||
cd "$ROOT" || exit 1
|
||||
|
||||
say() { printf '%s\n' "$*"; }
|
||||
ok() { say " [ok] $*"; }
|
||||
warn() { say " [warn] $*"; }
|
||||
err() { say " [error] $*"; }
|
||||
|
||||
OS="$(uname)"
|
||||
|
||||
say ""
|
||||
say " [yt-scraper] preparando el entorno en $ROOT ..."
|
||||
|
||||
# --- (1) Python >= 3.10 (se prefiere 3.12) ---
|
||||
PY=""
|
||||
for candidate in python3.12 python3.11 python3.10 python3 python; do
|
||||
command -v "$candidate" >/dev/null 2>&1 || continue
|
||||
ver="$("$candidate" -c 'import sys; print("%d.%d" % sys.version_info[:2])' 2>/dev/null || true)"
|
||||
case "$ver" in
|
||||
3.1[0-9] | 3.[2-9][0-9] | 4.*)
|
||||
PY="$(command -v "$candidate")"
|
||||
break
|
||||
;;
|
||||
esac
|
||||
done
|
||||
|
||||
if [ -z "$PY" ]; then
|
||||
err "no encontre Python >= 3.10 en el PATH."
|
||||
if [ "$OS" = "Darwin" ]; then
|
||||
say " instalalo con Homebrew: brew install [email protected]"
|
||||
else
|
||||
say " instalalo con el gestor de tu distro, p.ej.:"
|
||||
say " sudo apt install python3 python3-venv python3-pip"
|
||||
fi
|
||||
exit 1
|
||||
fi
|
||||
ok "Python encontrado: $PY ($("$PY" -c 'import sys; print(sys.version.split()[0])'))"
|
||||
|
||||
# --- (2) venv (crear si falta, reusar si ya existe) ---
|
||||
if [ ! -x ".venv/bin/python" ]; then
|
||||
say " [...] creando venv .venv/ ..."
|
||||
if ! "$PY" -m venv .venv; then
|
||||
err "no se pudo crear el venv (en Debian/Ubuntu falta: sudo apt install python3-venv)."
|
||||
exit 1
|
||||
fi
|
||||
else
|
||||
ok "venv .venv/ ya existe; reusandolo."
|
||||
fi
|
||||
VPY=".venv/bin/python"
|
||||
|
||||
# --- (3) dependencias (paquete + extras dev/web/analysis) ---
|
||||
say " [...] instalando dependencias: pip install -e \"[dev,web,analysis]\" ..."
|
||||
"$VPY" -m pip install --quiet --upgrade pip || true
|
||||
if ! "$VPY" -m pip install -e ".[dev,web,analysis]"; then
|
||||
err "fallo pip install; revisa el mensaje de arriba."
|
||||
exit 1
|
||||
fi
|
||||
ok "dependencias instaladas."
|
||||
|
||||
# --- (4) herramientas opcionales ---
|
||||
if command -v ffmpeg >/dev/null 2>&1; then
|
||||
ok "ffmpeg presente ($(command -v ffmpeg))."
|
||||
else
|
||||
warn "ffmpeg no encontrado (solo hace falta para descargar audio MP3)."
|
||||
if [ "$OS" = "Darwin" ]; then
|
||||
say " brew install ffmpeg"
|
||||
else
|
||||
say " sudo apt install ffmpeg # Debian/Ubuntu (dnf/pacman en otras distros)"
|
||||
fi
|
||||
fi
|
||||
|
||||
if command -v node >/dev/null 2>&1; then
|
||||
ok "node presente ($(command -v node)) — runtime JS para los retos de yt-dlp."
|
||||
else
|
||||
warn "node no encontrado (yt-dlp lo usa como runtime JS para algunos desafios)."
|
||||
if [ "$OS" = "Darwin" ]; then
|
||||
say " brew install node"
|
||||
else
|
||||
say " usa el gestor de tu distro o descargalo de https://nodejs.org"
|
||||
fi
|
||||
fi
|
||||
|
||||
# --- (5) configuracion inicial ---
|
||||
if [ -f "config.example.yaml" ] && [ ! -f "config.yaml" ]; then
|
||||
cp config.example.yaml config.yaml
|
||||
ok "config.yaml creado a partir de config.example.yaml (revisa canales y rutas)."
|
||||
elif [ -f "config.yaml" ]; then
|
||||
ok "config.yaml ya existe; no se toca."
|
||||
fi
|
||||
|
||||
# --- siguientes pasos ---
|
||||
say ""
|
||||
say " Listo. Siguientes pasos:"
|
||||
say " - arranca el servidor: bash scripts/start-server.sh (o: make serve)"
|
||||
say " - detene el servidor: bash scripts/stop-server.sh (o: make stop)"
|
||||
say " - o usa la CLI: .venv/bin/yt-scraper --help"
|
||||
say ""
|
||||
@@ -0,0 +1,173 @@
|
||||
# doctor.ps1 -- prepara el entorno para arrancar yt-scraper y delega en
|
||||
# start-server.ps1. Pensado para usuarios no tecnicos: doble clic y listo.
|
||||
#
|
||||
# Uso: powershell -File doctor.ps1 [-CheckOnly] [args para start-server]
|
||||
# -CheckOnly solo diagnostica e informa que haria; no instala ni abre nada.
|
||||
#
|
||||
# Exit codes: 0 ok | 10 sin Python instalable | 11 dependencias no reparadas
|
||||
# (cualquier otro codigo viene propagado de start-server.ps1)
|
||||
|
||||
param(
|
||||
[switch]$CheckOnly
|
||||
)
|
||||
|
||||
$ErrorActionPreference = 'Stop'
|
||||
$root = Split-Path -Parent $PSScriptRoot
|
||||
Set-Location -LiteralPath $root
|
||||
try { [Console]::OutputEncoding = [System.Text.Encoding]::UTF8 } catch {}
|
||||
|
||||
function Out-Line($msg, $color = $null) {
|
||||
if ($color) { Write-Host $msg -ForegroundColor $color }
|
||||
else { Write-Host $msg }
|
||||
[Console]::Out.Flush()
|
||||
}
|
||||
|
||||
function Test-PortOpen($port, $timeoutMs = 350) {
|
||||
$client = New-Object System.Net.Sockets.TcpClient
|
||||
try {
|
||||
$iar = $client.BeginConnect('127.0.0.1', [int]$port, $null, $null)
|
||||
if (-not $iar.AsyncWaitHandle.WaitOne($timeoutMs)) { return $false }
|
||||
$client.EndConnect($iar)
|
||||
return $true
|
||||
} catch { return $false }
|
||||
finally { $client.Close() }
|
||||
}
|
||||
|
||||
# Devuelve la ruta CONCRETA de python.exe si el candidato es un Python
|
||||
# >= 3.10 REAL; $null si no existe, es el stub de la Microsoft Store o es
|
||||
# demasiado viejo. Resuelve via sys.executable porque el launcher 'py' puede
|
||||
# cambiar su default al instalarse otro Python.
|
||||
function Resolve-PythonCandidate($tokens) {
|
||||
$name = $tokens[0]
|
||||
$cmd = Get-Command $name -ErrorAction SilentlyContinue
|
||||
if (-not $cmd) { return $null }
|
||||
# Stub de la Microsoft Store: vive en WindowsApps y no ejecuta nada.
|
||||
if ($cmd.Source -and $cmd.Source -like '*WindowsApps*') { return $null }
|
||||
try {
|
||||
$rest = @($tokens | Select-Object -Skip 1)
|
||||
$out = & $cmd.Source @($rest + @('-c', 'import sys; print(sys.executable)')) 2>&1
|
||||
if ($LASTEXITCODE -ne 0) { return $null }
|
||||
$exePath = (($out | Out-String).Trim() -split "`r?`n")[-1].Trim()
|
||||
if (-not ($exePath -like '*python*')) { return $null }
|
||||
$vout = & $exePath --version 2>&1
|
||||
if ($LASTEXITCODE -ne 0) { return $null }
|
||||
$s = ($vout | Out-String).Trim()
|
||||
if ($s -notmatch '^Python\s+(\d+)\.(\d+)') { return $null }
|
||||
if ([int]$Matches[1] -lt 3 -or ([int]$Matches[1] -eq 3 -and [int]$Matches[2] -lt 10)) { return $null }
|
||||
return $exePath
|
||||
} catch { return $null }
|
||||
}
|
||||
|
||||
# True si ese interprete ya tiene las dependencias del proyecto instaladas.
|
||||
function Test-Deps($exe) {
|
||||
try {
|
||||
& $exe -c 'import fastapi, uvicorn, sse_starlette, jinja2, yaml, yt_dlp, requests, slugify' *> $null
|
||||
return ($LASTEXITCODE -eq 0)
|
||||
} catch { return $false }
|
||||
}
|
||||
|
||||
Out-Line ""
|
||||
Out-Line " [yt-scraper] verificando tu sistema..." Cyan
|
||||
|
||||
# --- (1) Server ya vivo? Abrir navegador y terminar. -----------------------
|
||||
if (Test-Path '.run\server.info') {
|
||||
$parts = ((Get-Content '.run\server.info' -TotalCount 1) -split '\s+')
|
||||
if ($parts.Count -ge 1 -and $parts[0] -match '^\d+$' -and (Test-PortOpen ([int]$parts[0]) 350)) {
|
||||
Out-Line " [ok] El servidor ya estaba corriendo - abriendo..." Green
|
||||
Out-Line " http://127.0.0.1:$($parts[0])"
|
||||
if (-not $CheckOnly) { try { Start-Process "http://127.0.0.1:$($parts[0])" } catch {} }
|
||||
exit 0
|
||||
}
|
||||
}
|
||||
|
||||
# --- (2) Resolver Python ----------------------------------------------------
|
||||
# Maquinas con varios Pythons son comunes: se prefiere el primero que ademas
|
||||
# pasa el import-test de dependencias; si ninguno lo pasa, vale el primero
|
||||
# con version valida (la rama (3) instalara las deps ahi).
|
||||
$candidates = @(,@('py','-3')) + @(,@('python')) + @(,@('python3'))
|
||||
$pythonExe = $null
|
||||
foreach ($cand in $candidates) {
|
||||
$exe = Resolve-PythonCandidate $cand
|
||||
if (-not $exe) { continue }
|
||||
if (-not $pythonExe) { $pythonExe = $exe }
|
||||
if (Test-Deps $exe) { $pythonExe = $exe; break }
|
||||
}
|
||||
|
||||
if (-not $pythonExe) {
|
||||
Out-Line " [info] Python no esta instalado. Instalandolo automaticamente (~25 MB)..." Yellow
|
||||
if ($CheckOnly) {
|
||||
Out-Line " [check] AQUI: winget install Python.Python.3.12 + busqueda de ruta nueva" DarkGray
|
||||
exit 10
|
||||
}
|
||||
$winget = Get-Command winget -ErrorAction SilentlyContinue
|
||||
if (-not $winget) {
|
||||
Out-Line " [!] No pude instalar Python automaticamente (falta winget)." Red
|
||||
Out-Line " Abriendo la pagina de descarga de Python..." Gray
|
||||
Out-Line " Instalalo (marca 'Add python.exe to PATH') y vuelve a hacer doble clic." Gray
|
||||
try { Start-Process 'https://www.python.org/downloads/' } catch {}
|
||||
exit 10
|
||||
}
|
||||
& winget install --id Python.Python.3.12 --silent --accept-package-agreements --accept-source-agreements | Out-Null
|
||||
if ($LASTEXITCODE -ne 0) {
|
||||
Out-Line " [!] La instalacion de Python fallo (codigo $LASTEXITCODE). Reinstala manualmente:" Red
|
||||
Out-Line " https://www.python.org/downloads/" Gray
|
||||
exit 10
|
||||
}
|
||||
# El PATH de esta sesion no se refresca: sondear rutas conocidas.
|
||||
$fresh = @()
|
||||
$fresh += Get-ChildItem "$env:LOCALAPPDATA\Programs\Python\Python3*\python.exe" -ErrorAction SilentlyContinue
|
||||
$fresh += Get-ChildItem 'C:\Program Files\Python3*\python.exe' -ErrorAction SilentlyContinue
|
||||
foreach ($f in ($fresh | Sort-Object FullName -Descending)) {
|
||||
$pythonExe = Resolve-PythonCandidate @($f.FullName)
|
||||
if ($pythonExe) { break }
|
||||
}
|
||||
if (-not $pythonExe) { $pythonExe = Resolve-PythonCandidate @('py','-3') }
|
||||
if (-not $pythonExe) {
|
||||
Out-Line " [!] Se instalo Python pero no lo encuentro. Cierra esta ventana," Red
|
||||
Out-Line " abre una nueva e intenta de nuevo (el PATH se refresca al reabrir)." Gray
|
||||
exit 10
|
||||
}
|
||||
}
|
||||
Out-Line " [ok] Python encontrado: $pythonExe" DarkGray
|
||||
|
||||
# --- (3) Dependencias --------------------------------------------------------
|
||||
$depsOk = Test-Deps $pythonExe
|
||||
|
||||
if (-not $depsOk) {
|
||||
if ($CheckOnly) {
|
||||
Out-Line " [check] AQUI: instalar dependencias -> & '$pythonExe' -m pip install -e `".[web]`"" DarkGray
|
||||
exit 11
|
||||
}
|
||||
Out-Line " [info] Instalando las piezas que faltan por primera vez (puede tardar 1-2 min)..." Yellow
|
||||
try {
|
||||
& $pythonExe -m pip --version *> $null
|
||||
if ($LASTEXITCODE -ne 0) { & $pythonExe -m ensurepip --upgrade | Out-Null }
|
||||
} catch {
|
||||
& $pythonExe -m ensurepip --upgrade | Out-Null
|
||||
}
|
||||
& $pythonExe -m pip install -e ".[web]"
|
||||
if ($LASTEXITCODE -ne 0) {
|
||||
Out-Line " [!] No pude instalar las dependencias (revisa tu conexion a internet)" Red
|
||||
Out-Line " y vuelve a hacer doble clic en start-server.bat" Gray
|
||||
exit 11
|
||||
}
|
||||
}
|
||||
if ($CheckOnly) { Out-Line " [check] dependencias: OK" DarkGray }
|
||||
|
||||
# --- (4) ffmpeg (opcional: solo audio) --------------------------------------
|
||||
if (-not (Get-Command ffmpeg -ErrorAction SilentlyContinue)) {
|
||||
Out-Line " [aviso] La descarga de AUDIO no estara disponible (falta ffmpeg)." Yellow
|
||||
Out-Line " Todo lo demas funciona perfecto. Puedes ignorarlo." Gray
|
||||
}
|
||||
|
||||
# --- (5) Delegar al launcher -------------------------------------------------
|
||||
if ($CheckOnly) {
|
||||
Out-Line " [check] AQUI: lanzaria scripts\start-server.ps1 -PythonExe '$pythonExe'" DarkGray
|
||||
Out-Line ""
|
||||
Out-Line " Diagnostico completo. Todo listo para arrancar." Green
|
||||
exit 0
|
||||
}
|
||||
|
||||
Out-Line ""
|
||||
& (Join-Path $PSScriptRoot 'start-server.ps1') -PythonExe $pythonExe @args
|
||||
exit $LASTEXITCODE
|
||||
+176
-57
@@ -1,45 +1,117 @@
|
||||
param(
|
||||
# Ruta absoluta o nombre del interprete Python que lanza uvicorn.
|
||||
# doctor.ps1 la resuelve porque tras una instalacion fresca el PATH
|
||||
# de esta sesion aun no ve el Python nuevo.
|
||||
[string]$PythonExe = 'python'
|
||||
)
|
||||
|
||||
$ErrorActionPreference = 'Stop'
|
||||
$root = Split-Path -Parent $PSScriptRoot
|
||||
Set-Location -LiteralPath $root
|
||||
# Force UTF-8 so accented messages render correctly even when launched
|
||||
# from the .bat via `powershell -File` (PS5.1 uses the OEM codepage by default).
|
||||
try { [Console]::OutputEncoding = [System.Text.Encoding]::UTF8 } catch {}
|
||||
|
||||
# Robust browser opener — tries several Windows methods until one works.
|
||||
# All console output is force-flushed so the user sees live progress even when
|
||||
# the launcher is invoked from a .bat-launched powershell.exe window.
|
||||
function Out-Line($msg, $color = $null) {
|
||||
if ($color) { Write-Host $msg -ForegroundColor $color }
|
||||
else { Write-Host $msg }
|
||||
[Console]::Out.Flush()
|
||||
}
|
||||
|
||||
# Robust browser opener -- tries several Windows methods until one works.
|
||||
function Open-Browser($url) {
|
||||
$opened = $false
|
||||
# Method 1: Start-Process with the URL (default protocol handler)
|
||||
try { Start-Process -FilePath $url; $opened = $true } catch {}
|
||||
if (-not $opened) {
|
||||
# Method 2: explorer.exe with the URL (opens default browser)
|
||||
try { & explorer.exe $url; $opened = $true } catch {}
|
||||
}
|
||||
if (-not $opened) {
|
||||
# Method 3: cmd `start` builtin
|
||||
try { & cmd.exe /c start "" $url; $opened = $true } catch {}
|
||||
}
|
||||
if ($opened) { Write-Host " [ok] abriendo navegador -> $url" -ForegroundColor DarkGray }
|
||||
else { Write-Host " [warn] no se pudo abrir el navegador. Abre manualmente: $url" -ForegroundColor Yellow }
|
||||
if (-not $opened) { try { & explorer.exe $url; $opened = $true } catch {} }
|
||||
if (-not $opened) { try { & cmd.exe /c start "" $url; $opened = $true } catch {} }
|
||||
if ($opened) { Out-Line " [ok] navegador abierto -> $url" DarkGray }
|
||||
else { Out-Line " [warn] no se pudo abrir el navegador automaticamente" Yellow }
|
||||
return $opened
|
||||
}
|
||||
|
||||
Write-Host ""
|
||||
Write-Host " [yt-scraper] iniciando servidor local..." -ForegroundColor Cyan
|
||||
# Fast TCP-connect probe. Replaces Invoke-WebRequest (which on Windows
|
||||
# PowerShell 5.1 ignores -TimeoutSec for the connect phase and can hang
|
||||
# 3-15 s per failed attempt to a closed port -- making a typical port
|
||||
# scan silently take many minutes).
|
||||
function Test-PortOpen($port, $timeoutMs = 350) {
|
||||
$client = New-Object System.Net.Sockets.TcpClient
|
||||
try {
|
||||
$iar = $client.BeginConnect('127.0.0.1', $port, $null, $null)
|
||||
if (-not $iar.AsyncWaitHandle.WaitOne($timeoutMs)) { return $false }
|
||||
$client.EndConnect($iar)
|
||||
return $true
|
||||
} catch { return $false }
|
||||
finally { $client.Close() }
|
||||
}
|
||||
|
||||
function Get-ServerHealth($port, $timeoutMs = 800) {
|
||||
try {
|
||||
$r = Invoke-WebRequest -Uri "http://127.0.0.1:$port/healthz" -UseBasicParsing -TimeoutSec ([Math]::Max(1, [int]($timeoutMs/1000)))
|
||||
if ($r.Content -match 'ok') { return $true }
|
||||
} catch {}
|
||||
return $false
|
||||
}
|
||||
|
||||
function Is-PidAlive($processId) {
|
||||
if (-not $processId) { return $false }
|
||||
$p = Get-Process -Id ([int]$processId) -ErrorAction SilentlyContinue
|
||||
return $null -ne $p
|
||||
}
|
||||
|
||||
function Show-State($label, $obj) {
|
||||
Out-Line (" [{0}] puerto={1} pid={2} t={3:o}" -f $label, $obj.port, $obj.pid, (Get-Date))
|
||||
}
|
||||
|
||||
Out-Line ""
|
||||
Out-Line " [yt-scraper] iniciando servidor local..." Cyan
|
||||
|
||||
if (-not (Test-Path '.run')) { New-Item -ItemType Directory -Path '.run' | Out-Null }
|
||||
|
||||
# Already running?
|
||||
# --- (1) Check for an already-running server registered in server.info ---
|
||||
# A bound port is ground truth that a server is live. The cmd-wrapper PID
|
||||
# we recorded is a weak signal (cmd /c exits as soon as python takes over),
|
||||
# so we DON'T require the PID to be alive -- only that the port responds.
|
||||
$registeredPort = $null
|
||||
$registeredPid = $null
|
||||
$liveExisting = $false
|
||||
|
||||
if (Test-Path '.run\server.info') {
|
||||
$firstLine = (Get-Content '.run\server.info' -TotalCount 1)
|
||||
$existingPort = ($firstLine -split '\s+')[0]
|
||||
try {
|
||||
$r = Invoke-WebRequest -Uri "http://localhost:$existingPort/healthz" -UseBasicParsing -TimeoutSec 2
|
||||
if ($r.Content -match 'ok') {
|
||||
Write-Host " [ok] servidor ya activo en el puerto $existingPort" -ForegroundColor Green
|
||||
Open-Browser "http://127.0.0.1:$existingPort"
|
||||
exit 0
|
||||
$parts = $firstLine -split '\s+'
|
||||
if ($parts.Count -ge 2) {
|
||||
$registeredPort = $parts[0]
|
||||
$registeredPid = $parts[1]
|
||||
$portBusy = Test-PortOpen ([int]$registeredPort) 350
|
||||
$pidAlive = Is-PidAlive $registeredPid
|
||||
if ($portBusy) {
|
||||
# Port is bound and answering -- this IS the running server.
|
||||
# Even if the recorded cmd-wrapper PID died, python is alive.
|
||||
$liveExisting = $true
|
||||
} else {
|
||||
Out-Line " [info] Habia un registro viejo de otra sesion; limpiandolo..." Yellow
|
||||
Remove-Item '.run\server.info' -Force -ErrorAction SilentlyContinue
|
||||
if (-not $pidAlive) {
|
||||
# also nuke any stray python on that port, just in case
|
||||
Get-NetTCPConnection -State Listen -LocalPort ([int]$registeredPort) -ErrorAction SilentlyContinue |
|
||||
ForEach-Object { try { Stop-Process -Id $_.OwningProcess -Force -ErrorAction SilentlyContinue } catch {} }
|
||||
}
|
||||
}
|
||||
}
|
||||
} catch {}
|
||||
}
|
||||
|
||||
# Collect candidate free ports scanning from 8000.
|
||||
if ($liveExisting) {
|
||||
Out-Line " [ok] servidor ya activo en puerto $registeredPort" Green
|
||||
Out-Line " URL: http://localhost:$registeredPort"
|
||||
Out-Line " log: .run\server.log"
|
||||
Out-Line " stop: stop-server.bat"
|
||||
Open-Browser "http://127.0.0.1:$registeredPort" | Out-Null
|
||||
Out-Line ""
|
||||
exit 0
|
||||
}
|
||||
|
||||
# --- (2) Find free ports (fast TcpListener probe, like the original) ---
|
||||
$candidates = @()
|
||||
foreach ($p in 8000..8100) {
|
||||
try {
|
||||
@@ -50,53 +122,100 @@ foreach ($p in 8000..8100) {
|
||||
} catch {}
|
||||
}
|
||||
if ($candidates.Count -eq 0) {
|
||||
Write-Host " [error] no se encontro ningun puerto libre (8000-8100)" -ForegroundColor Red
|
||||
Out-Line " [error] no se encontro ningun puerto libre (8000-8100)" Red
|
||||
exit 2
|
||||
}
|
||||
|
||||
# Launch uvicorn DETACHED in its own persistent minimized console window, logging to .run/server.log.
|
||||
# The `cmd /c` window is an independent console: it survives this launcher exiting.
|
||||
# --- (3) Launch uvicorn on each candidate until one sticks ---
|
||||
$launched = $false
|
||||
$proc = $null
|
||||
$port = 0
|
||||
$proc = $null
|
||||
|
||||
foreach ($candidate in $candidates) {
|
||||
# clear previous log
|
||||
Remove-Item '.run\server.log' -Force -ErrorAction SilentlyContinue
|
||||
$cmdLine = '/c python -m uvicorn yt_scraper.webapp.app:app --host 127.0.0.1 --port ' + $candidate + ' --log-level info > .run\server.log 2>&1'
|
||||
$proc = Start-Process -FilePath 'cmd.exe' -ArgumentList $cmdLine -WorkingDirectory $root -WindowStyle Minimized -PassThru
|
||||
Set-Content -Path '.run\server.info' -Value "$candidate $($proc.Id) $(Get-Date -Format o)" -Encoding ASCII
|
||||
|
||||
$ok = $false
|
||||
for ($i = 0; $i -lt 40; $i++) {
|
||||
Start-Sleep -Milliseconds 500
|
||||
Remove-Item '.run\server.err.log' -Force -ErrorAction SilentlyContinue
|
||||
Out-Line " [...] Encendiendo el servidor (intento $candidate)..." Gray
|
||||
# Spawn directo de python: un wrapper cmd /c con la ruta entre comillas
|
||||
# dispara el quote-stripping de cmd cuando la ruta tiene espacios y el
|
||||
# comando muere en silencio. Start-Process con array de args no sufre eso,
|
||||
# y ademas server.info registra el PID real de python (no el del wrapper).
|
||||
try {
|
||||
$r = Invoke-WebRequest -Uri "http://localhost:$candidate/healthz" -UseBasicParsing -TimeoutSec 2
|
||||
if ($r.Content -match 'ok') { $ok = $true; break }
|
||||
} catch {}
|
||||
$proc = Start-Process -FilePath $PythonExe `
|
||||
-ArgumentList @('-m','uvicorn','yt_scraper.webapp.app:app','--host','127.0.0.1','--port',"$candidate",'--log-level','info') `
|
||||
-WorkingDirectory $root -WindowStyle Minimized -PassThru `
|
||||
-RedirectStandardOutput "$root\.run\server.log" `
|
||||
-RedirectStandardError "$root\.run\server.err.log"
|
||||
} catch {
|
||||
Out-Line " [warn] Start-Process fallo en $candidate : $($_.Exception.Message)" Yellow
|
||||
continue
|
||||
}
|
||||
if ($ok) { $port = $candidate; $launched = $true; break }
|
||||
Out-Line " El servidor corre en segundo plano. Registro: .run\server.log"
|
||||
|
||||
# failed on this port — show why from the log, then try the next candidate
|
||||
Stop-Process -Id $proc.Id -Force -ErrorAction SilentlyContinue
|
||||
Remove-Item '.run\server.info' -Force -ErrorAction SilentlyContinue
|
||||
# Fast TCP-poll for the bound port. Bound == uvicorn listening, which
|
||||
# is all the user needs to know "open this URL in a browser".
|
||||
$boundAt = $null
|
||||
$bound = $false
|
||||
$sw = [System.Diagnostics.Stopwatch]::StartNew()
|
||||
for ($i = 0; $i -lt 50; $i++) {
|
||||
Start-Sleep -Milliseconds 200
|
||||
if (Test-PortOpen $candidate 250) { $bound = $true; $boundAt = $sw.Elapsed; break }
|
||||
# show a tiny heartbeat every ~5 attempts
|
||||
if ($i -in 5,15,30,45) {
|
||||
Out-Line (" ... esperando puerto {0} ({1:0.0}s)" -f $candidate, $sw.Elapsed.TotalSeconds) DarkGray
|
||||
}
|
||||
}
|
||||
|
||||
if (-not $bound) {
|
||||
Out-Line " [warn] Este intento no respondio; probando el siguiente..." Yellow
|
||||
# Show the actual reason from the logs so the user can diagnose.
|
||||
foreach ($logName in @('.run\server.err.log', '.run\server.log')) {
|
||||
if (Test-Path $logName) {
|
||||
$tail = Get-Content $logName -Tail 6 -ErrorAction SilentlyContinue
|
||||
if ($tail) {
|
||||
Out-Line " (ultimas lineas de $logName):" DarkGray
|
||||
$tail | ForEach-Object { Out-Line " $_" DarkGray }
|
||||
}
|
||||
}
|
||||
}
|
||||
try { Stop-Process -Id $proc.Id -Force -ErrorAction SilentlyContinue } catch {}
|
||||
try { Get-NetTCPConnection -State Listen -LocalPort $candidate -ErrorAction SilentlyContinue | ForEach-Object { Stop-Process -Id $_.OwningProcess -Force -ErrorAction SilentlyContinue } } catch {}
|
||||
continue
|
||||
}
|
||||
|
||||
# Port is bound. Reveal the URL RIGHT NOW so the user sees progress.
|
||||
$port = $candidate
|
||||
Set-Content -Path '.run\server.info' -Value "$port $($proc.Id) $(Get-Date -Format o)" -Encoding ASCII
|
||||
|
||||
Out-Line " [ok] Servidor encendido." Green
|
||||
Out-Line " URL: http://localhost:$port"
|
||||
Out-Line " log: .run\server.log"
|
||||
Out-Line " stop: stop-server.bat"
|
||||
Out-Line " Confirmando que todo cargo bien..." DarkGray
|
||||
|
||||
Open-Browser "http://127.0.0.1:$port" | Out-Null
|
||||
|
||||
# Confirm /healthz (separate from TCP-bound so the URL appears first).
|
||||
$healthAt = $null
|
||||
for ($i = 0; $i -lt 25; $i++) {
|
||||
Start-Sleep -Milliseconds 200
|
||||
if (Get-ServerHealth $port 800) { $healthAt = $sw.Elapsed; break }
|
||||
}
|
||||
if ($healthAt) {
|
||||
Out-Line " [ok] /healthz OK (t=$([int]$healthAt.TotalSeconds)s)" Green
|
||||
} else {
|
||||
Out-Line " [warn] El servidor abrio pero tardo en responder; revisa .run\server.log" Yellow
|
||||
}
|
||||
Out-Line ""
|
||||
$launched = $true
|
||||
break
|
||||
}
|
||||
|
||||
if (-not $launched) {
|
||||
Write-Host " [error] no se pudo iniciar el servidor." -ForegroundColor Red
|
||||
Out-Line " [error] no se pudo iniciar el servidor en ninguno de los puertos candidatos." Red
|
||||
if (Test-Path '.run\server.log') {
|
||||
Write-Host " --- ultimo log del servidor ---" -ForegroundColor Yellow
|
||||
Get-Content '.run\server.log' -Tail 25 | ForEach-Object { Write-Host " $_" }
|
||||
Out-Line " --- ultimo log del servidor ---" Yellow
|
||||
Get-Content '.run\server.log' -Tail 25 | ForEach-Object { Out-Line " $_" }
|
||||
} else {
|
||||
Write-Host " instala dependencias: pip install -e `".[web]`"" -ForegroundColor Yellow
|
||||
Out-Line " instala dependencias: pip install -e `".[web]`"" Yellow
|
||||
}
|
||||
exit 3
|
||||
}
|
||||
|
||||
Open-Browser "http://127.0.0.1:$port"
|
||||
Write-Host ""
|
||||
Write-Host " [ok] servidor activo" -ForegroundColor Green
|
||||
Write-Host " URL: http://localhost:$port"
|
||||
Write-Host " PID: $($proc.Id) (ventana minimizada 'cmd' independiente)"
|
||||
Write-Host " log: .run\server.log"
|
||||
Write-Host " stop: stop-server.bat"
|
||||
Write-Host ""
|
||||
|
||||
@@ -0,0 +1,182 @@
|
||||
#!/usr/bin/env bash
|
||||
# start-server.sh — arranque del servidor web de yt-scraper en macOS/Linux.
|
||||
# Espejo de scripts/start-server.ps1 (Windows); mismo formato de .run/server.info
|
||||
# ("PUERTO PID FECHA"), asi ambos conviven sobre el mismo proyecto.
|
||||
#
|
||||
# Uso:
|
||||
# bash scripts/start-server.sh (o `make serve`)
|
||||
# PYTHON_EXE=/ruta/python3 bash scripts/start-server.sh # igual que -PythonExe en el .ps1
|
||||
#
|
||||
# Compatible con el bash 3.2 de macOS y con POSIX sh: sin arrays asociativos,
|
||||
# sin ${var,,}, sin mapfile.
|
||||
|
||||
set -u
|
||||
|
||||
# --- raiz del proyecto (carpeta padre de scripts/) ---
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)" || exit 1
|
||||
ROOT="$(cd "$SCRIPT_DIR/.." && pwd)" || exit 1
|
||||
cd "$ROOT" || exit 1
|
||||
|
||||
RUN_DIR=".run"
|
||||
INFO_FILE="$RUN_DIR/server.info"
|
||||
LOG_FILE="$RUN_DIR/server.log"
|
||||
ERR_LOG_FILE="$RUN_DIR/server.err.log"
|
||||
HOST="127.0.0.1"
|
||||
PORT_MIN=8000
|
||||
PORT_MAX=8100
|
||||
READY_TIMEOUT=30 # segundos maximos esperando /healthz
|
||||
|
||||
say() { printf '%s\n' "$*"; }
|
||||
|
||||
# --- interprete Python: venv del proyecto si existe, si no python3 ---
|
||||
if [ -n "${PYTHON_EXE:-}" ]; then
|
||||
PY="$PYTHON_EXE"
|
||||
elif [ -x ".venv/bin/python" ]; then
|
||||
PY=".venv/bin/python"
|
||||
else
|
||||
PY="$(command -v python3 || command -v python || true)"
|
||||
if [ -z "$PY" ]; then
|
||||
say " [error] no encontre python3. Instalalo con: brew install [email protected]"
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
|
||||
# --- GET a /healthz; exit 0 si responde 'ok' (curl, con fallback a python) ---
|
||||
healthz_ok() {
|
||||
# $1 = puerto
|
||||
if command -v curl >/dev/null 2>&1; then
|
||||
curl -fsS --max-time 2 "http://$HOST:$1/healthz" 2>/dev/null | grep -q ok
|
||||
else
|
||||
"$PY" -c "import urllib.request,sys; sys.exit(0 if b'ok' in urllib.request.urlopen('http://$HOST:$1/healthz', timeout=2).read() else 1)" >/dev/null 2>&1
|
||||
fi
|
||||
}
|
||||
|
||||
# --- abre el navegador segun el SO; si no puede, imprime la URL ---
|
||||
open_browser() {
|
||||
case "$(uname)" in
|
||||
Darwin)
|
||||
if command -v open >/dev/null 2>&1; then
|
||||
open "$1" && say " [ok] navegador abierto -> $1" && return 0
|
||||
fi
|
||||
;;
|
||||
*)
|
||||
if command -v xdg-open >/dev/null 2>&1; then
|
||||
xdg-open "$1" && say " [ok] navegador abierto -> $1" && return 0
|
||||
fi
|
||||
;;
|
||||
esac
|
||||
say " [info] abre esta URL en tu navegador: $1"
|
||||
}
|
||||
|
||||
mkdir -p "$RUN_DIR"
|
||||
|
||||
say ""
|
||||
say " [yt-scraper] iniciando servidor local..."
|
||||
|
||||
# --- (1) servidor ya registrado y vivo? -> reutilizarlo ---
|
||||
# El puerto respondiendo es la verdad: si /healthz contesta, ese ES el
|
||||
# servidor corriendo, aunque el PID registrado haya muerto (el proceso
|
||||
# padre puede terminar y dejar el hijo vivo).
|
||||
if [ -f "$INFO_FILE" ]; then
|
||||
reg_port=""
|
||||
first_line="$(head -n 1 "$INFO_FILE" 2>/dev/null || true)"
|
||||
# primera linea: "PUERTO PID FECHA"
|
||||
set -- $first_line
|
||||
reg_port="${1:-}"
|
||||
if [ -n "$reg_port" ] && healthz_ok "$reg_port"; then
|
||||
say " [ok] servidor ya activo en puerto $reg_port"
|
||||
say " URL: http://localhost:$reg_port"
|
||||
say " log: $LOG_FILE"
|
||||
say " stop: bash scripts/stop-server.sh"
|
||||
open_browser "http://$HOST:$reg_port"
|
||||
say ""
|
||||
exit 0
|
||||
fi
|
||||
say " [info] Habia un registro viejo de otra sesion; limpiandolo..."
|
||||
rm -f "$INFO_FILE"
|
||||
fi
|
||||
|
||||
# --- (2) puertos libres en 8000-8100 (prueba de bind con Python) ---
|
||||
FREE_PORTS="$("$PY" - <<'PYEOF' 2>/dev/null
|
||||
import socket
|
||||
free = []
|
||||
for port in range(8000, 8101):
|
||||
s = socket.socket()
|
||||
try:
|
||||
s.bind(("127.0.0.1", port))
|
||||
except OSError:
|
||||
s.close()
|
||||
continue
|
||||
s.close()
|
||||
free.append(str(port))
|
||||
if len(free) >= 8:
|
||||
break
|
||||
print(" ".join(free))
|
||||
PYEOF
|
||||
)"
|
||||
if [ -z "$FREE_PORTS" ]; then
|
||||
say " [error] no se encontro ningun puerto libre ($PORT_MIN-$PORT_MAX)"
|
||||
exit 2
|
||||
fi
|
||||
|
||||
# --- (3) lanzar uvicorn en cada candidato hasta que uno responda ---
|
||||
launched=0
|
||||
for PORT in $FREE_PORTS; do
|
||||
rm -f "$LOG_FILE" "$ERR_LOG_FILE"
|
||||
say " [...] Encendiendo el servidor (intento $PORT)..."
|
||||
nohup "$PY" -m uvicorn yt_scraper.webapp.app:app \
|
||||
--host "$HOST" --port "$PORT" --log-level info \
|
||||
>>"$LOG_FILE" 2>>"$ERR_LOG_FILE" &
|
||||
SERVER_PID=$!
|
||||
say " El servidor corre en segundo plano. Registro: $LOG_FILE"
|
||||
|
||||
# Esperar /healthz (max ~30 s); si el proceso muere, probar otro puerto.
|
||||
waited=0
|
||||
ready=0
|
||||
while [ "$waited" -lt "$READY_TIMEOUT" ]; do
|
||||
if ! kill -0 "$SERVER_PID" 2>/dev/null; then
|
||||
say " [warn] el proceso murio al arrancar; probando el siguiente puerto..."
|
||||
break
|
||||
fi
|
||||
if healthz_ok "$PORT"; then
|
||||
ready=1
|
||||
break
|
||||
fi
|
||||
sleep 1
|
||||
waited=$((waited + 1))
|
||||
done
|
||||
|
||||
if [ "$ready" -ne 1 ]; then
|
||||
# Mostrar el motivo probable del fallo para poder diagnosticarlo.
|
||||
if [ -f "$ERR_LOG_FILE" ]; then
|
||||
say " (ultimas lineas de $ERR_LOG_FILE):"
|
||||
tail -n 6 "$ERR_LOG_FILE" 2>/dev/null | sed 's/^/ /'
|
||||
fi
|
||||
kill "$SERVER_PID" 2>/dev/null
|
||||
continue
|
||||
fi
|
||||
|
||||
printf '%s %s %s\n' "$PORT" "$SERVER_PID" "$(date '+%Y-%m-%dT%H:%M:%S%z')" > "$INFO_FILE"
|
||||
say " [ok] Servidor encendido."
|
||||
say " URL: http://localhost:$PORT"
|
||||
say " log: $LOG_FILE"
|
||||
say " stop: bash scripts/stop-server.sh"
|
||||
open_browser "http://$HOST:$PORT"
|
||||
launched=1
|
||||
break
|
||||
done
|
||||
|
||||
if [ "$launched" -ne 1 ]; then
|
||||
say " [error] no se pudo iniciar el servidor en ninguno de los puertos candidatos."
|
||||
if [ -f "$LOG_FILE" ]; then
|
||||
say " --- ultimo log del servidor ---"
|
||||
tail -n 25 "$LOG_FILE" 2>/dev/null | sed 's/^/ /'
|
||||
else
|
||||
say ' instala dependencias: pip install -e ".[web]"'
|
||||
fi
|
||||
say ""
|
||||
exit 3
|
||||
fi
|
||||
|
||||
say ""
|
||||
exit 0
|
||||
+58
-21
@@ -1,11 +1,43 @@
|
||||
$ErrorActionPreference = 'Stop'
|
||||
$ErrorActionPreference = 'Continue'
|
||||
$root = Split-Path -Parent $PSScriptRoot
|
||||
Set-Location -LiteralPath $root
|
||||
|
||||
function Out-Line($msg, $color = $null) {
|
||||
if ($color) { Write-Host $msg -ForegroundColor $color }
|
||||
else { Write-Host $msg }
|
||||
[Console]::Out.Flush()
|
||||
}
|
||||
|
||||
Out-Line ""
|
||||
|
||||
function Test-PortOpen($port, $timeoutMs = 350) {
|
||||
$client = New-Object System.Net.Sockets.TcpClient
|
||||
try {
|
||||
$iar = $client.BeginConnect('127.0.0.1', $port, $null, $null)
|
||||
if (-not $iar.AsyncWaitHandle.WaitOne($timeoutMs)) { return $false }
|
||||
$client.EndConnect($iar)
|
||||
return $true
|
||||
} catch { return $false }
|
||||
finally { $client.Close() }
|
||||
}
|
||||
|
||||
function Kill-PortOwner($port) {
|
||||
Get-NetTCPConnection -State Listen -LocalPort ([int]$port) -ErrorAction SilentlyContinue |
|
||||
ForEach-Object {
|
||||
Out-Line " [kill] Cerrando un proceso que quedaba suelto..." Gray
|
||||
try { Stop-Process -Id $_.OwningProcess -Force -ErrorAction SilentlyContinue } catch {}
|
||||
}
|
||||
}
|
||||
|
||||
if (-not (Test-Path '.run\server.info')) {
|
||||
Write-Host ""
|
||||
Write-Host " [yt-scraper] no hay servidor registrado." -ForegroundColor Yellow
|
||||
Write-Host ""
|
||||
# Nothing registered, but maybe a stray uvicorn is listening. Sweep 8000-8100.
|
||||
Out-Line " [yt-scraper] Buscando servidores que quedaron sueltos..." Cyan
|
||||
$killed = 0
|
||||
foreach ($p in 8000..8100) {
|
||||
if (Test-PortOpen $p 200) { Kill-PortOwner $p; $killed++ }
|
||||
}
|
||||
if ($killed -eq 0) { Out-Line " [ok] nada que detener." Green } else { Out-Line " [ok] servidores sueltos terminados." Green }
|
||||
Out-Line ""
|
||||
exit 0
|
||||
}
|
||||
|
||||
@@ -14,33 +46,38 @@ $parts = $firstLine -split '\s+'
|
||||
$port = $parts[0]
|
||||
$pidv = $parts[1]
|
||||
|
||||
Write-Host ""
|
||||
Write-Host " [yt-scraper] deteniendo servidor (PID $pidv, puerto $port)..." -ForegroundColor Cyan
|
||||
Out-Line " [yt-scraper] Apagando el servidor..." Cyan
|
||||
|
||||
if ($pidv -and $pidv -ne 'pending') {
|
||||
# Kill the process tree (uvicorn + any workers).
|
||||
# Detect stale info: PID gone -> treat as "already stopped" but also kill anything on the port.
|
||||
$pidAlive = $false
|
||||
if ($pidv) {
|
||||
$proc = Get-Process -Id ([int]$pidv) -ErrorAction SilentlyContinue
|
||||
$pidAlive = $null -ne $proc
|
||||
}
|
||||
|
||||
if (-not $pidAlive) {
|
||||
Out-Line " [info] El servidor ya estaba apagado." Yellow
|
||||
Kill-PortOwner $port
|
||||
} else {
|
||||
try {
|
||||
Stop-Process -Id ([int]$pidv) -Force -ErrorAction Stop
|
||||
} catch {
|
||||
# fallback: taskkill
|
||||
& taskkill /PID $pidv /T /F 2>$null | Out-Null
|
||||
try { & taskkill /PID $pidv /T /F 2>$null | Out-Null } catch {}
|
||||
}
|
||||
}
|
||||
|
||||
# Confirm the port is freed.
|
||||
# Also kill any uvicorn still listening on the registered port (worker/respawn case).
|
||||
Kill-PortOwner $port
|
||||
|
||||
# Confirm port freed.
|
||||
$freed = $false
|
||||
for ($i = 0; $i -lt 10; $i++) {
|
||||
try {
|
||||
Invoke-WebRequest -Uri "http://localhost:$port/healthz" -UseBasicParsing -TimeoutSec 1 | Out-Null
|
||||
} catch { $freed = $true; break }
|
||||
for ($i = 0; $i -lt 15; $i++) {
|
||||
if (-not (Test-PortOpen $port 200)) { $freed = $true; break }
|
||||
Start-Sleep -Milliseconds 400
|
||||
}
|
||||
|
||||
Remove-Item '.run\server.info' -Force -ErrorAction SilentlyContinue
|
||||
|
||||
if ($freed) {
|
||||
Write-Host " [ok] puerto liberado, servidor detenido." -ForegroundColor Green
|
||||
} else {
|
||||
Write-Host " [warn] el puerto $port sigue ocupado; revisa procesos python." -ForegroundColor Yellow
|
||||
}
|
||||
Write-Host ""
|
||||
if ($freed) { Out-Line " [ok] puerto liberado, servidor detenido." Green }
|
||||
else { Out-Line " [warn] Algo sigue ocupando la conexion. Reinicia la PC si vuelve a pasar." Yellow }
|
||||
Out-Line ""
|
||||
|
||||
@@ -0,0 +1,91 @@
|
||||
#!/usr/bin/env bash
|
||||
# stop-server.sh — detiene el servidor web registrado en .run/server.info.
|
||||
# Espejo de scripts/stop-server.ps1 (Windows). Mensajes en espanol como el .ps1.
|
||||
#
|
||||
# Uso: bash scripts/stop-server.sh (o `make stop`)
|
||||
# Compatible con el bash 3.2 de macOS y con POSIX sh.
|
||||
|
||||
set -u
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)" || exit 1
|
||||
ROOT="$(cd "$SCRIPT_DIR/.." && pwd)" || exit 1
|
||||
cd "$ROOT" || exit 1
|
||||
|
||||
INFO_FILE=".run/server.info"
|
||||
|
||||
say() { printf '%s\n' "$*"; }
|
||||
|
||||
# --- algo sigue escuchando en el puerto? (nc, con fallback a python3) ---
|
||||
port_busy() {
|
||||
# $1 = puerto; exit 0 si hay un proceso aceptando conexiones ahi
|
||||
if command -v nc >/dev/null 2>&1; then
|
||||
nc -z 127.0.0.1 "$1" >/dev/null 2>&1
|
||||
elif command -v python3 >/dev/null 2>&1; then
|
||||
python3 - "$1" <<'PYEOF' >/dev/null 2>&1
|
||||
import socket, sys
|
||||
s = socket.socket()
|
||||
s.settimeout(1)
|
||||
try:
|
||||
s.connect(("127.0.0.1", int(sys.argv[1])))
|
||||
except OSError:
|
||||
sys.exit(1)
|
||||
PYEOF
|
||||
else
|
||||
return 1
|
||||
fi
|
||||
}
|
||||
|
||||
say ""
|
||||
|
||||
if [ ! -f "$INFO_FILE" ]; then
|
||||
say " [yt-scraper] no hay servidor registrado; nada que detener."
|
||||
say " [ok] nada que detener."
|
||||
say ""
|
||||
exit 0
|
||||
fi
|
||||
|
||||
first_line="$(head -n 1 "$INFO_FILE" 2>/dev/null || true)"
|
||||
# primera linea: "PUERTO PID FECHA"
|
||||
set -- $first_line
|
||||
PORT="${1:-}"
|
||||
PIDV="${2:-}"
|
||||
|
||||
say " [yt-scraper] Apagando el servidor..."
|
||||
|
||||
# PID ausente, no-numerico o ya muerto -> ya estaba apagado.
|
||||
if [ -z "$PIDV" ]; then
|
||||
say " [info] El servidor ya estaba apagado."
|
||||
else
|
||||
case "$PIDV" in
|
||||
*[!0-9]* | "")
|
||||
say " [info] registro corrupto; solo limpio el archivo."
|
||||
;;
|
||||
*)
|
||||
if ! kill -0 "$PIDV" 2>/dev/null; then
|
||||
say " [info] El servidor ya estaba apagado."
|
||||
else
|
||||
# kill amable, ~4 s de gracia y luego kill -9.
|
||||
kill "$PIDV" 2>/dev/null
|
||||
i=0
|
||||
while [ "$i" -lt 10 ] && kill -0 "$PIDV" 2>/dev/null; do
|
||||
sleep 0.4
|
||||
i=$((i + 1))
|
||||
done
|
||||
if kill -0 "$PIDV" 2>/dev/null; then
|
||||
say " [kill] no cerro solo; forzando (kill -9)..."
|
||||
kill -9 "$PIDV" 2>/dev/null
|
||||
sleep 0.5
|
||||
fi
|
||||
fi
|
||||
;;
|
||||
esac
|
||||
fi
|
||||
|
||||
rm -f "$INFO_FILE"
|
||||
|
||||
if [ -n "$PORT" ] && port_busy "$PORT"; then
|
||||
say " [warn] Algo sigue ocupando el puerto $PORT. Revisa con: lsof -i :$PORT"
|
||||
else
|
||||
say " [ok] puerto liberado, servidor detenido."
|
||||
fi
|
||||
say ""
|
||||
@@ -0,0 +1,50 @@
|
||||
"""Shared HTTP helpers for talking to YouTube CDNs.
|
||||
|
||||
Single source of truth for the headers we send to ``*.youtube.com`` endpoints
|
||||
(``ytimg.com``, ``ggpht.com``, ``youtube.com/api/timedtext``). Centralising
|
||||
them prevents the failures we hit when ``Referer: https://www.youtube.com/``
|
||||
was missing on avatar/thumbnail/subtitle fetches.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any, Mapping
|
||||
|
||||
import requests
|
||||
|
||||
|
||||
_YT_HEADERS: dict[str, str] = {
|
||||
# Chrome 124 on Windows. Mimics a real browser request; some Google CDN
|
||||
# endpoints (notably ``yt3.ggpht.com`` avatars and ``timedtext`` captions)
|
||||
# reject ``python-requests`` style UA-only calls with HTTP 403.
|
||||
"User-Agent": (
|
||||
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
|
||||
"(KHTML, like Gecko) Chrome/124.0 Safari/537.36"
|
||||
),
|
||||
"Referer": "https://www.youtube.com/",
|
||||
"Accept-Language": "en-US,en;q=0.9",
|
||||
# Image fetches want an Accept that lists image/* so the CDN can pick
|
||||
# the right encoded variant (avif/webp/etc).
|
||||
"Accept": (
|
||||
"image/avif,image/webp,image/png,image/jpeg,image/*,*/*;q=0.8"
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
def yt_headers(extra: Mapping[str, str] | None = None) -> dict[str, str]:
|
||||
"""Return a copy of the YouTube CDN headers merged with ``extra``."""
|
||||
if extra:
|
||||
return {**_YT_HEADERS, **dict(extra)}
|
||||
return dict(_YT_HEADERS)
|
||||
|
||||
|
||||
def yt_get(url: str, *, timeout: float = 20.0,
|
||||
headers: Mapping[str, str] | None = None,
|
||||
params: Mapping[str, Any] | None = None) -> requests.Response:
|
||||
"""GET ``url`` with the shared YouTube CDN headers.
|
||||
|
||||
Centralising this lets every call site (subtitle download, video
|
||||
thumbnail cache, channel avatar cache) inherit header updates from a
|
||||
single place.
|
||||
"""
|
||||
return requests.get(url, params=params, headers=yt_headers(headers),
|
||||
timeout=timeout)
|
||||
+166
-73
@@ -3,6 +3,7 @@ from __future__ import annotations
|
||||
import logging
|
||||
import shutil
|
||||
import sys
|
||||
import time
|
||||
from pathlib import Path
|
||||
from types import SimpleNamespace
|
||||
|
||||
@@ -11,13 +12,12 @@ from rich.console import Console
|
||||
from rich.progress import Progress, SpinnerColumn, TextColumn, BarColumn, TaskProgressColumn, TimeRemainingColumn
|
||||
from rich.table import Table
|
||||
|
||||
from .config import Config, load_config
|
||||
from .store import Store, VideoRef, VideoRow
|
||||
from .discover import discover_channel
|
||||
from .chapters import align_chapters, chapters_from_info, Chapter, Section
|
||||
from .parse import Segment
|
||||
from .render import build_filename_stem, render_markdown
|
||||
from .ratelimit import polite_sleep
|
||||
from .config import Config, load_config, parse_languages
|
||||
from .store import Store, order_pending
|
||||
from .discover import discover_channel, discover_incremental, extract_handle
|
||||
from .segments import seconds_to_ts
|
||||
from .render import safe_filename
|
||||
from .ratelimit import ThrottleGuard, configure_global_pacer, polite_sleep
|
||||
from .pipeline import process_video
|
||||
from .cookies import auto_import_dir, resolve_active_path
|
||||
|
||||
@@ -51,6 +51,7 @@ def cli(ctx, config_path, channel, all_channels, cookies, cookies_from_browser,
|
||||
cfg = load_config(cfg_path) if cfg_path.exists() else Config()
|
||||
if channel:
|
||||
cfg.channel_url = channel
|
||||
configure_global_pacer(cfg.delay.min_request_interval)
|
||||
store = Store(cfg.database_path_resolved)
|
||||
# import any loose cookie files into the vault (idempotent)
|
||||
try:
|
||||
@@ -72,7 +73,10 @@ def cli(ctx, config_path, channel, all_channels, cookies, cookies_from_browser,
|
||||
def _apply_filters(refs, since, no_shorts, no_live, min_duration, limit):
|
||||
filtered = list(refs)
|
||||
if since:
|
||||
filtered = [r for r in filtered if (r.upload_date or "") >= since.replace("-", "")]
|
||||
# Undated entries survive: yt-dlp's flat listing omits upload_date, so a
|
||||
# plain `>= since` would silently drop every video discovery found.
|
||||
cutoff = since.replace("-", "")
|
||||
filtered = [r for r in filtered if not r.upload_date or r.upload_date >= cutoff]
|
||||
if no_shorts:
|
||||
filtered = [r for r in filtered if "/shorts/" not in (r.url or "")]
|
||||
if no_live:
|
||||
@@ -95,20 +99,24 @@ def _apply_filters(refs, since, no_shorts, no_live, min_duration, limit):
|
||||
@click.option("--resume/--no-resume", default=True)
|
||||
@click.option("--dry-run", is_flag=True, default=False)
|
||||
@click.option("--reset-errors", is_flag=True, default=False)
|
||||
@click.option("--full", is_flag=True, default=False,
|
||||
help="Recorrer el canal entero en vez de solo los videos nuevos")
|
||||
@click.pass_obj
|
||||
def scrape(obj, limit, since, languages, no_auto, no_shorts, include_shorts, no_live, resume, dry_run, reset_errors):
|
||||
def scrape(obj, limit, since, languages, no_auto, no_shorts, include_shorts, no_live, resume, dry_run, reset_errors, full):
|
||||
"""Scrape completo: discovery + extraccion + markdown."""
|
||||
_run_scrape(obj, limit, since, languages, no_auto, no_shorts, include_shorts, no_live, resume, dry_run, reset_errors)
|
||||
_run_scrape(obj, limit, since, languages, no_auto, no_shorts, include_shorts, no_live, resume, dry_run, reset_errors, full)
|
||||
|
||||
|
||||
def _run_scrape(obj, limit, since, languages, no_auto, no_shorts, include_shorts, no_live, resume, dry_run, reset_errors):
|
||||
def _run_scrape(obj, limit, since, languages, no_auto, no_shorts, include_shorts, no_live, resume, dry_run, reset_errors, full=False):
|
||||
cfg: Config = obj.cfg
|
||||
store: Store = obj.store
|
||||
if not cfg.channel_url:
|
||||
console.print("[red]Error:[/red] falta la URL del canal. Usa --channel o config.yaml")
|
||||
sys.exit(1)
|
||||
if languages:
|
||||
cfg.languages = [l.strip() for l in languages.split(",") if l.strip()]
|
||||
# CLI flag accepts legacy CSV; per-language mode requires editing YAML.
|
||||
list_value = [l.strip() for l in languages.split(",") if l.strip()]
|
||||
cfg.languages = parse_languages(list_value, cfg.prefer_manual)
|
||||
if no_auto:
|
||||
cfg.prefer_manual = True
|
||||
if include_shorts:
|
||||
@@ -119,22 +127,52 @@ def _run_scrape(obj, limit, since, languages, no_auto, no_shorts, include_shorts
|
||||
cfg.include_live = False
|
||||
|
||||
console.print(f"[cyan]Canal:[/cyan] {cfg.channel_url}\n[cyan]DB:[/cyan] {cfg.database_path_resolved}")
|
||||
console.print("\n[bold blue]Paso 1:[/bold blue] Discovery")
|
||||
|
||||
# Incremental by default: only pull pages newer than the last video we have.
|
||||
known_channel = _known_channel_for(store, cfg.channel_url)
|
||||
incremental = cfg.sync.incremental and not full and bool(known_channel)
|
||||
known = store.known_video_ids(known_channel) if incremental else set()
|
||||
watermark = store.latest_upload_date(known_channel) if incremental else None
|
||||
|
||||
mode = (
|
||||
f"incremental (desde {_fmt_date(watermark)} en adelante)" if watermark
|
||||
else "incremental (solo videos nuevos)" if incremental
|
||||
else "completo"
|
||||
)
|
||||
console.print(f"\n[bold blue]Paso 1:[/bold blue] Discovery — [dim]{mode}[/dim]")
|
||||
try:
|
||||
with Progress(SpinnerColumn(), TextColumn("[progress.description]{task.description}"), transient=True) as prog:
|
||||
task = prog.add_task("Descubriendo videos...", total=None)
|
||||
channel_id, channel_name, refs = discover_channel(cfg.channel_url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests)
|
||||
result = discover_incremental(
|
||||
cfg.channel_url, known,
|
||||
sleep_subrequests=cfg.yt_dlp.sleep_subrequests,
|
||||
window=cfg.sync.window, max_window=cfg.sync.max_window,
|
||||
overlap=cfg.sync.overlap, since=watermark,
|
||||
)
|
||||
prog.update(task, completed=1, total=1)
|
||||
except Exception as exc:
|
||||
console.print(f"[red]Error en discovery:[/red] {exc}")
|
||||
sys.exit(1)
|
||||
|
||||
store.upsert_channel(channel_id, _extract_handle(cfg.channel_url), channel_name, len(refs))
|
||||
console.print(f"[green]Canal:[/green] {channel_name} ({channel_id}) — {len(refs)} videos")
|
||||
channel_id, channel_name, refs = result.channel_id, result.channel_name, result.refs
|
||||
if result.full_scan:
|
||||
store.upsert_channel(channel_id, extract_handle(cfg.channel_url), channel_name, len(refs))
|
||||
else:
|
||||
store.update_channel_meta(channel_id, name=channel_name)
|
||||
console.print(
|
||||
f"[green]Canal:[/green] {channel_name} ({channel_id}) — "
|
||||
f"{result.fetched} videos leidos en {result.passes} pasada(s), {result.new_count} nuevos"
|
||||
)
|
||||
if not result.caught_up:
|
||||
console.print(
|
||||
"[yellow]Aviso:[/yellow] se alcanzo el limite de ventana "
|
||||
f"({cfg.sync.max_window}). Usa --full si faltan videos antiguos."
|
||||
)
|
||||
|
||||
refs = _apply_filters(refs, since, not cfg.include_shorts, not cfg.include_live, cfg.min_duration_sec, limit)
|
||||
console.print(f"[yellow]Tras filtros:[/yellow] {len(refs)} videos")
|
||||
store.upsert_videos(refs)
|
||||
store.mark_channel_synced(channel_id)
|
||||
|
||||
if reset_errors:
|
||||
n = store.reset_errors(channel_id)
|
||||
@@ -145,8 +183,13 @@ def _run_scrape(obj, limit, since, languages, no_auto, no_shorts, include_shorts
|
||||
return
|
||||
|
||||
pending = store.get_pending(channel_id) if resume else store.get_all(channel_id)
|
||||
if result.full_scan:
|
||||
ref_ids = {r.video_id for r in refs}
|
||||
pending = [p for p in pending if p.video_id in ref_ids] if ref_ids else pending
|
||||
else:
|
||||
# Discovery only saw the newest slice; keep the older backlog reachable
|
||||
# but put this run's videos first so --limit still means "los mas nuevos".
|
||||
pending = order_pending(pending, refs)
|
||||
if limit:
|
||||
pending = pending[:limit]
|
||||
if not pending:
|
||||
@@ -162,17 +205,56 @@ def _run_scrape(obj, limit, since, languages, no_auto, no_shorts, include_shorts
|
||||
with Progress(SpinnerColumn(), TextColumn("[progress.description]{task.description}"),
|
||||
BarColumn(), TaskProgressColumn(), TimeRemainingColumn()) as progress:
|
||||
task = progress.add_task("Procesando", total=len(pending))
|
||||
guard = ThrottleGuard(
|
||||
threshold=cfg.delay.throttle_threshold,
|
||||
base=cfg.delay.backoff_base,
|
||||
cap=cfg.delay.backoff_cap,
|
||||
)
|
||||
for i, row in enumerate(pending):
|
||||
progress.update(task, description=f"{row.video_id} {(row.title or '')[:30]}", completed=i)
|
||||
process_video(row, cfg, store, channel_name, channel_id, cfg.channel_url,
|
||||
status = process_video(row, cfg, store, channel_name, channel_id, cfg.channel_url,
|
||||
cookies_file=cookie_path, cookies_from_browser=obj.cookies_from_browser)
|
||||
progress.advance(task)
|
||||
if status == "done":
|
||||
guard.note_success()
|
||||
else:
|
||||
failed = store.get_video(row.video_id)
|
||||
wait = guard.note_failure(getattr(failed, "error_msg", None) or status)
|
||||
if guard.tripped:
|
||||
# Everything still pending stays pending: that is what makes
|
||||
# this recoverable instead of 500 rows marked failed.
|
||||
console.print(f"\n[bold red]Detenido:[/bold red] {guard.tripped_reason}")
|
||||
console.print(f"[yellow]{len(pending) - i - 1} videos sin tocar, siguen pendientes.[/yellow]")
|
||||
break
|
||||
if wait > 0:
|
||||
console.print(f"[yellow]YouTube nos esta limitando; esperando {wait:.1f}s[/yellow]")
|
||||
time.sleep(wait)
|
||||
if i < len(pending) - 1:
|
||||
polite_sleep(cfg.delay.min_seconds, cfg.delay.max_seconds)
|
||||
console.print()
|
||||
_print_stats(store, channel_id)
|
||||
|
||||
|
||||
def _known_channel_for(store: Store, channel_url: str) -> str | None:
|
||||
"""Match a channel URL against an already-tracked channel, by handle or id.
|
||||
|
||||
Without a hit there is no local history to stop at, so discovery has to walk
|
||||
the whole channel — which is correct for a first run.
|
||||
"""
|
||||
if not channel_url:
|
||||
return None
|
||||
handle = extract_handle(channel_url).lstrip("@").lower()
|
||||
tail = channel_url.rstrip("/").split("/")[-1]
|
||||
for ch in store.list_channels():
|
||||
cid = ch.get("channel_id") or ""
|
||||
ch_handle = (ch.get("handle") or "").lstrip("@").lower()
|
||||
if handle and ch_handle == handle:
|
||||
return cid
|
||||
if cid and (cid == tail or f"/channel/{cid}" in channel_url):
|
||||
return cid
|
||||
return None
|
||||
|
||||
|
||||
def _print_dry_run(refs):
|
||||
table = Table(show_lines=False)
|
||||
table.add_column("Fecha", style="dim")
|
||||
@@ -186,6 +268,55 @@ def _print_dry_run(refs):
|
||||
console.print(f"... y {len(refs) - 50} mas")
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- recovery
|
||||
|
||||
@cli.command("reset")
|
||||
@click.option("--status", "statuses", multiple=True,
|
||||
type=click.Choice(Store.RETRYABLE_STATUSES),
|
||||
help="Estados a reintentar (por defecto: error y no_subtitles)")
|
||||
@click.option("--yes", is_flag=True, default=False, help="No preguntar")
|
||||
@click.pass_obj
|
||||
def reset_cmd(obj, statuses, yes):
|
||||
"""Devolver videos atascados en error/no_subtitles a 'pending' para reintentarlos."""
|
||||
store: Store = obj.store
|
||||
targets = tuple(statuses) if statuses else Store.RETRYABLE_STATUSES
|
||||
channels = _channel_targets(obj)
|
||||
channel_id = channels[0] if len(channels) == 1 else None
|
||||
|
||||
counts = store.retryable_counts(channel_id)
|
||||
affected = sum(counts.get(s, 0) for s in targets)
|
||||
if not affected:
|
||||
console.print("[green]No hay videos que reintentar.[/green]")
|
||||
return
|
||||
detail = ", ".join(f"{s}={counts.get(s, 0)}" for s in targets)
|
||||
scope = channel_id or "todos los canales"
|
||||
if not yes and not click.confirm(f"Reintentar {affected} videos ({detail}) en {scope}?"):
|
||||
return
|
||||
n = store.reset_videos(channel_id, targets)
|
||||
console.print(f"[yellow]{n}[/yellow] videos vueltos a 'pending'.")
|
||||
|
||||
|
||||
@cli.command("reconcile")
|
||||
@click.option("--prune", is_flag=True, default=False,
|
||||
help="Ademas, devolver a 'pending' los 'done' cuyo .md ya no existe")
|
||||
@click.pass_obj
|
||||
def reconcile_cmd(obj, prune):
|
||||
"""Re-escanear data/markdown y hacer que la DB coincida con el disco."""
|
||||
from .segments import reconcile_markdown
|
||||
cfg: Config = obj.cfg
|
||||
store: Store = obj.store
|
||||
result = reconcile_markdown(
|
||||
store, Path(cfg.output_dir_resolved),
|
||||
log=lambda m: console.print(f"[dim]{m}[/dim]"), prune=prune,
|
||||
)
|
||||
table = Table(title="Reconciliacion disco <-> DB")
|
||||
table.add_column("Metrica", style="bold")
|
||||
table.add_column("Valor", justify="right")
|
||||
for k, v in result.items():
|
||||
table.add_row(k, str(v))
|
||||
console.print(table)
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- search
|
||||
|
||||
@cli.command("search")
|
||||
@@ -248,7 +379,13 @@ def export_cmd(obj, fmt, out_dir):
|
||||
def audio_cmd(obj, limit, since, force):
|
||||
"""Descarga audio MP3 (requiere ffmpeg)."""
|
||||
if not shutil.which("ffmpeg"):
|
||||
console.print("[red]ffmpeg no encontrado.[/red] Instala: winget install ffmpeg")
|
||||
if sys.platform == "win32":
|
||||
hint = "winget install ffmpeg (o choco install ffmpeg)"
|
||||
elif sys.platform == "darwin":
|
||||
hint = "brew install ffmpeg"
|
||||
else:
|
||||
hint = "apt install ffmpeg (Debian/Ubuntu) / dnf install ffmpeg (Fedora) / pacman -S ffmpeg (Arch)"
|
||||
console.print(f"[red]ffmpeg no encontrado.[/red] Instala: {hint}")
|
||||
sys.exit(1)
|
||||
store: Store = obj.store
|
||||
cfg: Config = obj.cfg
|
||||
@@ -268,6 +405,7 @@ def audio_cmd(obj, limit, since, force):
|
||||
"outtmpl": str(out_dir / "%(title)s.%(ext)s"),
|
||||
"postprocessors": [{"key": "FFmpegExtractAudio", "preferredcodec": "mp3", "preferredquality": "128"}],
|
||||
"quiet": True, "no_warnings": True, "noprogress": True,
|
||||
"js_runtimes": {"node": {}, "deno": {}, "bun": {}, "quickjs": {}},
|
||||
}
|
||||
if cookie_path:
|
||||
ydl_opts["cookiefile"] = cookie_path
|
||||
@@ -275,7 +413,7 @@ def audio_cmd(obj, limit, since, force):
|
||||
ydl_opts["cookiesfrombrowser"] = (obj.cookies_from_browser,)
|
||||
with yt_dlp.YoutubeDL(ydl_opts) as ydl:
|
||||
for v in videos:
|
||||
target = out_dir / f"{_safe_filename(v.title or v.video_id)}.mp3"
|
||||
target = out_dir / f"{safe_filename(v.title or v.video_id)}.mp3"
|
||||
if target.exists() and not force:
|
||||
continue
|
||||
try:
|
||||
@@ -313,8 +451,8 @@ def channels_add(obj, url):
|
||||
cfg: Config = obj.cfg
|
||||
cfg.channel_url = url
|
||||
console.print("[cyan]Resolviendo canal...[/cyan]")
|
||||
channel_id, name, refs = discover_channel(url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests)
|
||||
obj.store.upsert_channel(channel_id, _extract_handle(url), name, len(refs))
|
||||
channel_id, name, avatar, refs = discover_channel(url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests)
|
||||
obj.store.upsert_channel(channel_id, extract_handle(url), name, len(refs), avatar=avatar)
|
||||
obj.store.upsert_videos(refs)
|
||||
console.print(f"[green]Added:[/green] {name} ({channel_id}) — {len(refs)} videos")
|
||||
|
||||
@@ -409,6 +547,7 @@ def re_render_cmd(obj, backfill):
|
||||
"""Regenerar Markdown desde segmentos almacenados."""
|
||||
store: Store = obj.store
|
||||
cfg: Config = obj.cfg
|
||||
from .pipeline import re_render_videos
|
||||
from .segments import backfill_from_markdown
|
||||
if backfill:
|
||||
md_root = Path(cfg.output_dir_resolved)
|
||||
@@ -419,26 +558,13 @@ def re_render_cmd(obj, backfill):
|
||||
if not videos:
|
||||
console.print("[yellow]No hay videos con segments_json. Usa --backfill.[/yellow]")
|
||||
return
|
||||
import json
|
||||
from .render import render_markdown
|
||||
console.print(f"[cyan]Re-renderizando {len(videos)} videos...[/cyan]")
|
||||
for v in videos:
|
||||
segs = [Segment(start=s["start"], end=s["end"], text=s["text"]) for s in json.loads(v.segments_json)]
|
||||
chapters = [Chapter(title=c["title"], start_time=c["start"], end_time=c.get("end", c["start"])) for c in json.loads(v.chapters_json or "[]")]
|
||||
sections = align_chapters(segs, chapters)
|
||||
context = {
|
||||
"video_id": v.video_id, "title": v.title or v.video_id, "channel_name": "",
|
||||
"channel_id": v.channel_id, "channel_url": "", "upload_date": v.upload_date or "",
|
||||
"duration": v.duration or 0, "url": v.url, "transcript_lang": v.transcript_lang or "",
|
||||
"transcript_src": v.transcript_src or "", "view_count": v.view_count, "like_count": v.like_count,
|
||||
"tags": _parse_tags(v.tags), "thumbnail": v.thumbnail or "", "description": v.description or "",
|
||||
"sections": sections,
|
||||
}
|
||||
stem = build_filename_stem(v.upload_date, v.title or v.video_id, cfg.filename_template)
|
||||
ch = store.get_channel(v.channel_id)
|
||||
out_subdir = Path(cfg.output_dir_resolved) / _safe_dirname((ch or {}).get("name") or "unknown")
|
||||
render_markdown(cfg.template_path_resolved, out_subdir, stem, context)
|
||||
console.print(f"[green]Re-render completo.[/green]")
|
||||
# Delegate to the pipeline implementation instead of a local loop: it
|
||||
# retires the .md a re-render supersedes (renamed titles used to leave
|
||||
# orphans), re-points markdown_path via mark_done and fills channel_name
|
||||
# from the channels table.
|
||||
n = sum(re_render_videos(store, cfg, cid) for cid in targets)
|
||||
console.print(f"[green]Re-render completo:[/green] {n} videos")
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- helpers
|
||||
@@ -472,12 +598,6 @@ def _title_for(store: Store, video_id: str) -> str | None:
|
||||
return v.title if v else None
|
||||
|
||||
|
||||
def _extract_handle(url: str) -> str:
|
||||
if "@" in url:
|
||||
return "@" + url.split("@", 1)[1].split("/", 1)[0]
|
||||
return ""
|
||||
|
||||
|
||||
def _fmt_date(d: str | None) -> str:
|
||||
if not d:
|
||||
return ""
|
||||
@@ -494,33 +614,6 @@ def _fmt_duration(seconds: int | None) -> str:
|
||||
return f"{h}:{m:02d}:{s:02d}" if h else f"{m}:{s:02d}"
|
||||
|
||||
|
||||
def seconds_to_ts(sec: float) -> str:
|
||||
total = int(sec)
|
||||
h, rem = divmod(total, 3600)
|
||||
m, s = divmod(rem, 60)
|
||||
return f"{h:d}:{m:02d}:{s:02d}" if h else f"{m:d}:{s:02d}"
|
||||
|
||||
|
||||
def _safe_dirname(name: str) -> str:
|
||||
safe = "".join(c for c in name if c not in r'\/:*?"<>|')
|
||||
return safe.strip().strip(".") or "unknown"
|
||||
|
||||
|
||||
def _safe_filename(name: str) -> str:
|
||||
safe = "".join(c for c in name if c not in r'\/:*?"<>|')
|
||||
return safe.strip().strip(".") or "untitled"
|
||||
|
||||
|
||||
def _parse_tags(tags_json: str | None) -> list[str]:
|
||||
if not tags_json:
|
||||
return []
|
||||
import json
|
||||
try:
|
||||
return json.loads(tags_json)
|
||||
except (json.JSONDecodeError, TypeError):
|
||||
return []
|
||||
|
||||
|
||||
def _parse_interval(s: str) -> float:
|
||||
s = s.strip().lower()
|
||||
if s.endswith("s"):
|
||||
|
||||
+115
-3
@@ -9,28 +9,79 @@ import yaml
|
||||
|
||||
@dataclass
|
||||
class DelayConfig:
|
||||
"""Pacing and failure handling for everything that touches YouTube.
|
||||
|
||||
`min_seconds`/`max_seconds` are the randomised gap between units of work
|
||||
(one video, one channel). `backoff_*` and `throttle_threshold` only come
|
||||
into play once YouTube starts refusing: the backoff pair feeds the
|
||||
truncated-exponential-with-jitter formula Google documents for its own
|
||||
APIs, and the threshold is how many consecutive rate-limit responses end
|
||||
the run instead of burning through the rest of the queue.
|
||||
"""
|
||||
|
||||
min_seconds: float = 1.5
|
||||
max_seconds: float = 3.5
|
||||
backoff_base: float = 2.0
|
||||
backoff_cap: float = 60.0
|
||||
throttle_threshold: int = 3
|
||||
# Minimum gap between *any* two requests to YouTube from this process,
|
||||
# including the ones the `/api/tools/*` endpoints make outside the job
|
||||
# runner. 0 disables the shared pacer and leaves only the per-unit sleeps.
|
||||
min_request_interval: float = 0.0
|
||||
# Bytes/sec ceiling for audio downloads (yt-dlp `ratelimit`). 0 = unlimited.
|
||||
audio_rate_limit: int = 0
|
||||
|
||||
|
||||
@dataclass
|
||||
class YtDlpConfig:
|
||||
retries: int = 10
|
||||
# Seconds between the individual HTTP requests inside one extraction.
|
||||
# Passed to yt-dlp as `sleep_interval_requests`; the name here is the
|
||||
# project's own and predates the discovery that the option this used to be
|
||||
# forwarded under (`sleep_subrequests`) does not exist in yt-dlp at all.
|
||||
sleep_subrequests: float = 2.0
|
||||
extractor_retries: int = 3
|
||||
socket_timeout: float = 30.0
|
||||
|
||||
|
||||
@dataclass
|
||||
class SyncConfig:
|
||||
"""Incremental channel sync — how far back a routine re-scan looks.
|
||||
|
||||
The /videos tab is reverse-chronological and every extra page is another
|
||||
request to YouTube, so a sync fetches `window` entries and stops as soon as
|
||||
it has seen `overlap` consecutive videos already in the DB. Only if the
|
||||
whole window turns out to be new does it widen (doubling up to `max_window`),
|
||||
which is the case where the channel really did publish a lot since last time.
|
||||
"""
|
||||
|
||||
incremental: bool = True
|
||||
window: int = 30
|
||||
max_window: int = 300
|
||||
overlap: int = 3
|
||||
|
||||
|
||||
@dataclass
|
||||
class Config:
|
||||
channel_url: str = ""
|
||||
languages: list[str] = field(default_factory=lambda: ["es", "en"])
|
||||
# Per-language subtitle preference. Each value is one of:
|
||||
# "manual" - only manually uploaded captions
|
||||
# "auto" - only YouTube-auto-generated captions
|
||||
# "any" - defer to the legacy `prefer_manual` flag
|
||||
# Legacy form: ``["es", "en"]`` - treated as ``{lang: <prefer_manual_default>}``.
|
||||
# Default to "any" (manual first, auto as fallback). A manual-only default
|
||||
# silently yields nothing on the many channels that publish only
|
||||
# auto-generated captions, and stores that as `no_subtitles`.
|
||||
languages: dict[str, str] = field(
|
||||
default_factory=lambda: {"es": "any", "en": "any"}
|
||||
)
|
||||
prefer_manual: bool = True
|
||||
include_shorts: bool = False
|
||||
include_live: bool = True
|
||||
min_duration_sec: int = 0
|
||||
delay: DelayConfig = field(default_factory=DelayConfig)
|
||||
yt_dlp: YtDlpConfig = field(default_factory=YtDlpConfig)
|
||||
sync: SyncConfig = field(default_factory=SyncConfig)
|
||||
database_path: str = "data/state.db"
|
||||
output_dir: str = "data/markdown"
|
||||
template_path: str = "templates/video.md.j2"
|
||||
@@ -49,6 +100,51 @@ class Config:
|
||||
return Path(self.template_path).resolve()
|
||||
|
||||
|
||||
# Modes accepted in the ``languages`` dict. Kept here so tests and CLI
|
||||
# share a single definition without importing the private extract constant.
|
||||
LANGUAGE_MODES = ("manual", "auto", "any")
|
||||
|
||||
|
||||
def keep_ref(cfg: Config, include_shorts: bool | None = None, no_live: bool | None = None):
|
||||
"""Predicate matching the shorts/live rules that decide what reaches the DB.
|
||||
|
||||
Lives next to the `Config` flags it reads so the policy has one home.
|
||||
Incremental discovery needs the same filter its stored ids were created
|
||||
under, otherwise the tail of a window is full of entries that can never be
|
||||
recognised as known and the window keeps widening for nothing. The optional
|
||||
overrides serve the job runner, whose per-run opts may disagree with the
|
||||
config the store was populated under.
|
||||
"""
|
||||
shorts = cfg.include_shorts if include_shorts is None else include_shorts
|
||||
skip_live = (not cfg.include_live) if no_live is None else no_live
|
||||
|
||||
def keep(r) -> bool: # VideoRef or VideoRow — both carry `.url`
|
||||
url = r.url or ""
|
||||
if not shorts and "/shorts/" in url:
|
||||
return False
|
||||
if skip_live and url.startswith("https://www.youtube.com/live/"):
|
||||
return False
|
||||
return True
|
||||
|
||||
return keep
|
||||
|
||||
|
||||
def parse_languages(raw: Any, prefer_manual: bool) -> dict[str, str]:
|
||||
"""Normalise legacy list / new dict / None into ``{lang: mode}``."""
|
||||
default = "manual" if prefer_manual else "auto"
|
||||
if raw is None:
|
||||
return {}
|
||||
if isinstance(raw, dict):
|
||||
out: dict[str, str] = {}
|
||||
for lang, mode in raw.items():
|
||||
m = str(mode).lower().strip()
|
||||
out[str(lang)] = m if m in LANGUAGE_MODES else "any"
|
||||
return out
|
||||
if isinstance(raw, (list, tuple)):
|
||||
return {str(l): default for l in raw}
|
||||
return {}
|
||||
|
||||
|
||||
def load_config(path: str | Path) -> Config:
|
||||
p = Path(path)
|
||||
if not p.exists():
|
||||
@@ -61,10 +157,15 @@ def load_config(path: str | Path) -> Config:
|
||||
def _build_config(raw: dict[str, Any]) -> Config:
|
||||
delay_raw = raw.get("delay") or {}
|
||||
ydl_raw = raw.get("yt_dlp") or {}
|
||||
sync_raw = raw.get("sync") or {}
|
||||
prefer_manual = bool(raw.get("prefer_manual", True))
|
||||
languages = parse_languages(raw.get("languages"), prefer_manual)
|
||||
if not languages:
|
||||
languages = {"es": "any", "en": "any"}
|
||||
return Config(
|
||||
channel_url=raw.get("channel_url", ""),
|
||||
languages=list(raw.get("languages", ["es", "en"])),
|
||||
prefer_manual=bool(raw.get("prefer_manual", True)),
|
||||
languages=languages,
|
||||
prefer_manual=prefer_manual,
|
||||
include_shorts=bool(raw.get("include_shorts", False)),
|
||||
include_live=bool(raw.get("include_live", True)),
|
||||
min_duration_sec=int(raw.get("min_duration_sec", 0)),
|
||||
@@ -73,10 +174,21 @@ def _build_config(raw: dict[str, Any]) -> Config:
|
||||
max_seconds=float(delay_raw.get("max_seconds", 3.5)),
|
||||
backoff_base=float(delay_raw.get("backoff_base", 2.0)),
|
||||
backoff_cap=float(delay_raw.get("backoff_cap", 60.0)),
|
||||
throttle_threshold=max(1, int(delay_raw.get("throttle_threshold", 3))),
|
||||
min_request_interval=max(0.0, float(delay_raw.get("min_request_interval", 0.0))),
|
||||
audio_rate_limit=max(0, int(delay_raw.get("audio_rate_limit", 0))),
|
||||
),
|
||||
yt_dlp=YtDlpConfig(
|
||||
retries=int(ydl_raw.get("retries", 10)),
|
||||
sleep_subrequests=float(ydl_raw.get("sleep_subrequests", 2.0)),
|
||||
extractor_retries=max(0, int(ydl_raw.get("extractor_retries", 3))),
|
||||
socket_timeout=float(ydl_raw.get("socket_timeout", 30.0)),
|
||||
),
|
||||
sync=SyncConfig(
|
||||
incremental=bool(sync_raw.get("incremental", True)),
|
||||
window=max(1, int(sync_raw.get("window", 30))),
|
||||
max_window=max(1, int(sync_raw.get("max_window", 300))),
|
||||
overlap=max(1, int(sync_raw.get("overlap", 3))),
|
||||
),
|
||||
database_path=raw.get("database_path", "data/state.db"),
|
||||
output_dir=raw.get("output_dir", "data/markdown"),
|
||||
|
||||
+238
-13
@@ -1,6 +1,7 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
import os
|
||||
import uuid
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
@@ -10,7 +11,12 @@ from .store import CookieRow, Store
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
# A logged-in YouTube session always sets this core set. A lone
|
||||
# `__Secure-3PSID` (partial extension export) is NOT a session — YouTube
|
||||
# treats the request as anonymous and members-only content stays locked,
|
||||
# which is how an "active membership" cookie once failed invisibly.
|
||||
SESSION_COOKIE_NAMES = {"SID", "SAPISID", "__Secure-3PSID", "SSID", "LOGIN_INFO", "HSID", "APISID"}
|
||||
FULL_SESSION_NAMES = {"SID", "HSID", "SSID"}
|
||||
|
||||
_DEFAULT_DIR = Path("cookies")
|
||||
|
||||
@@ -28,10 +34,16 @@ def parse_netscape(text: str) -> tuple[bool, dict]:
|
||||
"""
|
||||
lines = text.splitlines()
|
||||
expiries: list[int] = []
|
||||
session_expiries: list[int] = []
|
||||
names: set[str] = set()
|
||||
count = 0
|
||||
for line in lines:
|
||||
line = line.rstrip("\n")
|
||||
# "#HttpOnly_" is a data prefix, not a comment — and the session
|
||||
# cookies themselves (SID, HSID, ...) are HttpOnly, so skipping
|
||||
# these lines would silently strip the login out of an export.
|
||||
if line.startswith("#HttpOnly_"):
|
||||
line = line[len("#HttpOnly_"):]
|
||||
if not line.strip() or line.startswith("#"):
|
||||
continue
|
||||
parts = line.split("\t")
|
||||
@@ -48,14 +60,20 @@ def parse_netscape(text: str) -> tuple[bool, dict]:
|
||||
count += 1
|
||||
if expiry:
|
||||
expiries.append(expiry)
|
||||
if name in SESSION_COOKIE_NAMES:
|
||||
session_expiries.append(expiry)
|
||||
if count == 0:
|
||||
return False, {"count": 0, "has_session": False, "expires_at": None, "names": set()}
|
||||
has_session = bool(names & SESSION_COOKIE_NAMES)
|
||||
has_session = FULL_SESSION_NAMES <= names or "LOGIN_INFO" in names
|
||||
# What the user needs to know is when the LOGIN dies, not when the
|
||||
# earliest throwaway cookie (YSC and friends live hours) lapses — so
|
||||
# report the session cookies' own expiry, falling back to the file max.
|
||||
relevant = session_expiries or expiries
|
||||
expires_at = None
|
||||
if expiries:
|
||||
earliest = min(expiries)
|
||||
if earliest > 0:
|
||||
expires_at = datetime.fromtimestamp(earliest, tz=timezone.utc).isoformat(timespec="seconds")
|
||||
if relevant:
|
||||
latest = max(relevant)
|
||||
if latest > 0:
|
||||
expires_at = datetime.fromtimestamp(latest, tz=timezone.utc).isoformat(timespec="seconds")
|
||||
return True, {"count": count, "has_session": has_session, "expires_at": expires_at, "names": names}
|
||||
|
||||
|
||||
@@ -99,10 +117,15 @@ def import_text(store: Store, text: str, label: str, cookie_dir: str | Path | No
|
||||
def auto_import_dir(store: Store, dir_path: str | Path | None = None) -> int:
|
||||
"""Import any loose .txt Netscape files in dir that aren't tracked yet. Returns count."""
|
||||
cdir = cookies_dir(dir_path)
|
||||
tracked = {c.filename for c in store.list_cookies()}
|
||||
tracked = store.list_cookies()
|
||||
for c in tracked:
|
||||
if not (cdir / c.filename).exists():
|
||||
store.delete_cookie(c.id)
|
||||
|
||||
tracked_filenames = {c.filename for c in store.list_cookies()}
|
||||
n = 0
|
||||
for f in sorted(cdir.glob("*.txt")):
|
||||
if f.name in tracked:
|
||||
if f.name in tracked_filenames:
|
||||
continue
|
||||
ok, info = parse_netscape_file(f)
|
||||
if not ok:
|
||||
@@ -118,14 +141,210 @@ def auto_import_dir(store: Store, dir_path: str | Path | None = None) -> int:
|
||||
cookie_count=info["count"],
|
||||
)
|
||||
n += 1
|
||||
# activate first cookie if none active
|
||||
if not store.get_active_cookie():
|
||||
# activate first cookie if none active or active file missing
|
||||
active = store.get_active_cookie()
|
||||
if not active or not (cdir / active.filename).exists():
|
||||
cookies = store.list_cookies()
|
||||
if cookies:
|
||||
store.set_active_cookie(cookies[0].id)
|
||||
return n
|
||||
|
||||
|
||||
class BrowserCookieLockedError(RuntimeError):
|
||||
"""The browser's cookie database could not be read — it is running.
|
||||
|
||||
Chromium opens its Cookies file without sharing read access, so the
|
||||
browser (including its background/tray processes) must be fully closed
|
||||
for the extraction to succeed.
|
||||
"""
|
||||
|
||||
|
||||
# Where the Brave executable lives on a default install, in preference
|
||||
# order. The %VAR% placeholders only expand on Windows (elsewhere they stay
|
||||
# literal and simply never match), so one flat cross-platform list works.
|
||||
_BRAVE_EXE_CANDIDATES = (
|
||||
# Windows
|
||||
r"%ProgramFiles%\BraveSoftware\Brave-Browser\Application\brave.exe",
|
||||
r"%LocalAppData%\BraveSoftware\Brave-Browser\Application\brave.exe",
|
||||
# macOS (system-wide Applications and per-user ~/Applications)
|
||||
"/Applications/Brave Browser.app/Contents/MacOS/Brave Browser",
|
||||
"~/Applications/Brave Browser.app/Contents/MacOS/Brave Browser",
|
||||
# Linux (official deb/rpm packages, distro builds, snap, manual installs)
|
||||
"/usr/bin/brave-browser",
|
||||
"/usr/bin/brave",
|
||||
"/opt/brave.com/brave/brave-browser",
|
||||
"/snap/bin/brave",
|
||||
"/usr/local/bin/brave-browser",
|
||||
)
|
||||
|
||||
# Default profile ("User Data") directories per OS — same flat-list trick:
|
||||
# only the one for the current OS exists, the rest never match.
|
||||
_BRAVE_USER_DATA_CANDIDATES = (
|
||||
r"%LocalAppData%\BraveSoftware\Brave-Browser\User Data", # Windows
|
||||
"~/Library/Application Support/BraveSoftware/Brave-Browser", # macOS
|
||||
"~/.config/BraveSoftware/Brave-Browser", # Linux
|
||||
)
|
||||
|
||||
|
||||
def _expand_path(candidate: str) -> str:
|
||||
"""Expand %VAR% (Windows) and ~ (macOS/Linux) in a path candidate."""
|
||||
return os.path.expanduser(os.path.expandvars(candidate))
|
||||
|
||||
|
||||
def _find_brave() -> tuple[str | None, str | None]:
|
||||
"""(exe_path, user_data_dir) for a default Brave install on Windows/macOS/Linux."""
|
||||
exe = next(
|
||||
(p for p in (_expand_path(c) for c in _BRAVE_EXE_CANDIDATES) if os.path.isfile(p)),
|
||||
None,
|
||||
)
|
||||
ud = next(
|
||||
(p for p in (_expand_path(c) for c in _BRAVE_USER_DATA_CANDIDATES) if os.path.isdir(p)),
|
||||
None,
|
||||
)
|
||||
return exe, ud
|
||||
|
||||
|
||||
def _cdp_cookie(cookie: dict):
|
||||
"""DevTools cookie dict -> http.cookiejar.Cookie."""
|
||||
import http.cookiejar
|
||||
domain = cookie.get("domain") or ".youtube.com"
|
||||
http_only = bool(cookie.get("httpOnly"))
|
||||
return http.cookiejar.Cookie(
|
||||
version=0, name=cookie["name"], value=cookie.get("value") or "",
|
||||
port=None, port_specified=False,
|
||||
domain=domain, domain_specified=True, domain_initial_dot=domain.startswith("."),
|
||||
path=cookie.get("path") or "/", path_specified=True,
|
||||
secure=bool(cookie.get("secure")),
|
||||
expires=int(cookie.get("expires")) if cookie.get("expires") else None,
|
||||
discard=False, comment=None, comment_url=None,
|
||||
rest={"HttpOnly": None} if http_only else {},
|
||||
)
|
||||
|
||||
|
||||
def _extract_brave_cdp(timeout: float = 45.0) -> list:
|
||||
"""Launch Brave headless on its REAL profile and read decrypted cookies
|
||||
over DevTools.
|
||||
|
||||
Chromium 127+ encrypts new cookies "app-bound" (v20): only the browser
|
||||
itself can decrypt them, which is why yt-dlp's file-based extraction
|
||||
dies with "Failed to decrypt with DPAPI". Launching the browser with
|
||||
its user-data-dir passed EXPLICITLY on the command line keeps remote
|
||||
debugging allowed (Chromium 136+ blocks it for the implicit default
|
||||
dir), and the browser hands its own cookies over in plaintext. The
|
||||
browser must still be closed — the profile is single-writer.
|
||||
"""
|
||||
import json
|
||||
import socket
|
||||
import subprocess
|
||||
import time
|
||||
import urllib.request
|
||||
|
||||
exe, user_data = _find_brave()
|
||||
if not exe or not user_data:
|
||||
raise RuntimeError(
|
||||
"Brave installation not found (looked in the default install "
|
||||
"locations for Windows, macOS and Linux)"
|
||||
)
|
||||
|
||||
with socket.socket() as s:
|
||||
s.bind(("127.0.0.1", 0))
|
||||
port = s.getsockname()[1]
|
||||
|
||||
proc = subprocess.Popen(
|
||||
[
|
||||
exe, "--headless=new", f"--remote-debugging-port={port}",
|
||||
f"--user-data-dir={user_data}", "--no-first-run",
|
||||
"--no-default-browser-check", "about:blank",
|
||||
],
|
||||
stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL,
|
||||
)
|
||||
try:
|
||||
deadline = time.monotonic() + timeout
|
||||
version = None
|
||||
while time.monotonic() < deadline:
|
||||
try:
|
||||
version = json.load(urllib.request.urlopen(
|
||||
f"http://127.0.0.1:{port}/json/version", timeout=2))
|
||||
break
|
||||
except Exception:
|
||||
time.sleep(0.5)
|
||||
if version is None:
|
||||
# Most common cause: Brave is already running and the new
|
||||
# process just delegated to it without opening a debug port.
|
||||
raise BrowserCookieLockedError(
|
||||
"could not attach to Brave — close Brave completely "
|
||||
"(including background processes) and try again"
|
||||
)
|
||||
|
||||
from websockets.sync.client import connect
|
||||
with connect(version["webSocketDebuggerUrl"], open_timeout=10) as ws:
|
||||
ws.send(json.dumps({"id": 1, "method": "Storage.getCookies"}))
|
||||
reply = json.loads(ws.recv())
|
||||
return [_cdp_cookie(c) for c in reply.get("result", {}).get("cookies", [])]
|
||||
finally:
|
||||
proc.kill()
|
||||
|
||||
|
||||
|
||||
def import_from_browser(
|
||||
store: Store,
|
||||
browser: str = "brave",
|
||||
profile: str | None = None,
|
||||
label: str | None = None,
|
||||
cookie_dir: str | Path | None = None,
|
||||
) -> str:
|
||||
"""Pull youtube.com cookies straight from a local browser profile.
|
||||
|
||||
Tries yt-dlp's file-based extraction first; on Chromium 127+ profiles
|
||||
whose cookies are app-bound (v20 — "Failed to decrypt with DPAPI"),
|
||||
falls back to launching the browser itself headless and reading the
|
||||
decrypted cookies over DevTools. Only youtube.com cookies are kept and
|
||||
stored in the vault as a normal Netscape file, so activation, expiry
|
||||
tracking and every extraction path work unchanged.
|
||||
"""
|
||||
from yt_dlp.cookies import extract_cookies_from_browser
|
||||
|
||||
cookies = None
|
||||
try:
|
||||
cookies = extract_cookies_from_browser(browser, profile or None)
|
||||
except Exception as exc:
|
||||
msg = str(exc)
|
||||
if "Could not copy Chrome cookie database" in msg or "database is locked" in msg:
|
||||
raise BrowserCookieLockedError(
|
||||
f"could not read {browser}'s cookie database — close {browser} completely "
|
||||
"(including background processes) and try again"
|
||||
) from exc
|
||||
# App-bound (v20) cookies: only the browser can decrypt them.
|
||||
if browser == "brave" and ("decrypt with DPAPI" in msg or "decrypt" in msg.lower()):
|
||||
log.info("Brave cookies are app-bound; falling back to DevTools extraction")
|
||||
cookies = _extract_brave_cdp()
|
||||
else:
|
||||
raise
|
||||
|
||||
lines = []
|
||||
for c in cookies:
|
||||
if "youtube.com" not in (c.domain or ""):
|
||||
continue
|
||||
# Netscape format; the #HttpOnly_ prefix is stripped again on parse.
|
||||
prefix = "#HttpOnly_" if c.has_nonstandard_attr("httponly") else ""
|
||||
domain = c.domain or ".youtube.com"
|
||||
include_subdomains = "TRUE" if domain.startswith(".") else "FALSE"
|
||||
expiry = int(c.expires) if c.expires else 0
|
||||
lines.append(
|
||||
f"{prefix}{domain}\t{include_subdomains}\t{c.path or '/'}\t"
|
||||
f"{'TRUE' if c.secure else 'FALSE'}\t{expiry}\t{c.name}\t{c.value}"
|
||||
)
|
||||
if not lines:
|
||||
raise ValueError(f"no youtube.com cookies found in {browser} (not logged in?)")
|
||||
|
||||
text = (
|
||||
"# Netscape HTTP Cookie File\n"
|
||||
f"# Extracted from {browser} profile {profile or 'default'}\n"
|
||||
+ "\n".join(lines) + "\n"
|
||||
)
|
||||
return import_text(store, text, label=label or f"{browser}", cookie_dir=cookie_dir)
|
||||
|
||||
|
||||
def set_active(store: Store, cookie_id: str) -> None:
|
||||
store.set_active_cookie(cookie_id)
|
||||
|
||||
@@ -143,12 +362,18 @@ def delete(store: Store, cookie_id: str, cookie_dir: str | Path | None = None) -
|
||||
|
||||
|
||||
def resolve_active_path(store: Store, cookie_dir: str | Path | None = None) -> str | None:
|
||||
row = store.get_active_cookie()
|
||||
if not row:
|
||||
return None
|
||||
cdir = cookies_dir(cookie_dir)
|
||||
row = store.get_active_cookie()
|
||||
if row:
|
||||
path = cdir / row.filename
|
||||
return str(path) if path.exists() else None
|
||||
if path.exists():
|
||||
return str(path)
|
||||
for c in store.list_cookies():
|
||||
path = cdir / c.filename
|
||||
if path.exists():
|
||||
store.set_active_cookie(c.id)
|
||||
return str(path)
|
||||
return None
|
||||
|
||||
|
||||
def is_expired(row: CookieRow) -> bool:
|
||||
|
||||
+244
-8
@@ -1,31 +1,76 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
from typing import Any
|
||||
from dataclasses import dataclass, field
|
||||
from datetime import datetime, timezone
|
||||
from typing import Any, Callable, Collection
|
||||
|
||||
import yt_dlp
|
||||
|
||||
from .ratelimit import GLOBAL_PACER, ydl_throttle_opts
|
||||
from .store import VideoRef
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
|
||||
def discover_channel(channel_url: str, sleep_subrequests: float = 2.0) -> tuple[str, str, list[VideoRef]]:
|
||||
"""Returns (channel_id, channel_name, video_refs)."""
|
||||
def _entry_upload_date(entry: dict[str, Any]) -> tuple[str | None, int]:
|
||||
"""(YYYYMMDD | None, approx_flag) para un entry flat.
|
||||
|
||||
Con `youtubetab:approximate_date` yt-dlp llena `timestamp` parseando el
|
||||
texto relativo que YouTube ya manda en el listado ("hace 3 semanas"): la
|
||||
fecha sale gratis, con la precision del texto (dia para recientes, mas
|
||||
gruesa para antiguos). Un upload_date crudo siempre gana: es exacto.
|
||||
"""
|
||||
exact = entry.get("upload_date")
|
||||
if exact:
|
||||
return str(exact), 0
|
||||
ts = entry.get("timestamp")
|
||||
if not ts:
|
||||
return None, 0
|
||||
return datetime.fromtimestamp(int(ts), tz=timezone.utc).strftime("%Y%m%d"), 1
|
||||
|
||||
# Defaults for incremental sync; `SyncConfig` in config.py is the tunable copy.
|
||||
DEFAULT_SYNC_WINDOW = 30
|
||||
DEFAULT_MAX_WINDOW = 300
|
||||
DEFAULT_OVERLAP = 3
|
||||
|
||||
|
||||
def discover_channel(
|
||||
channel_url: str,
|
||||
sleep_subrequests: float = 2.0,
|
||||
limit: int | None = None,
|
||||
) -> tuple[str, str, str | None, list[VideoRef]]:
|
||||
"""Returns (channel_id, channel_name, avatar_url, video_refs).
|
||||
|
||||
`limit` caps how many entries yt-dlp pulls off the playlist. It is not a
|
||||
post-filter: yt-dlp stops requesting continuation pages once it has enough,
|
||||
so a small limit is the difference between one request and dozens.
|
||||
"""
|
||||
ydl_opts: dict[str, Any] = {
|
||||
"extract_flat": "in_playlist",
|
||||
"quiet": True,
|
||||
"no_warnings": True,
|
||||
"skip_download": True,
|
||||
"extract_flat_args": None,
|
||||
"sleep_subrequests": sleep_subrequests,
|
||||
# Parse the relative time text ("3 weeks ago") YouTube already includes
|
||||
# in the listing into an approximate `timestamp` per entry. Rides the
|
||||
# same requests: zero extra calls.
|
||||
"extractor_args": {"youtubetab": {"approximate_date": ["true"]}},
|
||||
**ydl_throttle_opts(sleep_subrequests),
|
||||
}
|
||||
if limit:
|
||||
# Only honoured because extract_info processes the result below. With
|
||||
# `process=False` yt-dlp hands back a lazy generator that ignores
|
||||
# playlistend, and walking it paginates the entire channel: measured on
|
||||
# a 2564-video channel, 2 requests versus 86.
|
||||
ydl_opts["playlistend"] = int(limit)
|
||||
|
||||
GLOBAL_PACER.wait()
|
||||
with yt_dlp.YoutubeDL(ydl_opts) as ydl:
|
||||
info = ydl.extract_info(channel_url, download=False)
|
||||
|
||||
channel_id = info.get("channel_id") or info.get("uploader_id") or info.get("id") or "unknown"
|
||||
channel_name = info.get("channel") or info.get("title") or info.get("uploader") or "Unknown"
|
||||
avatar = _pick_channel_avatar(info)
|
||||
|
||||
entries = _flatten_entries(info)
|
||||
refs: list[VideoRef] = []
|
||||
@@ -36,22 +81,213 @@ def discover_channel(channel_url: str, sleep_subrequests: float = 2.0) -> tuple[
|
||||
if entry.get("_type") == "playlist":
|
||||
continue
|
||||
url = entry.get("url") or f"https://www.youtube.com/watch?v={video_id}"
|
||||
upload_date = entry.get("upload_date")
|
||||
upload_date, date_approx = _entry_upload_date(entry)
|
||||
duration = entry.get("duration")
|
||||
title = entry.get("title") or video_id
|
||||
# Flat entries report availability ("subscriber_only" for members-only
|
||||
# content), so a video we could never fetch is identifiable from the
|
||||
# listing itself instead of costing a failed extraction to discover.
|
||||
availability = entry.get("availability")
|
||||
refs.append(
|
||||
VideoRef(
|
||||
video_id=str(video_id),
|
||||
channel_id=str(channel_id),
|
||||
title=str(title),
|
||||
url=str(url),
|
||||
upload_date=str(upload_date) if upload_date else None,
|
||||
upload_date=upload_date,
|
||||
date_approx=date_approx,
|
||||
duration=int(duration) if duration else None,
|
||||
availability=str(availability) if availability else None,
|
||||
# The tab is reverse-chronological and flat entries carry no
|
||||
# upload_date, so this index is the only thing that says how old
|
||||
# an un-extracted video is. Recorded verbatim; later filtering
|
||||
# (shorts/live, `since`) leaves gaps but never reorders.
|
||||
position=len(refs),
|
||||
)
|
||||
)
|
||||
|
||||
log.info("Discovered %d videos on channel %s (%s)", len(refs), channel_name, channel_id)
|
||||
return str(channel_id), str(channel_name), refs
|
||||
return str(channel_id), str(channel_name), avatar, refs
|
||||
|
||||
|
||||
@dataclass
|
||||
class IncrementalDiscovery:
|
||||
"""Outcome of a windowed channel sync."""
|
||||
|
||||
channel_id: str
|
||||
channel_name: str
|
||||
avatar: str | None
|
||||
refs: list[VideoRef] = field(default_factory=list) # window fetched, newest first
|
||||
new_refs: list[VideoRef] = field(default_factory=list) # subset absent from known_ids
|
||||
fetched: int = 0 # entries yt-dlp actually returned on the last pass
|
||||
passes: int = 0 # flat extractions performed
|
||||
window: int = 0 # playlistend used on the last pass
|
||||
caught_up: bool = False # reached videos we already had (or the channel's end)
|
||||
exhausted: bool = False # the window covered the entire channel
|
||||
full_scan: bool = False # we deliberately walked everything
|
||||
|
||||
@property
|
||||
def new_count(self) -> int:
|
||||
return len(self.new_refs)
|
||||
|
||||
|
||||
def discover_incremental(
|
||||
channel_url: str,
|
||||
known_ids: Collection[str],
|
||||
*,
|
||||
sleep_subrequests: float = 2.0,
|
||||
window: int = DEFAULT_SYNC_WINDOW,
|
||||
max_window: int = DEFAULT_MAX_WINDOW,
|
||||
overlap: int = DEFAULT_OVERLAP,
|
||||
since: str | None = None,
|
||||
keep: Callable[[VideoRef], bool] | None = None,
|
||||
) -> IncrementalDiscovery:
|
||||
"""Fetch only the newest slice of a channel instead of paginating all of it.
|
||||
|
||||
A channel's /videos tab is reverse-chronological, so once `overlap`
|
||||
consecutive entries are ones we already have, everything older is already in
|
||||
the DB and further pages buy nothing but requests against YouTube's limits.
|
||||
|
||||
`since` (YYYYMMDD) is a second, best-effort stop condition for the case where
|
||||
yt-dlp does report upload dates — the YouTube tab extractor usually does not
|
||||
populate them in flat mode, which is why the id overlap is what actually
|
||||
terminates the walk.
|
||||
|
||||
`keep` filters each window before the overlap test, so it must match whatever
|
||||
filter decided which videos reached the DB (shorts/live); otherwise the tail
|
||||
would be full of entries that could never be "known" and the window would
|
||||
widen pointlessly.
|
||||
"""
|
||||
known = set(known_ids)
|
||||
overlap = max(1, int(overlap))
|
||||
window = max(1, int(window))
|
||||
max_window = max(window, int(max_window))
|
||||
|
||||
def _apply(raw: list[VideoRef]) -> list[VideoRef]:
|
||||
return [r for r in raw if keep(r)] if keep else list(raw)
|
||||
|
||||
# No prior state means there is no boundary to find — walk the whole channel.
|
||||
if not known:
|
||||
cid, name, avatar, raw = discover_channel(channel_url, sleep_subrequests=sleep_subrequests)
|
||||
refs = _apply(raw)
|
||||
log.info("Full discovery of %s: %d videos", channel_url, len(refs))
|
||||
return IncrementalDiscovery(
|
||||
channel_id=cid, channel_name=name, avatar=avatar,
|
||||
refs=refs, new_refs=list(refs), fetched=len(raw), passes=1,
|
||||
window=len(raw), caught_up=True, exhausted=True, full_scan=True,
|
||||
)
|
||||
|
||||
size = window
|
||||
passes = 0
|
||||
while True:
|
||||
cid, name, avatar, raw = discover_channel(
|
||||
channel_url, sleep_subrequests=sleep_subrequests, limit=size
|
||||
)
|
||||
passes += 1
|
||||
exhausted = len(raw) < size
|
||||
refs = _apply(raw)
|
||||
tail = refs[-overlap:]
|
||||
|
||||
caught_up = exhausted or (bool(tail) and all(r.video_id in known for r in tail))
|
||||
if not caught_up and since and tail:
|
||||
dated = [r.upload_date for r in tail if r.upload_date]
|
||||
caught_up = bool(dated) and all(d < since for d in dated)
|
||||
|
||||
if caught_up or size >= max_window:
|
||||
new_refs = [r for r in refs if r.video_id not in known]
|
||||
log.info(
|
||||
"Incremental sync of %s: %d fetched over %d pass(es), %d new%s",
|
||||
channel_url, len(raw), passes, len(new_refs),
|
||||
"" if caught_up else " (window ceiling hit; older videos not checked)",
|
||||
)
|
||||
return IncrementalDiscovery(
|
||||
channel_id=cid, channel_name=name, avatar=avatar,
|
||||
refs=refs, new_refs=new_refs, fetched=len(raw), passes=passes,
|
||||
window=size, caught_up=caught_up, exhausted=exhausted,
|
||||
)
|
||||
|
||||
# Everything in the window was new: the channel published more than we
|
||||
# looked at, so widen and try again rather than miss uploads.
|
||||
size = min(size * 2, max_window)
|
||||
|
||||
|
||||
def extract_handle(url: str) -> str:
|
||||
"""@handle embedded in a channel URL, or "" for /channel/<id> URLs.
|
||||
|
||||
Every caller that records a channel (CLI, webapp add-channel, jobs, watch)
|
||||
normalises the handle the same way; this is that one shared definition.
|
||||
"""
|
||||
if "@" in url:
|
||||
return "@" + url.split("@", 1)[1].split("/", 1)[0]
|
||||
return ""
|
||||
|
||||
|
||||
def _pick_channel_avatar(info: dict[str, Any]) -> str | None:
|
||||
"""Best-effort channel avatar URL from a yt-dlp channel info dict.
|
||||
|
||||
Only sources that ACTUALLY point to an image are considered. Notably we do
|
||||
NOT fall back to `channel_url` or `thumbnail` (those are page URLs / banners).
|
||||
"""
|
||||
ths = info.get("thumbnails") or []
|
||||
if isinstance(ths, list):
|
||||
for t in reversed(ths):
|
||||
url = (t.get("url") if isinstance(t, dict) else None)
|
||||
if isinstance(url, str) and url:
|
||||
return url
|
||||
for key in ("avatar", "channel_icon"):
|
||||
v = info.get(key)
|
||||
if isinstance(v, str) and v.startswith("http"):
|
||||
return v
|
||||
if isinstance(v, dict):
|
||||
u = v.get("url")
|
||||
if isinstance(u, str) and u:
|
||||
return u
|
||||
return None
|
||||
|
||||
|
||||
_CHANNEL_PATH_TAILS = (
|
||||
"/videos", "/shorts", "/streams", "/featured", "/playlists",
|
||||
"/community", "/about", "/channels",
|
||||
)
|
||||
|
||||
|
||||
def _channel_root_url(channel_url: str) -> str:
|
||||
"""Strip a tab suffix from a YouTube channel URL so yt-dlp extracts the
|
||||
channel home (where the avatar reliably lives), not a tab."""
|
||||
u = (channel_url or "").rstrip("/")
|
||||
for tail in _CHANNEL_PATH_TAILS:
|
||||
if u.endswith(tail):
|
||||
u = u[: -len(tail)]
|
||||
break
|
||||
return u or channel_url
|
||||
|
||||
|
||||
def deep_channel_avatar(channel_url: str, sleep_subrequests: float = 2.0) -> str | None:
|
||||
"""Robust avatar recovery via a yt-dlp call on the channel root URL.
|
||||
|
||||
`playlistend` and `extract_flat` are load-bearing, not tuning. Without them
|
||||
this asked yt-dlp to fully extract every video the channel has ever
|
||||
published in order to read one image URL: yt-dlp redirects a bare channel
|
||||
URL back to /videos, then walks /videos, /streams and /shorts, and
|
||||
`download=False` suppresses only the media download, not the extraction.
|
||||
Measured against a 2564-video channel it was still going at 735 requests
|
||||
when the measurement aborted it; the form below costs 4.
|
||||
"""
|
||||
ydl_opts: dict[str, Any] = {
|
||||
"quiet": True, "no_warnings": True, "skip_download": True,
|
||||
"extract_flat": "in_playlist",
|
||||
"playlistend": 1,
|
||||
**ydl_throttle_opts(sleep_subrequests),
|
||||
}
|
||||
target = _channel_root_url(channel_url)
|
||||
# yt-dlp probes /videos, /streams and /shorts on a bare channel URL.
|
||||
GLOBAL_PACER.wait(cost=3)
|
||||
try:
|
||||
with yt_dlp.YoutubeDL(ydl_opts) as ydl:
|
||||
info = ydl.extract_info(target, download=False)
|
||||
except Exception:
|
||||
return None
|
||||
return _pick_channel_avatar(info or {})
|
||||
|
||||
|
||||
def _flatten_entries(info: dict[str, Any]) -> list[dict[str, Any]]:
|
||||
|
||||
+366
-28
@@ -1,13 +1,19 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import logging
|
||||
import os
|
||||
import re
|
||||
import urllib.request
|
||||
from dataclasses import dataclass
|
||||
from typing import Any
|
||||
from pathlib import Path
|
||||
from typing import Any, Mapping
|
||||
|
||||
import requests
|
||||
import yt_dlp
|
||||
|
||||
from ._yt_http import yt_get
|
||||
from .parse import Segment, parse_auto_dump
|
||||
from .ratelimit import GLOBAL_PACER, ydl_throttle_opts
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
@@ -26,66 +32,327 @@ class VideoData:
|
||||
segments: list[Segment]
|
||||
subtitle: SubtitlePick | None
|
||||
has_chapters: bool
|
||||
# Why `segments` came back empty. "No transcript" has several very
|
||||
# different causes — the video genuinely has no captions, the language
|
||||
# policy rejected the tracks that do exist, or the download was throttled —
|
||||
# and collapsing them into one terminal status hides recoverable failures.
|
||||
skip_reason: str | None = None
|
||||
|
||||
|
||||
_LANGUAGE_MODES = ("manual", "auto", "any")
|
||||
|
||||
|
||||
def _sources_for(mode: str, prefer_manual: bool, manual: dict, auto: dict) -> list[tuple[str, dict]]:
|
||||
"""Return ordered list of (label, tracks_dict) to try for ``mode``.
|
||||
|
||||
``any`` defers to the legacy :data:`prefer_manual` global default.
|
||||
``manual`` / ``auto`` force the track family even if the other has
|
||||
a higher-priority language elsewhere in the iteration.
|
||||
"""
|
||||
if mode == "manual":
|
||||
return [("manual", manual)]
|
||||
if mode == "auto":
|
||||
return [("auto", auto)]
|
||||
if prefer_manual:
|
||||
return [("manual", manual), ("auto", auto)]
|
||||
return [("auto", auto), ("manual", manual)]
|
||||
|
||||
|
||||
def extract_video(
|
||||
video_url: str,
|
||||
languages: list[str],
|
||||
languages: Mapping[str, str] | list[str],
|
||||
retries: int = 10,
|
||||
sleep_subrequests: float = 2.0,
|
||||
prefer_manual: bool = True,
|
||||
cookies_file: str | None = None,
|
||||
cookies_from_browser: str | None = None,
|
||||
extractor_retries: int = 3,
|
||||
socket_timeout: float = 30.0,
|
||||
) -> VideoData:
|
||||
"""Run yt-dlp on ``video_url`` and pull the preferred subtitle track.
|
||||
|
||||
``languages`` may be either a list (legacy, every entry uses
|
||||
``prefer_manual``) or a mapping ``{lang: mode}`` where ``mode`` is one
|
||||
of ``"manual"``, ``"auto"`` or ``"any"``. The mapping form is the
|
||||
preferred interface because it lets you mix per-language policies such
|
||||
as ``{"en": "manual", "es": "auto", "pt": "any"}``.
|
||||
"""
|
||||
languages_dict = _coerce_languages(languages, prefer_manual)
|
||||
|
||||
ydl_opts: dict[str, Any] = {
|
||||
"writesubtitles": True,
|
||||
"writeautomaticsub": True,
|
||||
"subtitleslangs": languages,
|
||||
"subtitleslangs": list(languages_dict.keys()),
|
||||
"skip_download": True,
|
||||
"quiet": True,
|
||||
"no_warnings": True,
|
||||
"retries": retries,
|
||||
"sleep_subrequests": sleep_subrequests,
|
||||
"noprogress": True,
|
||||
# Never let yt-dlp probe formats: it costs one HTTP request per format
|
||||
# and we only ever want captions and metadata.
|
||||
"check_formats": None,
|
||||
**ydl_throttle_opts(
|
||||
sleep_subrequests,
|
||||
extractor_retries=extractor_retries,
|
||||
socket_timeout=socket_timeout,
|
||||
),
|
||||
}
|
||||
if cookies_file:
|
||||
ydl_opts["cookiefile"] = cookies_file
|
||||
if cookies_from_browser:
|
||||
ydl_opts["cookiesfrombrowser"] = (cookies_from_browser,)
|
||||
|
||||
# One video extraction is two requests: the watch page and the InnerTube
|
||||
# player call. yt-dlp spaces them itself; the pacer needs to know they exist.
|
||||
GLOBAL_PACER.wait(cost=2)
|
||||
try:
|
||||
with yt_dlp.YoutubeDL(ydl_opts) as ydl:
|
||||
info = ydl.extract_info(video_url, download=False)
|
||||
except Exception:
|
||||
# Logged-in sessions on current YouTube increasingly end here
|
||||
# ("The page needs to be reloaded" / format-availability failures).
|
||||
# The session itself is usually fine — the watch page still hands
|
||||
# metadata and caption tracks to a plain cookie'd GET — so try that
|
||||
# before giving up. Without cookies there is nothing to fall back to.
|
||||
if cookies_file:
|
||||
fallback = extract_via_watch_page(video_url, cookies_file, languages_dict, prefer_manual)
|
||||
if fallback is not None:
|
||||
return fallback
|
||||
raise
|
||||
|
||||
pick = pick_subtitle(info, languages, prefer_manual)
|
||||
pick = pick_subtitle(info, languages_dict, prefer_manual)
|
||||
segments: list[Segment] = []
|
||||
skip_reason: str | None = None
|
||||
if pick:
|
||||
raw = _download_subtitle(pick.url)
|
||||
raw, dl_error = _download_subtitle(pick.url)
|
||||
if raw:
|
||||
segments = parse_auto_dump(raw)
|
||||
if not segments:
|
||||
log.warning("Could not parse subtitle for %s (format=%s)", video_url, pick.ext)
|
||||
skip_reason = f"subtitle downloaded but parsed empty (lang={pick.lang}, format={pick.ext})"
|
||||
else:
|
||||
skip_reason = (
|
||||
f"subtitle track found (lang={pick.lang}, {pick.source}) but the download failed "
|
||||
f"— usually throttling; retry later [{dl_error}]"
|
||||
)
|
||||
else:
|
||||
skip_reason = describe_missing_subtitle(info, languages_dict)
|
||||
|
||||
has_chapters = bool(info.get("chapters"))
|
||||
return VideoData(info=info, segments=segments, subtitle=pick, has_chapters=has_chapters)
|
||||
return VideoData(
|
||||
info=info, segments=segments, subtitle=pick,
|
||||
has_chapters=has_chapters, skip_reason=skip_reason,
|
||||
)
|
||||
|
||||
|
||||
def pick_subtitle(info: dict[str, Any], languages: list[str], prefer_manual: bool = True) -> SubtitlePick | None:
|
||||
_WATCH_PAGE_UA = (
|
||||
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
|
||||
"(KHTML, like Gecko) Chrome/152.0.7977.64 Safari/537.36"
|
||||
)
|
||||
|
||||
# Precompiled: the fallback can fire on every video of a throttled batch, and
|
||||
# re-compiling per call showed up under those runs.
|
||||
_INITIAL_PLAYER_RE = re.compile(r"ytInitialPlayerResponse\s*=\s*(\{.+?\})\s*;")
|
||||
|
||||
# One opener per (cookie file, mtime): the jar parse is per-call work that is
|
||||
# pure waste inside a batch. Keyed on mtime so a re-imported cookie file under
|
||||
# the same path still gets a fresh jar; only the newest entry is kept.
|
||||
_OPENER_CACHE: dict[tuple[str, float], urllib.request.OpenerDirector] = {}
|
||||
|
||||
|
||||
def _session_urlopen(cookies_file: str, url: str, *, timeout: float = 20.0):
|
||||
path = str(Path(cookies_file).resolve())
|
||||
try:
|
||||
mtime = os.path.getmtime(path)
|
||||
except OSError:
|
||||
mtime = -1.0
|
||||
key = (path, mtime)
|
||||
opener = _OPENER_CACHE.get(key)
|
||||
if opener is None:
|
||||
jar = yt_dlp.cookies.YoutubeDLCookieJar(cookies_file)
|
||||
jar.load(ignore_discard=True, ignore_expires=True)
|
||||
opener = urllib.request.build_opener(urllib.request.HTTPCookieProcessor(jar))
|
||||
opener.addheaders = [
|
||||
("User-Agent", _WATCH_PAGE_UA),
|
||||
("Accept-Language", "es-ES,es;q=0.9,en;q=0.8"),
|
||||
]
|
||||
_OPENER_CACHE.clear()
|
||||
_OPENER_CACHE[key] = opener
|
||||
return opener.open(url, timeout=timeout)
|
||||
|
||||
|
||||
def extract_via_watch_page(
|
||||
video_url: str,
|
||||
cookies_file: str,
|
||||
languages: Mapping[str, str],
|
||||
prefer_manual: bool = True,
|
||||
) -> VideoData | None:
|
||||
"""Session-cookie fallback: scrape the watch page directly.
|
||||
|
||||
yt-dlp's InnerTube clients reject logged-in sessions that lack a PO
|
||||
token (playability "The page needs to be reloaded") or return no
|
||||
formats/captions, which kills cookie-authenticated videos — members
|
||||
being the case this exists for. The plain watch page served to the
|
||||
logged-in browser still carries `ytInitialPlayerResponse` with
|
||||
metadata and caption tracks, so GET it with the vault cookie and
|
||||
reuse the normal subtitle picker. Returns None when the page holds
|
||||
no caption tracks at all, so callers keep their own error semantics.
|
||||
"""
|
||||
GLOBAL_PACER.wait()
|
||||
html = _session_urlopen(cookies_file, video_url).read().decode("utf-8", "replace")
|
||||
m = _INITIAL_PLAYER_RE.search(html)
|
||||
if not m:
|
||||
log.warning("watch-page fallback: no ytInitialPlayerResponse for %s", video_url)
|
||||
return None
|
||||
try:
|
||||
pr = json.loads(m.group(1))
|
||||
except json.JSONDecodeError:
|
||||
log.warning("watch-page fallback: unparseable player response for %s", video_url)
|
||||
return None
|
||||
|
||||
status = (pr.get("playabilityStatus") or {}).get("status")
|
||||
if status != "OK":
|
||||
reason = (pr.get("playabilityStatus") or {}).get("reason") or status
|
||||
raise RuntimeError(f"watch-page fallback: video not playable ({reason})")
|
||||
|
||||
details = pr.get("videoDetails") or {}
|
||||
micro = (pr.get("microformat") or {}).get("playerMicroformatRenderer") or {}
|
||||
tracks = (
|
||||
(pr.get("captions") or {}).get("playerCaptionsTracklistRenderer") or {}
|
||||
).get("captionTracks") or []
|
||||
if not tracks:
|
||||
return None
|
||||
|
||||
# Reuse pick_subtitle by shaping the tracks as an info dict.
|
||||
info: dict[str, Any] = {
|
||||
"title": details.get("title"),
|
||||
"channel": details.get("author"),
|
||||
"duration": int(details["lengthSeconds"]) if str(details.get("lengthSeconds", "")).isdigit() else None,
|
||||
"view_count": int(details["viewCount"]) if str(details.get("viewCount", "")).isdigit() else None,
|
||||
"description": details.get("shortDescription") or "",
|
||||
"tags": details.get("keywords") or [],
|
||||
"thumbnail": (details.get("thumbnail") or {}).get("thumbnails", [{}])[-1].get("url"),
|
||||
"upload_date": (micro.get("publishDate") or micro.get("uploadDate") or "").replace("-", "") or None,
|
||||
# The watch page carries no availability signal; leaving it unset
|
||||
# keeps the (more informed) discovery value in the store.
|
||||
"availability": None,
|
||||
"subtitles": {},
|
||||
"automatic_captions": {},
|
||||
}
|
||||
for t in tracks:
|
||||
base = t.get("baseUrl") or ""
|
||||
if not base:
|
||||
continue
|
||||
entry = [{"ext": "json3", "url": base + ("&" if "?" in base else "?") + "fmt=json3"}]
|
||||
if t.get("kind") == "asr":
|
||||
info["automatic_captions"].setdefault(t.get("languageCode", ""), []).extend(entry)
|
||||
else:
|
||||
info["subtitles"].setdefault(t.get("languageCode", ""), []).extend(entry)
|
||||
|
||||
pick = pick_subtitle(info, languages, prefer_manual)
|
||||
segments: list[Segment] = []
|
||||
skip_reason: str | None = None
|
||||
if pick:
|
||||
try:
|
||||
GLOBAL_PACER.wait()
|
||||
raw = _session_urlopen(cookies_file, pick.url).read().decode("utf-8", "replace")
|
||||
segments = parse_auto_dump(raw)
|
||||
if not segments:
|
||||
skip_reason = "subtitle downloaded but parsed empty (watch-page fallback)"
|
||||
except Exception as exc: # pylint: disable=broad-except
|
||||
skip_reason = f"caption download failed via watch-page fallback: {exc}"
|
||||
else:
|
||||
skip_reason = describe_missing_subtitle(info, languages)
|
||||
|
||||
return VideoData(
|
||||
info=info, segments=segments, subtitle=pick,
|
||||
has_chapters=bool(info.get("chapters")), skip_reason=skip_reason,
|
||||
)
|
||||
|
||||
|
||||
def describe_missing_subtitle(info: dict[str, Any], languages: Mapping[str, str]) -> str:
|
||||
"""Explain why no track matched, distinguishing 'none exist' from 'policy rejected them'.
|
||||
|
||||
A channel that only publishes auto-generated captions scanned under a
|
||||
manual-only policy yields nothing — which is a config problem, not a
|
||||
property of the video, and the message has to say so.
|
||||
"""
|
||||
manual = {k: v for k, v in (info.get("subtitles") or {}).items() if v}
|
||||
auto = {k: v for k, v in (info.get("automatic_captions") or {}).items() if v}
|
||||
if not manual and not auto:
|
||||
return "no caption tracks published for this video"
|
||||
|
||||
wanted = ", ".join(f"{lang}={mode}" for lang, mode in languages.items()) or "(none configured)"
|
||||
modes = {str(m).lower() for m in languages.values()}
|
||||
parts = [f"no track matched the language policy ({wanted})"]
|
||||
parts.append(f"available: {len(manual)} manual, {len(auto)} auto")
|
||||
if auto and not manual and modes == {"manual"}:
|
||||
parts.append(
|
||||
"this video has ONLY auto-generated captions — set the language mode "
|
||||
"to 'any' or 'auto' to use them"
|
||||
)
|
||||
return "; ".join(parts)
|
||||
|
||||
|
||||
def _coerce_languages(languages: Mapping[str, str] | list[str] | None,
|
||||
prefer_manual: bool) -> dict[str, str]:
|
||||
"""Normalise legacy list / new dict / None into ``{lang: mode}``."""
|
||||
default = "manual" if prefer_manual else "auto"
|
||||
if languages is None:
|
||||
return {}
|
||||
if isinstance(languages, Mapping):
|
||||
out: dict[str, str] = {}
|
||||
for lang, mode in languages.items():
|
||||
m = str(mode).lower().strip()
|
||||
if m not in _LANGUAGE_MODES:
|
||||
m = "any"
|
||||
out[str(lang)] = m
|
||||
return out
|
||||
if isinstance(languages, (list, tuple)):
|
||||
return {str(l): default for l in languages}
|
||||
return {}
|
||||
|
||||
|
||||
def pick_subtitle(info: dict[str, Any],
|
||||
languages: Mapping[str, str],
|
||||
prefer_manual: bool = True) -> SubtitlePick | None:
|
||||
"""Pick the best subtitle track for ``info`` honouring per-language mode.
|
||||
|
||||
See :func:`extract_video` for the ``languages`` schema. ``prefer_manual``
|
||||
is only consulted for entries whose mode is ``"any"``.
|
||||
"""
|
||||
manual = info.get("subtitles") or {}
|
||||
auto = info.get("automatic_captions") or {}
|
||||
|
||||
ordered_sources: list[tuple[str, dict[str, Any]]]
|
||||
if prefer_manual:
|
||||
ordered_sources = [("manual", manual), ("auto", auto)]
|
||||
else:
|
||||
ordered_sources = [("auto", auto), ("manual", manual)]
|
||||
if isinstance(languages, (list, tuple)):
|
||||
# legacy path: convert on the fly
|
||||
default = "manual" if prefer_manual else "auto"
|
||||
languages = {l: default for l in languages}
|
||||
|
||||
for source_label, tracks in ordered_sources:
|
||||
for lang in languages:
|
||||
# Config order is a preference between languages we can read, not an
|
||||
# instruction to accept a machine translation when the real transcript is
|
||||
# sitting right there. An English channel scanned under {es, es-419, en}
|
||||
# was yielding Spanish auto-translations of English speech.
|
||||
#
|
||||
# Two passes rather than a reorder. The reorder alone needed to know the
|
||||
# spoken language, and when neither an `-orig` key nor `info["language"]`
|
||||
# was present it silently fell back to config order and reintroduced the
|
||||
# bug. Rejecting translations outright in the first pass needs no such
|
||||
# knowledge: whatever language it lands on, it is the one actually spoken.
|
||||
ordered = list(languages.items())
|
||||
spoken = original_language(info)
|
||||
if spoken and any(_normalize_lang(l) == spoken for l, _ in ordered):
|
||||
ordered.sort(key=lambda kv: _normalize_lang(kv[0]) != spoken)
|
||||
|
||||
for allow_translations in (False, True):
|
||||
for lang, mode in ordered:
|
||||
if mode not in _LANGUAGE_MODES:
|
||||
mode = "any"
|
||||
sources = _sources_for(mode, prefer_manual, manual, auto)
|
||||
normalized = _normalize_lang(lang)
|
||||
for track_lang, formats in tracks.items():
|
||||
if _normalize_lang(track_lang) != normalized:
|
||||
continue
|
||||
if not formats:
|
||||
for source_label, tracks in sources:
|
||||
for track_lang, formats in _ordered_tracks(tracks, normalized):
|
||||
if not allow_translations and _is_translation(track_lang, formats):
|
||||
continue
|
||||
pick = _pick_best_format(formats)
|
||||
if pick:
|
||||
@@ -93,7 +360,7 @@ def pick_subtitle(info: dict[str, Any], languages: list[str], prefer_manual: boo
|
||||
url=pick["url"],
|
||||
ext=pick["ext"],
|
||||
lang=track_lang,
|
||||
source="manual" if source_label == "manual" else "auto",
|
||||
source=source_label,
|
||||
)
|
||||
return None
|
||||
|
||||
@@ -115,11 +382,82 @@ def _normalize_lang(code: str) -> str:
|
||||
return base
|
||||
|
||||
|
||||
def _download_subtitle(url: str) -> str | None:
|
||||
try:
|
||||
resp = requests.get(url, timeout=15, headers={"User-Agent": "Mozilla/5.0"})
|
||||
resp.raise_for_status()
|
||||
return resp.text
|
||||
except requests.RequestException as exc:
|
||||
log.error("Failed to download subtitle from %s: %s", url, exc)
|
||||
def _is_original_track(code: str) -> bool:
|
||||
"""True for YouTube's original-ASR track, which it suffixes with ``-orig``.
|
||||
|
||||
YouTube publishes the speech-recognised track as ``<lang>-orig`` and then a
|
||||
long tail of machine translations keyed by bare language code — including a
|
||||
translation *into the video's own language*. So on a Spanish video both
|
||||
``es-orig`` and ``es`` exist, and only the first is the real transcript.
|
||||
"""
|
||||
return code.replace("_", "-").lower().endswith("-orig")
|
||||
|
||||
|
||||
def original_language(info: dict[str, Any]) -> str | None:
|
||||
"""The language actually spoken in the video, normalised, or None.
|
||||
|
||||
Prefers the ``-orig`` track that YouTube itself publishes over ``info`` keys,
|
||||
because the ``-orig`` suffix is direct evidence from the caption list while
|
||||
``language`` is metadata that YouTube localises along with the title.
|
||||
"""
|
||||
for code in (info.get("automatic_captions") or {}):
|
||||
if _is_original_track(code):
|
||||
return _normalize_lang(code)
|
||||
for key in ("language", "original_language"):
|
||||
val = info.get(key)
|
||||
if isinstance(val, str) and val:
|
||||
return _normalize_lang(val)
|
||||
return None
|
||||
|
||||
|
||||
def _is_translation(code: str, formats: list[dict[str, Any]] | None) -> bool:
|
||||
"""True when this track is YouTube machine-translating some other track.
|
||||
|
||||
The caption URL says so outright: yt-dlp builds a translated track by
|
||||
appending ``tlang=`` to the base track's URL and omits it when the target
|
||||
equals the source language. That is direct evidence, unlike the ``-orig``
|
||||
naming convention, and it is what lets us reject a translation even for a
|
||||
video whose spoken language we could not otherwise determine.
|
||||
"""
|
||||
if _is_original_track(code):
|
||||
return False
|
||||
for fmt in formats or []:
|
||||
url = fmt.get("url") or ""
|
||||
if "tlang=" in url:
|
||||
return True
|
||||
return False
|
||||
|
||||
|
||||
def _ordered_tracks(tracks: dict, normalized: str) -> list[tuple[str, Any]]:
|
||||
"""Tracks matching `normalized`, original-ASR first.
|
||||
|
||||
Without this the picker took whichever key yt-dlp happened to list first.
|
||||
That silently returned the right thing on Spanish channels (``es-orig``
|
||||
sorts before ``es``) and the wrong thing everywhere else.
|
||||
"""
|
||||
matches = [(c, f) for c, f in tracks.items() if _normalize_lang(c) == normalized and f]
|
||||
matches.sort(key=lambda kv: not _is_original_track(kv[0]))
|
||||
return matches
|
||||
|
||||
|
||||
def _download_subtitle(url: str, *, timeout: float = 15.0) -> tuple[str | None, str | None]:
|
||||
"""Fetch a caption track. Returns (text, error_description).
|
||||
|
||||
The error text is returned rather than only logged because a 429 here is
|
||||
how YouTube throttling most often shows up on this path, and the circuit
|
||||
breaker upstream can only see it if it survives into the stored reason.
|
||||
"""
|
||||
try:
|
||||
GLOBAL_PACER.wait()
|
||||
resp = yt_get(url, timeout=timeout)
|
||||
resp.raise_for_status()
|
||||
return resp.text, None
|
||||
except Exception as exc: # pylint: disable=broad-except
|
||||
log.error("Failed to download subtitle from %s: %s", url, exc)
|
||||
# Lead with a normalised "HTTP Error <status>" token. Downstream, the
|
||||
# throttle detector has to recognise a 429 here, and depending on the
|
||||
# prose is fragile: a 429 served without a reason phrase (routine over
|
||||
# HTTP/2) says nothing about "too many requests".
|
||||
status = getattr(getattr(exc, "response", None), "status_code", None)
|
||||
prefix = f"HTTP Error {status}: " if status else ""
|
||||
return None, f"{prefix}{type(exc).__name__}: {exc}"
|
||||
|
||||
+48
-15
@@ -4,11 +4,11 @@ import logging
|
||||
import time
|
||||
from typing import Callable
|
||||
|
||||
from .config import Config
|
||||
from .config import Config, keep_ref
|
||||
from .cookies import resolve_active_path
|
||||
from .discover import discover_channel
|
||||
from .discover import discover_incremental, extract_handle
|
||||
from .pipeline import process_video
|
||||
from .ratelimit import polite_sleep
|
||||
from .ratelimit import ThrottleGuard, polite_sleep
|
||||
from .store import Store
|
||||
|
||||
|
||||
@@ -65,13 +65,34 @@ def _run_once(
|
||||
if on_log:
|
||||
on_log(msg)
|
||||
|
||||
channel_id_found, channel_name, refs = discover_channel(
|
||||
cfg.channel_url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests
|
||||
# A watch loop re-runs on an interval; walking the whole channel every tick
|
||||
# is exactly what burns through YouTube's tolerance. Only look at what is
|
||||
# newer than the last video we already have.
|
||||
known = store.known_video_ids(channel_id) if (channel_id and cfg.sync.incremental) else set()
|
||||
result = discover_incremental(
|
||||
cfg.channel_url,
|
||||
known,
|
||||
sleep_subrequests=cfg.yt_dlp.sleep_subrequests,
|
||||
window=cfg.sync.window,
|
||||
max_window=cfg.sync.max_window,
|
||||
overlap=cfg.sync.overlap,
|
||||
since=store.latest_upload_date(channel_id) if known else None,
|
||||
keep=keep_ref(cfg),
|
||||
)
|
||||
target_channel = channel_id or channel_id_found
|
||||
store.upsert_channel(target_channel, _extract_handle(cfg.channel_url), channel_name, len(refs))
|
||||
channel_name, refs = result.channel_name, result.refs
|
||||
target_channel = channel_id or result.channel_id
|
||||
if result.full_scan:
|
||||
store.upsert_channel(
|
||||
target_channel, extract_handle(cfg.channel_url), channel_name, len(refs), avatar=result.avatar
|
||||
)
|
||||
else:
|
||||
store.update_channel_meta(target_channel, name=channel_name, avatar=result.avatar)
|
||||
store.upsert_videos(refs)
|
||||
_emit(f"watch: discovered {len(refs)} videos on {channel_name}")
|
||||
store.mark_channel_synced(target_channel)
|
||||
_emit(
|
||||
f"watch: scanned {result.fetched} newest video(s) on {channel_name} — "
|
||||
f"{result.new_count} new"
|
||||
)
|
||||
|
||||
pending = store.get_pending(target_channel)
|
||||
if not pending:
|
||||
@@ -86,17 +107,29 @@ def _run_once(
|
||||
cookie_path = vault
|
||||
_emit(f"watch: using active vault cookie {vault}")
|
||||
|
||||
guard = ThrottleGuard(
|
||||
threshold=cfg.delay.throttle_threshold,
|
||||
base=cfg.delay.backoff_base,
|
||||
cap=cfg.delay.backoff_cap,
|
||||
)
|
||||
for row in pending:
|
||||
process_video(
|
||||
status = process_video(
|
||||
row, cfg, store, channel_name, target_channel, cfg.channel_url,
|
||||
cookies_file=cookie_path,
|
||||
cookies_from_browser=cookies_from_browser,
|
||||
on_log=on_log,
|
||||
)
|
||||
if status == "done":
|
||||
guard.note_success()
|
||||
else:
|
||||
failed = store.get_video(row.video_id)
|
||||
wait = guard.note_failure(getattr(failed, "error_msg", None) or status)
|
||||
if guard.tripped:
|
||||
# Unattended loop: stop this pass rather than spend the whole
|
||||
# backlog against a throttled session. The next tick retries.
|
||||
_emit(f"watch: {guard.tripped_reason} — stopping this pass")
|
||||
return
|
||||
if wait > 0:
|
||||
_emit(f"watch: throttled, waiting {wait:.1f}s")
|
||||
time.sleep(wait)
|
||||
polite_sleep(cfg.delay.min_seconds, cfg.delay.max_seconds)
|
||||
|
||||
|
||||
def _extract_handle(url: str) -> str:
|
||||
if "@" in url:
|
||||
return "@" + url.split("@", 1)[1].split("/", 1)[0]
|
||||
return ""
|
||||
|
||||
+133
-42
@@ -8,13 +8,84 @@ from typing import Callable
|
||||
from .chapters import align_chapters, chapters_from_info
|
||||
from .config import Config
|
||||
from .extract import extract_video
|
||||
from .render import build_filename_stem, render_markdown
|
||||
from .ratelimit import is_rate_limited
|
||||
from .render import build_filename_stem, render_markdown, safe_dirname
|
||||
from .store import Store, VideoRow
|
||||
from ._yt_http import yt_get
|
||||
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
|
||||
def _render_and_retire(
|
||||
cfg: Config,
|
||||
video_id: str,
|
||||
channel_name: str,
|
||||
context: dict,
|
||||
previous: str | None,
|
||||
) -> tuple[Path, str]:
|
||||
"""Write the note for `video_id` and delete whatever file it used to own.
|
||||
|
||||
Identity is the video id; the *filename* is derived from the title, and
|
||||
titles are not stable. YouTube serves them localised — the same video came
|
||||
back as "La controversia de Claude Fable 5" on one pass and "The Claude
|
||||
Fable controversy 5" on the next — and creators rename videos outright.
|
||||
Re-rendering under a new stem without retiring the old path leaves an
|
||||
orphan the database no longer references, which is how a library ends up
|
||||
with more notes than videos.
|
||||
|
||||
Shared by `process_video` and `re_render_videos` because they previously
|
||||
each built the stem their own way and drifted: one used the raw compact
|
||||
date, the other the normalised one, and the result was 94 files for 61 rows.
|
||||
"""
|
||||
stem = build_filename_stem(
|
||||
upload_date=context["upload_date"],
|
||||
title=context["title"],
|
||||
template=cfg.filename_template,
|
||||
video_id=video_id,
|
||||
)
|
||||
out_subdir = Path(cfg.output_dir_resolved) / safe_dirname(channel_name)
|
||||
md_path = render_markdown(cfg.template_path_resolved, out_subdir, stem, context)
|
||||
|
||||
out_root = Path(cfg.output_dir_resolved).parent
|
||||
try:
|
||||
rel = md_path.relative_to(out_root) if md_path.is_relative_to(out_root) else md_path
|
||||
except ValueError:
|
||||
rel = md_path
|
||||
rel_str = str(rel)
|
||||
|
||||
if previous and previous != rel_str:
|
||||
old = out_root / previous
|
||||
try:
|
||||
if old.exists() and old.resolve() != md_path.resolve():
|
||||
old.unlink()
|
||||
log.info("retired superseded note %s -> %s", previous, rel_str)
|
||||
except OSError as exc:
|
||||
log.warning("could not remove superseded note %s: %s", old, exc)
|
||||
return md_path, rel_str
|
||||
|
||||
|
||||
def _store_metadata(store: Store, row: VideoRow, info: dict) -> None:
|
||||
"""Persist everything the extraction learned that is not the transcript.
|
||||
|
||||
Split out of the render path so it can run before the no-transcript exit:
|
||||
these fields are already paid for by the time we know whether captions
|
||||
downloaded, and `update_video_metadata` only writes the keys it is given.
|
||||
"""
|
||||
store.update_video_metadata(
|
||||
row.video_id,
|
||||
view_count=info.get("view_count"),
|
||||
like_count=info.get("like_count"),
|
||||
tags=info.get("tags") or None,
|
||||
thumbnail=info.get("thumbnail"),
|
||||
description=(info.get("description") or "").strip() or None,
|
||||
)
|
||||
# discovery in flat mode reports no upload_date, so this is where it lands
|
||||
ud = info.get("upload_date") or row.upload_date
|
||||
if ud:
|
||||
store.set_upload_date(row.video_id, ud.replace("-", "") if "-" in str(ud) else str(ud))
|
||||
|
||||
|
||||
def process_video(
|
||||
row: VideoRow,
|
||||
cfg: Config,
|
||||
@@ -42,23 +113,46 @@ def process_video(
|
||||
prefer_manual=cfg.prefer_manual,
|
||||
cookies_file=cookies_file,
|
||||
cookies_from_browser=cookies_from_browser,
|
||||
extractor_retries=cfg.yt_dlp.extractor_retries,
|
||||
socket_timeout=cfg.yt_dlp.socket_timeout,
|
||||
)
|
||||
except Exception as exc:
|
||||
log.error("Error extrayendo %s: %s", row.video_id, exc)
|
||||
store.mark_error(row.video_id, str(exc))
|
||||
return "error"
|
||||
|
||||
info = data.info
|
||||
|
||||
# Metadata is persisted BEFORE the no-transcript exit. The extraction already
|
||||
# cost its requests and the info dict is in hand; discarding it because the
|
||||
# separate caption fetch failed means a retry re-spends them for data we
|
||||
# already had. Measured after a throttling incident: five rows left with
|
||||
# upload_date, view_count, description and thumbnail all NULL.
|
||||
# Agrupadas en una transaccion: una conexion/commit en vez de tres.
|
||||
with store.transaction():
|
||||
store.set_availability(row.video_id, info.get("availability"))
|
||||
_store_metadata(store, row, info)
|
||||
|
||||
if not data.segments:
|
||||
log.warning("Sin transcripcion para %s", row.video_id)
|
||||
store.mark_status(row.video_id, "no_subtitles")
|
||||
reason = data.skip_reason or "no transcript"
|
||||
log.warning("Sin transcripcion para %s: %s", row.video_id, reason)
|
||||
_emit(f"{row.video_id}: {reason}")
|
||||
# A throttled caption fetch is not "this video has no captions". Both
|
||||
# states are retryable, but only `error` is honest about the cause, and
|
||||
# `no_subtitles` counts are what tell you a channel publishes none.
|
||||
# Which of the three requests YouTube refused should not decide this.
|
||||
if is_rate_limited(reason):
|
||||
store.mark_error(row.video_id, reason)
|
||||
return "error"
|
||||
store.mark_status(row.video_id, "no_subtitles", reason)
|
||||
return "no_subtitles"
|
||||
|
||||
chapters = chapters_from_info(data.info)
|
||||
sections = align_chapters(data.segments, chapters)
|
||||
|
||||
info = data.info
|
||||
|
||||
# persist segments + rich metadata to DB (for search, stats, webapp)
|
||||
# persist segments + rich metadata to DB (for search, stats, webapp);
|
||||
# una transaccion: delete+inserts+update atomicos y un solo commit
|
||||
with store.transaction():
|
||||
store.store_segments(row.video_id, data.segments)
|
||||
seg_json = json.dumps(
|
||||
[{"start": s.start, "end": s.end, "text": s.text} for s in data.segments],
|
||||
@@ -70,20 +164,10 @@ def process_video(
|
||||
)
|
||||
store.update_video_metadata(
|
||||
row.video_id,
|
||||
view_count=info.get("view_count"),
|
||||
like_count=info.get("like_count"),
|
||||
tags=info.get("tags") or None,
|
||||
thumbnail=info.get("thumbnail"),
|
||||
description=(info.get("description") or "").strip() or None,
|
||||
chapters_json=ch_json,
|
||||
segments_json=seg_json,
|
||||
)
|
||||
|
||||
# ensure upload_date is populated (discovery sometimes lacks it)
|
||||
ud = info.get("upload_date") or row.upload_date
|
||||
if ud:
|
||||
store.set_upload_date(row.video_id, ud.replace("-", "") if "-" in str(ud) else str(ud))
|
||||
|
||||
context = {
|
||||
"video_id": row.video_id,
|
||||
"title": info.get("title") or row.title or row.video_id,
|
||||
@@ -103,22 +187,12 @@ def process_video(
|
||||
"sections": sections,
|
||||
}
|
||||
|
||||
stem = build_filename_stem(
|
||||
upload_date=context["upload_date"],
|
||||
title=context["title"],
|
||||
template=cfg.filename_template,
|
||||
md_path, rel_str = _render_and_retire(
|
||||
cfg, row.video_id, channel_name, context, row.markdown_path
|
||||
)
|
||||
out_subdir = Path(cfg.output_dir_resolved) / _safe_dirname(channel_name)
|
||||
md_path = render_markdown(cfg.template_path_resolved, out_subdir, stem, context)
|
||||
|
||||
out_root = Path(cfg.output_dir_resolved).parent
|
||||
try:
|
||||
rel = md_path.relative_to(out_root) if md_path.is_relative_to(out_root) else md_path
|
||||
except ValueError:
|
||||
rel = md_path
|
||||
store.mark_done(
|
||||
row.video_id,
|
||||
str(rel),
|
||||
rel_str,
|
||||
data.subtitle.lang if data.subtitle else None,
|
||||
data.subtitle.source if data.subtitle else None,
|
||||
data.has_chapters,
|
||||
@@ -135,11 +209,6 @@ def _normalize_date(d: str | None) -> str:
|
||||
return d
|
||||
|
||||
|
||||
def _safe_dirname(name: str) -> str:
|
||||
safe = "".join(c for c in name if c not in r'\/:*?"<>|')
|
||||
return safe.strip().strip(".") or "unknown"
|
||||
|
||||
|
||||
def thumbnail_url_for(video_row) -> str:
|
||||
"""Thumbnail URL for a video: stored URL, else the canonical YouTube one derived from its id."""
|
||||
if video_row and getattr(video_row, "thumbnail", None):
|
||||
@@ -150,20 +219,38 @@ def thumbnail_url_for(video_row) -> str:
|
||||
|
||||
def cache_thumbnail(store, video_id: str, thumbnails_dir) -> bool:
|
||||
"""Download + cache a video's thumbnail to <thumbnails_dir>/<video_id>.jpg. Returns True on success/existing."""
|
||||
import requests as _requests
|
||||
out = Path(thumbnails_dir) / f"{video_id}.jpg"
|
||||
if out.exists():
|
||||
if out.exists() and out.stat().st_size > 0:
|
||||
return True
|
||||
v = store.get_video(video_id)
|
||||
url = thumbnail_url_for(v) if v else f"https://i.ytimg.com/vi/{video_id}/hqdefault.jpg"
|
||||
try:
|
||||
r = _requests.get(url, timeout=15, headers={"User-Agent": "Mozilla/5.0"})
|
||||
r = yt_get(url, timeout=20.0)
|
||||
r.raise_for_status()
|
||||
if r.content:
|
||||
out.parent.mkdir(parents=True, exist_ok=True)
|
||||
out.write_bytes(r.content)
|
||||
return True
|
||||
except Exception:
|
||||
return False
|
||||
return False
|
||||
|
||||
|
||||
def cache_channel_avatar(channel_id: str, url: str, avatars_dir) -> bool:
|
||||
"""Download + cache a channel avatar to <avatars_dir>/<channel_id>.jpg. Returns True on success/existing."""
|
||||
out = Path(avatars_dir) / f"{channel_id}.jpg"
|
||||
out.parent.mkdir(parents=True, exist_ok=True)
|
||||
if out.exists() and out.stat().st_size > 0:
|
||||
return True
|
||||
try:
|
||||
r = yt_get(url, timeout=20.0)
|
||||
r.raise_for_status()
|
||||
if r.content:
|
||||
out.write_bytes(r.content)
|
||||
return True
|
||||
except Exception:
|
||||
pass
|
||||
return False
|
||||
|
||||
|
||||
def re_render_videos(store: Store, cfg: Config, channel_id: str | None = None) -> int:
|
||||
@@ -171,7 +258,6 @@ def re_render_videos(store: Store, cfg: Config, channel_id: str | None = None) -
|
||||
import json
|
||||
from .chapters import align_chapters, Chapter
|
||||
from .parse import Segment
|
||||
from .render import build_filename_stem, render_markdown
|
||||
|
||||
videos = [v for v in store.get_all(channel_id) if v.status == "done" and v.segments_json]
|
||||
n = 0
|
||||
@@ -193,14 +279,19 @@ def re_render_videos(store: Store, cfg: Config, channel_id: str | None = None) -
|
||||
context = {
|
||||
"video_id": v.video_id, "title": v.title or v.video_id,
|
||||
"channel_name": ch.get("name") or "", "channel_id": v.channel_id, "channel_url": "",
|
||||
"upload_date": v.upload_date or "", "duration": v.duration or 0, "url": v.url,
|
||||
# Must match process_video's normalisation (pipeline.py:96): the
|
||||
# frontmatter is a bidirectional contract that backfill re-reads.
|
||||
"upload_date": _normalize_date(v.upload_date), "duration": v.duration or 0, "url": v.url,
|
||||
"transcript_lang": v.transcript_lang or "", "transcript_src": v.transcript_src or "",
|
||||
"view_count": v.view_count, "like_count": v.like_count, "tags": tags,
|
||||
"thumbnail": v.thumbnail or "", "description": v.description or "", "sections": sections,
|
||||
}
|
||||
stem = build_filename_stem(v.upload_date, v.title or v.video_id, cfg.filename_template)
|
||||
out_subdir = Path(cfg.output_dir_resolved) / _safe_dirname(ch.get("name") or "unknown")
|
||||
render_markdown(cfg.template_path_resolved, out_subdir, stem, context)
|
||||
md_path, rel_str = _render_and_retire(
|
||||
cfg, v.video_id, ch.get("name") or "unknown", context, v.markdown_path
|
||||
)
|
||||
# mark_done is the only writer of markdown_path; without it the DB
|
||||
# keeps pointing at the pre-render file.
|
||||
store.mark_done(v.video_id, rel_str, v.transcript_lang, v.transcript_src, bool(chapters))
|
||||
n += 1
|
||||
except Exception as exc:
|
||||
log.warning("re-render failed for %s: %s", v.video_id, exc)
|
||||
|
||||
+280
-4
@@ -1,14 +1,290 @@
|
||||
"""Politeness primitives shared by every path that reaches YouTube.
|
||||
|
||||
Three separate concerns live here, and they are not interchangeable:
|
||||
|
||||
* `Pacer` spaces requests out so a burst never leaves this process. It is
|
||||
process-global on purpose — the JobManager serialises *jobs*, but the
|
||||
`/api/tools/*` endpoints run outside it, so without a shared pacer two
|
||||
consumers can hammer YouTube while each believes it is being polite.
|
||||
* `backoff_delay` is what to wait *after* a failure. It follows the algorithm
|
||||
Google documents for its own APIs: `min(base * 2**n + jitter, cap)`, with the
|
||||
jitter redrawn on every attempt so retries from concurrent clients do not
|
||||
re-synchronise into waves.
|
||||
* `ThrottleGuard` decides when to stop trying. YouTube's throttle is a session
|
||||
ban of up to an hour; once it lands, every further request is both useless
|
||||
and harmful. The guard is what turns "511 videos marked failed" into "stopped
|
||||
after 3, kept them retryable".
|
||||
|
||||
Deliberately absent: any retry of the request itself. yt-dlp's YouTube
|
||||
extractor explicitly excludes 403/429 from its own RetryManager, so a throttled
|
||||
call fails once and comes back here — retrying it in a tight loop is the exact
|
||||
behaviour that earns the ban in the first place.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import random
|
||||
import re
|
||||
import threading
|
||||
import time
|
||||
from dataclasses import dataclass, field
|
||||
from typing import Callable
|
||||
|
||||
# Substrings that mean "YouTube is refusing because of *rate*, not content".
|
||||
#
|
||||
# Kept distinct from `Store.PERMANENT_ERROR_PATTERNS`: those describe videos we
|
||||
# will never get (members-only, deleted). These describe videos we could get if
|
||||
# we asked more slowly, so they must stay retryable and must trip the breaker.
|
||||
_THROTTLE_SIGNS = (
|
||||
"rate-limited by youtube",
|
||||
"this content isn't available, try again later",
|
||||
"sign in to confirm you're not a bot",
|
||||
"sign in to confirm your age", # same bot-wall, different copy
|
||||
"http error 429",
|
||||
"too many requests",
|
||||
"ratelimitexceeded",
|
||||
"userratelimitexceeded",
|
||||
"temporarily blocked",
|
||||
)
|
||||
|
||||
# 403 with a quota reason is NOT the same failure: Google documents it as
|
||||
# daily and explicitly warns retries may not work for hours (AIP-194). Treated
|
||||
# as fatal-for-now rather than as something to back off and retry into.
|
||||
_QUOTA_SIGNS = (
|
||||
"quotaexceeded",
|
||||
"dailylimitexceeded",
|
||||
)
|
||||
|
||||
# Two spellings reach us and they are not the same string:
|
||||
# yt-dlp -> "HTTP Error 429: Too Many Requests"
|
||||
# requests -> "HTTPError: 429 Client Error: Too Many Requests for url: ..."
|
||||
# The original pattern only matched the first, so the second was detected purely
|
||||
# by its "too many requests" prose — and a 429 served without a reason phrase
|
||||
# (routine over HTTP/2) carries no such prose and slipped through the breaker.
|
||||
_HTTP_STATUS = re.compile(
|
||||
r"HTTP\s*Error[:\s]+(\d{3})"
|
||||
r"|(\d{3})\s+(?:Client|Server)\s+Error",
|
||||
re.IGNORECASE,
|
||||
)
|
||||
|
||||
#: Statuses Google documents as retryable, and which mean "slow down" here.
|
||||
_RETRYABLE_STATUSES = frozenset({"408", "429"})
|
||||
|
||||
|
||||
def is_rate_limited(error: object) -> bool:
|
||||
"""True when `error` looks like YouTube throttling us rather than a bad video."""
|
||||
text = str(error or "").lower()
|
||||
if any(sign in text for sign in _THROTTLE_SIGNS):
|
||||
return True
|
||||
for m in _HTTP_STATUS.finditer(text):
|
||||
status = m.group(1) or m.group(2)
|
||||
if status in _RETRYABLE_STATUSES:
|
||||
return True
|
||||
return False
|
||||
|
||||
|
||||
def is_quota_exhausted(error: object) -> bool:
|
||||
"""True for Google's daily-quota refusals, which backing off will not fix."""
|
||||
text = str(error or "").lower()
|
||||
return any(sign in text for sign in _QUOTA_SIGNS)
|
||||
|
||||
|
||||
def polite_sleep(min_s: float = 1.5, max_s: float = 3.5) -> None:
|
||||
delay = random.uniform(min_s, max_s)
|
||||
time.sleep(delay)
|
||||
"""Randomised pause between units of work (one video, one channel)."""
|
||||
if max_s < min_s:
|
||||
max_s = min_s
|
||||
time.sleep(random.uniform(min_s, max_s))
|
||||
|
||||
|
||||
def backoff_sleep(attempt: int, base: float = 2.0, cap: float = 60.0) -> None:
|
||||
delay = min(cap, base * (2 ** attempt) + random.uniform(0, 1))
|
||||
def backoff_delay(attempt: int, base: float = 2.0, cap: float = 60.0) -> float:
|
||||
"""Truncated exponential backoff with full jitter, per Google's retry guidance.
|
||||
|
||||
`attempt` is 0-indexed, so the first wait after a failure is ~`base`.
|
||||
Returns the delay instead of sleeping so callers can log it and tests can
|
||||
assert on it without spending real seconds.
|
||||
"""
|
||||
if attempt < 0:
|
||||
attempt = 0
|
||||
# 2**attempt overflows into absurd floats long before it matters; clamp the
|
||||
# exponent so a runaway counter can't turn into an OverflowError.
|
||||
exponent = min(attempt, 32)
|
||||
return min(cap, base * (2**exponent) + random.uniform(0, 1))
|
||||
|
||||
|
||||
def backoff_sleep(attempt: int, base: float = 2.0, cap: float = 60.0) -> float:
|
||||
"""Sleep `backoff_delay(...)` and return how long it waited."""
|
||||
delay = backoff_delay(attempt, base, cap)
|
||||
time.sleep(delay)
|
||||
return delay
|
||||
|
||||
|
||||
class Pacer:
|
||||
"""Enforces a minimum gap between requests, process-wide and thread-safe.
|
||||
|
||||
`wait()` blocks only for the remainder of the gap, so a slow caller never
|
||||
pays twice: if the previous request already took longer than `min_interval`
|
||||
there is nothing left to wait for.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
min_interval: float = 0.0,
|
||||
*,
|
||||
clock: Callable[[], float] = time.monotonic,
|
||||
sleeper: Callable[[float], None] = time.sleep,
|
||||
) -> None:
|
||||
self.min_interval = max(0.0, float(min_interval))
|
||||
self._clock = clock
|
||||
self._sleeper = sleeper
|
||||
self._lock = threading.Lock()
|
||||
self._next_at = 0.0
|
||||
|
||||
def configure(self, min_interval: float) -> None:
|
||||
with self._lock:
|
||||
self.min_interval = max(0.0, float(min_interval))
|
||||
|
||||
def wait(self, cost: int = 1) -> float:
|
||||
"""Block until the next request is allowed. Returns seconds actually slept.
|
||||
|
||||
`cost` is how many HTTP requests the caller is about to make. It matters
|
||||
because the caller is usually yt-dlp: one `extract_info` on a video is
|
||||
two requests (watch page + InnerTube player), and charging it as one
|
||||
made the pacer under-count by half. yt-dlp spaces those two internally
|
||||
via `sleep_interval_requests`; `cost` is what keeps this pacer's idea of
|
||||
the budget honest about them.
|
||||
"""
|
||||
cost = max(1, int(cost))
|
||||
with self._lock:
|
||||
if self.min_interval <= 0:
|
||||
return 0.0
|
||||
now = self._clock()
|
||||
delay = self._next_at - now
|
||||
if delay < 0:
|
||||
delay = 0.0
|
||||
# Reserve the slots before releasing the lock so concurrent callers
|
||||
# queue up behind each other instead of all reading the same `now`.
|
||||
self._next_at = now + delay + self.min_interval * cost
|
||||
if delay > 0:
|
||||
self._sleeper(delay)
|
||||
return delay
|
||||
|
||||
def penalise(self, seconds: float) -> None:
|
||||
"""Push the next allowed request out by `seconds` (used after a 429)."""
|
||||
if seconds <= 0:
|
||||
return
|
||||
with self._lock:
|
||||
self._next_at = max(self._next_at, self._clock() + seconds)
|
||||
|
||||
|
||||
#: Shared by `_yt_http.yt_get` and every yt-dlp entry point. Idle (0.0) until
|
||||
#: something calls `configure_global_pacer`, so importing this module never
|
||||
#: slows down a test suite that does not opt in.
|
||||
GLOBAL_PACER = Pacer(0.0)
|
||||
|
||||
|
||||
def configure_global_pacer(min_interval: float) -> None:
|
||||
GLOBAL_PACER.configure(min_interval)
|
||||
|
||||
|
||||
def ydl_throttle_opts(
|
||||
sleep_requests: float = 2.0,
|
||||
*,
|
||||
extractor_retries: int = 3,
|
||||
socket_timeout: float = 30.0,
|
||||
js_runtimes: dict[str, dict] | None = None,
|
||||
) -> dict[str, object]:
|
||||
"""The politeness half of every `ydl_opts` dict in this project.
|
||||
|
||||
Exists because the option names are easy to get subtly wrong, and yt-dlp
|
||||
silently ignores keys it does not recognise — this project shipped
|
||||
`sleep_subrequests` (not a real option) for its entire history, so nothing
|
||||
ever slept between the sub-requests of an extraction.
|
||||
|
||||
Only `sleep_interval_requests` throttles *extraction*; `sleep_interval` and
|
||||
`max_sleep_interval` fire in the file downloader and never trigger under
|
||||
`skip_download`, which is every metadata path here.
|
||||
"""
|
||||
return {
|
||||
# Sleeps inside InfoExtractor._request_webpage, i.e. before each HTTP
|
||||
# call an extractor makes: watch page, InnerTube player/browse, and the
|
||||
# continuation pages of a channel tab.
|
||||
"sleep_interval_requests": max(0.0, float(sleep_requests)),
|
||||
# `retries` governs the downloader; extraction retries are this one.
|
||||
# Note yt-dlp's YouTube extractor refuses to retry 403/429 at all, so
|
||||
# this only covers transient 5xx and network errors.
|
||||
"extractor_retries": int(extractor_retries),
|
||||
"socket_timeout": float(socket_timeout),
|
||||
"js_runtimes": js_runtimes if js_runtimes is not None else {"node": {}, "deno": {}, "bun": {}, "quickjs": {}},
|
||||
}
|
||||
|
||||
|
||||
@dataclass
|
||||
class ThrottleGuard:
|
||||
"""Circuit breaker for a batch of YouTube work.
|
||||
|
||||
Feed it every outcome. It answers one question — "should this run keep
|
||||
going?" — and, while it still says yes, how long to wait first.
|
||||
|
||||
The counter is *consecutive*: an isolated throttled video between successes
|
||||
is noise, three in a row means the session is banned and everything after
|
||||
it will fail too. Only the consecutive form distinguishes those.
|
||||
"""
|
||||
|
||||
threshold: int = 3
|
||||
base: float = 2.0
|
||||
cap: float = 60.0
|
||||
|
||||
consecutive: int = 0
|
||||
throttled_total: int = 0
|
||||
tripped_reason: str | None = None
|
||||
_waits: list[float] = field(default_factory=list)
|
||||
|
||||
@property
|
||||
def tripped(self) -> bool:
|
||||
return self.tripped_reason is not None
|
||||
|
||||
def note_success(self) -> None:
|
||||
"""A request went through: the session is healthy again."""
|
||||
self.consecutive = 0
|
||||
|
||||
def note_failure(self, error: object) -> float:
|
||||
"""Record a failed unit of work. Returns seconds the caller should wait.
|
||||
|
||||
Non-throttle failures (a private video, a parse error) reset the
|
||||
consecutive counter: they say nothing about our request rate, and
|
||||
letting them accumulate would trip the breaker on a channel that simply
|
||||
has a few dead videos.
|
||||
"""
|
||||
if is_quota_exhausted(error):
|
||||
self.throttled_total += 1
|
||||
self.tripped_reason = (
|
||||
"YouTube reports the quota is exhausted; Google documents this as "
|
||||
"daily, so retrying now cannot succeed"
|
||||
)
|
||||
return 0.0
|
||||
|
||||
if not is_rate_limited(error):
|
||||
self.consecutive = 0
|
||||
return 0.0
|
||||
|
||||
self.throttled_total += 1
|
||||
self.consecutive += 1
|
||||
if self.consecutive >= self.threshold:
|
||||
self.tripped_reason = (
|
||||
f"{self.consecutive} consecutive rate-limit responses from YouTube; "
|
||||
"the session is throttled (YouTube states up to an hour) and further "
|
||||
"requests would only mark healthy videos as failed"
|
||||
)
|
||||
return 0.0
|
||||
|
||||
delay = backoff_delay(self.consecutive - 1, self.base, self.cap)
|
||||
self._waits.append(delay)
|
||||
GLOBAL_PACER.penalise(delay)
|
||||
return delay
|
||||
|
||||
def summary(self) -> str:
|
||||
if self.tripped_reason:
|
||||
return f"stopped: {self.tripped_reason}"
|
||||
if self.throttled_total:
|
||||
return f"{self.throttled_total} throttled response(s) absorbed by backoff"
|
||||
return "no throttling seen"
|
||||
|
||||
@@ -32,7 +32,16 @@ def to_json(value) -> str:
|
||||
return json.dumps(value, ensure_ascii=False)
|
||||
|
||||
|
||||
# One Environment per template directory, cached for the process lifetime:
|
||||
# building it (loader + filters) per rendered note was the dominant cost of
|
||||
# batch renders. Jinja's own per-env template cache keeps the compiled
|
||||
# Template, and its default auto_reload still picks up on-disk edits.
|
||||
_ENV_CACHE: dict[Path, Environment] = {}
|
||||
|
||||
|
||||
def _make_env(template_dir: Path) -> Environment:
|
||||
env = _ENV_CACHE.get(template_dir)
|
||||
if env is None:
|
||||
env = Environment(
|
||||
loader=FileSystemLoader(str(template_dir)),
|
||||
autoescape=select_autoescape(disabled_extensions=("j2", "txt")),
|
||||
@@ -42,6 +51,7 @@ def _make_env(template_dir: Path) -> Environment:
|
||||
env.filters["format_timestamp"] = format_timestamp
|
||||
env.filters["quote_yaml"] = quote_yaml
|
||||
env.filters["to_json"] = to_json
|
||||
_ENV_CACHE[template_dir] = env
|
||||
return env
|
||||
|
||||
|
||||
@@ -63,10 +73,44 @@ def render_markdown(
|
||||
return out_file
|
||||
|
||||
|
||||
def build_filename_stem(upload_date: str | None, title: str, template: str = "{upload_date}_{slug}") -> str:
|
||||
def safe_dirname(name: str | None) -> str:
|
||||
"""Channel name -> markdown subdirectory name, shared by every .md writer.
|
||||
|
||||
Strips the characters Windows forbids in a path segment, then leading and
|
||||
trailing spaces/dots (also illegal there), falling back to "unknown" so the
|
||||
output tree never grows a nameless root. Callers with a bare filename want
|
||||
`safe_filename` instead ("untitled" fallback).
|
||||
"""
|
||||
safe = "".join(c for c in (name or "") if c not in r'\/:*?"<>|')
|
||||
return safe.strip().strip(".") or "unknown"
|
||||
|
||||
|
||||
def safe_filename(name: str) -> str:
|
||||
"""Free-text (title) -> file-legal stem; "untitled" when nothing survives."""
|
||||
safe = "".join(c for c in name if c not in r'\/:*?"<>|')
|
||||
return safe.strip().strip(".") or "untitled"
|
||||
|
||||
|
||||
def build_filename_stem(
|
||||
upload_date: str | None,
|
||||
title: str,
|
||||
template: str = "{upload_date}_{slug}",
|
||||
video_id: str | None = None,
|
||||
) -> str:
|
||||
"""Filename for a video's note, from `template`.
|
||||
|
||||
`{video_id}` is offered because the other two variables are unstable:
|
||||
YouTube serves titles localised (the same video came back Spanish on one
|
||||
pass and English on the next) and creators rename things. A template
|
||||
including the id makes the file identifiable from disk alone, without
|
||||
consulting the database. It is opt-in — the default is unchanged so
|
||||
existing libraries keep their filenames.
|
||||
"""
|
||||
from slugify import slugify
|
||||
date_part = upload_date or "unknown-date"
|
||||
slug = slugify(title, max_length=60) or "untitled"
|
||||
stem = template.format(upload_date=date_part, slug=slug, title=title)
|
||||
stem = template.format(
|
||||
upload_date=date_part, slug=slug, title=title, video_id=video_id or "",
|
||||
)
|
||||
safe = "".join(c for c in stem if c not in r'\/:*?"<>|')
|
||||
return safe.strip().strip(".")
|
||||
|
||||
+114
-10
@@ -8,7 +8,7 @@ from pathlib import Path
|
||||
from typing import Callable
|
||||
|
||||
from .parse import Segment
|
||||
from .store import SearchHit, Store
|
||||
from .store import Store
|
||||
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
@@ -91,14 +91,6 @@ def _strip_quotes(value: str) -> str:
|
||||
return v
|
||||
|
||||
|
||||
def store_segments(store: Store, video_id: str, segments: list[Segment]) -> None:
|
||||
store.store_segments(video_id, segments)
|
||||
|
||||
|
||||
def search(store: Store, query: str, channel_id: str | None = None, limit: int = 50) -> list[SearchHit]:
|
||||
return store.search_segments(query, channel_id=channel_id, limit=limit)
|
||||
|
||||
|
||||
def backfill_from_markdown(
|
||||
store: Store,
|
||||
md_root: Path,
|
||||
@@ -122,7 +114,9 @@ def backfill_from_markdown(
|
||||
for md_path in md_files:
|
||||
try:
|
||||
text = md_path.read_text(encoding="utf-8")
|
||||
except OSError as exc:
|
||||
except (OSError, UnicodeDecodeError) as exc:
|
||||
# UnicodeDecodeError is a ValueError, not an OSError — letting it
|
||||
# escape aborted the loop and silently skipped every later file.
|
||||
_log(f"backfill: skip unreadable {md_path}: {exc}")
|
||||
continue
|
||||
parsed = parse_markdown(text)
|
||||
@@ -170,6 +164,116 @@ def backfill_from_markdown(
|
||||
return n
|
||||
|
||||
|
||||
def reconcile_markdown(
|
||||
store: Store,
|
||||
md_root: Path,
|
||||
log: Callable[[str], None] | None = None,
|
||||
*,
|
||||
prune: bool = False,
|
||||
) -> dict[str, int]:
|
||||
"""Make the DB agree with what is actually on disk.
|
||||
|
||||
`backfill_from_markdown` fills in segments and metadata but never touches
|
||||
`status` or `markdown_path`, so a video whose .md exists can sit at
|
||||
`error`/`no_subtitles`/`pending` forever and the UI keeps showing a failure
|
||||
for work that is already done. This walks the markdown tree and repairs:
|
||||
|
||||
- a row with a real .md but a non-done status -> marked done
|
||||
|
||||
`prune=True` additionally sends `done` rows whose .md has disappeared back
|
||||
to pending. That direction is opt-in because it is destructive when aimed
|
||||
at the wrong root: pointed at an empty or unrelated markdown tree it would
|
||||
demote every finished video in the database. It is also skipped outright
|
||||
when the tree contains no .md at all, which is never a real "everything was
|
||||
deleted" state — it means the root is wrong.
|
||||
|
||||
Returns counts so the caller can report what changed. Idempotent.
|
||||
"""
|
||||
def _log(msg: str) -> None:
|
||||
if log:
|
||||
log(msg)
|
||||
else:
|
||||
logging.getLogger(__name__).info(msg)
|
||||
|
||||
md_root = Path(md_root)
|
||||
data_root = md_root.parent
|
||||
out = {
|
||||
"scanned": 0, "repaired_done": 0, "orphan_md": 0,
|
||||
"missing_md": 0, "backfilled": 0, "stale_dupe": 0,
|
||||
}
|
||||
if not md_root.exists():
|
||||
_log(f"reconcile: markdown root not found: {md_root}")
|
||||
return out
|
||||
|
||||
out["backfilled"] = backfill_from_markdown(store, md_root, log=log)
|
||||
|
||||
seen: dict[str, Path] = {}
|
||||
for md_path in sorted(md_root.rglob("*.md")):
|
||||
out["scanned"] += 1
|
||||
try:
|
||||
text = md_path.read_text(encoding="utf-8")
|
||||
except (OSError, UnicodeDecodeError) as exc:
|
||||
_log(f"reconcile: skip unreadable {md_path}: {exc}")
|
||||
continue
|
||||
meta = parse_markdown(text).metadata
|
||||
video_id = meta.get("video_id")
|
||||
if not video_id:
|
||||
continue
|
||||
row = store.get_video(video_id)
|
||||
if not row:
|
||||
out["orphan_md"] += 1
|
||||
continue
|
||||
rel = md_path.relative_to(data_root).as_posix()
|
||||
|
||||
# A second .md for a video the DB already resolves elsewhere. Older
|
||||
# re-renders built the filename from a differently-formatted date, so
|
||||
# they wrote a sibling file the DB never learned about; it is dead
|
||||
# weight that every later scan has to wade through.
|
||||
canonical = (row.markdown_path or "").replace("\\", "/")
|
||||
if video_id in seen or (row.status == "done" and canonical and canonical != rel):
|
||||
out["stale_dupe"] += 1
|
||||
if prune:
|
||||
try:
|
||||
md_path.unlink()
|
||||
_log(f"reconcile: removed stale duplicate {rel}")
|
||||
except OSError as exc:
|
||||
_log(f"reconcile: could not remove {rel}: {exc}")
|
||||
else:
|
||||
_log(f"reconcile: stale duplicate (use prune to delete): {rel}")
|
||||
continue
|
||||
|
||||
seen[video_id] = md_path
|
||||
if row.status != "done" or not row.markdown_path:
|
||||
store.mark_done(
|
||||
video_id, rel,
|
||||
meta.get("transcript_lang") or row.transcript_lang,
|
||||
meta.get("transcript_src") or row.transcript_src,
|
||||
bool(meta.get("has_chapters")) or bool(row.has_chapters),
|
||||
)
|
||||
out["repaired_done"] += 1
|
||||
_log(f"reconcile: {video_id} had a .md on disk but status={row.status} -> done")
|
||||
|
||||
# The other direction is destructive, so it needs both an explicit opt-in
|
||||
# and evidence that we are looking at a real markdown tree.
|
||||
if prune and out["scanned"]:
|
||||
for row in store.get_all():
|
||||
if row.status != "done" or row.video_id in seen:
|
||||
continue
|
||||
path = data_root / row.markdown_path if row.markdown_path else None
|
||||
if path is None or not path.exists():
|
||||
store.mark_status(row.video_id, "pending", "markdown file missing on disk")
|
||||
out["missing_md"] += 1
|
||||
_log(f"reconcile: {row.video_id} marked done but .md is gone -> pending")
|
||||
elif prune:
|
||||
_log("reconcile: markdown tree is empty — refusing to prune (wrong root?)")
|
||||
|
||||
_log(
|
||||
"reconcile: scanned {scanned} .md, repaired {repaired_done}, "
|
||||
"re-queued {missing_md}, orphans {orphan_md}, stale duplicates {stale_dupe}".format(**out)
|
||||
)
|
||||
return out
|
||||
|
||||
|
||||
def _to_int(value: str | None) -> int | None:
|
||||
if value is None:
|
||||
return None
|
||||
|
||||
+684
-74
@@ -2,6 +2,7 @@ from __future__ import annotations
|
||||
|
||||
import json
|
||||
import sqlite3
|
||||
import threading
|
||||
from contextlib import contextmanager
|
||||
from dataclasses import dataclass
|
||||
from datetime import datetime, timezone
|
||||
@@ -49,9 +50,48 @@ _VIDEO_COLUMNS: dict[str, str] = {
|
||||
"description": "TEXT",
|
||||
"chapters_json": "TEXT",
|
||||
"segments_json": "TEXT",
|
||||
"availability": "TEXT",
|
||||
# Position in the channel's reverse-chronological /videos tab (higher =
|
||||
# newer). The only recency signal discovery produces: yt-dlp's flat listing
|
||||
# reports no upload_date for YouTube entries, so a video that has never been
|
||||
# extracted has no date to sort by.
|
||||
"channel_seq": "INTEGER",
|
||||
# 1 cuando upload_date viene del discovery aproximado (texto relativo de
|
||||
# YouTube: "hace 3 semanas"), 0/NULL cuando es exacto (extraccion).
|
||||
"upload_date_approx": "INTEGER DEFAULT 0",
|
||||
"video_download_status": "TEXT DEFAULT 'not_downloaded'",
|
||||
"video_path": "TEXT",
|
||||
"video_filename": "TEXT",
|
||||
"video_size": "INTEGER",
|
||||
"video_downloaded_at": "TEXT",
|
||||
"video_error": "TEXT",
|
||||
}
|
||||
|
||||
_CHANNEL_COLUMNS: dict[str, str] = {
|
||||
"avatar": "TEXT",
|
||||
# Incremental-sync watermark: when we last looked, and the newest upload
|
||||
# date we know of. `last_video_date` is the "desde aqui en adelante" mark.
|
||||
"last_synced_at": "TEXT",
|
||||
"last_video_date": "TEXT",
|
||||
}
|
||||
|
||||
_EXTRA_SCHEMA = """
|
||||
CREATE INDEX IF NOT EXISTS idx_videos_channel_seq ON videos(channel_id, channel_seq);
|
||||
|
||||
-- SORT_DATE_SQL runs two per-channel lookups for every undated row, and undated
|
||||
-- is the majority of a library until it is fully scraped (4541 of 4959 rows on
|
||||
-- the real one). Both are PARTIAL and covering, indexing only the dated rows —
|
||||
-- which is what the lookups are hunting for. Without the partial predicate the
|
||||
-- "nearest dated video above me" search walks every row in between checking the
|
||||
-- table for a date it will not find: 2500 rows deep into a channel whose 8
|
||||
-- dated videos all sit at the top, that is quadratic, and it measured 576 ms
|
||||
-- per page against 8 ms with these. Keep the WHERE clauses spelled exactly as
|
||||
-- the queries spell them or SQLite will not consider the index.
|
||||
CREATE INDEX IF NOT EXISTS idx_videos_dated_seq ON videos(channel_id, channel_seq, upload_date)
|
||||
WHERE upload_date IS NOT NULL AND upload_date <> '';
|
||||
CREATE INDEX IF NOT EXISTS idx_videos_dated ON videos(channel_id, upload_date)
|
||||
WHERE upload_date IS NOT NULL AND upload_date <> '';
|
||||
|
||||
CREATE TABLE IF NOT EXISTS transcript_segments (
|
||||
video_id TEXT NOT NULL,
|
||||
idx INTEGER NOT NULL,
|
||||
@@ -92,6 +132,143 @@ CREATE TABLE IF NOT EXISTS scrape_jobs (
|
||||
"""
|
||||
|
||||
|
||||
# yt-dlp's availability enum (see yt_dlp/extractor/common.py:414):
|
||||
# 'private' | 'premium_only' | 'subscriber_only' | 'needs_auth' | 'unlisted' | 'public'.
|
||||
# Only these four mean "we cannot fetch it" — `unlisted` downloads perfectly
|
||||
# well and must NOT be treated as blocked.
|
||||
BLOCKING_AVAILABILITY = {
|
||||
"subscriber_only": "members_only",
|
||||
"premium_only": "premium_only",
|
||||
"private": "private",
|
||||
"needs_auth": "needs_auth",
|
||||
}
|
||||
|
||||
|
||||
# Sorts below every real YYYYMMDD. Reached only when a video has no date and
|
||||
# neither does anything else in its channel, i.e. we have zero evidence about
|
||||
# when it was published. Such a video does not get to outrank videos we do know
|
||||
# something about; within its channel `channel_seq` still orders it correctly.
|
||||
# Stripped before it reaches the UI — it is a rank, not a date.
|
||||
NO_DATE_SENTINEL = "00000000"
|
||||
|
||||
# The date a video is ordered by.
|
||||
#
|
||||
# Chronological order here has to mean what it means on YouTube: newest upload
|
||||
# first, whether or not we have scraped the video. The obstacle is that
|
||||
# discovery cannot supply `upload_date` — yt-dlp's flat listing does not report
|
||||
# one for YouTube entries — so every video without a .md also has a NULL date.
|
||||
# Measured on the real library: 4541 of 4959 rows.
|
||||
#
|
||||
# `channel_seq` is the video's position in the channel's reverse-chronological
|
||||
# /videos tab, which gives an exact within-channel order and a defensible date:
|
||||
#
|
||||
# 1. its own upload_date, once an extraction has learned it;
|
||||
# 2. else the date of the nearest video ABOVE it in the channel that has one —
|
||||
# it was published no earlier than that, and ties break by rank, so it
|
||||
# lands in the slot YouTube would give it;
|
||||
# 3. else the newest date known anywhere in its channel. This is the run at
|
||||
# the very top of a channel, above every dated video: it is newer than all
|
||||
# of them (rank settles that) but claiming more would be inventing a date,
|
||||
# and it used to let a wholly un-scraped channel take over page one;
|
||||
# 4. else nothing is known at all — see NO_DATE_SENTINEL.
|
||||
#
|
||||
# The fallback this replaced was `discovered_at`, which dated every un-scraped
|
||||
# video "today" and pinned the entire backlog above everything else.
|
||||
#
|
||||
# The `upload_date IS NOT NULL AND upload_date <> ''` spelling is load-bearing:
|
||||
# it is what makes the partial indexes above applicable.
|
||||
SORT_DATE_SQL = f"""COALESCE(
|
||||
NULLIF(videos.upload_date, ''),
|
||||
(SELECT v2.upload_date FROM videos v2
|
||||
WHERE v2.channel_id = videos.channel_id
|
||||
AND v2.channel_seq > COALESCE(videos.channel_seq, -1)
|
||||
AND v2.upload_date IS NOT NULL AND v2.upload_date <> ''
|
||||
ORDER BY v2.channel_seq ASC LIMIT 1),
|
||||
(SELECT MAX(v3.upload_date) FROM videos v3
|
||||
WHERE v3.channel_id = videos.channel_id
|
||||
AND v3.upload_date IS NOT NULL AND v3.upload_date <> ''),
|
||||
'{NO_DATE_SENTINEL}')"""
|
||||
|
||||
# Ordering references the `sort_date` alias rather than repeating the subquery,
|
||||
# so every query that uses `_order_clause` must select `SORT_DATE_SQL AS
|
||||
# sort_date`.
|
||||
#
|
||||
# The tiebreak is `channel_id` THEN `channel_seq`, in that order. `channel_seq`
|
||||
# is a per-channel counter whose maximum is that channel's video count, so
|
||||
# comparing it ACROSS channels just ranks by catalogue size: measured on the
|
||||
# real library, 92% of adjacent pairs tie on sort_date, and in the 20250419 tie
|
||||
# the whole of one channel preceded the whole of another purely because 575 >
|
||||
# 476. Grouping by channel first keeps each channel's block contiguous and its
|
||||
# internal order — the part that has to match YouTube — untouched. `video_id`
|
||||
# then makes the order total, without which LIMIT/OFFSET paging can repeat or
|
||||
# skip rows between pages.
|
||||
_NEWEST_FIRST = (
|
||||
"sort_date DESC, videos.channel_id, videos.channel_seq DESC, videos.video_id DESC"
|
||||
)
|
||||
|
||||
# Ascending needs the unknown-date rows pushed out explicitly. Descending gets
|
||||
# it for free — NO_DATE_SENTINEL sorts below every real date — but that is the
|
||||
# same reason it sorts FIRST under ASC, which made "oldest" open with the 212
|
||||
# videos of a channel nothing has ever extracted, ahead of a genuine 2017 upload.
|
||||
# "We do not know" is not "the beginning of time"; it belongs at the end either way.
|
||||
_OLDEST_FIRST = (
|
||||
f"(sort_date = '{NO_DATE_SENTINEL}'), "
|
||||
"sort_date ASC, videos.channel_id, videos.channel_seq ASC, videos.video_id ASC"
|
||||
)
|
||||
|
||||
|
||||
def _nulls_last(column: str, direction: str) -> str:
|
||||
"""`ORDER BY` fragment that keeps NULLs at the bottom either way.
|
||||
|
||||
SQLite only accepts NULLS LAST from 3.30; the boolean-first form works on
|
||||
every version, and a video with no view count should not outrank one with a
|
||||
known count just because the column is empty.
|
||||
"""
|
||||
return f"videos.{column} IS NULL, videos.{column} {direction}, {_NEWEST_FIRST}"
|
||||
|
||||
|
||||
def _order_clause(sort: str | None) -> str:
|
||||
"""Map an API sort key to SQL. Unknown keys fall back to newest-first.
|
||||
|
||||
Both the bare and suffixed spellings are accepted because the web UI sends
|
||||
`upload_date` / `view_count` while the CLI and older callers send
|
||||
`upload_date_desc` / `views_desc`; the mismatch used to drop every non-date
|
||||
sort onto a raw `upload_date DESC` that ignored the inference above.
|
||||
"""
|
||||
by_views = _nulls_last("view_count", "DESC")
|
||||
by_likes = _nulls_last("like_count", "DESC")
|
||||
return {
|
||||
"upload_date": _NEWEST_FIRST,
|
||||
"upload_date_desc": _NEWEST_FIRST,
|
||||
"newest": _NEWEST_FIRST,
|
||||
"upload_date_asc": _OLDEST_FIRST,
|
||||
"oldest": _OLDEST_FIRST,
|
||||
"duration": _nulls_last("duration", "DESC"),
|
||||
"duration_desc": _nulls_last("duration", "DESC"),
|
||||
"duration_asc": _nulls_last("duration", "ASC"),
|
||||
"view_count": by_views,
|
||||
"view_count_desc": by_views,
|
||||
"views_desc": by_views,
|
||||
"like_count": by_likes,
|
||||
"like_count_desc": by_likes,
|
||||
"likes_desc": by_likes,
|
||||
"title": f"videos.title COLLATE NOCASE ASC, {_NEWEST_FIRST}",
|
||||
"title_asc": f"videos.title COLLATE NOCASE ASC, {_NEWEST_FIRST}",
|
||||
# Por nombre de canal (no por UUID): el subquery consulta channels,
|
||||
# que tiene decenas de filas y PK sobre channel_id.
|
||||
"channel": (
|
||||
"(SELECT c.name FROM channels c WHERE c.channel_id = videos.channel_id) "
|
||||
"COLLATE NOCASE ASC, videos.channel_id, videos.channel_seq DESC, videos.video_id DESC"
|
||||
),
|
||||
# Agrupa por ciclo de vida: hecho, luego pendientes, luego los que
|
||||
# necesitan atencion (sin subs), errores al final.
|
||||
"status": (
|
||||
"CASE videos.status WHEN 'done' THEN 0 WHEN 'pending' THEN 1 "
|
||||
"WHEN 'no_subtitles' THEN 2 ELSE 3 END ASC, " + _NEWEST_FIRST
|
||||
),
|
||||
}.get((sort or "").strip(), _NEWEST_FIRST)
|
||||
|
||||
|
||||
@dataclass
|
||||
class VideoRef:
|
||||
video_id: str
|
||||
@@ -100,6 +277,17 @@ class VideoRef:
|
||||
url: str
|
||||
upload_date: str | None = None
|
||||
duration: int | None = None
|
||||
# Reported by flat discovery, so a members-only video is known before we
|
||||
# ever spend an extraction attempt on it.
|
||||
availability: str | None = None
|
||||
# 0-based index in the listing this ref came from (0 = newest). Discovery
|
||||
# walks the /videos tab in reverse-chronological order, so this is the
|
||||
# chronological rank of a video we have no upload_date for yet.
|
||||
position: int | None = None
|
||||
# 1 si upload_date es aproximado (derivado del texto relativo del listado),
|
||||
# 0 si es exacto o no hay fecha. Al final: los llamadores posicionales
|
||||
# existentes terminan en (upload_date, duration) y no deben desplazarse.
|
||||
date_approx: int = 0
|
||||
|
||||
|
||||
@dataclass
|
||||
@@ -123,6 +311,41 @@ class VideoRow:
|
||||
description: str | None = None
|
||||
chapters_json: str | None = None
|
||||
segments_json: str | None = None
|
||||
availability: str | None = None
|
||||
channel_seq: int | None = None
|
||||
# 1 cuando upload_date es aproximado (discovery), 0/NULL si es exacto.
|
||||
upload_date_approx: int | None = None
|
||||
video_download_status: str | None = None
|
||||
video_path: str | None = None
|
||||
video_filename: str | None = None
|
||||
video_size: int | None = None
|
||||
video_downloaded_at: str | None = None
|
||||
video_error: str | None = None
|
||||
# Date the row was ordered by. Equals `upload_date` when it is known; for a
|
||||
# video discovery has not extracted yet it is inferred from `channel_seq`
|
||||
# (see `query_videos`). Only populated by queries that compute it.
|
||||
sort_date: str | None = None
|
||||
|
||||
@property
|
||||
def block_reason(self) -> str | None:
|
||||
"""Why this video can never be fetched, or None if it can.
|
||||
|
||||
Prefers the discovery signal (known before any attempt) and falls back
|
||||
to the recorded error for rows burned in before availability existed.
|
||||
"""
|
||||
blocked = BLOCKING_AVAILABILITY.get((self.availability or "").lower())
|
||||
if blocked:
|
||||
return blocked
|
||||
msg = (self.error_msg or "").lower()
|
||||
if not msg:
|
||||
return None
|
||||
if "members-only" in msg or "join this channel to get access" in msg:
|
||||
return "members_only"
|
||||
if "private video" in msg:
|
||||
return "private"
|
||||
if "has been removed" in msg or "has been terminated" in msg:
|
||||
return "removed"
|
||||
return None
|
||||
|
||||
|
||||
@dataclass
|
||||
@@ -171,10 +394,30 @@ class JobRow:
|
||||
last_error: str | None
|
||||
|
||||
|
||||
def order_pending(rows: list[VideoRow], refs: list[VideoRef]) -> list[VideoRow]:
|
||||
"""Newest-window-first ordering for the pending queue.
|
||||
|
||||
`get_pending` is ordered by discovery time, which used to coincide with
|
||||
newest-first because discovery saw the whole channel at once. With windowed
|
||||
sync that no longer holds, so put the videos from this run's window (the
|
||||
`refs` discovery just returned, newest first) at the front and keep the rest
|
||||
of the backlog behind them. Shared by the CLI scrape and the webapp's
|
||||
channel jobs, which must agree on what `--limit` / `limit` mean.
|
||||
"""
|
||||
by_id = {r.video_id: r for r in rows}
|
||||
ordered = [by_id.pop(x.video_id) for x in refs if x.video_id in by_id]
|
||||
ordered.extend(by_id.values())
|
||||
return ordered
|
||||
|
||||
|
||||
class Store:
|
||||
def __init__(self, db_path: str | Path):
|
||||
self.db_path = Path(db_path)
|
||||
self.db_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
# Pila de transacciones ambientales, por hilo: permite agrupar varias
|
||||
# llamadas a Store en una sola conexion/commit sin pasar la conexion
|
||||
# como parametro a cada metodo.
|
||||
self._local = threading.local()
|
||||
self._init_schema()
|
||||
|
||||
def _connect(self) -> sqlite3.Connection:
|
||||
@@ -192,10 +435,58 @@ class Store:
|
||||
for col, coltype in _VIDEO_COLUMNS.items():
|
||||
if col not in existing:
|
||||
conn.execute(f"ALTER TABLE videos ADD COLUMN {col} {coltype}")
|
||||
ch_existing = {row["name"] for row in conn.execute("PRAGMA table_info(channels)")}
|
||||
for col, coltype in _CHANNEL_COLUMNS.items():
|
||||
if col not in ch_existing:
|
||||
conn.execute(f"ALTER TABLE channels ADD COLUMN {col} {coltype}")
|
||||
conn.executescript(_EXTRA_SCHEMA)
|
||||
# Unconditional in spirit, not in cost: a row can also arrive
|
||||
# unranked afterwards, and an unranked row is displayed in the
|
||||
# wrong place rather than merely in an arbitrary one. But
|
||||
# _rank_unranked scans the whole videos table, and Store is built
|
||||
# on every webapp import — so gate it behind the same predicate it
|
||||
# matches on, which is a single indexed probe.
|
||||
if conn.execute(
|
||||
"SELECT 1 FROM videos WHERE channel_seq IS NULL LIMIT 1"
|
||||
).fetchone() is not None:
|
||||
_rank_unranked(conn)
|
||||
|
||||
@contextmanager
|
||||
def transaction(self):
|
||||
"""Agrupa varias escrituras de Store en una sola conexion y un commit.
|
||||
|
||||
`process_video` costaba ~6 connect/commit por video (un fsync cada
|
||||
uno en WAL). Dentro de este bloque, toda llamada a Store hecha desde
|
||||
el MISMO hilo se une a la conexion ambiental; un error revierte el
|
||||
grupo entero, que es la atomicidad por video que se quiere de todos
|
||||
modos.
|
||||
"""
|
||||
conn = self._connect()
|
||||
stack: list[sqlite3.Connection] = getattr(self._local, "stack", None) or []
|
||||
self._local.stack = stack
|
||||
stack.append(conn)
|
||||
try:
|
||||
yield conn
|
||||
conn.commit()
|
||||
except BaseException:
|
||||
conn.rollback()
|
||||
raise
|
||||
finally:
|
||||
stack.remove(conn)
|
||||
conn.close()
|
||||
|
||||
@contextmanager
|
||||
def _cursor(self) -> Iterator[sqlite3.Cursor]:
|
||||
# Dentro de transaction(): misma conexion y sin commit intermedio —
|
||||
# el bloque externo decide cuando el trabajo se vuelve durable.
|
||||
stack: list[sqlite3.Connection] = getattr(self._local, "stack", None) or []
|
||||
if stack:
|
||||
cur = stack[-1].cursor()
|
||||
try:
|
||||
yield cur
|
||||
finally:
|
||||
cur.close()
|
||||
return
|
||||
conn = self._connect()
|
||||
try:
|
||||
yield conn.cursor()
|
||||
@@ -205,18 +496,19 @@ class Store:
|
||||
|
||||
# ------------------------------------------------------------------ channels
|
||||
|
||||
def upsert_channel(self, channel_id: str, handle: str | None, name: str | None, video_count: int = 0) -> None:
|
||||
def upsert_channel(self, channel_id: str, handle: str | None, name: str | None, video_count: int = 0, avatar: str | None = None) -> None:
|
||||
now = _now_iso()
|
||||
with self._cursor() as cur:
|
||||
cur.execute(
|
||||
"""INSERT INTO channels (channel_id, handle, name, last_scraped, video_count)
|
||||
VALUES (?, ?, ?, ?, ?)
|
||||
"""INSERT INTO channels (channel_id, handle, name, last_scraped, video_count, avatar)
|
||||
VALUES (?, ?, ?, ?, ?, ?)
|
||||
ON CONFLICT(channel_id) DO UPDATE SET
|
||||
handle = excluded.handle,
|
||||
name = excluded.name,
|
||||
last_scraped = excluded.last_scraped,
|
||||
video_count = excluded.video_count""",
|
||||
(channel_id, handle, name, now, video_count),
|
||||
video_count = excluded.video_count,
|
||||
avatar = COALESCE(excluded.avatar, channels.avatar)""",
|
||||
(channel_id, handle, name, now, video_count, avatar),
|
||||
)
|
||||
|
||||
def list_channels(self) -> list[dict]:
|
||||
@@ -237,30 +529,172 @@ class Store:
|
||||
cur.execute("DELETE FROM videos WHERE channel_id = ?", (channel_id,))
|
||||
cur.execute("DELETE FROM channels WHERE channel_id = ?", (channel_id,))
|
||||
|
||||
def known_video_ids(self, channel_id: str) -> set[str]:
|
||||
"""Every video id already recorded for a channel — the boundary an
|
||||
incremental discovery walks back to."""
|
||||
with self._cursor() as cur:
|
||||
cur.execute("SELECT video_id FROM videos WHERE channel_id = ?", (channel_id,))
|
||||
return {row["video_id"] for row in cur.fetchall()}
|
||||
|
||||
def latest_upload_date(self, channel_id: str) -> str | None:
|
||||
"""Newest known upload date (YYYYMMDD) for a channel, or None."""
|
||||
with self._cursor() as cur:
|
||||
cur.execute(
|
||||
"SELECT MAX(upload_date) AS d FROM videos "
|
||||
"WHERE channel_id = ? AND upload_date IS NOT NULL AND upload_date <> ''",
|
||||
(channel_id,),
|
||||
)
|
||||
row = cur.fetchone()
|
||||
return row["d"] if row and row["d"] else None
|
||||
|
||||
def mark_channel_synced(self, channel_id: str) -> dict[str, Any]:
|
||||
"""Refresh a channel's counters from what is actually stored.
|
||||
|
||||
Incremental discovery only ever sees the newest slice, so `video_count`
|
||||
has to be recounted here — passing len(refs) would shrink an 848-video
|
||||
channel to the size of the window.
|
||||
"""
|
||||
now = _now_iso()
|
||||
with self._cursor() as cur:
|
||||
cur.execute("SELECT COUNT(*) AS n FROM videos WHERE channel_id = ?", (channel_id,))
|
||||
count = int(cur.fetchone()["n"])
|
||||
cur.execute(
|
||||
"SELECT MAX(upload_date) AS d FROM videos "
|
||||
"WHERE channel_id = ? AND upload_date IS NOT NULL AND upload_date <> ''",
|
||||
(channel_id,),
|
||||
)
|
||||
row = cur.fetchone()
|
||||
last_date = row["d"] if row and row["d"] else None
|
||||
cur.execute(
|
||||
"""UPDATE channels
|
||||
SET video_count = ?, last_scraped = ?, last_synced_at = ?,
|
||||
last_video_date = COALESCE(?, last_video_date)
|
||||
WHERE channel_id = ?""",
|
||||
(count, now, now, last_date, channel_id),
|
||||
)
|
||||
return {"video_count": count, "last_video_date": last_date, "last_synced_at": now}
|
||||
|
||||
def update_channel_meta(self, channel_id: str, *, name: str | None = None, avatar: str | None = None) -> bool:
|
||||
"""Update only the fields explicitly passed. Preserves last_scraped and video_count."""
|
||||
sets: list[str] = []
|
||||
params: list[Any] = []
|
||||
if name is not None:
|
||||
sets.append("name = ?"); params.append(name)
|
||||
if avatar is not None:
|
||||
sets.append("avatar = ?"); params.append(avatar)
|
||||
if not sets:
|
||||
return False
|
||||
params.append(channel_id)
|
||||
with self._cursor() as cur:
|
||||
cur.execute(f"UPDATE channels SET {', '.join(sets)} WHERE channel_id = ?", params)
|
||||
return cur.rowcount > 0
|
||||
|
||||
# ------------------------------------------------------------------ videos
|
||||
|
||||
def upsert_videos(self, refs: list[VideoRef]) -> int:
|
||||
"""Insert/refresh discovered videos and re-rank the channel's recency order.
|
||||
|
||||
`refs` arrive in /videos-tab order (newest first). That order is the only
|
||||
chronological signal discovery yields — the flat listing carries no
|
||||
upload_date — so it is recorded as `channel_seq` and is what lets the UI
|
||||
place a video that has never been extracted where YouTube would show it.
|
||||
|
||||
The window is ranked ABOVE the channel's current maximum rather than from
|
||||
zero: a sync only fetches the newest slice, and everything it did not
|
||||
fetch is by construction older than everything it did. Lifting the window
|
||||
keeps both halves consistently ordered without re-walking the channel.
|
||||
"""
|
||||
now = _now_iso()
|
||||
inserted = 0
|
||||
with self._cursor() as cur:
|
||||
for r in refs:
|
||||
incoming = {r.video_id for r in refs}
|
||||
existing: set[str] = set()
|
||||
if incoming:
|
||||
placeholders = ",".join("?" for _ in incoming)
|
||||
cur.execute(
|
||||
"""INSERT INTO videos
|
||||
(video_id, channel_id, title, url, upload_date, duration, status, discovered_at)
|
||||
VALUES (?, ?, ?, ?, ?, ?, 'pending', ?)
|
||||
ON CONFLICT(video_id) DO UPDATE SET
|
||||
title = excluded.title,
|
||||
upload_date = excluded.upload_date,
|
||||
duration = excluded.duration""",
|
||||
(r.video_id, r.channel_id, r.title, r.url, r.upload_date, r.duration, now),
|
||||
f"SELECT video_id FROM videos WHERE video_id IN ({placeholders})",
|
||||
list(incoming),
|
||||
)
|
||||
if cur.rowcount > 0:
|
||||
existing = {row["video_id"] for row in cur.fetchall()}
|
||||
|
||||
by_channel: dict[str, list[VideoRef]] = {}
|
||||
for r in refs:
|
||||
by_channel.setdefault(r.channel_id, []).append(r)
|
||||
seqs: dict[str, int] = {}
|
||||
for channel_id, group in by_channel.items():
|
||||
# Honour an explicit position when discovery set one; otherwise
|
||||
# the list order is the listing order.
|
||||
ordered = sorted(
|
||||
enumerate(group),
|
||||
key=lambda pair: pair[1].position if pair[1].position is not None else pair[0],
|
||||
)
|
||||
cur.execute(
|
||||
"SELECT COALESCE(MAX(channel_seq), 0) AS m FROM videos WHERE channel_id = ?",
|
||||
(channel_id,),
|
||||
)
|
||||
base = int(cur.fetchone()["m"] or 0)
|
||||
width = len(ordered)
|
||||
for rank, (_, r) in enumerate(ordered):
|
||||
seqs[r.video_id] = base + width - rank
|
||||
|
||||
# One executemany instead of a prepared-statement round per ref:
|
||||
# discovery hands over hundreds of rows at once and the statement
|
||||
# text is identical for all of them.
|
||||
cur.executemany(
|
||||
"""INSERT INTO videos
|
||||
(video_id, channel_id, title, url, upload_date,
|
||||
upload_date_approx, duration, availability,
|
||||
channel_seq, status, discovered_at)
|
||||
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, 'pending', ?)
|
||||
ON CONFLICT(video_id) DO UPDATE SET
|
||||
title = COALESCE(excluded.title, videos.title),
|
||||
upload_date = CASE
|
||||
WHEN excluded.upload_date IS NULL THEN videos.upload_date
|
||||
WHEN videos.upload_date IS NULL THEN excluded.upload_date
|
||||
WHEN COALESCE(videos.upload_date_approx, 0) = 1
|
||||
THEN excluded.upload_date
|
||||
ELSE videos.upload_date END,
|
||||
upload_date_approx = CASE
|
||||
WHEN videos.upload_date IS NOT NULL
|
||||
AND COALESCE(videos.upload_date_approx, 0) = 0
|
||||
THEN videos.upload_date_approx
|
||||
ELSE COALESCE(excluded.upload_date_approx,
|
||||
videos.upload_date_approx) END,
|
||||
duration = COALESCE(excluded.duration, videos.duration),
|
||||
availability = COALESCE(excluded.availability, videos.availability),
|
||||
channel_seq = COALESCE(excluded.channel_seq, videos.channel_seq)""",
|
||||
[
|
||||
(r.video_id, r.channel_id, r.title, r.url, r.upload_date,
|
||||
int(r.date_approx or 0), r.duration, r.availability,
|
||||
seqs.get(r.video_id), now)
|
||||
for r in refs
|
||||
],
|
||||
)
|
||||
for r in refs:
|
||||
if r.video_id not in existing:
|
||||
inserted += 1
|
||||
existing.add(r.video_id)
|
||||
return inserted
|
||||
|
||||
def get_pending(self, channel_id: str | None = None, limit: int | None = None) -> list[VideoRow]:
|
||||
def get_pending(
|
||||
self,
|
||||
channel_id: str | None = None,
|
||||
limit: int | None = None,
|
||||
*,
|
||||
include_blocked: bool = False,
|
||||
) -> list[VideoRow]:
|
||||
"""Pending videos, excluding ones discovery already told us we cannot
|
||||
fetch (members-only, premium, private). Bulk runs should not spend
|
||||
requests on those; an explicit per-video "Process" click still can,
|
||||
which is what makes the membership case recoverable.
|
||||
"""
|
||||
sql = "SELECT * FROM videos WHERE status = 'pending'"
|
||||
params: list[Any] = []
|
||||
if not include_blocked:
|
||||
blocking = sorted(BLOCKING_AVAILABILITY)
|
||||
marks = ",".join("?" for _ in blocking)
|
||||
sql += f" AND (availability IS NULL OR availability NOT IN ({marks}))"
|
||||
params.extend(blocking)
|
||||
if channel_id:
|
||||
sql += " AND channel_id = ?"
|
||||
params.append(channel_id)
|
||||
@@ -300,6 +734,7 @@ class Store:
|
||||
sort: str = "upload_date_desc",
|
||||
page: int = 1,
|
||||
size: int = 50,
|
||||
blocked: bool | None = None,
|
||||
) -> tuple[list[VideoRow], int]:
|
||||
where: list[str] = []
|
||||
params: list[Any] = []
|
||||
@@ -307,6 +742,21 @@ class Store:
|
||||
where.append("channel_id = ?"); params.append(channel_id)
|
||||
if status:
|
||||
where.append("status = ?"); params.append(status)
|
||||
if blocked is not None:
|
||||
# Match on the discovery signal or on the recorded error, so rows
|
||||
# burned in before `availability` existed are still findable.
|
||||
marks = ",".join("?" for _ in sorted(BLOCKING_AVAILABILITY))
|
||||
# COALESCE is load-bearing: with a NULL availability the IN test is
|
||||
# NULL, and NOT(NULL) is NULL, so the negated branch would silently
|
||||
# return zero rows instead of "everything fetchable".
|
||||
expr = (
|
||||
f"(COALESCE(availability, '') IN ({marks}) "
|
||||
"OR COALESCE(error_msg, '') LIKE '%members-only%' "
|
||||
"OR COALESCE(error_msg, '') LIKE '%Join this channel to get access%' "
|
||||
"OR COALESCE(error_msg, '') LIKE '%Private video%')"
|
||||
)
|
||||
where.append(expr if blocked else f"NOT {expr}")
|
||||
params.extend(sorted(BLOCKING_AVAILABILITY))
|
||||
if date_from:
|
||||
where.append("upload_date >= ?"); params.append(date_from.replace("-", ""))
|
||||
if date_to:
|
||||
@@ -317,19 +767,14 @@ class Store:
|
||||
where.append("(title LIKE ? OR description LIKE ?)")
|
||||
params.extend([f"%{q}%", f"%{q}%"])
|
||||
clause = ("WHERE " + " AND ".join(where)) if where else ""
|
||||
order = {
|
||||
"upload_date_desc": "upload_date DESC",
|
||||
"upload_date_asc": "upload_date ASC",
|
||||
"duration_desc": "duration DESC",
|
||||
"views_desc": "view_count DESC",
|
||||
"title_asc": "title ASC",
|
||||
}.get(sort, "upload_date DESC")
|
||||
order = _order_clause(sort)
|
||||
offset = max(0, (page - 1) * size)
|
||||
with self._cursor() as cur:
|
||||
cur.execute(f"SELECT COUNT(*) AS n FROM videos {clause}", params)
|
||||
total = cur.fetchone()["n"]
|
||||
cur.execute(
|
||||
f"SELECT * FROM videos {clause} ORDER BY {order} LIMIT ? OFFSET ?",
|
||||
f"SELECT videos.*, {SORT_DATE_SQL} AS sort_date FROM videos {clause} "
|
||||
f"ORDER BY {order} LIMIT ? OFFSET ?",
|
||||
[*params, size, offset],
|
||||
)
|
||||
rows = [_row_to_videorow(r) for r in cur.fetchall()]
|
||||
@@ -392,20 +837,70 @@ class Store:
|
||||
(markdown_path, transcript_lang, transcript_src, int(has_chapters), now, video_id),
|
||||
)
|
||||
|
||||
def mark_status(self, video_id: str, status: str) -> None:
|
||||
def mark_status(self, video_id: str, status: str, reason: str | None = None) -> None:
|
||||
"""Set a video's status, optionally recording why.
|
||||
|
||||
`no_subtitles` used to be stored with error_msg=NULL, which left no way
|
||||
to tell "this video has no captions" apart from "the language policy
|
||||
rejected the captions it does have" — the second is recoverable.
|
||||
"""
|
||||
now = _now_iso()
|
||||
with self._cursor() as cur:
|
||||
cur.execute(
|
||||
"UPDATE videos SET status = ?, processed_at = ? WHERE video_id = ?",
|
||||
(status, now, video_id),
|
||||
"UPDATE videos SET status = ?, error_msg = ?, processed_at = ? WHERE video_id = ?",
|
||||
# 2000, not 500: a throttled caption fetch stores the timedtext
|
||||
# URL, and at 500 the cut landed twenty characters before the
|
||||
# `tlang=` parameter — the one token proving the request was for
|
||||
# a machine translation rather than the real transcript.
|
||||
(status, reason[:2000] if reason else None, now, video_id),
|
||||
)
|
||||
|
||||
def update_video_download(
|
||||
self,
|
||||
video_id: str,
|
||||
status: str,
|
||||
*,
|
||||
path: str | None = None,
|
||||
filename: str | None = None,
|
||||
size: int | None = None,
|
||||
error: str | None = None,
|
||||
) -> None:
|
||||
"""Persist media-download state independently from transcript state."""
|
||||
now = _now_iso() if status == "done" else None
|
||||
with self._cursor() as cur:
|
||||
cur.execute(
|
||||
"""UPDATE videos SET
|
||||
video_download_status = ?,
|
||||
video_path = COALESCE(?, video_path),
|
||||
video_filename = COALESCE(?, video_filename),
|
||||
video_size = COALESCE(?, video_size),
|
||||
video_downloaded_at = COALESCE(?, video_downloaded_at),
|
||||
video_error = ?
|
||||
WHERE video_id = ?""",
|
||||
(status, path, filename, size, now, error[:2000] if error else None, video_id),
|
||||
)
|
||||
|
||||
def set_availability(self, video_id: str, availability: str | None) -> None:
|
||||
"""Record what a full extraction learned; more authoritative than the
|
||||
flat listing, which omits the field for most entries."""
|
||||
if not availability:
|
||||
return
|
||||
with self._cursor() as cur:
|
||||
cur.execute(
|
||||
"UPDATE videos SET availability = ? WHERE video_id = ?",
|
||||
(str(availability), video_id),
|
||||
)
|
||||
|
||||
def set_upload_date(self, video_id: str, upload_date: str | None) -> None:
|
||||
if not upload_date:
|
||||
return
|
||||
with self._cursor() as cur:
|
||||
# La fecha de la extraccion es exacta: sobrescribe una aproximada
|
||||
# del discovery y apaga su bandera; nunca degrada una real.
|
||||
cur.execute(
|
||||
"UPDATE videos SET upload_date = ? WHERE video_id = ? AND upload_date IS NULL",
|
||||
"""UPDATE videos SET upload_date = ?, upload_date_approx = 0
|
||||
WHERE video_id = ?
|
||||
AND (upload_date IS NULL OR COALESCE(upload_date_approx, 0) = 1)""",
|
||||
(upload_date, video_id),
|
||||
)
|
||||
|
||||
@@ -610,66 +1105,100 @@ class Store:
|
||||
def dashboard(self) -> dict:
|
||||
with self._cursor() as cur:
|
||||
channels = [dict(r) for r in cur.execute("SELECT * FROM channels ORDER BY name").fetchall()]
|
||||
# One GROUP BY instead of one aggregate query per channel (N+1).
|
||||
# Channels with zero videos are absent from the grouping and get
|
||||
# the same zeros/NULLs the per-channel query used to return.
|
||||
agg = {
|
||||
r["channel_id"]: r
|
||||
for r in cur.execute(
|
||||
"SELECT channel_id, COUNT(*) AS n, COALESCE(SUM(duration),0) AS dur, "
|
||||
"MIN(upload_date) AS mind, MAX(upload_date) AS maxd, "
|
||||
"COALESCE(SUM(view_count),0) AS views, COALESCE(SUM(like_count),0) AS likes "
|
||||
"FROM videos GROUP BY channel_id"
|
||||
).fetchall()
|
||||
}
|
||||
for ch in channels:
|
||||
cid = ch["channel_id"]
|
||||
row = cur.execute(
|
||||
"SELECT COUNT(*) AS n, COALESCE(SUM(duration),0) AS dur, MIN(upload_date) AS mind, MAX(upload_date) AS maxd, COALESCE(SUM(view_count),0) AS views, COALESCE(SUM(like_count),0) AS likes FROM videos WHERE channel_id = ?",
|
||||
(cid,),
|
||||
).fetchone()
|
||||
ch["video_count_db"] = row["n"]
|
||||
ch["total_duration"] = row["dur"]
|
||||
ch["date_min"] = row["mind"]
|
||||
ch["date_max"] = row["maxd"]
|
||||
ch["total_views"] = row["views"]
|
||||
ch["total_likes"] = row["likes"]
|
||||
row = agg.get(ch["channel_id"])
|
||||
ch["video_count_db"] = row["n"] if row else 0
|
||||
ch["total_duration"] = row["dur"] if row else 0
|
||||
ch["date_min"] = row["mind"] if row else None
|
||||
ch["date_max"] = row["maxd"] if row else None
|
||||
ch["total_views"] = row["views"] if row else 0
|
||||
ch["total_likes"] = row["likes"] if row else 0
|
||||
status_breakdown = {
|
||||
r["status"]: r["n"]
|
||||
for r in cur.execute("SELECT status, COUNT(*) AS n FROM videos GROUP BY status").fetchall()
|
||||
}
|
||||
# uploads over time (by month)
|
||||
uploads = [
|
||||
{"month": r["m"], "count": r["n"]}
|
||||
# top tags (tags is JSON array text). Counted in SQL via json_each
|
||||
# instead of loading every tags string into Python; json_valid
|
||||
# skips rows the old json.loads/except path also skipped, and the
|
||||
# TRIM/LOWER/empty-filter mirrors the per-tag normalisation.
|
||||
tag_counts = {
|
||||
r["tag"]: r["n"]
|
||||
for r in cur.execute(
|
||||
"SELECT substr(upload_date,1,6) AS m, COUNT(*) AS n FROM videos WHERE upload_date IS NOT NULL GROUP BY m ORDER BY m"
|
||||
"SELECT LOWER(TRIM(je.value)) AS tag, COUNT(*) AS n "
|
||||
"FROM videos v, json_each(v.tags) je "
|
||||
"WHERE json_valid(v.tags) AND TRIM(je.value) <> '' "
|
||||
"GROUP BY tag ORDER BY n DESC LIMIT 20"
|
||||
).fetchall()
|
||||
]
|
||||
# duration histogram (buckets)
|
||||
hist = [
|
||||
{"bucket": r["b"], "count": r["n"]}
|
||||
for r in cur.execute(
|
||||
"""SELECT
|
||||
CASE WHEN duration < 300 THEN '<5m'
|
||||
WHEN duration < 600 THEN '5-10m'
|
||||
WHEN duration < 1200 THEN '10-20m'
|
||||
WHEN duration < 2400 THEN '20-40m'
|
||||
ELSE '40m+' END AS b,
|
||||
COUNT(*) AS n
|
||||
FROM videos WHERE duration IS NOT NULL GROUP BY b"""
|
||||
).fetchall()
|
||||
]
|
||||
# top tags (tags is JSON array text)
|
||||
tag_rows = cur.execute("SELECT tags FROM videos WHERE tags IS NOT NULL AND tags != '[]'").fetchall()
|
||||
tag_counts: dict[str, int] = {}
|
||||
for tr in tag_rows:
|
||||
try:
|
||||
for t in json.loads(tr["tags"]):
|
||||
t = (t or "").strip().lower()
|
||||
if t:
|
||||
tag_counts[t] = tag_counts.get(t, 0) + 1
|
||||
except (json.JSONDecodeError, TypeError):
|
||||
continue
|
||||
top_tags = sorted(tag_counts.items(), key=lambda x: x[1], reverse=True)[:20]
|
||||
}
|
||||
top_tags = sorted(tag_counts.items(), key=lambda x: x[1], reverse=True)
|
||||
return {
|
||||
"channels": channels,
|
||||
"status_breakdown": status_breakdown,
|
||||
"uploads_over_time": uploads,
|
||||
"duration_histogram": hist,
|
||||
"top_tags": [{"tag": t, "count": c} for t, c in top_tags],
|
||||
}
|
||||
|
||||
RETRYABLE_STATUSES = ("error", "no_subtitles")
|
||||
|
||||
# Failures that will never resolve by trying again: paying for a membership
|
||||
# or the video coming back from the dead are not retry outcomes. Retrying
|
||||
# them just spends requests against the rate limit that the videos which
|
||||
# CAN succeed need. Kept narrow on purpose — YouTube's throttling message
|
||||
# ("rate-limited ... try again later") is retryable and must not match here.
|
||||
PERMANENT_ERROR_PATTERNS = (
|
||||
"%members-only%",
|
||||
"%Join this channel to get access%",
|
||||
"%Private video%",
|
||||
"%This video has been removed%",
|
||||
"%video has been terminated%",
|
||||
)
|
||||
|
||||
def _permanent_sql(self, negate: bool = True) -> str:
|
||||
clause = " OR ".join("error_msg LIKE ?" for _ in self.PERMANENT_ERROR_PATTERNS)
|
||||
return f"NOT (error_msg IS NOT NULL AND ({clause}))" if negate else f"(error_msg IS NOT NULL AND ({clause}))"
|
||||
|
||||
def reset_errors(self, channel_id: str | None = None) -> int:
|
||||
sql = "UPDATE videos SET status = 'pending', error_msg = NULL WHERE status = 'error'"
|
||||
params: list[Any] = []
|
||||
return self.reset_videos(channel_id, statuses=("error",))
|
||||
|
||||
def reset_videos(
|
||||
self,
|
||||
channel_id: str | None = None,
|
||||
statuses: tuple[str, ...] | list[str] = ("error",),
|
||||
*,
|
||||
keep_done: bool = True,
|
||||
include_permanent: bool = False,
|
||||
) -> int:
|
||||
"""Send videos in the given terminal statuses back to `pending`.
|
||||
|
||||
`no_subtitles` has to be resettable, not just `error`: it is recorded
|
||||
whenever the language policy matched nothing, so a config fix is
|
||||
worthless if the affected rows can never be retried. `done` is never
|
||||
reset here — re-running finished work is what burns rate limits.
|
||||
"""
|
||||
allowed = [s for s in statuses if s in self.RETRYABLE_STATUSES or not keep_done]
|
||||
allowed = [s for s in allowed if s != "done"]
|
||||
if not allowed:
|
||||
return 0
|
||||
placeholders = ",".join("?" for _ in allowed)
|
||||
sql = (
|
||||
f"UPDATE videos SET status = 'pending', error_msg = NULL "
|
||||
f"WHERE status IN ({placeholders})"
|
||||
)
|
||||
params: list[Any] = list(allowed)
|
||||
if not include_permanent:
|
||||
sql += f" AND {self._permanent_sql()}"
|
||||
params.extend(self.PERMANENT_ERROR_PATTERNS)
|
||||
if channel_id:
|
||||
sql += " AND channel_id = ?"
|
||||
params.append(channel_id)
|
||||
@@ -677,6 +1206,31 @@ class Store:
|
||||
cur.execute(sql, params)
|
||||
return cur.rowcount
|
||||
|
||||
def retryable_counts(self, channel_id: str | None = None) -> dict[str, int]:
|
||||
"""Videos stuck in each retryable status, plus how many are permanently
|
||||
blocked. The retry button must not promise to fix members-only videos.
|
||||
"""
|
||||
placeholders = ",".join("?" for _ in self.RETRYABLE_STATUSES)
|
||||
base = f"FROM videos WHERE status IN ({placeholders})"
|
||||
base_params: list[Any] = list(self.RETRYABLE_STATUSES)
|
||||
tail = ""
|
||||
if channel_id:
|
||||
tail = " AND channel_id = ?"
|
||||
with self._cursor() as cur:
|
||||
cur.execute(
|
||||
f"SELECT status, COUNT(*) n {base} AND {self._permanent_sql()}{tail} GROUP BY status",
|
||||
base_params + list(self.PERMANENT_ERROR_PATTERNS) + ([channel_id] if channel_id else []),
|
||||
)
|
||||
out = {s: 0 for s in self.RETRYABLE_STATUSES}
|
||||
for row in cur.fetchall():
|
||||
out[row["status"]] = int(row["n"])
|
||||
cur.execute(
|
||||
f"SELECT COUNT(*) n {base} AND {self._permanent_sql(negate=False)}{tail}",
|
||||
base_params + list(self.PERMANENT_ERROR_PATTERNS) + ([channel_id] if channel_id else []),
|
||||
)
|
||||
out["permanent"] = int(cur.fetchone()["n"])
|
||||
return out
|
||||
|
||||
|
||||
def _sanitize_fts(query: str) -> str:
|
||||
# Build a safe AND FTS5 query from whitespace-separated terms.
|
||||
@@ -695,6 +1249,52 @@ def _now_iso() -> str:
|
||||
return datetime.now(timezone.utc).isoformat(timespec="seconds")
|
||||
|
||||
|
||||
def _rank_unranked(conn: sqlite3.Connection) -> int:
|
||||
"""Give every row that has no recency rank one, above its channel's maximum.
|
||||
|
||||
Covers two cases with the same rule. On a database that predates the column
|
||||
every row is unranked, and this is the initial seed. Afterwards a row can
|
||||
still arrive unranked from a process running the pre-`channel_seq` code — a
|
||||
long-lived server that has not been restarted since the migration — which is
|
||||
exactly what happened on the live library: three videos discovered after the
|
||||
migration, all NULL.
|
||||
|
||||
An unranked row is not merely unordered, it is actively misplaced:
|
||||
`COALESCE(channel_seq, -1)` in SORT_DATE_SQL makes rule 2 pick the OLDEST
|
||||
dated video in the channel, so a video discovery has only just found — one
|
||||
of the channel's newest — is shown at the very bottom. Measured: two
|
||||
brand-new Alex Hormozi uploads displayed with sort_date 20180720, second to
|
||||
last of 513.
|
||||
|
||||
The ordering is `discovered_at DESC, rowid ASC`, which is the order the rows
|
||||
were actually learned in: `upsert_videos` stamps one timestamp per discovery
|
||||
batch, discovery only ever adds ids newer than everything already stored, and
|
||||
within a batch the insert order is the channel's /videos tab —
|
||||
reverse-chronological. It is a reconstruction, not an observation; every
|
||||
later sync overwrites the slice it touches with the real thing.
|
||||
|
||||
Idempotent: with nothing unranked it does no writes at all.
|
||||
"""
|
||||
by_channel: dict[str, list[str]] = {}
|
||||
for row in conn.execute(
|
||||
"SELECT video_id, channel_id FROM videos WHERE channel_seq IS NULL "
|
||||
"ORDER BY channel_id, discovered_at DESC, rowid ASC"
|
||||
):
|
||||
by_channel.setdefault(row["channel_id"], []).append(row["video_id"])
|
||||
if not by_channel:
|
||||
return 0
|
||||
updates: list[tuple[int, str]] = []
|
||||
for channel_id, ids in by_channel.items():
|
||||
row = conn.execute(
|
||||
"SELECT COALESCE(MAX(channel_seq), 0) AS m FROM videos WHERE channel_id = ?",
|
||||
(channel_id,),
|
||||
).fetchone()
|
||||
base = int(row["m"] or 0)
|
||||
updates.extend((base + len(ids) - i, vid) for i, vid in enumerate(ids))
|
||||
conn.executemany("UPDATE videos SET channel_seq = ? WHERE video_id = ?", updates)
|
||||
return len(updates)
|
||||
|
||||
|
||||
def _row_to_videorow(row: sqlite3.Row) -> VideoRow:
|
||||
keys = row.keys()
|
||||
return VideoRow(
|
||||
@@ -717,6 +1317,16 @@ def _row_to_videorow(row: sqlite3.Row) -> VideoRow:
|
||||
description=row["description"] if "description" in keys else None,
|
||||
chapters_json=row["chapters_json"] if "chapters_json" in keys else None,
|
||||
segments_json=row["segments_json"] if "segments_json" in keys else None,
|
||||
availability=row["availability"] if "availability" in keys else None,
|
||||
channel_seq=row["channel_seq"] if "channel_seq" in keys else None,
|
||||
upload_date_approx=row["upload_date_approx"] if "upload_date_approx" in keys else None,
|
||||
video_download_status=row["video_download_status"] if "video_download_status" in keys else None,
|
||||
video_path=row["video_path"] if "video_path" in keys else None,
|
||||
video_filename=row["video_filename"] if "video_filename" in keys else None,
|
||||
video_size=row["video_size"] if "video_size" in keys else None,
|
||||
video_downloaded_at=row["video_downloaded_at"] if "video_downloaded_at" in keys else None,
|
||||
video_error=row["video_error"] if "video_error" in keys else None,
|
||||
sort_date=row["sort_date"] if "sort_date" in keys else None,
|
||||
)
|
||||
|
||||
|
||||
|
||||
+323
-30
@@ -1,6 +1,7 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from concurrent.futures import ThreadPoolExecutor
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
@@ -12,8 +13,22 @@ from .. import analysis as analysis_mod
|
||||
from .. import cookies as cookies_mod
|
||||
from .. import export as export_mod
|
||||
from ..config import Config
|
||||
from ..discover import extract_handle
|
||||
from ..ratelimit import polite_sleep
|
||||
from ..render import safe_dirname
|
||||
from .. import store as store_mod
|
||||
from ..store import Store
|
||||
|
||||
#: How many thumbnails a *implicit* whole-channel request may fetch. The UI
|
||||
#: pulls the rest lazily through /api/thumbnails/{id} as rows scroll into view,
|
||||
#: so this only bounds the eager burst that follows "Add channel".
|
||||
THUMBNAIL_AUTO_LIMIT = 60
|
||||
|
||||
#: Thumbnail fetches go to the i.ytimg.com CDN (not youtube.com), so the
|
||||
#: Pacer does not apply to them; a small pool turns ~60 sequential round
|
||||
#: trips into ~10 batches without hammering the CDN.
|
||||
THUMBNAIL_WORKERS = 6
|
||||
|
||||
|
||||
def build_router(store: Store, cfg: Config, jobs) -> APIRouter:
|
||||
r = APIRouter(prefix="/api")
|
||||
@@ -28,16 +43,33 @@ def build_router(store: Store, cfg: Config, jobs) -> APIRouter:
|
||||
def channels_list():
|
||||
return {"items": store.list_channels()}
|
||||
|
||||
# Sync `def`, not `async def`, on purpose: the body blocks on yt-dlp for as
|
||||
# long as the channel takes to page. FastAPI runs `def` handlers on the
|
||||
# threadpool, whereas an `async def` would hold the event loop and freeze
|
||||
# every other request — including the SSE stream of a running job.
|
||||
@r.post("/channels")
|
||||
async def add_channel(payload: dict):
|
||||
from ..discover import discover_channel
|
||||
def add_channel(payload: dict):
|
||||
from ..discover import discover_channel, deep_channel_avatar
|
||||
from ..pipeline import cache_channel_avatar
|
||||
url = (payload or {}).get("url")
|
||||
if not url:
|
||||
raise HTTPException(400, "url required")
|
||||
cid, name, refs = discover_channel(url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests)
|
||||
store.upsert_channel(cid, _handle(url), name, len(refs))
|
||||
cid, name, avatar, refs = discover_channel(url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests)
|
||||
avatar_cached = False
|
||||
if not avatar:
|
||||
avatar = deep_channel_avatar(url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests)
|
||||
store.upsert_channel(cid, extract_handle(url), name, len(refs), avatar=avatar)
|
||||
store.upsert_videos(refs)
|
||||
return {"channel_id": cid, "name": name, "video_count": len(refs)}
|
||||
if avatar:
|
||||
avatars_dir = Path(cfg.output_dir_resolved).parent / "avatars"
|
||||
avatar_cached = cache_channel_avatar(cid, avatar, avatars_dir)
|
||||
return {
|
||||
"channel_id": cid,
|
||||
"name": name,
|
||||
"video_count": len(refs),
|
||||
"avatar": avatar,
|
||||
"avatar_cached": avatar_cached,
|
||||
}
|
||||
|
||||
@r.delete("/channels/{channel_id}")
|
||||
def del_channel(channel_id: str):
|
||||
@@ -52,8 +84,11 @@ def build_router(store: Store, cfg: Config, jobs) -> APIRouter:
|
||||
date_to: str | None = Query(None, alias="to"),
|
||||
min_dur: int | None = None, q: str | None = None,
|
||||
sort: str = "upload_date_desc", page: int = 1, size: int = 50,
|
||||
blocked: bool | None = None,
|
||||
):
|
||||
rows, total = store.query_videos(channel, status, date_from, date_to, min_dur, q, sort, page, size)
|
||||
rows, total = store.query_videos(
|
||||
channel, status, date_from, date_to, min_dur, q, sort, page, size, blocked=blocked,
|
||||
)
|
||||
return {"items": [_video_dict(v) for v in rows], "total": total, "page": page, "size": size}
|
||||
|
||||
@r.get("/videos/{video_id}")
|
||||
@@ -140,10 +175,18 @@ def build_router(store: Store, cfg: Config, jobs) -> APIRouter:
|
||||
return {"deleted": n}
|
||||
|
||||
@r.delete("/scrape/{job_id}")
|
||||
def delete_or_cancel_job(job_id: str):
|
||||
def delete_or_cancel_job(job_id: str, action: str = "cancel"):
|
||||
job = store.get_job(job_id)
|
||||
if not job:
|
||||
raise HTTPException(404, "job not found")
|
||||
# action=purge is unconditional: removes the row regardless of status.
|
||||
# Used by the "Recent jobs" table, where the user's intent is to clear
|
||||
# history (including rows that survived a server restart as zombies).
|
||||
# action=cancel (default) only works for terminal jobs by clearing the
|
||||
# row; for live jobs it asks the JobManager to stop them cooperatively.
|
||||
if action == "purge":
|
||||
store.delete_job(job_id)
|
||||
return {"deleted": job_id}
|
||||
if job.status in store.TERMINAL_STATUSES:
|
||||
store.delete_job(job_id)
|
||||
return {"deleted": job_id}
|
||||
@@ -194,6 +237,31 @@ def build_router(store: Store, cfg: Config, jobs) -> APIRouter:
|
||||
cookies_mod.set_active(store, imported[0])
|
||||
return {"imported": imported}
|
||||
|
||||
@r.post("/cookies/from-browser")
|
||||
def cookies_from_browser(payload: dict):
|
||||
"""Extract youtube.com cookies from a local browser (default: Brave).
|
||||
|
||||
Requires the browser to be fully closed; Chromium keeps its cookie
|
||||
database locked while running.
|
||||
"""
|
||||
payload = payload or {}
|
||||
browser = (payload.get("browser") or "brave").strip().lower()
|
||||
profile = payload.get("profile") or None
|
||||
try:
|
||||
cid = cookies_mod.import_from_browser(store, browser=browser, profile=profile)
|
||||
except cookies_mod.BrowserCookieLockedError as exc:
|
||||
raise HTTPException(409, str(exc))
|
||||
except Exception as exc:
|
||||
raise HTTPException(400, str(exc))
|
||||
# The user imported it precisely to use it.
|
||||
cookies_mod.set_active(store, cid)
|
||||
row = store.get_cookie(cid)
|
||||
return {
|
||||
"imported": [cid], "active": cid, "browser": browser,
|
||||
"cookie_count": row.cookie_count if row else None,
|
||||
"has_session": row.has_session if row else None,
|
||||
}
|
||||
|
||||
@r.post("/cookies/{cookie_id}/activate")
|
||||
def cookies_activate(cookie_id: str):
|
||||
cookies_mod.set_active(store, cookie_id)
|
||||
@@ -254,8 +322,11 @@ def build_router(store: Store, cfg: Config, jobs) -> APIRouter:
|
||||
"database": _folder_info(Path(cfg.database_path_resolved).parent),
|
||||
}}
|
||||
|
||||
# `def`, not `async def`: same reason as add_channel — the body blocks on
|
||||
# SQLite reads and a subprocess/os.startfile call, which would hold the
|
||||
# event loop if this were async.
|
||||
@r.post("/folders/open")
|
||||
async def open_folder(payload: dict):
|
||||
def open_folder(payload: dict):
|
||||
kind = (payload or {}).get("kind")
|
||||
channel_id = (payload or {}).get("channel_id")
|
||||
data_root = Path(cfg.output_dir_resolved).parent
|
||||
@@ -264,7 +335,7 @@ def build_router(store: Store, cfg: Config, jobs) -> APIRouter:
|
||||
if channel_id:
|
||||
ch = store.get_channel(channel_id) or {}
|
||||
if ch.get("name"):
|
||||
target = Path(cfg.output_dir_resolved) / _safe_dir(ch["name"])
|
||||
target = Path(cfg.output_dir_resolved) / safe_dirname(ch["name"])
|
||||
elif kind == "exports":
|
||||
target = data_root / "exports"
|
||||
elif kind == "audio":
|
||||
@@ -340,6 +411,17 @@ def build_router(store: Store, cfg: Config, jobs) -> APIRouter:
|
||||
raise HTTPException(404, "markdown file missing on disk")
|
||||
return FileResponse(str(p), filename=p.name, media_type="text/markdown")
|
||||
|
||||
@r.post("/videos/{video_id}/open-markdown")
|
||||
def open_video_markdown(video_id: str):
|
||||
v = store.get_video(video_id)
|
||||
if not v or not v.markdown_path:
|
||||
raise HTTPException(404, "markdown not generated yet")
|
||||
p = Path(cfg.output_dir_resolved).parent / v.markdown_path
|
||||
if not p.exists():
|
||||
raise HTTPException(404, "markdown file missing on disk")
|
||||
_open_in_os(p)
|
||||
return {"opened": str(p)}
|
||||
|
||||
@r.api_route("/videos/{video_id}/audio", methods=["GET", "HEAD"])
|
||||
def video_audio(video_id: str, request: Request):
|
||||
# audio is stored as data/audio/<video_id>.mp3 (see jobs._run_audio outtmpl)
|
||||
@@ -348,21 +430,54 @@ def build_router(store: Store, cfg: Config, jobs) -> APIRouter:
|
||||
raise HTTPException(404, "audio not downloaded yet")
|
||||
return FileResponse(str(p), filename=f"{video_id}.mp3", media_type="audio/mpeg")
|
||||
|
||||
@r.api_route("/videos/{video_id}/media", methods=["GET", "HEAD"])
|
||||
def video_media(video_id: str, request: Request, inline: bool = False):
|
||||
v = store.get_video(video_id)
|
||||
if not v or v.video_download_status != "done" or not v.video_path:
|
||||
raise HTTPException(404, "video not downloaded yet")
|
||||
root = Path(cfg.output_dir_resolved).parent.resolve()
|
||||
p = (root / v.video_path).resolve()
|
||||
try:
|
||||
p.relative_to(root)
|
||||
except ValueError:
|
||||
raise HTTPException(500, "invalid stored video path")
|
||||
if not p.exists():
|
||||
raise HTTPException(404, "video file missing on disk")
|
||||
return FileResponse(
|
||||
str(p), filename=v.video_filename or p.name, media_type="video/webm",
|
||||
content_disposition_type="inline" if inline else "attachment",
|
||||
)
|
||||
|
||||
# -------------------------------------------------- thumbnails (local cache)
|
||||
@r.post("/tools/thumbnails")
|
||||
async def download_thumbnails(payload: dict):
|
||||
def download_thumbnails(payload: dict):
|
||||
from ..pipeline import cache_thumbnail
|
||||
video_ids = (payload or {}).get("video_ids") or []
|
||||
channel_id = (payload or {}).get("channel_id")
|
||||
explicit = bool(video_ids)
|
||||
if channel_id and not video_ids:
|
||||
video_ids = [v.video_id for v in store.get_all(channel_id)]
|
||||
# Adding a channel fires this from the frontend with just a channel_id,
|
||||
# which used to mean one CDN request per video in the catalog — ~860 in
|
||||
# a burst for a large channel, on top of the discovery that just ran.
|
||||
# Cap the implicit form; an explicit list of ids is the user asking.
|
||||
limit = int((payload or {}).get("limit") or 0)
|
||||
if not explicit:
|
||||
limit = limit or THUMBNAIL_AUTO_LIMIT
|
||||
skipped = 0
|
||||
if limit and len(video_ids) > limit:
|
||||
skipped = len(video_ids) - limit
|
||||
video_ids = video_ids[:limit]
|
||||
out_dir = Path(cfg.output_dir_resolved).parent / "thumbnails"
|
||||
out_dir.mkdir(parents=True, exist_ok=True)
|
||||
n = 0
|
||||
for vid in video_ids:
|
||||
if cache_thumbnail(store, vid, out_dir):
|
||||
n += 1
|
||||
return {"downloaded": n, "dir": str(out_dir)}
|
||||
# CDN fetches, one per video, each independent: run them on a small
|
||||
# thread pool instead of sequentially. cache_thumbnail owns its own
|
||||
# SQLite connection and writes one file per id, so this is safe;
|
||||
# failures still count as False exactly as before.
|
||||
with ThreadPoolExecutor(max_workers=THUMBNAIL_WORKERS) as pool:
|
||||
results = list(pool.map(lambda vid: cache_thumbnail(store, vid, out_dir), video_ids))
|
||||
n = sum(1 for ok in results if ok)
|
||||
return {"downloaded": n, "dir": str(out_dir), "skipped": skipped}
|
||||
|
||||
@r.get("/thumbnails/{video_id}")
|
||||
def serve_thumbnail(video_id: str):
|
||||
@@ -375,6 +490,152 @@ def build_router(store: Store, cfg: Config, jobs) -> APIRouter:
|
||||
url = thumbnail_url_for(v) if v else f"https://i.ytimg.com/vi/{video_id}/hqdefault.jpg"
|
||||
return RedirectResponse(url=url, status_code=302)
|
||||
|
||||
# -------------------------------------------------- freshness / recovery
|
||||
@r.get("/videos-retryable")
|
||||
def videos_retryable(channel: str | None = None):
|
||||
"""How many videos are stuck in a retryable terminal status.
|
||||
|
||||
`total` excludes the permanently blocked ones (members-only, private,
|
||||
removed) so the retry button never promises to fix them.
|
||||
"""
|
||||
counts = store.retryable_counts(channel)
|
||||
return {
|
||||
"counts": counts,
|
||||
"total": sum(counts.get(s, 0) for s in Store.RETRYABLE_STATUSES),
|
||||
"permanent": counts.get("permanent", 0),
|
||||
}
|
||||
|
||||
# `def`, not `async def`: the UPDATE below blocks on SQLite; run it on the
|
||||
# threadpool like the other blocking handlers.
|
||||
@r.post("/videos/reset")
|
||||
def reset_videos(payload: dict):
|
||||
"""Send `error` / `no_subtitles` videos back to pending so they can be retried.
|
||||
|
||||
`no_subtitles` is resettable on purpose: it is recorded whenever the
|
||||
language policy matched nothing, so fixing the config is useless if the
|
||||
affected rows stay terminal.
|
||||
"""
|
||||
body = payload or {}
|
||||
channel_id = body.get("channel_id")
|
||||
statuses = body.get("statuses") or ["error", "no_subtitles"]
|
||||
bad = [s for s in statuses if s not in Store.RETRYABLE_STATUSES]
|
||||
if bad:
|
||||
raise HTTPException(400, f"not retryable: {bad}")
|
||||
n = store.reset_videos(channel_id, tuple(statuses))
|
||||
return {"reset": n, "statuses": statuses, "channel_id": channel_id}
|
||||
|
||||
@r.post("/tools/reconcile")
|
||||
def reconcile(payload: dict | None = None):
|
||||
"""Re-scan data/markdown and make the DB agree with disk, on demand.
|
||||
|
||||
Without this the markdown tree is only read at server startup, so any
|
||||
.md produced afterwards is invisible until a restart.
|
||||
"""
|
||||
from ..segments import reconcile_markdown
|
||||
# prune is opt-in: it demotes `done` rows whose .md vanished, which is
|
||||
# destructive if the markdown root is ever misconfigured.
|
||||
prune = bool((payload or {}).get("prune"))
|
||||
return reconcile_markdown(store, Path(cfg.output_dir_resolved), prune=prune)
|
||||
|
||||
# -------------------------------------------------- channel avatars (local cache)
|
||||
@r.post("/tools/sync-channels")
|
||||
def sync_channels(payload: dict):
|
||||
"""Re-scrape each tracked channel to refresh metadata + avatar (with deep fallback)."""
|
||||
from ..discover import discover_channel, deep_channel_avatar
|
||||
from ..pipeline import cache_channel_avatar
|
||||
channel_id = (payload or {}).get("channel_id")
|
||||
channels = [store.get_channel(channel_id)] if channel_id else store.list_channels()
|
||||
channels = [c for c in channels if c]
|
||||
avatars_dir = Path(cfg.output_dir_resolved).parent / "avatars"
|
||||
avatars_dir.mkdir(parents=True, exist_ok=True)
|
||||
n_synced = 0
|
||||
n_avatars = 0
|
||||
n_recovered = 0
|
||||
errors: list[dict] = []
|
||||
for i, c in enumerate(channels):
|
||||
# This loop runs outside the JobManager, so nothing else is spacing
|
||||
# it out; syncing every channel used to be one burst.
|
||||
if i:
|
||||
polite_sleep(cfg.delay.min_seconds, cfg.delay.max_seconds)
|
||||
cid = c["channel_id"]
|
||||
handle = (c.get("handle") or "").lstrip("@")
|
||||
ch_url = (
|
||||
f"https://www.youtube.com/@{handle}/videos" if handle
|
||||
else f"https://www.youtube.com/channel/{cid}"
|
||||
)
|
||||
try:
|
||||
# Metadata + avatar live on the channel object, not the video
|
||||
# list — one entry is enough and costs a single request instead
|
||||
# of paginating the entire channel.
|
||||
_new_id, new_name, avatar, _refs = discover_channel(
|
||||
ch_url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests, limit=1,
|
||||
)
|
||||
except Exception as exc:
|
||||
errors.append({"channel_id": cid, "name": c.get("name"), "error": str(exc)})
|
||||
continue
|
||||
if avatar is None:
|
||||
avatar = deep_channel_avatar(ch_url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests)
|
||||
if avatar:
|
||||
n_recovered += 1
|
||||
store.update_channel_meta(cid, name=new_name, avatar=avatar)
|
||||
if avatar and cache_channel_avatar(cid, avatar, avatars_dir):
|
||||
n_avatars += 1
|
||||
n_synced += 1
|
||||
return {
|
||||
"synced": n_synced,
|
||||
"avatars_cached": n_avatars,
|
||||
"avatars_recovered_via_deep_fallback": n_recovered,
|
||||
"errors": errors,
|
||||
}
|
||||
|
||||
@r.post("/tools/avatars")
|
||||
def download_avatars(payload: dict):
|
||||
from ..pipeline import cache_channel_avatar
|
||||
from ..discover import discover_channel
|
||||
channel_id = (payload or {}).get("channel_id")
|
||||
channels = [store.get_channel(channel_id)] if channel_id else store.list_channels()
|
||||
channels = [c for c in channels if c]
|
||||
out_dir = Path(cfg.output_dir_resolved).parent / "avatars"
|
||||
out_dir.mkdir(parents=True, exist_ok=True)
|
||||
n = 0
|
||||
for i, c in enumerate(channels):
|
||||
if i:
|
||||
polite_sleep(cfg.delay.min_seconds, cfg.delay.max_seconds)
|
||||
cid = c["channel_id"]
|
||||
url = c.get("avatar")
|
||||
if not url:
|
||||
handle = (c.get("handle") or "").lstrip("@")
|
||||
ch_url = f"https://www.youtube.com/@{handle}/videos" if handle else f"https://www.youtube.com/channel/{cid}"
|
||||
try:
|
||||
# Avatar only — no reason to walk the video list.
|
||||
_, _, avatar, _ = discover_channel(
|
||||
ch_url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests, limit=1,
|
||||
)
|
||||
except Exception:
|
||||
avatar = None
|
||||
if avatar:
|
||||
store.upsert_channel(cid, c.get("handle"), c.get("name"), c.get("video_count") or 0, avatar=avatar)
|
||||
url = avatar
|
||||
if url and cache_channel_avatar(cid, url, out_dir):
|
||||
n += 1
|
||||
return {"downloaded": n, "dir": str(out_dir)}
|
||||
|
||||
@r.get("/avatars/{channel_id}")
|
||||
def serve_avatar(channel_id: str):
|
||||
from fastapi.responses import RedirectResponse
|
||||
from ..pipeline import cache_channel_avatar
|
||||
local = Path(cfg.output_dir_resolved).parent / "avatars" / f"{channel_id}.jpg"
|
||||
if local.exists():
|
||||
return FileResponse(str(local), media_type="image/jpeg")
|
||||
c = store.get_channel(channel_id) or {}
|
||||
url = c.get("avatar")
|
||||
if not url:
|
||||
raise HTTPException(404, "no avatar for channel")
|
||||
avatars_dir = Path(cfg.output_dir_resolved).parent / "avatars"
|
||||
if cache_channel_avatar(channel_id, url, avatars_dir) and local.exists():
|
||||
return FileResponse(str(local), media_type="image/jpeg")
|
||||
return RedirectResponse(url=url, status_code=302)
|
||||
|
||||
# -------------------------------------------------- clip (transcript segment)
|
||||
@r.get("/clip/{video_id}")
|
||||
def clip_video(video_id: str, frm: str = Query("0:00", alias="from"), to: str = Query("", alias="to")):
|
||||
@@ -411,8 +672,10 @@ def build_router(store: Store, cfg: Config, jobs) -> APIRouter:
|
||||
"top_tags": d["top_tags"], "totals": totals}
|
||||
|
||||
# -------------------------------------------------- audio (job)
|
||||
# `def`, not `async def`: no await here, and the SQLite reads below would
|
||||
# block the event loop.
|
||||
@r.post("/tools/audio")
|
||||
async def tools_audio(payload: dict):
|
||||
def tools_audio(payload: dict):
|
||||
video_ids = (payload or {}).get("video_ids") or []
|
||||
channel_id = (payload or {}).get("channel_id")
|
||||
if not video_ids and not channel_id:
|
||||
@@ -423,24 +686,65 @@ def build_router(store: Store, cfg: Config, jobs) -> APIRouter:
|
||||
job_id = jobs.enqueue(channel_id, opts)
|
||||
return {"job_id": job_id}
|
||||
|
||||
# `def`, not `async def`: same as tools_audio — SQLite reads + filesystem
|
||||
# stats only, no await.
|
||||
@r.post("/tools/video")
|
||||
def tools_video(payload: dict):
|
||||
video_ids = (payload or {}).get("video_ids") or []
|
||||
if len(video_ids) != 1:
|
||||
raise HTTPException(400, "exactly one video_id required")
|
||||
video_id = str(video_ids[0])
|
||||
v = store.get_video(video_id)
|
||||
if not v:
|
||||
raise HTTPException(404, "video not found")
|
||||
md_path = Path(cfg.output_dir_resolved).parent / v.markdown_path if v.markdown_path else None
|
||||
if v.status != "done" or not v.markdown_path or not md_path or not md_path.exists():
|
||||
raise HTTPException(409, "markdown must be generated and present before downloading video")
|
||||
opts = {"mode": "video", "video_ids": [video_id]}
|
||||
job_id = jobs.enqueue(v.channel_id, opts)
|
||||
return {"job_id": job_id}
|
||||
|
||||
return r
|
||||
|
||||
|
||||
def _video_dict(v) -> dict:
|
||||
import json as _json
|
||||
tags = []
|
||||
if v.tags:
|
||||
try:
|
||||
tags = _json.loads(v.tags)
|
||||
tags = json.loads(v.tags)
|
||||
except Exception:
|
||||
tags = []
|
||||
return {
|
||||
"video_id": v.video_id, "channel_id": v.channel_id, "title": v.title, "url": v.url,
|
||||
"upload_date": v.upload_date, "duration": v.duration, "status": v.status,
|
||||
"upload_date": v.upload_date,
|
||||
"upload_date_approx": bool(v.upload_date_approx),
|
||||
"duration": v.duration, "status": v.status,
|
||||
"transcript_lang": v.transcript_lang, "has_chapters": bool(v.has_chapters),
|
||||
"view_count": v.view_count, "like_count": v.like_count, "tags": tags,
|
||||
"thumbnail": v.thumbnail, "markdown_path": v.markdown_path,
|
||||
"video_download_status": v.video_download_status or "not_downloaded",
|
||||
"video_path": v.video_path, "video_filename": v.video_filename,
|
||||
"video_size": v.video_size, "video_downloaded_at": v.video_downloaded_at,
|
||||
"video_error": v.video_error,
|
||||
# The date the row was ordered by. For a video discovery has not
|
||||
# extracted yet there is no real upload_date, so this is inferred from
|
||||
# its position in the channel listing — flagged, never passed off as
|
||||
# exact. The sentinel means "newer than anything dated in this channel",
|
||||
# which is a rank, not a date, so it does not reach the client.
|
||||
"sort_date": None if v.sort_date == store_mod.NO_DATE_SENTINEL else v.sort_date,
|
||||
# date_estimated = "la fecha que ves no es exacta": tanto la inferida
|
||||
# por rank como la aproximada del discovery (texto relativo de YouTube).
|
||||
"date_estimated": bool(
|
||||
(v.sort_date and not v.upload_date and v.sort_date != store_mod.NO_DATE_SENTINEL)
|
||||
or v.upload_date_approx
|
||||
),
|
||||
"error_msg": v.error_msg,
|
||||
"description": v.description,
|
||||
"availability": v.availability,
|
||||
# None when fetchable; otherwise members_only / premium_only / private /
|
||||
# needs_auth / removed. Derived, so it stays correct for rows recorded
|
||||
# before `availability` existed.
|
||||
"block_reason": v.block_reason,
|
||||
}
|
||||
|
||||
|
||||
@@ -461,22 +765,11 @@ def _loads(s):
|
||||
return {}
|
||||
|
||||
|
||||
def _handle(url: str) -> str:
|
||||
if "@" in url:
|
||||
return "@" + url.split("@", 1)[1].split("/", 1)[0]
|
||||
return ""
|
||||
|
||||
|
||||
def _folder_info(path: Path) -> dict:
|
||||
path = Path(path)
|
||||
return {"path": str(path.resolve()), "exists": path.exists()}
|
||||
|
||||
|
||||
def _safe_dir(name: str) -> str:
|
||||
safe = "".join(c for c in (name or "") if c not in r'\/:*?"<>|')
|
||||
return (safe.strip().strip(".") or "unknown")
|
||||
|
||||
|
||||
def _open_in_os(path: Path) -> None:
|
||||
import os
|
||||
import sys
|
||||
|
||||
@@ -1,5 +1,7 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
import threading
|
||||
from pathlib import Path
|
||||
|
||||
from fastapi import FastAPI
|
||||
@@ -8,7 +10,8 @@ from fastapi.staticfiles import StaticFiles
|
||||
|
||||
from ..config import Config, load_config
|
||||
from ..cookies import auto_import_dir
|
||||
from ..segments import backfill_from_markdown
|
||||
from ..ratelimit import configure_global_pacer
|
||||
from ..segments import reconcile_markdown
|
||||
from ..store import Store
|
||||
from .api import build_router
|
||||
from .jobs import JobManager
|
||||
@@ -23,24 +26,66 @@ def create_app(cfg: Config | None = None, db_path: str | Path | None = None) ->
|
||||
cfg.database_path = str(db_path)
|
||||
store = Store(cfg.database_path_resolved)
|
||||
|
||||
# The job runner serialises jobs, but /api/tools/* endpoints run outside it
|
||||
# on the threadpool. The pacer is the only thing that stops those two from
|
||||
# hitting YouTube simultaneously, so it has to be armed before any router.
|
||||
configure_global_pacer(cfg.delay.min_request_interval)
|
||||
|
||||
# auto-import any loose cookies into the vault
|
||||
try:
|
||||
auto_import_dir(store)
|
||||
except Exception:
|
||||
pass
|
||||
# auto-backfill segments from existing markdown (idempotent)
|
||||
# uvicorn only configures its own loggers, so the root logger has no
|
||||
# handler and everything yt_scraper logs at startup vanishes.
|
||||
pkg_log = logging.getLogger("yt_scraper")
|
||||
if not pkg_log.handlers:
|
||||
handler = logging.StreamHandler()
|
||||
handler.setFormatter(logging.Formatter("%(levelname)s %(name)s: %(message)s"))
|
||||
pkg_log.addHandler(handler)
|
||||
pkg_log.setLevel(logging.INFO)
|
||||
|
||||
# Reconcile, not just backfill: backfill_from_markdown populates segments
|
||||
# and metadata but never touches `status`/`markdown_path`, so a video whose
|
||||
# .md is already on disk would keep showing a failure after every restart.
|
||||
#
|
||||
# It walks and parses EVERY .md in the tree, which on a large library
|
||||
# blocked uvicorn boot for the whole scan — so it runs on a daemon thread
|
||||
# and the server answers immediately (SQLite/WAL absorbs the concurrent
|
||||
# writes; /api/tools/reconcile still re-runs it on demand). healthz
|
||||
# reports whether the initial pass has finished.
|
||||
reconcile_done = threading.Event()
|
||||
|
||||
def _startup_reconcile() -> None:
|
||||
try:
|
||||
md_root = cfg.output_dir_resolved
|
||||
if Path(md_root).exists():
|
||||
backfill_from_markdown(store, Path(md_root))
|
||||
md_root = Path(cfg.output_dir_resolved)
|
||||
if md_root.exists():
|
||||
reconcile_markdown(store, md_root, log=pkg_log.info)
|
||||
except Exception:
|
||||
pass
|
||||
pkg_log.exception("startup reconcile failed")
|
||||
finally:
|
||||
reconcile_done.set()
|
||||
|
||||
threading.Thread(target=_startup_reconcile, name="startup-reconcile", daemon=True).start()
|
||||
|
||||
app = FastAPI(title="yt-scraper platform", version="1.0.0")
|
||||
|
||||
# StaticFiles/FileResponse send etag + last-modified but no Cache-Control,
|
||||
# so a browser may heuristically cache an old app.js across updates and
|
||||
# run yesterday's JS against a new backend. "no-cache" forces
|
||||
# revalidation on every load (cheap 304s on a local app) while keeping
|
||||
# the etag benefits.
|
||||
@app.middleware("http")
|
||||
async def _revalidate_shell(request, call_next):
|
||||
response = await call_next(request)
|
||||
path = request.url.path
|
||||
if path == "/" or path.startswith("/static/"):
|
||||
response.headers["Cache-Control"] = "no-cache"
|
||||
return response
|
||||
|
||||
@app.get("/healthz")
|
||||
def healthz():
|
||||
return {"status": "ok"}
|
||||
return {"status": "ok", "reconcile_done": reconcile_done.is_set()}
|
||||
|
||||
jobs = JobManager(store, cfg)
|
||||
app.state.jobs = jobs
|
||||
|
||||
+452
-40
@@ -6,14 +6,15 @@ import threading
|
||||
import time
|
||||
import uuid
|
||||
from collections import deque
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
from ..config import Config, load_config
|
||||
from ..config import Config, keep_ref, load_config, parse_languages
|
||||
from ..cookies import resolve_active_path
|
||||
from ..discover import discover_channel
|
||||
from ..discover import discover_incremental, extract_handle
|
||||
from ..pipeline import process_video
|
||||
from ..ratelimit import polite_sleep
|
||||
from ..store import Store
|
||||
from ..ratelimit import ThrottleGuard, polite_sleep
|
||||
from ..store import Store, order_pending
|
||||
|
||||
|
||||
class JobManager:
|
||||
@@ -89,12 +90,18 @@ class JobManager:
|
||||
if not job:
|
||||
return
|
||||
opts = json.loads(job.opts_json) if job.opts_json else {}
|
||||
if opts.get("video_ids"):
|
||||
self._run_batch(job_id, opts)
|
||||
if opts.get("mode") == "discover":
|
||||
self._run_discovery(job_id, opts)
|
||||
return
|
||||
if opts.get("mode") == "audio":
|
||||
self._run_audio(job_id, opts)
|
||||
return
|
||||
if opts.get("mode") == "video":
|
||||
self._run_video(job_id, opts)
|
||||
return
|
||||
if opts.get("video_ids"):
|
||||
self._run_batch(job_id, opts)
|
||||
return
|
||||
self._run_channel(job_id, opts)
|
||||
|
||||
def _resolve_channel_for_video(self, video_id: str) -> tuple[str, str, str]:
|
||||
@@ -107,19 +114,85 @@ class JobManager:
|
||||
url = f"https://www.youtube.com/@{handle}/videos" if handle else row.url
|
||||
return (name, row.channel_id, url)
|
||||
|
||||
# ---- throttling ------------------------------------------------------
|
||||
#
|
||||
# Every loop that touches YouTube shares these three helpers. Before them,
|
||||
# a throttled session was invisible to the runner: it kept going and turned
|
||||
# one rate-limit into hundreds of videos marked failed (the incident that
|
||||
# left 511 rows in `no_subtitles` and 347 in `error`, all of them saying
|
||||
# "rate-limited by YouTube"). Now the run stops and everything it has not
|
||||
# reached stays `pending`, which is the retryable state.
|
||||
|
||||
def _new_guard(self) -> ThrottleGuard:
|
||||
d = self.cfg.delay
|
||||
return ThrottleGuard(
|
||||
threshold=d.throttle_threshold,
|
||||
base=d.backoff_base,
|
||||
cap=d.backoff_cap,
|
||||
)
|
||||
|
||||
def _note_outcome(self, job_id: str, guard: ThrottleGuard, video_id: str, status: str) -> bool:
|
||||
"""Feed one video's result to the breaker. Returns False to stop the run.
|
||||
|
||||
`process_video` returns only a status string, so the reason is read back
|
||||
from the row: the message it stored is the one place the throttling
|
||||
signature survives.
|
||||
"""
|
||||
if status == "done":
|
||||
guard.note_success()
|
||||
return True
|
||||
row = self.store.get_video(video_id)
|
||||
wait = guard.note_failure(getattr(row, "error_msg", None) or status)
|
||||
if guard.tripped:
|
||||
return False
|
||||
if wait > 0:
|
||||
self._emit(job_id, "log", {
|
||||
"msg": f"YouTube is throttling us — waiting {wait:.1f}s before the next video"
|
||||
})
|
||||
time.sleep(wait)
|
||||
return True
|
||||
|
||||
def _stop_throttled(
|
||||
self, job_id: str, guard: ThrottleGuard, completed: int, total: int,
|
||||
extra: dict[str, Any] | None = None,
|
||||
) -> None:
|
||||
msg = guard.tripped_reason or "stopped by the throttling circuit breaker"
|
||||
remaining = max(0, total - completed)
|
||||
detail = (
|
||||
f"{msg}. Stopped after {completed}/{total}; the remaining {remaining} "
|
||||
"were left untouched and are still pending."
|
||||
)
|
||||
self.store.update_job(
|
||||
job_id, status="error", completed=completed, last_error=detail, finished=True
|
||||
)
|
||||
self._emit(job_id, "log", {"msg": f"ABORTED — {detail}"})
|
||||
self._emit(job_id, "error", {
|
||||
"message": detail, "throttled": True,
|
||||
"completed": completed, "total": total, **(extra or {}),
|
||||
})
|
||||
|
||||
def _run_batch(self, job_id: str, opts: dict[str, Any]) -> None:
|
||||
from ..pipeline import process_video as _process_video, cache_thumbnail
|
||||
"""Generate the .md for an explicit list of videos — and nothing else.
|
||||
|
||||
Thumbnails are deliberately NOT fetched here. They are already cached by
|
||||
the channel-level paths (add-channel, the Thumbnails tool), and
|
||||
/api/thumbnails/{id} redirects to the CDN for anything missing, so
|
||||
piggybacking them on a .md batch only spent extra requests per video for
|
||||
an image the UI could already display.
|
||||
"""
|
||||
from ..pipeline import process_video as _process_video
|
||||
video_ids = opts.get("video_ids") or []
|
||||
cookie_path = opts.get("cookies_file") or resolve_active_path(self.store)
|
||||
thumb_dir = Path(self.cfg.output_dir_resolved).parent / "thumbnails"
|
||||
total = len(video_ids)
|
||||
self.store.update_job(job_id, status="running", total=total, completed=0)
|
||||
self._emit(job_id, "progress", {"completed": 0, "total": total})
|
||||
self._emit(job_id, "log", {"msg": f"processing {total} videos (.md + thumbnails)"})
|
||||
self._emit(job_id, "log", {"msg": f"generating .md for {total} videos"})
|
||||
cfg = _clone_config(self.cfg)
|
||||
if opts.get("languages"):
|
||||
cfg.languages = opts["languages"]
|
||||
cfg.languages = parse_languages(opts["languages"], cfg.prefer_manual)
|
||||
completed = 0
|
||||
outcomes = {"processed": 0, "no_subtitles": 0, "errors": 0}
|
||||
guard = self._new_guard()
|
||||
for vid in video_ids:
|
||||
if job_id in self._cancel:
|
||||
self.store.update_job(job_id, status="cancelled", completed=completed, finished=True)
|
||||
@@ -128,15 +201,17 @@ class JobManager:
|
||||
row = self.store.get_video(vid)
|
||||
if not row:
|
||||
self._emit(job_id, "log", {"msg": f"skip unknown {vid}"})
|
||||
outcomes["errors"] += 1
|
||||
completed += 1
|
||||
self.store.update_job(job_id, completed=completed)
|
||||
continue
|
||||
# UX optimization: skip re-extracting videos whose .md already exists — just ensure thumbnail cached.
|
||||
# UX optimization: a video whose .md is already on disk costs nothing
|
||||
# to "download" — skip the extraction entirely.
|
||||
md_path = Path(cfg.output_dir_resolved).parent / row.markdown_path if row.markdown_path else None
|
||||
if row.status == "done" and row.markdown_path and md_path and md_path.exists():
|
||||
cache_thumbnail(self.store, vid, thumb_dir)
|
||||
status = "done"
|
||||
self._emit(job_id, "log", {"msg": f"cached {vid} (md already present)"})
|
||||
extracted = False
|
||||
self._emit(job_id, "log", {"msg": f"skip {vid} (.md already present)"})
|
||||
else:
|
||||
channel_name, channel_id, channel_url = self._resolve_channel_for_video(vid)
|
||||
status = _process_video(
|
||||
@@ -144,12 +219,28 @@ class JobManager:
|
||||
cookies_file=cookie_path,
|
||||
on_log=lambda m: self._emit(job_id, "log", {"msg": m}),
|
||||
)
|
||||
cache_thumbnail(self.store, vid, thumb_dir)
|
||||
extracted = True
|
||||
if status == "done":
|
||||
outcomes["processed"] += 1
|
||||
elif status == "no_subtitles":
|
||||
outcomes["no_subtitles"] += 1
|
||||
self._emit(job_id, "log", {"msg": f"{vid}: no transcript; .md not generated"})
|
||||
else:
|
||||
outcomes["errors"] += 1
|
||||
self._emit(job_id, "log", {"msg": f"{vid}: processing failed; .md not generated"})
|
||||
completed += 1
|
||||
self.store.update_job(job_id, completed=completed)
|
||||
self._emit(job_id, "progress", {"completed": completed, "total": total, "video_id": vid, "status": status})
|
||||
if not self._note_outcome(job_id, guard, vid, status):
|
||||
self._stop_throttled(job_id, guard, completed, total, outcomes)
|
||||
return
|
||||
# Only pace when we actually hit YouTube. A batch of already-rendered
|
||||
# videos costs nothing but a disk check, and sleeping through it
|
||||
# would make multi-select feel broken for no benefit.
|
||||
if extracted and completed < total:
|
||||
polite_sleep(cfg.delay.min_seconds, cfg.delay.max_seconds)
|
||||
self.store.update_job(job_id, status="done", completed=completed, finished=True)
|
||||
self._emit(job_id, "done", {"completed": completed, "total": total})
|
||||
self._emit(job_id, "done", {"completed": completed, "total": total, **outcomes, "throttling": guard.summary()})
|
||||
|
||||
def _run_audio(self, job_id: str, opts: dict[str, Any]) -> None:
|
||||
import yt_dlp
|
||||
@@ -165,15 +256,29 @@ class JobManager:
|
||||
total = len(videos)
|
||||
self.store.update_job(job_id, status="running", total=total, completed=0)
|
||||
self._emit(job_id, "log", {"msg": f"downloading {total} audio tracks -> {out_dir}"})
|
||||
cfg = self.cfg
|
||||
ydl_opts = {
|
||||
"format": "bestaudio/best",
|
||||
"outtmpl": str(out_dir / "%(id)s.%(ext)s"),
|
||||
"postprocessors": [{"key": "FFmpegExtractAudio", "preferredcodec": "mp3", "preferredquality": "128"}],
|
||||
"quiet": True, "no_warnings": True, "noprogress": True,
|
||||
# This is the only path that downloads media, so it is also the only
|
||||
# one where yt-dlp's own download-side throttles actually fire.
|
||||
"sleep_interval": cfg.delay.min_seconds,
|
||||
"max_sleep_interval": cfg.delay.max_seconds,
|
||||
"sleep_interval_requests": cfg.yt_dlp.sleep_subrequests,
|
||||
"retries": cfg.yt_dlp.retries,
|
||||
"socket_timeout": 30.0,
|
||||
"js_runtimes": {"node": {}, "deno": {}, "bun": {}, "quickjs": {}},
|
||||
}
|
||||
if cfg.delay.audio_rate_limit:
|
||||
ydl_opts["ratelimit"] = cfg.delay.audio_rate_limit
|
||||
if cookie_path:
|
||||
ydl_opts["cookiefile"] = cookie_path
|
||||
completed = 0
|
||||
guard = self._new_guard()
|
||||
# One YoutubeDL for the whole batch: keeps the connection pool, the
|
||||
# cookie jar and the resolved player JS alive across videos.
|
||||
with yt_dlp.YoutubeDL(ydl_opts) as ydl:
|
||||
for v in videos:
|
||||
if job_id in self._cancel:
|
||||
@@ -183,13 +288,132 @@ class JobManager:
|
||||
try:
|
||||
ydl.download([v.url])
|
||||
self._emit(job_id, "log", {"msg": f"OK {v.video_id}"})
|
||||
guard.note_success()
|
||||
except Exception as exc:
|
||||
self._emit(job_id, "log", {"msg": f"FAIL {v.video_id}: {exc}"})
|
||||
wait = guard.note_failure(exc)
|
||||
if guard.tripped:
|
||||
self._stop_throttled(job_id, guard, completed, total)
|
||||
return
|
||||
if wait > 0:
|
||||
self._emit(job_id, "log", {
|
||||
"msg": f"YouTube is throttling us — waiting {wait:.1f}s before the next track"
|
||||
})
|
||||
time.sleep(wait)
|
||||
completed += 1
|
||||
self.store.update_job(job_id, completed=completed)
|
||||
self._emit(job_id, "progress", {"completed": completed, "total": total, "video_id": v.video_id})
|
||||
self.store.update_job(job_id, status="done", completed=completed, finished=True)
|
||||
self._emit(job_id, "done", {"completed": completed, "total": total})
|
||||
self._emit(job_id, "done", {"completed": completed, "total": total, "throttling": guard.summary()})
|
||||
|
||||
def _run_video(self, job_id: str, opts: dict[str, Any]) -> None:
|
||||
"""Download one validated video as a Chromium-friendly WebM."""
|
||||
import yt_dlp
|
||||
|
||||
video_id = (opts.get("video_ids") or [None])[0]
|
||||
row = self.store.get_video(video_id) if video_id else None
|
||||
if not row:
|
||||
self._fail_video_job(job_id, "video not found")
|
||||
return
|
||||
|
||||
root = Path(self.cfg.output_dir_resolved).parent
|
||||
md_path = root / row.markdown_path if row.markdown_path else None
|
||||
if row.status != "done" or not row.markdown_path or not md_path or not md_path.exists():
|
||||
self._fail_video_job(job_id, "markdown must be generated and present before downloading video", video_id)
|
||||
return
|
||||
|
||||
out_dir = root / "videos" / video_id
|
||||
out_dir.mkdir(parents=True, exist_ok=True)
|
||||
filename = row.video_filename or _safe_video_filename(row.title or video_id)
|
||||
target = out_dir / filename
|
||||
if target.suffix.lower() != ".webm":
|
||||
target = target.with_suffix(".webm")
|
||||
rel_path = target.relative_to(root).as_posix()
|
||||
|
||||
if target.exists() and target.stat().st_size > 0:
|
||||
self.store.update_video_download(video_id, "done", path=rel_path, filename=target.name, size=target.stat().st_size)
|
||||
self.store.update_job(job_id, status="done", total=1, completed=1, finished=True)
|
||||
self._emit(job_id, "progress", {"completed": 1, "total": 1, "video_id": video_id, "status": "done", "skipped": True})
|
||||
self._emit(job_id, "done", {"completed": 1, "total": 1, "skipped": True})
|
||||
return
|
||||
|
||||
cookie_path = opts.get("cookies_file") or resolve_active_path(self.store)
|
||||
restricted = bool(row.block_reason or row.availability in {
|
||||
"subscriber_only", "premium_only", "private", "needs_auth",
|
||||
})
|
||||
self.store.update_video_download(video_id, "queued", filename=target.name, error=None)
|
||||
self.store.update_job(job_id, status="running", total=1, completed=0)
|
||||
self._emit(job_id, "log", {"msg": f"downloading {video_id} -> {target.name}"})
|
||||
|
||||
def hook(data: dict[str, Any]) -> None:
|
||||
if data.get("status") not in ("downloading", "finished"):
|
||||
return
|
||||
downloaded = int(data.get("downloaded_bytes") or 0)
|
||||
total = int(data.get("total_bytes") or data.get("total_bytes_estimate") or 0)
|
||||
percent = round(downloaded * 100 / total, 1) if total else None
|
||||
self._emit(job_id, "progress", {
|
||||
"completed": 0, "total": 1, "video_id": video_id,
|
||||
"status": data.get("status"), "downloaded_bytes": downloaded,
|
||||
"total_bytes": total, "percent": percent,
|
||||
"speed": data.get("speed"), "eta": data.get("eta"),
|
||||
})
|
||||
|
||||
self.store.update_video_download(video_id, "downloading", filename=target.name, error=None)
|
||||
ydl_opts = {
|
||||
# Prefer YouTube's HLS VP9 + Opus pair. Chromium can play the
|
||||
# resulting WebM directly, and HLS avoids the recurring mid-range
|
||||
# 403s seen on long HTTPS media requests.
|
||||
"format": "bestvideo[height<=1080][protocol^=m3u8_native][vcodec^=vp09]+bestaudio[acodec^=opus]/bestvideo[height<=1080][protocol^=m3u8_native]+bestaudio[acodec^=opus]/bestvideo[height<=1080][ext=webm]+bestaudio[ext=webm]",
|
||||
"merge_output_format": "webm",
|
||||
"outtmpl": str(out_dir / f"{filename.rsplit('.', 1)[0]}.%(ext)s"),
|
||||
"quiet": True, "no_warnings": True, "noprogress": True,
|
||||
"progress_hooks": [hook],
|
||||
"continuedl": True,
|
||||
# Current YouTube extraction requires an external JS runtime and
|
||||
# the EJS challenge scripts. Node is available in the supported
|
||||
# local browser setup and is installed by yt-dlp[default].
|
||||
"js_runtimes": {"node": {}},
|
||||
"fragment_retries": 10,
|
||||
"retries": self.cfg.yt_dlp.retries,
|
||||
"extractor_retries": self.cfg.yt_dlp.extractor_retries,
|
||||
"sleep_interval": self.cfg.delay.min_seconds,
|
||||
"max_sleep_interval": self.cfg.delay.max_seconds,
|
||||
"sleep_interval_requests": self.cfg.yt_dlp.sleep_subrequests,
|
||||
"socket_timeout": self.cfg.yt_dlp.socket_timeout,
|
||||
"extractor_args": {"youtube": {"player_client": ["default", "web"]}},
|
||||
}
|
||||
if cookie_path and restricted:
|
||||
ydl_opts["cookiefile"] = cookie_path
|
||||
|
||||
try:
|
||||
with yt_dlp.YoutubeDL(ydl_opts) as ydl:
|
||||
ydl.download([row.url])
|
||||
if not target.exists():
|
||||
# yt-dlp may retain the requested stem but choose a different
|
||||
# extension when the merge was skipped; locate only this ID's
|
||||
# directory and accept the generated WebM as the canonical file.
|
||||
candidates = sorted(out_dir.glob("*.webm"), key=lambda p: p.stat().st_mtime, reverse=True)
|
||||
if candidates:
|
||||
target = candidates[0]
|
||||
if not target.exists() or target.stat().st_size <= 0:
|
||||
raise RuntimeError("yt-dlp finished without producing a WebM file")
|
||||
rel_path = target.relative_to(root).as_posix()
|
||||
size = target.stat().st_size
|
||||
self.store.update_video_download(video_id, "done", path=rel_path, filename=target.name, size=size, error=None)
|
||||
self.store.update_job(job_id, status="done", total=1, completed=1, finished=True)
|
||||
self._emit(job_id, "progress", {"completed": 1, "total": 1, "video_id": video_id, "status": "done", "percent": 100, "total_bytes": size})
|
||||
self._emit(job_id, "done", {"completed": 1, "total": 1, "video_id": video_id, "size": size})
|
||||
except Exception as exc:
|
||||
msg = str(exc)
|
||||
self.store.update_video_download(video_id, "error", filename=target.name, error=msg)
|
||||
self.store.update_job(job_id, status="error", total=1, completed=0, last_error=msg, finished=True)
|
||||
self._emit(job_id, "error", {"message": msg, "video_id": video_id})
|
||||
|
||||
def _fail_video_job(self, job_id: str, message: str, video_id: str | None = None) -> None:
|
||||
if video_id:
|
||||
self.store.update_video_download(video_id, "error", error=message)
|
||||
self.store.update_job(job_id, status="error", total=1, completed=0, last_error=message, finished=True)
|
||||
self._emit(job_id, "error", {"message": message, **({"video_id": video_id} if video_id else {})})
|
||||
|
||||
def _run_channel(self, job_id: str, opts: dict[str, Any]) -> None:
|
||||
job = self.store.get_job(job_id)
|
||||
@@ -204,37 +428,75 @@ class JobManager:
|
||||
return
|
||||
cfg.channel_url = channel_url
|
||||
|
||||
self.store.update_job(job_id, status="running")
|
||||
self._emit(job_id, "log", {"msg": f"discovering {channel_url}"})
|
||||
try:
|
||||
_channel_id, channel_name, refs = discover_channel(channel_url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests)
|
||||
except Exception as exc:
|
||||
self.store.update_job(job_id, status="error", last_error=str(exc), finished=True)
|
||||
self._emit(job_id, "error", {"message": f"discovery failed: {exc}"})
|
||||
return
|
||||
|
||||
limit = opts.get("limit")
|
||||
since = opts.get("since")
|
||||
languages = opts.get("languages")
|
||||
include_shorts = opts.get("include_shorts", cfg.include_shorts)
|
||||
no_live = opts.get("no_live", not cfg.include_live)
|
||||
cookie_override = opts.get("cookies_file")
|
||||
full = bool(opts.get("full"))
|
||||
|
||||
sync = cfg.sync
|
||||
incremental = sync.incremental and not full and bool(channel_id)
|
||||
known = self.store.known_video_ids(channel_id) if incremental else set()
|
||||
|
||||
self.store.update_job(job_id, status="running")
|
||||
self._emit(job_id, "log", {
|
||||
"msg": f"discovering {channel_url}"
|
||||
+ (f" (incremental — newest {sync.window} first)" if known else " (full)")
|
||||
})
|
||||
try:
|
||||
result = discover_incremental(
|
||||
channel_url,
|
||||
known,
|
||||
sleep_subrequests=cfg.yt_dlp.sleep_subrequests,
|
||||
window=sync.window,
|
||||
max_window=sync.max_window,
|
||||
overlap=sync.overlap,
|
||||
since=self.store.latest_upload_date(channel_id) if incremental else None,
|
||||
keep=keep_ref(cfg, include_shorts, no_live),
|
||||
)
|
||||
except Exception as exc:
|
||||
self.store.update_job(job_id, status="error", last_error=str(exc), finished=True)
|
||||
self._emit(job_id, "error", {"message": f"discovery failed: {exc}"})
|
||||
return
|
||||
|
||||
_channel_id, channel_name, avatar = result.channel_id, result.channel_name, result.avatar
|
||||
refs = result.refs
|
||||
self._emit(job_id, "log", {
|
||||
"msg": f"fetched {result.fetched} entries in {result.passes} pass(es), {result.new_count} new"
|
||||
})
|
||||
|
||||
if languages:
|
||||
cfg.languages = languages
|
||||
cfg.languages = parse_languages(languages, cfg.prefer_manual)
|
||||
if since:
|
||||
refs = [r for r in refs if (r.upload_date or "") >= since.replace("-", "")]
|
||||
if not include_shorts:
|
||||
refs = [r for r in refs if "/shorts/" not in (r.url or "")]
|
||||
if no_live:
|
||||
refs = [r for r in refs if not (r.url or "").startswith("https://www.youtube.com/live/")]
|
||||
if limit:
|
||||
refs = refs[: int(limit)]
|
||||
# Keep undated entries: flat discovery does not report upload_date,
|
||||
# so `>= since` on a missing date would discard the whole channel.
|
||||
cutoff = since.replace("-", "")
|
||||
refs = [r for r in refs if not r.upload_date or r.upload_date >= cutoff]
|
||||
|
||||
self.store.upsert_channel(_channel_id, _handle(channel_url), channel_name, len(refs))
|
||||
if not avatar:
|
||||
from ..discover import deep_channel_avatar
|
||||
avatar = deep_channel_avatar(channel_url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests)
|
||||
if result.full_scan:
|
||||
self.store.upsert_channel(_channel_id, extract_handle(channel_url), channel_name, len(refs), avatar=avatar)
|
||||
else:
|
||||
self.store.update_channel_meta(_channel_id, name=channel_name, avatar=avatar)
|
||||
if avatar:
|
||||
from ..pipeline import cache_channel_avatar
|
||||
cache_channel_avatar(_channel_id, avatar, Path(self.cfg.output_dir_resolved).parent / "avatars")
|
||||
self.store.upsert_videos(refs)
|
||||
self.store.mark_channel_synced(_channel_id)
|
||||
|
||||
pending = [r for r in self.store.get_pending(_channel_id) if r.video_id in {x.video_id for x in refs}]
|
||||
# Process the freshly-seen window first, then the older backlog, so a
|
||||
# `limit` still means "the newest N" now that discovery stops early.
|
||||
pending = order_pending(self.store.get_pending(_channel_id), refs)
|
||||
pending = [r for r in pending if keep_ref(cfg, include_shorts, no_live)(r)]
|
||||
if since:
|
||||
cutoff = since.replace("-", "")
|
||||
pending = [r for r in pending if not r.upload_date or r.upload_date >= cutoff]
|
||||
if limit:
|
||||
pending = pending[: int(limit)]
|
||||
total = len(pending)
|
||||
self.store.update_job(job_id, total=total)
|
||||
self._emit(job_id, "progress", {"completed": 0, "total": total})
|
||||
@@ -242,6 +504,8 @@ class JobManager:
|
||||
|
||||
cookie_path = cookie_override or resolve_active_path(self.store)
|
||||
completed = 0
|
||||
outcomes = {"processed": 0, "no_subtitles": 0, "errors": 0}
|
||||
guard = self._new_guard()
|
||||
for row in pending:
|
||||
if job_id in self._cancel:
|
||||
self.store.update_job(job_id, status="cancelled", completed=completed, finished=True)
|
||||
@@ -252,14 +516,158 @@ class JobManager:
|
||||
cookies_file=cookie_path,
|
||||
on_log=lambda m: self._emit(job_id, "log", {"msg": m}),
|
||||
)
|
||||
if status == "done":
|
||||
outcomes["processed"] += 1
|
||||
elif status == "no_subtitles":
|
||||
outcomes["no_subtitles"] += 1
|
||||
self._emit(job_id, "log", {"msg": f"{row.video_id}: no transcript; .md not generated"})
|
||||
else:
|
||||
outcomes["errors"] += 1
|
||||
self._emit(job_id, "log", {"msg": f"{row.video_id}: processing failed; .md not generated"})
|
||||
completed += 1
|
||||
self.store.update_job(job_id, completed=completed)
|
||||
self._emit(job_id, "progress", {"completed": completed, "total": total, "video_id": row.video_id, "status": status})
|
||||
if not self._note_outcome(job_id, guard, row.video_id, status):
|
||||
self._stop_throttled(job_id, guard, completed, total, outcomes)
|
||||
return
|
||||
if completed < total:
|
||||
polite_sleep(cfg.delay.min_seconds, cfg.delay.max_seconds)
|
||||
|
||||
self.store.update_job(job_id, status="done", completed=completed, finished=True)
|
||||
self._emit(job_id, "done", {"completed": completed, "total": total})
|
||||
self._emit(job_id, "done", {"completed": completed, "total": total, **outcomes, "throttling": guard.summary()})
|
||||
|
||||
def _run_discovery(self, job_id: str, opts: dict[str, Any] | None = None) -> None:
|
||||
"""Refresh the catalog without extracting or downloading video content."""
|
||||
job = self.store.get_job(job_id)
|
||||
if not job:
|
||||
return
|
||||
opts = opts or {}
|
||||
full = bool(opts.get("full"))
|
||||
|
||||
if job.channel_id:
|
||||
channel = self.store.get_channel(job.channel_id)
|
||||
channels = [channel] if channel else []
|
||||
else:
|
||||
channels = self.store.list_channels()
|
||||
channels = [c for c in channels if c]
|
||||
total = len(channels)
|
||||
self.store.update_job(job_id, status="running", total=total, completed=0)
|
||||
mode = "full rescan" if full else "incremental (recent uploads only)"
|
||||
self._emit(job_id, "log", {"msg": f"investigating {total} channel(s) — {mode}"})
|
||||
|
||||
totals = {"new_videos": 0, "known_videos": 0, "errors": 0, "fetched": 0}
|
||||
completed = 0
|
||||
guard = self._new_guard()
|
||||
for channel in channels:
|
||||
if job_id in self._cancel:
|
||||
self.store.update_job(job_id, status="cancelled", completed=completed, finished=True)
|
||||
self._emit(job_id, "cancelled", {"completed": completed, "total": total})
|
||||
return
|
||||
|
||||
# Pace between channels too: each one is a fresh burst of
|
||||
# continuation requests, and a catalog refresh over every channel
|
||||
# used to fire them back to back.
|
||||
if completed:
|
||||
polite_sleep(self.cfg.delay.min_seconds, self.cfg.delay.max_seconds)
|
||||
|
||||
channel_id = channel["channel_id"]
|
||||
try:
|
||||
result = self._discover_catalog_channel(channel, full=full)
|
||||
except Exception as exc:
|
||||
wait = guard.note_failure(exc)
|
||||
if guard.tripped:
|
||||
self._stop_throttled(job_id, guard, completed, total, totals)
|
||||
return
|
||||
if wait > 0:
|
||||
self._emit(job_id, "log", {
|
||||
"msg": f"YouTube is throttling us — waiting {wait:.1f}s before the next channel"
|
||||
})
|
||||
time.sleep(wait)
|
||||
if job.channel_id:
|
||||
self.store.update_job(job_id, status="error", completed=completed, last_error=str(exc), finished=True)
|
||||
self._emit(job_id, "error", {"message": f"discovery failed: {exc}"})
|
||||
return
|
||||
totals["errors"] += 1
|
||||
self._emit(job_id, "log", {"msg": f"discovery failed for {channel.get('name') or channel_id}: {exc}"})
|
||||
else:
|
||||
guard.note_success()
|
||||
totals["new_videos"] += result["new_videos"]
|
||||
totals["known_videos"] += result["known_videos"]
|
||||
totals["fetched"] += result["fetched"]
|
||||
name = channel.get("name") or channel_id
|
||||
scope = (
|
||||
f"scanned {result['fetched']} newest"
|
||||
if result["incremental"] else f"scanned all {result['fetched']}"
|
||||
)
|
||||
if not result["caught_up"]:
|
||||
scope += " (hit window ceiling — run a full rescan if videos are missing)"
|
||||
self._emit(job_id, "log", {"msg": f"{name}: {scope}, {result['new_videos']} new"})
|
||||
self._emit(
|
||||
job_id,
|
||||
"progress",
|
||||
{
|
||||
"channel_id": channel_id,
|
||||
"new_videos": result["new_videos"],
|
||||
"known_videos": result["known_videos"],
|
||||
"fetched": result["fetched"],
|
||||
"incremental": result["incremental"],
|
||||
"caught_up": result["caught_up"],
|
||||
"last_video_date": result["last_video_date"],
|
||||
"completed": completed + 1,
|
||||
"total": total,
|
||||
},
|
||||
)
|
||||
completed += 1
|
||||
self.store.update_job(job_id, completed=completed)
|
||||
|
||||
self.store.update_job(job_id, status="done", completed=completed, finished=True)
|
||||
self._emit(job_id, "done", {"completed": completed, "total": total, **totals})
|
||||
|
||||
def _discover_catalog_channel(self, channel: dict, *, full: bool = False) -> dict[str, Any]:
|
||||
"""Refresh one channel's catalog.
|
||||
|
||||
Incremental by default: only the newest slice of the channel is fetched,
|
||||
stopping at the first run of videos we already have. `full` forces the
|
||||
old behaviour (walk every page) for when the local catalog is suspect.
|
||||
"""
|
||||
channel_id = channel["channel_id"]
|
||||
channel_url = _resolve_channel_url(self.store, self.cfg, channel_id)
|
||||
if not channel_url:
|
||||
raise RuntimeError("no channel url")
|
||||
|
||||
sync = self.cfg.sync
|
||||
incremental = sync.incremental and not full
|
||||
known = self.store.known_video_ids(channel_id) if incremental else set()
|
||||
|
||||
result = discover_incremental(
|
||||
channel_url,
|
||||
known,
|
||||
sleep_subrequests=self.cfg.yt_dlp.sleep_subrequests,
|
||||
window=sync.window,
|
||||
max_window=sync.max_window,
|
||||
overlap=sync.overlap,
|
||||
since=self.store.latest_upload_date(channel_id) if incremental else None,
|
||||
keep=keep_ref(self.cfg),
|
||||
)
|
||||
refs = result.refs
|
||||
|
||||
if result.full_scan:
|
||||
self.store.upsert_channel(result.channel_id, extract_handle(channel_url), result.channel_name, len(refs))
|
||||
new_videos = self.store.upsert_videos(refs)
|
||||
else:
|
||||
self.store.update_channel_meta(result.channel_id, name=result.channel_name)
|
||||
new_videos = self.store.upsert_videos(refs)
|
||||
marks = self.store.mark_channel_synced(result.channel_id)
|
||||
|
||||
return {
|
||||
"new_videos": new_videos,
|
||||
"known_videos": len({r.video_id for r in refs}) - new_videos,
|
||||
"fetched": result.fetched,
|
||||
"passes": result.passes,
|
||||
"incremental": not result.full_scan,
|
||||
"caught_up": result.caught_up,
|
||||
"last_video_date": marks["last_video_date"],
|
||||
}
|
||||
|
||||
|
||||
def _clone_config(cfg: Config) -> Config:
|
||||
@@ -279,7 +687,11 @@ def _resolve_channel_url(store: Store, cfg: Config, channel_id: str | None) -> s
|
||||
return f"https://www.youtube.com/channel/{channel_id}"
|
||||
|
||||
|
||||
def _handle(url: str) -> str:
|
||||
if "@" in url:
|
||||
return "@" + url.split("@", 1)[1].split("/", 1)[0]
|
||||
return ""
|
||||
def _safe_video_filename(title: str) -> str:
|
||||
"""Keep the displayed title while making a valid, bounded Windows name."""
|
||||
invalid = set(r'\\/:*?"<>|')
|
||||
safe = "".join("_" if c in invalid or ord(c) < 32 else c for c in title)
|
||||
safe = " ".join(safe.strip().split())
|
||||
safe = safe.rstrip(" .") or "video"
|
||||
# Leave room for the id directory and yt-dlp's temporary suffixes.
|
||||
return safe[:180] + ".webm"
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -21,10 +21,13 @@
|
||||
<link rel="stylesheet" href="/static/styles.css" />
|
||||
</head>
|
||||
<body class="min-h-screen relative">
|
||||
<div id="root" x-cloak x-data="platform()" x-init="init()" class="relative z-10 flex min-h-screen">
|
||||
<div id="root" x-cloak x-data="platform()" x-init="init()" @keydown.escape.window="onGlobalEscape($event)" class="relative z-10 flex min-h-screen">
|
||||
|
||||
<button class="fixed top-4 left-4 z-40 md:hidden btn btn-ghost !p-2" @click="toggleSidebar()" :aria-label="sidebarOpen ? 'Close navigation' : 'Open navigation'"><svg x-show="!sidebarOpen" class="w-5 h-5" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><path stroke-linecap="round" d="M4 6h16M4 12h16M4 18h16"/></svg><svg x-show="sidebarOpen" class="w-5 h-5" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><path stroke-linecap="round" d="M6 6l12 12M18 6L6 18"/></svg></button>
|
||||
<div x-show="sidebarOpen" x-transition.opacity @click="closeSidebar()" class="fixed inset-0 z-20 bg-black/60 md:hidden"></div>
|
||||
|
||||
<!-- ============ SIDEBAR ============ -->
|
||||
<aside class="w-64 shrink-0 border-r border-zinc-800/80 bg-zinc-950/70 backdrop-blur-xl flex flex-col fixed inset-y-0 left-0 z-30">
|
||||
<aside class="w-64 shrink-0 border-r border-zinc-800/80 bg-zinc-950/70 backdrop-blur-xl flex flex-col fixed inset-y-0 left-0 z-30 transform transition-transform md:translate-x-0" :class="sidebarOpen ? 'translate-x-0' : '-translate-x-full'" x-transition>
|
||||
<div class="px-5 py-5 flex items-center gap-3 border-b border-zinc-800/70">
|
||||
<div class="w-9 h-9 rounded-xl accent-grad flex items-center justify-center font-bold text-white shadow-lg shadow-rose-500/20">
|
||||
<svg viewBox="0 0 24 24" class="w-5 h-5" fill="currentColor"><path d="M10 15.5v-7l6 3.5-6 3.5zM21 6.5c0-1.1-.9-2-2-2H5c-1.1 0-2 .9-2 2v11c0 1.1.9 2 2 2h14c1.1 0 2-.9 2-2v-11z"/></svg>
|
||||
@@ -46,14 +49,23 @@
|
||||
|
||||
<div class="px-4 py-4 border-t border-zinc-800/70 text-[0.7rem] text-zinc-500">
|
||||
<div class="flex items-center gap-2">
|
||||
<span class="w-2 h-2 rounded-full" :class="health ? 'bg-emerald-500' : 'bg-zinc-600'"></span>
|
||||
<span x-text="health ? 'backend online' : 'checking…'"></span>
|
||||
<span class="w-2 h-2 rounded-full" :class="health ? 'bg-emerald-500' : (healthFails ? 'bg-rose-500 animate-pulse' : 'bg-zinc-600')"></span>
|
||||
<span x-text="health ? 'backend online' : (healthFails ? 'backend unreachable' : 'checking…')"></span>
|
||||
</div>
|
||||
</div>
|
||||
</aside>
|
||||
|
||||
<!-- ============ MAIN ============ -->
|
||||
<main class="flex-1 ml-64 min-w-0">
|
||||
<main class="flex-1 ml-0 md:ml-64 min-w-0">
|
||||
<div x-show="!health && healthFails" x-transition.opacity class="bg-rose-500/15 border-b border-rose-500/40 text-rose-200" role="alert" aria-live="polite">
|
||||
<div class="max-w-7xl mx-auto px-8 py-2.5 flex items-center justify-between text-sm">
|
||||
<div class="flex items-center gap-2">
|
||||
<svg class="w-4 h-4 shrink-0" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><path stroke-linecap="round" stroke-linejoin="round" d="M12 9v2m0 4h.01M4.93 4.93l14.14 14.14M12 2a10 10 0 100 20 10 10 0 000-20z"/></svg>
|
||||
<span><span x-text="healthFails"></span> failed health check(s). The backend may be down; requests will retry automatically.</span>
|
||||
</div>
|
||||
<button class="text-rose-200 hover:text-white underline text-xs" @click="checkHealth()">Retry now</button>
|
||||
</div>
|
||||
</div>
|
||||
<div class="max-w-7xl mx-auto px-8 py-8">
|
||||
|
||||
<!-- ===== Dashboard ===== -->
|
||||
@@ -74,19 +86,23 @@
|
||||
<div class="grid grid-cols-2 md:grid-cols-4 gap-4">
|
||||
<div class="glass p-4">
|
||||
<div class="text-xs text-zinc-500 uppercase tracking-wider">Channels</div>
|
||||
<div class="mt-1 text-2xl font-bold text-white font-mono" x-text="dash.channels.length"></div>
|
||||
<template x-if="loading.dashboard"><div class="skeleton mt-1 h-7 w-12" aria-hidden="true"></div></template>
|
||||
<template x-if="!loading.dashboard"><div class="mt-1 text-2xl font-bold text-white font-mono" x-text="dash.channels.length"></div></template>
|
||||
</div>
|
||||
<div class="glass p-4">
|
||||
<div class="text-xs text-zinc-500 uppercase tracking-wider">Videos</div>
|
||||
<div class="mt-1 text-2xl font-bold text-white font-mono" x-text="fmtNum(dash.channels.reduce((s,c)=>s+(c.video_count||0),0))"></div>
|
||||
<template x-if="loading.dashboard"><div class="skeleton mt-1 h-7 w-16" aria-hidden="true"></div></template>
|
||||
<template x-if="!loading.dashboard"><div class="mt-1 text-2xl font-bold text-white font-mono" x-text="fmtNum(dash.channels.reduce((s,c)=>s+(c.video_count||0),0))"></div></template>
|
||||
</div>
|
||||
<div class="glass p-4">
|
||||
<div class="text-xs text-zinc-500 uppercase tracking-wider">Total Views</div>
|
||||
<div class="mt-1 text-2xl font-bold text-white font-mono" x-text="fmtNum(dash.channels.reduce((s,c)=>s+(c.total_views||0),0))"></div>
|
||||
<template x-if="loading.dashboard"><div class="skeleton mt-1 h-7 w-20" aria-hidden="true"></div></template>
|
||||
<template x-if="!loading.dashboard"><div class="mt-1 text-2xl font-bold text-white font-mono" x-text="fmtNum(dash.channels.reduce((s,c)=>s+(c.total_views||0),0))"></div></template>
|
||||
</div>
|
||||
<div class="glass p-4">
|
||||
<div class="text-xs text-zinc-500 uppercase tracking-wider">Total Duration</div>
|
||||
<div class="mt-1 text-2xl font-bold text-white font-mono" x-text="humanDur(dash.channels.reduce((s,c)=>s+(c.total_duration||0),0))"></div>
|
||||
<template x-if="loading.dashboard"><div class="skeleton mt-1 h-7 w-20" aria-hidden="true"></div></template>
|
||||
<template x-if="!loading.dashboard"><div class="mt-1 text-2xl font-bold text-white font-mono" x-text="humanDur(dash.channels.reduce((s,c)=>s+(c.total_duration||0),0))"></div></template>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
@@ -107,12 +123,18 @@
|
||||
</template>
|
||||
<div class="grid md:grid-cols-2 gap-4">
|
||||
<template x-for="c in dash.channels" :key="c.channel_id">
|
||||
<div class="glass p-5 hover:border-rose-500/40 transition-colors">
|
||||
<div class="glass p-5 hover:border-rose-500/40 transition-colors cursor-pointer focus:outline-none focus:ring-2 focus:ring-rose-500/40" role="button" tabindex="0" @click="openChannel(c.channel_id)" @keydown.enter.prevent="openChannel(c.channel_id)" @keydown.space.prevent="openChannel(c.channel_id)" :aria-label="'Open videos for ' + (c.name || c.channel_id)">
|
||||
<div class="flex items-start justify-between gap-3">
|
||||
<div class="flex items-center gap-3 min-w-0">
|
||||
<div class="relative w-10 h-10 shrink-0">
|
||||
<img class="avatar w-10 h-10 shrink-0" :src="c.avatar || '/api/avatars/'+c.channel_id" :alt="c.name" :data-fb-src="'/api/avatars/'+c.channel_id" onerror="if(!this.dataset.swapped){this.dataset.swapped='1'; this.src=this.dataset.fbSrc||''}else{this.style.visibility='hidden'; if(this.nextElementSibling) this.nextElementSibling.style.display='inline-flex'}" />
|
||||
<span class="avatar avatar-initial rounded-full w-10 h-10 text-sm absolute inset-0 hidden" :text-content="(c.name||c.channel_id||'?').charAt(0)" aria-hidden="true"></span>
|
||||
</div>
|
||||
<div class="min-w-0">
|
||||
<div class="text-white font-semibold truncate" x-text="c.name"></div>
|
||||
<div class="text-xs text-zinc-500 truncate" x-text="c.handle || c.channel_id"></div>
|
||||
</div>
|
||||
</div>
|
||||
<div class="text-right shrink-0">
|
||||
<div class="text-lg font-bold text-white font-mono" x-text="fmtNum(c.video_count)"></div>
|
||||
<div class="text-[0.65rem] text-zinc-500 uppercase">videos</div>
|
||||
@@ -127,29 +149,32 @@
|
||||
</template>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<!-- charts -->
|
||||
<div class="grid lg:grid-cols-2 gap-4">
|
||||
<div class="glass p-5">
|
||||
<h3 class="text-sm font-semibold text-zinc-200 mb-3">Uploads over time</h3>
|
||||
<div class="h-64"><canvas id="chart-uploads"></canvas></div>
|
||||
</div>
|
||||
<div class="glass p-5">
|
||||
<h3 class="text-sm font-semibold text-zinc-200 mb-3">Duration distribution</h3>
|
||||
<div class="h-64"><canvas id="chart-duration"></canvas></div>
|
||||
</div>
|
||||
<div class="glass p-5 lg:col-span-2">
|
||||
<h3 class="text-sm font-semibold text-zinc-200 mb-3">Top tags</h3>
|
||||
<div class="h-64"><canvas id="chart-tags"></canvas></div>
|
||||
</div>
|
||||
</div>
|
||||
</section>
|
||||
</template>
|
||||
|
||||
<!-- ===== Channels ===== -->
|
||||
<template x-if="view==='channels'">
|
||||
<section class="space-y-6" x-transition.opacity>
|
||||
<header><h1 class="text-2xl font-bold text-white">Channels</h1><p class="text-sm text-zinc-500">Add a channel URL to start tracking</p></header>
|
||||
<header class="flex items-end justify-between gap-3 flex-wrap">
|
||||
<div>
|
||||
<h1 class="text-2xl font-bold text-white">Channels</h1>
|
||||
<p class="text-sm text-zinc-500">Add a channel URL to start tracking</p>
|
||||
</div>
|
||||
<div class="flex items-center gap-2">
|
||||
<span x-show="sync.lastResult" class="text-xs transition-opacity" :class="sync.lastResultKind === 'error' ? 'text-rose-400' : 'text-emerald-400'" x-text="sync.lastResult"></span>
|
||||
<button class="btn accent-grad btn-primary" :disabled="jobActive() || channels.items.length===0" @click="startDiscovery()" title="Scan only what was uploaded since the last video we have">
|
||||
<svg x-show="jobActive() && scrape.kind==='discovery'" class="spin w-4 h-4" viewBox="0 0 24 24" fill="none"><circle cx="12" cy="12" r="9" stroke="currentColor" stroke-width="3" stroke-dasharray="40 20"/></svg>
|
||||
Investigate all
|
||||
</button>
|
||||
<button class="btn btn-ghost" :disabled="jobActive() || channels.items.length===0" @click="startDiscovery(null, true)" title="Re-read every page of every channel — slow, only if videos are missing">
|
||||
Full rescan
|
||||
</button>
|
||||
<button class="btn btn-ghost" :disabled="sync.busy==='all' || channels.items.length===0" @click="syncChannels()">
|
||||
<svg x-show="sync.busy==='all'" class="spin w-4 h-4" viewBox="0 0 24 24" fill="none"><circle cx="12" cy="12" r="9" stroke="currentColor" stroke-width="3" stroke-dasharray="40 20"/></svg>
|
||||
Sync channels
|
||||
</button>
|
||||
</div>
|
||||
</header>
|
||||
|
||||
<form class="glass p-4 flex gap-2" @submit.prevent="addChannel()">
|
||||
<input class="field flex-1" type="text" placeholder="https://www.youtube.com/@handle" x-model="forms.channelUrl" />
|
||||
@@ -162,20 +187,59 @@
|
||||
|
||||
<div class="glass overflow-hidden">
|
||||
<table class="tbl">
|
||||
<thead><tr><th>Name</th><th>Handle</th><th>Videos</th><th>Pending</th><th>Last scraped</th><th class="text-right">Actions</th></tr></thead>
|
||||
<thead><tr><th>Name</th><th>Handle</th><th>Videos</th><th>Pending</th><th>Latest video</th><th>Last scraped</th><th class="text-right">Actions</th></tr></thead>
|
||||
<tbody>
|
||||
<template x-if="loading.channels"><tr><td colspan="6" class="text-center text-zinc-500 py-8">loading…</td></tr></template>
|
||||
<template x-if="!loading.channels && channels.items.length===0"><tr><td colspan="6" class="text-center text-zinc-500 py-8">No channels tracked yet.</td></tr></template>
|
||||
<template x-if="loading.channels">
|
||||
<template x-for="i in 5" :key="i">
|
||||
<tr>
|
||||
<td>
|
||||
<template x-if="i===1"><span class="sr-only" role="status">Loading channels…</span></template>
|
||||
<div class="flex items-center gap-2" aria-hidden="true">
|
||||
<div class="skeleton w-7 h-7 rounded-full"></div>
|
||||
<div class="skeleton h-4 w-28"></div>
|
||||
</div>
|
||||
</td>
|
||||
<td><div class="skeleton h-3.5 w-20" aria-hidden="true"></div></td>
|
||||
<td><div class="skeleton h-3.5 w-10" aria-hidden="true"></div></td>
|
||||
<td><div class="skeleton h-5 w-20 rounded-full" aria-hidden="true"></div></td>
|
||||
<td><div class="skeleton h-3 w-16" aria-hidden="true"></div></td>
|
||||
<td><div class="skeleton h-3 w-16" aria-hidden="true"></div></td>
|
||||
<td>
|
||||
<div class="flex items-center justify-end gap-1" aria-hidden="true">
|
||||
<div class="skeleton h-5 w-14"></div>
|
||||
<div class="skeleton h-5 w-12"></div>
|
||||
<div class="skeleton h-5 w-12"></div>
|
||||
</div>
|
||||
</td>
|
||||
</tr>
|
||||
</template>
|
||||
</template>
|
||||
<template x-if="!loading.channels && channels.items.length===0"><tr><td colspan="7" class="text-center text-zinc-500 py-8">No channels tracked yet.</td></tr></template>
|
||||
<template x-for="c in channels.items" :key="c.channel_id">
|
||||
<tr class="row-hover">
|
||||
<td><button class="text-white font-medium hover:text-rose-400 transition-colors text-left" @click="openChannel(c.channel_id)" x-text="c.name"></button></td>
|
||||
<td>
|
||||
<div class="flex items-center gap-2">
|
||||
<div class="relative w-7 h-7 shrink-0">
|
||||
<img class="avatar w-7 h-7 shrink-0" :src="c.avatar || '/api/avatars/'+c.channel_id" :alt="c.name" :data-fb-src="'/api/avatars/'+c.channel_id" onerror="if(!this.dataset.swapped){this.dataset.swapped='1'; this.src=this.dataset.fbSrc||''}else{this.style.visibility='hidden'; if(this.nextElementSibling) this.nextElementSibling.style.display='inline-flex'}" />
|
||||
<span class="avatar avatar-initial rounded-full w-7 h-7 text-xs absolute inset-0 hidden" :text-content="(c.name||c.channel_id||'?').charAt(0)" aria-hidden="true"></span>
|
||||
</div>
|
||||
<button class="text-white font-medium hover:text-rose-400 transition-colors text-left" @click="openChannel(c.channel_id)" x-text="c.name"></button>
|
||||
</div>
|
||||
</td>
|
||||
<td class="text-zinc-400" x-text="c.handle || '—'"></td>
|
||||
<td><button class="font-mono text-zinc-300 hover:text-rose-400 transition-colors" @click="openChannel(c.channel_id)" x-text="fmtNum(c.video_count)"></button></td>
|
||||
<td>
|
||||
<span class="pill" :class="(channels.pending[c.channel_id]||0) > 0 ? 'st-pending' : 'st-skipped'" x-text="(channels.pending[c.channel_id]||0) + ' pending'"></span>
|
||||
</td>
|
||||
<td class="text-zinc-400 text-xs font-mono" x-text="fmtDate(c.last_video_date)" title="Newest upload we know of — incremental scans start here"></td>
|
||||
<td class="text-zinc-400 text-xs" x-text="c.last_scraped ? fmtDate(c.last_scraped) : '—'"></td>
|
||||
<td class="text-right whitespace-nowrap">
|
||||
<button class="btn accent-grad btn-primary !py-1 !px-2 text-xs" @click="startDiscovery(c.channel_id)" :disabled="jobActive()" title="Find videos uploaded since the latest one we have">
|
||||
<svg x-show="jobActive() && scrape.kind==='discovery' && scrape.targetChannel===c.channel_id" class="spin w-3 h-3" viewBox="0 0 24 24" fill="none"><circle cx="12" cy="12" r="9" stroke="currentColor" stroke-width="3" stroke-dasharray="40 20"></circle></svg>
|
||||
Investigate
|
||||
</button>
|
||||
<button class="btn btn-ghost !py-1 !px-2 text-xs" @click="startDiscovery(c.channel_id, true)" :disabled="jobActive()" title="Re-read every page of this channel">Full</button>
|
||||
<button class="btn btn-ghost !py-1 !px-2 text-xs" @click="syncChannels(c.channel_id)" :disabled="sync.busy===c.channel_id || sync.busy==='all'" title="Re-scrape channel info + avatar">Sync</button>
|
||||
<button class="btn btn-ghost !py-1 !px-2 text-xs" @click="openChannel(c.channel_id)">Videos</button>
|
||||
<button class="btn btn-ghost !py-1 !px-2 text-xs" @click="openFolder('markdown', c.channel_id)" title="Open .md folder">.md</button>
|
||||
<button class="btn accent-grad btn-primary !py-1 !px-2 text-xs" @click="downloadChannelPending(c.channel_id)" title="Process pending videos to .md" :disabled="(channels.pending[c.channel_id]||0)===0">Pending .md</button>
|
||||
@@ -192,30 +256,84 @@
|
||||
<!-- ===== Videos ===== -->
|
||||
<template x-if="view==='videos'">
|
||||
<section class="space-y-5" x-transition.opacity>
|
||||
<header><h1 class="text-2xl font-bold text-white">Videos</h1><p class="text-sm text-zinc-500">Browse and filter the catalog</p></header>
|
||||
<header class="flex items-end justify-between gap-3 flex-wrap">
|
||||
<div><h1 class="text-2xl font-bold text-white">Videos</h1><p class="text-sm text-zinc-500">Browse and filter the catalog</p></div>
|
||||
<div class="flex items-center gap-2 flex-wrap">
|
||||
<button class="btn btn-ghost" :disabled="loading.reconcile" @click="reconcile()" title="Re-scan data/markdown and repair any status that disagrees with what is on disk">
|
||||
<svg x-show="loading.reconcile" class="spin w-4 h-4" viewBox="0 0 24 24" fill="none"><circle cx="12" cy="12" r="9" stroke="currentColor" stroke-width="3" stroke-dasharray="40 20"/></svg>
|
||||
Sync from disk
|
||||
</button>
|
||||
<button class="btn btn-ghost" x-show="retryable.total > 0" :disabled="loading.reset || jobActive()" @click="resetVideos()"
|
||||
:title="'Send ' + retryable.total + ' failed/skipped videos back to pending so they can be retried'">
|
||||
<svg x-show="loading.reset" class="spin w-4 h-4" viewBox="0 0 24 24" fill="none"><circle cx="12" cy="12" r="9" stroke="currentColor" stroke-width="3" stroke-dasharray="40 20"/></svg>
|
||||
<span x-text="'Retry ' + retryable.total + ' stuck'"></span>
|
||||
</button>
|
||||
<button class="btn accent-grad btn-primary" :disabled="jobActive() || channels.items.length===0" @click="startDiscovery(filters.channel || null)">
|
||||
<svg x-show="jobActive() && scrape.kind==='discovery'" class="spin w-4 h-4" viewBox="0 0 24 24" fill="none"><circle cx="12" cy="12" r="9" stroke="currentColor" stroke-width="3" stroke-dasharray="40 20"/></svg>
|
||||
<span x-text="filters.channel ? 'Investigate channel' : 'Investigate all'"></span>
|
||||
</button>
|
||||
</div>
|
||||
</header>
|
||||
|
||||
<div x-show="retryable.total > 0" class="glass px-4 py-3 text-sm text-zinc-300 flex items-start gap-3">
|
||||
<span class="text-rose-400 font-bold shrink-0">!</span>
|
||||
<div>
|
||||
<span x-text="retryable.total"></span> video(s) sit in a terminal status
|
||||
(<span x-text="(retryable.counts.no_subtitles||0) + ' no_subtitles, ' + (retryable.counts.error||0) + ' error'"></span>)
|
||||
and will never be picked up by “Pending .md”.
|
||||
<span class="text-zinc-500">“no_subtitles” often just means the language policy rejected the auto-generated captions — hover a status badge for the recorded reason.</span>
|
||||
<span x-show="retryable.permanent > 0" class="block mt-1 text-zinc-500">
|
||||
A further <span class="text-zinc-300" x-text="retryable.permanent"></span> are permanently blocked
|
||||
(members-only, private or removed) and are excluded from retries — no amount of re-running will fetch them.
|
||||
</span>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div class="glass p-4 grid md:grid-cols-6 gap-3">
|
||||
<select class="field" x-model="filters.channel" @change="loadVideos(1)"><option value="">All channels</option><template x-for="c in channels.items" :key="c.channel_id"><option :value="c.channel_id" x-text="c.name"></option></template></select>
|
||||
<select class="field" x-model="filters.status" @change="loadVideos(1)">
|
||||
<option value="">Any status</option>
|
||||
<option value="done">Done</option><option value="pending">Pending</option><option value="error">Error</option><option value="skipped">Skipped</option><option value="new">New</option>
|
||||
<option value="done">Done</option>
|
||||
<option value="pending">Pending</option>
|
||||
<option value="no_subtitles">No subtitles</option>
|
||||
<option value="error">Error</option>
|
||||
<option value="__blocked__">Locked (members-only…)</option>
|
||||
</select>
|
||||
<input class="field" type="date" x-model="filters.from" @change="loadVideos(1)" />
|
||||
<input class="field" type="date" x-model="filters.to" @change="loadVideos(1)" />
|
||||
<input class="field" type="number" min="0" step="60" placeholder="min sec" x-model="filters.min_dur" @change="loadVideos(1)" />
|
||||
<input class="field" type="number" min="0" step="30" placeholder="min duration (s)" title="Hide videos shorter than this many seconds" x-model="filters.min_dur" @change="loadVideos(1)" aria-label="Minimum duration in seconds" />
|
||||
<div class="flex gap-2">
|
||||
<input class="field flex-1" type="text" placeholder="search title…" x-model="filters.q" @keydown.enter="loadVideos(1)" />
|
||||
<button class="btn btn-ghost" @click="loadVideos(1)">Go</button>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<!-- active filters: visible at a glance and one click to remove, so
|
||||
a channel filter set from the dashboard never feels like a trap -->
|
||||
<div x-show="filters.channel || filters.status" x-cloak class="flex flex-wrap items-center gap-2">
|
||||
<span class="text-xs text-zinc-500">Showing</span>
|
||||
<template x-if="filters.channel">
|
||||
<button class="filter-chip" @click="clearFilter('channel')" title="Clear this filter and show all channels">
|
||||
<span class="max-w-56 truncate" x-text="chanName(filters.channel)"></span>
|
||||
<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" aria-hidden="true"><path stroke-linecap="round" d="M6 6l12 12M18 6L6 18"/></svg>
|
||||
</button>
|
||||
</template>
|
||||
<template x-if="filters.status">
|
||||
<button class="filter-chip" @click="clearFilter('status')" title="Clear this status filter">
|
||||
<span x-text="filters.status === '__blocked__' ? 'locked videos' : filters.status + ' videos'"></span>
|
||||
<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" aria-hidden="true"><path stroke-linecap="round" d="M6 6l12 12M18 6L6 18"/></svg>
|
||||
</button>
|
||||
</template>
|
||||
<button class="text-xs text-zinc-400 hover:text-white underline underline-offset-2 transition-colors" @click="clearAllFilters()">Clear all</button>
|
||||
</div>
|
||||
|
||||
<!-- bulk action bar -->
|
||||
<template x-if="videos.selected.length > 0">
|
||||
<div class="bulk-bar glass p-3 flex flex-wrap items-center gap-2">
|
||||
<span class="pill accent-grad btn-primary font-bold" x-text="videos.selected.length + ' selected'"></span>
|
||||
<button class="btn accent-grad btn-primary !py-1.5 !px-3 text-xs" @click="downloadMdBatch()" title="Download .md + thumbnails for the selection">
|
||||
<button class="btn accent-grad btn-primary !py-1.5 !px-3 text-xs" @click="downloadMdBatch()" title="Generate the .md notes for the selection. Thumbnails are not touched — they already come from the channel scrape.">
|
||||
<svg class="w-3.5 h-3.5" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><path stroke-linecap="round" stroke-linejoin="round" d="M12 4v12m0 0l-4-4m4 4l4-4M4 18v2h16v-2"/></svg>
|
||||
Download .md + thumbs
|
||||
Download .md
|
||||
</button>
|
||||
<button class="btn btn-ghost !py-1.5 !px-3 text-xs" @click="downloadAudio()">
|
||||
<svg class="w-3.5 h-3.5" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><path stroke-linecap="round" stroke-linejoin="round" d="M9 18V6l10-2v12M9 18a3 3 0 11-6 0 3 3 0 016 0zm10-2a3 3 0 11-6 0 3 3 0 016 0z"/></svg>
|
||||
@@ -232,26 +350,69 @@
|
||||
<thead><tr>
|
||||
<th class="w-8"><input type="checkbox" class="accent-rose-500" :checked="allSelected" @change="toggleSelectAll()" title="Select page" /></th>
|
||||
<th class="w-24">Thumb</th><th>Title</th><th>Channel</th>
|
||||
<th><select class="bg-transparent text-zinc-500 uppercase text-[0.7rem] tracking-wider font-semibold" x-model="filters.sort" @change="loadVideos(videos.page)">
|
||||
<option value="upload_date">Uploaded</option><option value="duration">Duration</option><option value="view_count">Views</option><option value="like_count">Likes</option>
|
||||
</select></th>
|
||||
<!-- No @click on the th: the select bubbles up to it, so a plain
|
||||
click on the dropdown used to cycle the sort as well. -->
|
||||
<th class="select-none">
|
||||
<select class="bg-transparent text-zinc-300 uppercase text-[0.7rem] tracking-wider font-semibold cursor-pointer hover:text-white focus:outline-none" x-model="filters.sort" @change="loadVideos(1)" aria-label="Sort videos" title="Sort order — newest upload first by default">
|
||||
<option value="upload_date">Newest ▼</option>
|
||||
<option value="oldest">Oldest ▲</option>
|
||||
<option value="view_count">Views ▼</option>
|
||||
<option value="like_count">Likes ▼</option>
|
||||
<option value="duration">Duration ▼</option>
|
||||
<option value="title">Title A→Z</option>
|
||||
<option value="channel">Channel A→Z</option>
|
||||
<option value="status">Status</option>
|
||||
</select>
|
||||
</th>
|
||||
<th>Duration</th><th>Views</th><th>Status</th><th class="text-right">Actions</th>
|
||||
</tr></thead>
|
||||
<tbody>
|
||||
<template x-if="loading.videos"><tr><td colspan="9" class="text-center text-zinc-500 py-10">loading…</td></tr></template>
|
||||
<template x-if="!loading.videos && videos.items.length===0"><tr><td colspan="9" class="text-center text-zinc-500 py-10">No videos match these filters.</td></tr></template>
|
||||
<!-- skeleton rows mirror the real columns; the sr-only status rides row 1 so each table has exactly one live region -->
|
||||
<template x-if="loading.videos">
|
||||
<template x-for="i in 8" :key="i">
|
||||
<tr>
|
||||
<td>
|
||||
<template x-if="i===1"><span class="sr-only" role="status">Loading videos…</span></template>
|
||||
<div class="skeleton w-4 h-4 rounded" aria-hidden="true"></div>
|
||||
</td>
|
||||
<td><div class="skeleton w-20 h-12 rounded-lg" aria-hidden="true"></div></td>
|
||||
<td><div class="skeleton h-4 w-full max-w-md" aria-hidden="true"></div></td>
|
||||
<td><div class="skeleton h-3 w-24" aria-hidden="true"></div></td>
|
||||
<td><div class="skeleton h-3 w-16" aria-hidden="true"></div></td>
|
||||
<td><div class="skeleton h-3 w-12" aria-hidden="true"></div></td>
|
||||
<td><div class="skeleton h-3 w-14" aria-hidden="true"></div></td>
|
||||
<td><div class="skeleton h-5 w-16 rounded-full" aria-hidden="true"></div></td>
|
||||
<td>
|
||||
<div class="flex items-center justify-end gap-1" aria-hidden="true">
|
||||
<div class="skeleton h-5 w-12"></div>
|
||||
<div class="skeleton h-5 w-10"></div>
|
||||
</div>
|
||||
</td>
|
||||
</tr>
|
||||
</template>
|
||||
</template>
|
||||
<template x-if="!loading.videos && videos.items.length===0"><tr><td colspan="9" class="text-center text-zinc-500 py-10"><div>No videos match these filters.</div><button class="btn btn-ghost mt-3" @click.stop="clearAllFilters()">Clear filters</button></td></tr></template>
|
||||
<template x-for="v in videos.items" :key="v.video_id">
|
||||
<tr class="row-hover" :class="isSelected(v.video_id) ? 'row-selected' : ''" @click="openVideo(v.video_id)">
|
||||
<td @click.stop><input type="checkbox" class="accent-rose-500" :checked="isSelected(v.video_id)" @change="toggleSelect(v.video_id)" /></td>
|
||||
<td><img class="thumb w-20 h-12" :src="'/api/thumbnails/'+v.video_id" :alt="v.title" onerror="this.style.visibility='hidden'" /></td>
|
||||
<td class="text-zinc-100 max-w-md truncate" x-text="v.title"></td>
|
||||
<td class="text-zinc-400 text-xs" x-text="chanName(v.channel_id)"></td>
|
||||
<td class="text-zinc-400 font-mono text-xs" x-text="fmtDate(v.upload_date)"></td>
|
||||
<td class="font-mono text-xs" :class="v.upload_date ? 'text-zinc-400' : 'text-zinc-600 italic'" x-text="videoDate(v)" :title="videoDateTitle(v)"></td>
|
||||
<td class="font-mono text-zinc-300 text-xs" x-text="ts(v.duration)"></td>
|
||||
<td class="font-mono text-zinc-300 text-xs" x-text="fmtNum(v.view_count)"></td>
|
||||
<td><span class="pill" :class="statusClass(v.status)" x-text="v.status"></span></td>
|
||||
<td>
|
||||
<div class="flex items-center gap-1 flex-wrap">
|
||||
<span class="pill" :class="statusClass(v.status)" x-text="v.status" :title="v.error_msg || v.status"></span>
|
||||
<span x-show="isBlocked(v)" class="pill st-locked" :title="blockTitle(v)">
|
||||
<svg class="w-3 h-3" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" aria-hidden="true"><rect x="4" y="10" width="16" height="10" rx="2"/><path d="M8 10V7a4 4 0 1 1 8 0v3"/></svg>
|
||||
<span x-text="blockLabel(v)"></span>
|
||||
</span>
|
||||
</div>
|
||||
</td>
|
||||
<td class="text-right whitespace-nowrap" @click.stop>
|
||||
<button class="btn btn-ghost !py-1 !px-2 text-xs" :class="v.status==='done' ? 'md-done' : 'md-todo'" @click="v.status==='done' ? downloadMd(v.video_id) : processOne(v.video_id)" :title="v.status==='done' ? 'Download .md' : 'Process to .md'" x-text="v.status==='done' ? '.md' : 'Process'"></button>
|
||||
<button x-show="!isBlocked(v) || v.status==='done'" class="btn btn-ghost !py-1 !px-2 text-xs" :class="v.status==='done' ? 'md-done' : 'md-todo'" @click="v.status==='done' ? downloadMd(v.video_id) : processOne(v.video_id)" :title="v.status==='done' ? 'Download .md' : 'Process to .md'" x-text="v.status==='done' ? '.md' : 'Process'"></button>
|
||||
<button x-show="isBlocked(v) && v.status!=='done'" class="btn btn-ghost !py-1 !px-2 text-xs opacity-50" @click="processOne(v.video_id)" :title="blockTitle(v) + ' Click anyway if you have since bought the membership and refreshed your cookies.'">Locked</button>
|
||||
<button class="btn btn-ghost !py-1 !px-2 text-xs" @click="openClip(v.video_id)" title="Extract transcript clip">Clip</button>
|
||||
</td>
|
||||
</tr>
|
||||
@@ -261,20 +422,28 @@
|
||||
</div>
|
||||
|
||||
<!-- pagination -->
|
||||
<div class="flex items-center justify-between text-sm">
|
||||
<div class="text-zinc-500" x-text="videos.total ? (videos.size*(videos.page-1)+1)+'–'+Math.min(videos.size*videos.page, videos.total)+' of '+videos.total : '0 results'"></div>
|
||||
<div class="flex gap-2">
|
||||
<button class="btn btn-ghost" :disabled="videos.page<=1" @click="loadVideos(videos.page-1)">Prev</button>
|
||||
<button class="btn btn-ghost" :disabled="videos.size*videos.page >= videos.total" @click="loadVideos(videos.page+1)">Next</button>
|
||||
<div class="flex items-center justify-between gap-3 text-sm flex-wrap">
|
||||
<div class="text-zinc-500"><span x-text="videos.total ? (videos.size*(videos.page-1)+1)+'–'+Math.min(videos.size*videos.page, videos.total)+' of '+videos.total : '0 results'"></span><span class="ml-3" x-text="'Page '+(videos.total ? videos.page : 0)+' of '+pageCount"></span></div>
|
||||
<div class="flex items-center gap-2">
|
||||
<label class="text-zinc-500 text-xs" for="videos-size">Per page</label><select id="videos-size" class="field !py-1.5 !px-2 w-20" x-model.number="videos.size" @change="localStorage.setItem('videos-size', videos.size); loadVideos(1)"><template x-for="n in [10,25,50,100]" :key="n"><option :value="n" x-text="n"></option></template></select>
|
||||
<button class="btn btn-ghost !py-1.5 !px-2" :disabled="videos.page<=1" @click="loadVideos(videos.page-1); window.scrollTo({top:0,behavior:'smooth'})">Prev</button>
|
||||
<template x-for="item in pageNumbers" :key="item"><button x-show="item !== '…'" class="btn btn-ghost !py-1.5 !px-2 min-w-8" :class="item===videos.page ? 'accent-grad btn-primary' : ''" @click="loadVideos(item); window.scrollTo({top:0,behavior:'smooth'})" x-text="item"></button><span x-show="item === '…'" class="px-1 text-zinc-500">…</span></template>
|
||||
<button class="btn btn-ghost !py-1.5 !px-2" :disabled="videos.page>=pageCount" @click="loadVideos(videos.page+1); window.scrollTo({top:0,behavior:'smooth'})">Next</button>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
</section>
|
||||
</template>
|
||||
|
||||
<!-- ===== Detail ===== -->
|
||||
<template x-if="view==='detail'">
|
||||
<section class="space-y-5" x-transition.opacity>
|
||||
<button class="btn btn-ghost" @click="setView('videos')">← Back to videos</button>
|
||||
<div class="flex items-center justify-between gap-3 flex-wrap">
|
||||
<button class="btn btn-ghost" @click="goBack()" title="Go back (or press Esc)">
|
||||
<svg class="w-4 h-4 shrink-0" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" aria-hidden="true"><path stroke-linecap="round" stroke-linejoin="round" d="M15 19l-7-7 7-7"/></svg>
|
||||
<span x-text="backLabel"></span>
|
||||
</button>
|
||||
</div>
|
||||
<template x-if="loading.detail"><div class="glass p-10 text-center text-zinc-500">loading…</div></template>
|
||||
<template x-if="!loading.detail && detail.video">
|
||||
<div class="space-y-5">
|
||||
@@ -288,19 +457,88 @@
|
||||
<div><span class="text-zinc-500">Duration </span><span class="text-zinc-200 font-mono" x-text="ts(detail.video.duration)"></span></div>
|
||||
<div><span class="text-zinc-500">Views </span><span class="text-zinc-200 font-mono" x-text="fmtNum(detail.video.view_count)"></span></div>
|
||||
<div><span class="text-zinc-500">Likes </span><span class="text-zinc-200 font-mono" x-text="fmtNum(detail.video.like_count)"></span></div>
|
||||
<div><span class="pill" :class="statusClass(detail.video.status)" x-text="detail.video.status"></span></div>
|
||||
<div class="flex items-center gap-1">
|
||||
<span class="pill" :class="statusClass(detail.video.status)" x-text="detail.video.status"></span>
|
||||
<span x-show="isBlocked(detail.video)" class="pill st-locked" :title="blockTitle(detail.video)">
|
||||
<svg class="w-3 h-3" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" aria-hidden="true"><rect x="4" y="10" width="16" height="10" rx="2"/><path d="M8 10V7a4 4 0 1 1 8 0v3"/></svg>
|
||||
<span x-text="blockLabel(detail.video)"></span>
|
||||
</span>
|
||||
</div>
|
||||
</div>
|
||||
<div x-show="isBlocked(detail.video)" class="mt-3 text-xs text-amber-300/90" x-text="blockTitle(detail.video)"></div>
|
||||
<div x-show="detail.video.status==='no_subtitles' || detail.video.status==='error'" class="mt-3 text-xs text-amber-300/90">
|
||||
No <span class="font-mono">.md</span> was generated.
|
||||
<span x-show="detail.video.error_msg" class="block mt-1 text-zinc-400 font-mono" x-text="detail.video.error_msg"></span>
|
||||
<span x-show="!detail.video.error_msg" class="block mt-1 text-zinc-500">No reason was recorded (processed before reasons were stored). Re-process to find out why.</span>
|
||||
</div>
|
||||
<div x-show="detail.video.status==='error' && detail.video.error_msg" class="mt-3 text-xs text-rose-300" x-text="detail.video.error_msg"></div>
|
||||
<div class="flex flex-wrap gap-2 mt-4">
|
||||
<button class="btn accent-grad btn-primary !py-1.5 !px-3 text-xs" @click="detail.video.status==='done' ? downloadMd(detail.video.video_id) : processOne(detail.video.video_id)" x-text="detail.video.status==='done' ? 'Download .md' : 'Process to .md'"></button>
|
||||
<button class="btn accent-grad btn-primary !py-1.5 !px-3 text-xs" @click="startDiscovery(detail.video.channel_id)" :disabled="jobActive()" title="Find recent videos from this channel without downloading transcripts">Investigate channel</button>
|
||||
<button x-show="!isBlocked(detail.video) || detail.video.status==='done'" class="btn accent-grad btn-primary !py-1.5 !px-3 text-xs" @click="detail.video.status==='done' ? downloadMd(detail.video.video_id) : processOne(detail.video.video_id)" x-text="detail.video.status==='done' ? 'Download .md' : 'Process to .md'"></button>
|
||||
<button class="btn btn-ghost !py-1.5 !px-3 text-xs" @click="copyMd(detail.video.video_id)" x-show="detail.video && detail.video.status==='done'" title="Copiar todo el contenido del .md al portapapeles">
|
||||
<svg class="w-3.5 h-3.5" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><rect x="9" y="9" width="13" height="13" rx="2" ry="2"/><path d="M5 15H4a2 2 0 0 1-2-2V4a2 2 0 0 1 2-2h9a2 2 0 0 1 2 2v1"/></svg>
|
||||
<span>Copiar .md</span>
|
||||
</button>
|
||||
<button class="btn btn-ghost !py-1.5 !px-3 text-xs" @click="openMd(detail.video.video_id)" x-show="detail.video && detail.video.status==='done'" title="Abrir archivo .md en tu aplicación predeterminada (Obsidian, VS Code, etc.)">
|
||||
<svg class="w-3.5 h-3.5" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><path d="M18 13v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/><polyline points="15 3 21 3 21 9"/><line x1="10" y1="14" x2="21" y2="3"/></svg>
|
||||
<span>Abrir .md</span>
|
||||
</button>
|
||||
<button x-show="isBlocked(detail.video) && detail.video.status!=='done'" class="btn btn-ghost !py-1.5 !px-3 text-xs opacity-60" @click="processOne(detail.video.video_id)" :title="blockTitle(detail.video) + ' Click anyway if you have since bought the membership and refreshed your cookies.'">Try anyway</button>
|
||||
<button class="btn btn-ghost !py-1.5 !px-3 text-xs" @click="downloadAudioOne(detail.video.video_id)" x-show="detail.video.status==='done'" x-text="detail.hasAudio ? 'Re-download audio' : 'Download audio'">
|
||||
<svg class="w-3.5 h-3.5" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><path stroke-linecap="round" stroke-linejoin="round" d="M9 18V6l10-2v12M9 18a3 3 0 11-6 0 3 3 0 016 0zm10-2a3 3 0 11-6 0 3 3 0 016 0z"/></svg>
|
||||
</button>
|
||||
<button class="btn accent-grad btn-primary !py-1.5 !px-3 text-xs" @click="downloadVideoOne(detail.video.video_id)" x-show="detail.video.status==='done' && detail.video.video_download_status!=='done'" :disabled="detail.video.video_download_status==='queued' || detail.video.video_download_status==='downloading'">
|
||||
<svg x-show="detail.video.video_download_status==='queued' || detail.video.video_download_status==='downloading'" class="spin w-3.5 h-3.5" viewBox="0 0 24 24" fill="none"><circle cx="12" cy="12" r="9" stroke="currentColor" stroke-width="3" stroke-dasharray="40 20"/></svg>
|
||||
<span x-text="detail.video.video_download_status==='queued' || detail.video.video_download_status==='downloading' ? 'Downloading video…' : 'Download video · 1080p WebM'"></span>
|
||||
</button>
|
||||
<a class="btn btn-ghost !py-1.5 !px-3 text-xs" x-show="detail.video.video_download_status==='done'" :href="'/api/videos/'+detail.video.video_id+'/media'" :download="detail.video.video_filename || ''">Download WebM</a>
|
||||
<button class="btn btn-ghost !py-1.5 !px-3 text-xs" @click="openClip(detail.video.video_id)">Clip</button>
|
||||
<template x-if="detail.video.url"><a class="btn btn-ghost !py-1.5 !px-3 text-xs" :href="detail.video.url" target="_blank" rel="noopener">Open on YouTube ↗</a></template>
|
||||
</div>
|
||||
<div class="flex flex-wrap gap-1.5 mt-3" x-show="(detail.video.tags||[]).length">
|
||||
<template x-for="t in (detail.video.tags||[]).slice(0,8)" :key="t"><span class="pill st-skipped" x-text="t"></span></template>
|
||||
</div>
|
||||
<div x-show="detail.video.description" x-data="{ descOpen: false }" class="mt-3 pt-3 border-t border-zinc-800/70">
|
||||
<button class="flex items-center gap-2 text-xs text-zinc-400 hover:text-rose-400 transition-colors w-full text-left" @click="descOpen = !descOpen" :aria-expanded="descOpen ? 'true' : 'false'" title="Click to toggle description">
|
||||
<svg :class="descOpen ? 'rotate-90' : ''" class="w-3 h-3 transition-transform shrink-0" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5"><path stroke-linecap="round" stroke-linejoin="round" d="M9 5l7 7-7 7"/></svg>
|
||||
<span class="uppercase tracking-wider font-semibold">Description</span>
|
||||
</button>
|
||||
<div x-show="descOpen" x-transition.opacity class="mt-3 text-sm text-zinc-300 whitespace-pre-wrap leading-relaxed max-h-72 overflow-y-auto pr-1" x-text="detail.video.description"></div>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<!-- local cinema player; the browser supplies the decoder while this shell supplies the experience -->
|
||||
<div class="cinema-card" x-show="detail.video && detail.video.video_download_status==='done'">
|
||||
<div class="cinema-stage">
|
||||
<video x-ref="videoPlayer" class="cinema-video" controls playsinline preload="metadata"
|
||||
:src="'/api/videos/'+detail.video.video_id+'/media?inline=1'"
|
||||
@error="onVideoError()"></video>
|
||||
<div class="cinema-error" x-show="detail.videoError">
|
||||
<div class="text-white font-semibold">Este navegador no pudo reproducir el WebM directamente.</div>
|
||||
<div class="text-sm text-zinc-400 mt-1">Puedes descargar el archivo y abrirlo con VLC u otro reproductor local.</div>
|
||||
<a class="btn btn-ghost mt-3" :href="'/api/videos/'+detail.video.video_id+'/media'" :download="detail.video.video_filename || ''">Descargar WebM</a>
|
||||
</div>
|
||||
</div>
|
||||
<div class="cinema-meta">
|
||||
<div>
|
||||
<div class="text-white font-semibold" x-text="detail.video.title"></div>
|
||||
<div class="text-xs text-zinc-500 mt-1">1080p · WebM · <span x-text="fmtBytes(detail.video.video_size)"></span></div>
|
||||
</div>
|
||||
<button class="btn btn-ghost !py-1.5 !px-3 text-xs" @click="openVideoPlayer()">Play</button>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div class="glass p-4" x-show="detail.video && (detail.video.video_download_status==='queued' || detail.video.video_download_status==='downloading')">
|
||||
<div class="flex items-center justify-between gap-3 text-sm">
|
||||
<span class="text-zinc-300">Downloading video…</span>
|
||||
<span class="font-mono text-zinc-400" x-text="detail.video.video_progress && detail.video.video_progress.percent != null ? detail.video.video_progress.percent + '%' : 'preparing' "></span>
|
||||
</div>
|
||||
<div class="progress-track mt-3"><div class="progress-fill" :style="'width:' + ((detail.video.video_progress && detail.video.video_progress.percent) || 0) + '%' "></div></div>
|
||||
<div class="text-xs text-zinc-500 mt-2" x-show="detail.video.video_progress">
|
||||
<span x-text="fmtBytes(detail.video.video_progress.downloaded_bytes)"></span>
|
||||
<span x-show="detail.video.video_progress.total_bytes"> / <span x-text="fmtBytes(detail.video.video_progress.total_bytes)"></span></span>
|
||||
<span class="ml-2" x-show="detail.video.video_progress.speed">· <span x-text="fmtSpeed(detail.video.video_progress.speed)"></span></span>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
@@ -372,12 +610,21 @@
|
||||
<section class="space-y-5" x-transition.opacity>
|
||||
<header><h1 class="text-2xl font-bold text-white">Search transcripts</h1><p class="text-sm text-zinc-500">Find moments across all videos</p></header>
|
||||
<form class="glass p-4 flex flex-col md:flex-row gap-2" @submit.prevent="loadSearch()">
|
||||
<input class="field flex-1" type="text" placeholder="search query…" x-model="search.q" />
|
||||
<select class="field md:w-56" x-model="search.channel"><option value="">All channels</option><template x-for="c in channels.items" :key="c.channel_id"><option :value="c.channel_id" x-text="c.name"></option></template></select>
|
||||
<button class="btn accent-grad btn-primary" :disabled="loading.search">Search</button>
|
||||
<input x-ref="searchInput" class="field flex-1" type="text" placeholder="search query… (Press Enter or /)" x-model="search.q" aria-label="Search query" />
|
||||
<select class="field md:w-56" x-model="search.channel" @change="if(search.ran) loadSearch()" aria-label="Filter by channel"><option value="">All channels</option><template x-for="c in channels.items" :key="c.channel_id"><option :value="c.channel_id" x-text="c.name"></option></template></select>
|
||||
<button class="btn accent-grad btn-primary" :disabled="loading.search || !search.q.trim()" @click="loadSearch()">Search</button>
|
||||
<button class="btn btn-ghost" type="button" @click="resetSearch()" x-show="search.ran || search.q || search.channel" title="Clear search">Clear</button>
|
||||
</form>
|
||||
<template x-if="loading.search"><div class="glass p-10 text-center text-zinc-500">searching…</div></template>
|
||||
<template x-if="!loading.search && search.ran && search.items.length===0"><div class="glass p-10 text-center text-zinc-500">No matches found.</div></template>
|
||||
<template x-if="!loading.search && search.ran && search.items.length===0">
|
||||
<div class="glass p-10 text-center">
|
||||
<div class="text-zinc-500 mb-3">No matches for <span class="text-zinc-300 font-mono" x-text="search.q"></span><span x-show="search.channel"> in <span class="text-zinc-300" x-text="chanName(search.channel)"></span></span>.</div>
|
||||
<button class="btn btn-ghost" @click="resetSearch()">Clear filters</button>
|
||||
</div>
|
||||
</template>
|
||||
<template x-if="!loading.search && !search.ran">
|
||||
<div class="glass p-10 text-center text-zinc-500">Type a query and press <kbd class="kbd">Enter</kbd> to search transcripts across all channels. Press / anywhere to jump to this search.<br><span class="text-xs text-zinc-500">Tip: try a phrase in its original language for best matches.</span></div>
|
||||
</template>
|
||||
<div class="space-y-3">
|
||||
<template x-for="r in search.items" :key="r.video_id+r.start_sec">
|
||||
<div class="glass p-4 hover:border-rose-500/40 transition-colors cursor-pointer" @click="openVideo(r.video_id)">
|
||||
@@ -457,12 +704,13 @@
|
||||
</div>
|
||||
</div>
|
||||
<div>
|
||||
<label class="block text-xs text-zinc-500 uppercase tracking-wider mb-1">Languages <span class="text-zinc-600 normal-case">(comma-sep)</span></label>
|
||||
<label class="block text-xs text-zinc-500 uppercase tracking-wider mb-1">Languages <span class="text-zinc-500 normal-case">(comma-sep)</span></label>
|
||||
<input class="field" type="text" placeholder="en, es" x-model="scrape.form.languages" />
|
||||
</div>
|
||||
<div class="flex gap-6 pt-1">
|
||||
<label class="flex items-center gap-2 text-sm text-zinc-300 cursor-pointer"><input type="checkbox" class="accent-rose-500" x-model="scrape.form.include_shorts" /> Include shorts</label>
|
||||
<label class="flex items-center gap-2 text-sm text-zinc-300 cursor-pointer"><input type="checkbox" class="accent-rose-500" x-model="scrape.form.no_live" /> Skip live</label>
|
||||
<label class="flex items-center gap-2 text-sm text-zinc-300 cursor-pointer" title="Off: discovery reads only what was uploaded since the latest video we have"><input type="checkbox" class="accent-rose-500" x-model="scrape.form.full" /> Full channel rescan</label>
|
||||
</div>
|
||||
<button class="btn accent-grad btn-primary w-full" :disabled="loading.startScrape || !scrape.form.channel_id">
|
||||
<svg x-show="loading.startScrape" class="spin w-4 h-4" viewBox="0 0 24 24" fill="none"><circle cx="12" cy="12" r="9" stroke="currentColor" stroke-width="3" stroke-dasharray="40 20"/></svg>
|
||||
@@ -478,7 +726,7 @@
|
||||
<span class="pill" :class="scrape.error ? 'st-error' : (scrape.done ? 'st-done' : 'st-running')" x-text="scrape.error ? 'error' : (scrape.done ? 'done' : 'running')"></span>
|
||||
</template>
|
||||
</div>
|
||||
<template x-if="!scrape.jobId"><div class="flex-1 flex items-center justify-center text-zinc-600 text-sm">No active job. Start one to see live progress.</div></template>
|
||||
<template x-if="!scrape.jobId"><div class="flex-1 flex items-center justify-center text-zinc-500 text-sm">No active job. Start one to see live progress.</div></template>
|
||||
<template x-if="scrape.jobId">
|
||||
<div class="mt-4 space-y-3 flex-1 flex flex-col">
|
||||
<div class="flex justify-between text-xs text-zinc-500"><span>Progress</span><span class="font-mono" x-text="(scrape.progress.completed||0)+' / '+(scrape.progress.total||0)"></span></div>
|
||||
@@ -500,17 +748,33 @@
|
||||
<div class="px-4 py-3 border-b border-zinc-800/70 flex items-center justify-between">
|
||||
<div class="flex flex-col">
|
||||
<h3 class="text-sm font-semibold text-zinc-200">Recent jobs</h3>
|
||||
<span class="text-[0.65rem] text-zinc-600 mt-0.5">Clearing history removes job records only — downloaded .md files stay.</span>
|
||||
<span class="text-[0.65rem] text-zinc-500 mt-0.5">Removes job records only — downloaded .md files and video statuses (done / error / no_subtitles) are NOT affected. To re-queue stuck videos, use "Retry X stuck" on the Videos tab.</span>
|
||||
</div>
|
||||
<div class="flex gap-1.5">
|
||||
<button class="btn btn-ghost !py-1 !px-2 text-xs" @click="loadJobs()">Refresh</button>
|
||||
<button class="btn btn-danger !py-1 !px-2 text-xs" @click="clearJobHistory()" :disabled="scrape.jobs.length === 0">Clear history</button>
|
||||
<button class="btn btn-danger !py-1 !px-2 text-xs" @click="clearJobHistory()" :disabled="scrape.jobs.length === 0" title="Remove finished job records from this list. Does not change video statuses.">Clear history</button>
|
||||
</div>
|
||||
</div>
|
||||
<table class="tbl">
|
||||
<thead><tr><th>Channel</th><th>Status</th><th>Progress</th><th>Started</th><th>Finished</th><th></th></tr></thead>
|
||||
<tbody>
|
||||
<template x-if="loading.jobs"><tr><td colspan="6" class="text-center text-zinc-500 py-6">loading…</td></tr></template>
|
||||
<template x-if="loading.jobs">
|
||||
<template x-for="i in 4" :key="i">
|
||||
<tr>
|
||||
<td>
|
||||
<template x-if="i===1"><span class="sr-only" role="status">Loading jobs…</span></template>
|
||||
<div class="skeleton h-3.5 w-28" aria-hidden="true"></div>
|
||||
</td>
|
||||
<td><div class="skeleton h-5 w-16 rounded-full" aria-hidden="true"></div></td>
|
||||
<td><div class="skeleton h-3 w-12" aria-hidden="true"></div></td>
|
||||
<td><div class="skeleton h-3 w-24" aria-hidden="true"></div></td>
|
||||
<td><div class="skeleton h-3 w-24" aria-hidden="true"></div></td>
|
||||
<td>
|
||||
<div class="flex justify-end" aria-hidden="true"><div class="skeleton h-5 w-12"></div></div>
|
||||
</td>
|
||||
</tr>
|
||||
</template>
|
||||
</template>
|
||||
<template x-if="!loading.jobs && scrape.jobs.length===0"><tr><td colspan="6" class="text-center text-zinc-500 py-6">No jobs yet.</td></tr></template>
|
||||
<template x-for="j in scrape.jobs" :key="j.id">
|
||||
<tr>
|
||||
@@ -519,7 +783,7 @@
|
||||
<td class="font-mono text-xs text-zinc-400" x-text="(j.completed||0)+' / '+(j.total||0)"></td>
|
||||
<td class="font-mono text-xs text-zinc-500" x-text="j.started_at ? fmtDateTime(j.started_at) : '—'"></td>
|
||||
<td class="font-mono text-xs text-zinc-500" x-text="j.finished_at ? fmtDateTime(j.finished_at) : '—'"></td>
|
||||
<td class="text-right"><button class="btn !py-1 !px-2 text-xs" :class="jobIsTerminal(j) ? 'btn-ghost' : 'btn-danger'" @click="deleteJob(j.id)" x-text="jobActionLabel(j)" :title="jobIsTerminal(j) ? 'Remove this job from history (downloads stay)' : 'Cancel this running job'"></button></td>
|
||||
<td class="text-right"><button class="btn btn-ghost !py-1 !px-2 text-xs" @click="deleteJob(j.id)" x-text="'Clear'" :title="jobIsTerminal(j) ? 'Remove this finished job from history (downloads stay)' : 'Force-remove this row (the job itself is being managed elsewhere or never started)'"></button></td>
|
||||
</tr>
|
||||
</template>
|
||||
</tbody>
|
||||
@@ -533,8 +797,19 @@
|
||||
<section class="space-y-5" x-transition.opacity>
|
||||
<header><h1 class="text-2xl font-bold text-white">Cookie vault</h1><p class="text-sm text-zinc-500">Upload YouTube cookie files to authenticate</p></header>
|
||||
|
||||
<div class="glass px-4 py-3 flex items-center justify-between gap-3">
|
||||
<div>
|
||||
<div class="text-sm text-zinc-200 font-medium">Import from Brave</div>
|
||||
<div class="text-xs text-zinc-500">Pulls your logged-in YouTube session straight from the browser's cookie store — no extensions or manual exports. <span class="text-amber-400/90">Close Brave first</span>: it keeps its cookie database locked while running.</div>
|
||||
</div>
|
||||
<button class="btn !py-1.5 !px-3 text-xs shrink-0" :disabled="loading.cookies" @click="importBrowserCookies()">
|
||||
<span x-show="!loading.cookies">Import from Brave</span>
|
||||
<span x-show="loading.cookies">Importing…</span>
|
||||
</button>
|
||||
</div>
|
||||
|
||||
<div class="dropzone p-8 text-center transition-all" :class="cookies.drag ? 'drag' : ''"
|
||||
@click="$refs.cookieFile.click()"
|
||||
@click="ref('cookieFile') && ref('cookieFile').click()"
|
||||
@dragover.prevent="cookies.drag=true"
|
||||
@dragleave.prevent="cookies.drag=false"
|
||||
@drop.prevent="handleDrop($event)">
|
||||
@@ -553,7 +828,29 @@
|
||||
<table class="tbl">
|
||||
<thead><tr><th>Label</th><th>File</th><th>Expires</th><th>Session</th><th>Active</th><th>Test</th><th></th></tr></thead>
|
||||
<tbody>
|
||||
<template x-if="loading.cookies"><tr><td colspan="7" class="text-center text-zinc-500 py-6">loading…</td></tr></template>
|
||||
<template x-if="loading.cookies">
|
||||
<template x-for="i in 4" :key="i">
|
||||
<tr>
|
||||
<td>
|
||||
<template x-if="i===1"><span class="sr-only" role="status">Loading cookies…</span></template>
|
||||
<div class="skeleton h-4 w-24" aria-hidden="true"></div>
|
||||
</td>
|
||||
<td><div class="skeleton h-3 w-40" aria-hidden="true"></div></td>
|
||||
<td><div class="skeleton h-3 w-28" aria-hidden="true"></div></td>
|
||||
<td><div class="skeleton h-5 w-16 rounded-full" aria-hidden="true"></div></td>
|
||||
<td><div class="skeleton w-4 h-4 rounded-full" aria-hidden="true"></div></td>
|
||||
<td>
|
||||
<div class="flex items-center gap-2" aria-hidden="true">
|
||||
<div class="skeleton h-5 w-10"></div>
|
||||
<div class="skeleton h-3 w-16"></div>
|
||||
</div>
|
||||
</td>
|
||||
<td>
|
||||
<div class="flex justify-end" aria-hidden="true"><div class="skeleton h-5 w-14"></div></div>
|
||||
</td>
|
||||
</tr>
|
||||
</template>
|
||||
</template>
|
||||
<template x-if="!loading.cookies && cookies.items.length===0"><tr><td colspan="7" class="text-center text-zinc-500 py-6">Vault is empty.</td></tr></template>
|
||||
<template x-for="c in cookies.items" :key="c.id">
|
||||
<tr>
|
||||
@@ -647,6 +944,21 @@
|
||||
</button>
|
||||
</div>
|
||||
|
||||
<!-- Channel avatars -->
|
||||
<div class="tool-card glass p-5">
|
||||
<div class="tool-ico"><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><path stroke-linecap="round" stroke-linejoin="round" d="M17.982 18.725A7.488 7.488 0 0012 15.75a7.488 7.488 0 00-5.982 2.975m11.963 0a9 9 0 10-11.963 0m11.963 0A8.966 8.966 0 0112 21a8.966 8.966 0 01-5.982-2.275M15 9.75a3 3 0 11-6 0 3 3 0 016 0z"/></svg></div>
|
||||
<h3 class="text-sm font-semibold text-zinc-100">Download avatars</h3>
|
||||
<p class="text-xs text-zinc-500 mt-1 mb-2">Scrape + cache channel profile images locally.</p>
|
||||
<select class="field !text-xs !py-1.5 mb-2" x-model="tools.avatarScope">
|
||||
<option value="all">All tracked channels</option>
|
||||
<template x-for="c in channels.items" :key="c.channel_id"><option :value="c.channel_id" x-text="c.name"></option></template>
|
||||
</select>
|
||||
<button class="btn btn-ghost w-full" :disabled="tools.busy==='avatars'" @click="downloadAvatarsScope(tools.avatarScope)">
|
||||
<svg x-show="tools.busy==='avatars'" class="spin w-4 h-4" viewBox="0 0 24 24" fill="none"><circle cx="12" cy="12" r="9" stroke="currentColor" stroke-width="3" stroke-dasharray="40 20"/></svg>
|
||||
Download
|
||||
</button>
|
||||
</div>
|
||||
|
||||
<!-- Audio -->
|
||||
<div class="tool-card glass p-5">
|
||||
<div class="tool-ico"><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><path stroke-linecap="round" stroke-linejoin="round" d="M9 18V6l10-2v12M9 18a3 3 0 11-6 0 3 3 0 016 0zm10-2a3 3 0 11-6 0 3 3 0 016 0z"/></svg></div>
|
||||
@@ -787,7 +1099,7 @@
|
||||
|
||||
<!-- ============ FLOATING JOB WIDGET ============ -->
|
||||
<template x-if="scrape.jobId && view !== 'scrape' && !jobWidget.dismissed">
|
||||
<div class="job-widget" x-transition.opacity>
|
||||
<div class="job-widget" x-transition.opacity @mouseenter="_bumpAutoHide()" @focusin="_bumpAutoHide()">
|
||||
<template x-if="!jobWidget.collapsed">
|
||||
<div class="space-y-2.5">
|
||||
<div class="flex items-center justify-between gap-3">
|
||||
@@ -796,7 +1108,7 @@
|
||||
<span class="text-xs font-semibold text-white uppercase tracking-wider" x-text="scrape.kind ? scrape.kind + ' job' : 'job'"></span>
|
||||
<span class="pill !text-[0.6rem] !py-0" :class="scrape.error ? 'st-error' : (scrape.done ? 'st-done' : 'st-running')" x-text="scrape.error ? 'error' : (scrape.done ? 'done' : 'running')"></span>
|
||||
</div>
|
||||
<button class="jw-x" @click="closeJobWidget()" :title="(!scrape.done && !scrape.error) ? 'Minimize' : 'Dismiss'">✕</button>
|
||||
<button class="jw-x" @click="closeJobWidget()" :title="(!scrape.done && !scrape.error) ? 'Minimize' : 'Dismiss'" :aria-label="(!scrape.done && !scrape.error) ? 'Minimize job widget' : 'Dismiss job widget'">✕</button>
|
||||
</div>
|
||||
<div class="flex justify-between text-[0.7rem] text-zinc-500">
|
||||
<span>Progress</span>
|
||||
@@ -832,7 +1144,7 @@
|
||||
<div class="modal-card w-full max-w-2xl" @click.stop>
|
||||
<div class="modal-head">
|
||||
<h2 class="text-base font-semibold text-white">Clip extractor</h2>
|
||||
<button class="jw-x" @click="closeClip()">✕</button>
|
||||
<button class="jw-x" @click="closeClip()" aria-label="Close clip extractor" title="Close (Esc)">✕</button>
|
||||
</div>
|
||||
<div class="p-5 space-y-3">
|
||||
<div>
|
||||
@@ -882,7 +1194,7 @@
|
||||
<div class="modal-card w-full max-w-3xl" @click.stop>
|
||||
<div class="modal-head">
|
||||
<h2 class="text-base font-semibold text-white">Stats</h2>
|
||||
<button class="jw-x" @click="closeStats()">✕</button>
|
||||
<button class="jw-x" @click="closeStats()" aria-label="Close stats panel" title="Close (Esc)">✕</button>
|
||||
</div>
|
||||
<div class="p-5 space-y-5">
|
||||
<template x-if="stats.loading"><div class="text-zinc-500 text-sm text-center py-8">crunching…</div></template>
|
||||
@@ -935,6 +1247,25 @@
|
||||
</div>
|
||||
</template>
|
||||
|
||||
<!-- ============ IN-APP CONFIRM (replaces window.confirm) ============ -->
|
||||
<template x-if="confirmBox.open">
|
||||
<div class="modal-backdrop" @click.self="resolveConfirm(false)" x-transition.opacity>
|
||||
<div class="modal-card w-full max-w-md" @click.stop role="alertdialog" aria-modal="true" :aria-label="confirmBox.title">
|
||||
<div class="modal-head">
|
||||
<h2 class="text-base font-semibold text-white" x-text="confirmBox.title"></h2>
|
||||
<button class="jw-x" @click="resolveConfirm(false)" aria-label="Close dialog" title="Close (Esc)">✕</button>
|
||||
</div>
|
||||
<div class="p-5 space-y-4">
|
||||
<p class="text-sm text-zinc-300 leading-relaxed whitespace-pre-line" x-text="confirmBox.message"></p>
|
||||
<div class="flex gap-2 justify-end">
|
||||
<button class="btn btn-ghost flex-1" @click="resolveConfirm(false)" x-init="$el.focus()" x-text="confirmBox.cancelLabel"></button>
|
||||
<button class="btn flex-1" :class="confirmBox.danger ? 'btn-danger' : 'accent-grad btn-primary'" @click="resolveConfirm(true)" x-text="confirmBox.confirmLabel"></button>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
</template>
|
||||
|
||||
</div>
|
||||
</body>
|
||||
</html>
|
||||
|
||||
@@ -69,6 +69,10 @@ body::before {
|
||||
.nav-item .nav-ico { width: 1.1rem; height: 1.1rem; opacity: 0.9; }
|
||||
|
||||
/* ---- form controls ---- */
|
||||
/* El popup nativo de <select> se pinta con el esquema del UA: sin esto abre
|
||||
blanco con texto claro heredado y queda ilegible sobre la UI oscura. */
|
||||
select { color-scheme: dark; }
|
||||
option { background-color: #18181b; color: #e4e4e7; }
|
||||
.field {
|
||||
background: #09090b;
|
||||
border: 1px solid #27272a;
|
||||
@@ -103,6 +107,22 @@ select.field { appearance: none; background-image: linear-gradient(45deg, transp
|
||||
border: 1px solid transparent;
|
||||
}
|
||||
|
||||
/* ---- active-filter chips (videos) ---- */
|
||||
/* Removable summary of the filters constraining the list: rose tint marks
|
||||
them as the accent-colored "current context", the ✕ is the affordance. */
|
||||
.filter-chip {
|
||||
display: inline-flex; align-items: center; gap: 0.4rem;
|
||||
padding: 0.25rem 0.65rem; border-radius: 999px;
|
||||
background: rgba(244, 63, 94, 0.10);
|
||||
border: 1px solid rgba(244, 63, 94, 0.35);
|
||||
color: #fecdd3; font-size: 0.75rem; font-weight: 600;
|
||||
cursor: pointer; transition: all 0.15s ease;
|
||||
}
|
||||
.filter-chip:hover { background: rgba(244, 63, 94, 0.20); color: #fff; border-color: rgba(244, 63, 94, 0.55); }
|
||||
.filter-chip:active { transform: translateY(1px); }
|
||||
.filter-chip svg { width: 0.75rem; height: 0.75rem; opacity: 0.7; flex-shrink: 0; }
|
||||
.filter-chip:hover svg { opacity: 1; }
|
||||
|
||||
/* ---- table ---- */
|
||||
.tbl { width: 100%; border-collapse: separate; border-spacing: 0; }
|
||||
.tbl th { text-align: left; font-size: 0.7rem; text-transform: uppercase; letter-spacing: 0.05em; color: #71717a; font-weight: 600; padding: 0.6rem 0.75rem; border-bottom: 1px solid #27272a; }
|
||||
@@ -116,6 +136,9 @@ select.field { appearance: none; background-image: linear-gradient(45deg, transp
|
||||
.st-running{ background: rgba(59, 130, 246, 0.12); color: #60a5fa; border-color: rgba(59, 130, 246, 0.35); }
|
||||
.st-skipped{ background: rgba(161, 161, 170, 0.12); color: #a1a1aa; border-color: rgba(161, 161, 170, 0.35); }
|
||||
.st-new { background: rgba(168, 85, 247, 0.12); color: #c084fc; border-color: rgba(168, 85, 247, 0.35); }
|
||||
/* Gated content: a dashed border reads as "not available to you" rather than
|
||||
as one more processing state, so it never gets confused with st-pending. */
|
||||
.st-locked { background: rgba(217, 119, 6, 0.10); color: #fbbf24; border-color: rgba(217, 119, 6, 0.45); border-style: dashed; }
|
||||
|
||||
/* ---- progress bar ---- */
|
||||
.progress-track { background: #18181b; border-radius: 999px; overflow: hidden; height: 0.6rem; border: 1px solid #27272a; }
|
||||
@@ -136,6 +159,20 @@ select.field { appearance: none; background-image: linear-gradient(45deg, transp
|
||||
.spin { animation: spin 0.9s linear infinite; }
|
||||
@keyframes spin { to { transform: rotate(360deg); } }
|
||||
|
||||
/* ---- skeleton loaders ---- */
|
||||
/* One class only: size/shape is varied per element with Tailwind utilities. */
|
||||
.skeleton {
|
||||
background: linear-gradient(90deg, #18181b 25%, #27272a 50%, #18181b 75%);
|
||||
background-size: 200% 100%;
|
||||
animation: skeleton-shimmer 1.4s ease infinite;
|
||||
border-radius: 0.375rem;
|
||||
}
|
||||
@keyframes skeleton-shimmer {
|
||||
0% { background-position: 200% 0; }
|
||||
100% { background-position: -200% 0; }
|
||||
}
|
||||
@media (prefers-reduced-motion: reduce) { .skeleton { animation: none; } }
|
||||
|
||||
/* ---- wordcloud ---- */
|
||||
.cloud-word { display: inline-block; margin: 0.35rem 0.4rem; line-height: 1; cursor: default; transition: opacity 0.15s ease; }
|
||||
.cloud-word:hover { opacity: 0.75; }
|
||||
@@ -151,6 +188,15 @@ select.field { appearance: none; background-image: linear-gradient(45deg, transp
|
||||
/* thumbnail */
|
||||
.thumb { border-radius: 0.5rem; object-fit: cover; background: #18181b; }
|
||||
|
||||
/* inline keyboard hint */
|
||||
.kbd { display: inline-block; padding: 0.05rem 0.4rem; font-family: ui-monospace, "JetBrains Mono", "Cascadia Code", monospace; font-size: 0.7rem; color: #d4d4d8; background: #18181b; border: 1px solid #3f3f46; border-radius: 0.3rem; box-shadow: inset 0 -1px 0 #27272a; }
|
||||
|
||||
/* initial-letter fallback for avatars / thumbnails */
|
||||
.avatar-initial { background: linear-gradient(135deg, #27272a, #3f3f46); color: #d4d4d8; font-weight: 600; display: inline-flex; align-items: center; justify-content: center; text-transform: uppercase; user-select: none; }
|
||||
|
||||
/* channel avatar */
|
||||
.avatar { border-radius: 9999px; object-fit: cover; background: #18181b; border: 1px solid #27272a; }
|
||||
|
||||
/* fade-up transition for views */
|
||||
.fade-enter-active { transition: all 0.22s ease; }
|
||||
.fade-enter-from { opacity: 0; transform: translateY(6px); }
|
||||
@@ -240,6 +286,26 @@ select.field { appearance: none; background-image: linear-gradient(45deg, transp
|
||||
.audio-bar { filter: invert(0.92) hue-rotate(170deg) sepia(0.15); height: 36px; }
|
||||
.audio-bar::-webkit-media-controls-panel { background: rgba(24,24,27,0.85); }
|
||||
|
||||
/* ===== local cinema player ===== */
|
||||
.cinema-card {
|
||||
background: #030305; border: 1px solid #27272a; border-radius: 1rem;
|
||||
overflow: hidden; box-shadow: 0 24px 80px rgba(0,0,0,0.45);
|
||||
}
|
||||
.cinema-stage {
|
||||
position: relative; aspect-ratio: 16 / 9; background: #000;
|
||||
display: flex; align-items: center; justify-content: center;
|
||||
}
|
||||
.cinema-video { width: 100%; height: 100%; object-fit: contain; background: #000; }
|
||||
.cinema-error {
|
||||
position: absolute; inset: 0; display: flex; flex-direction: column;
|
||||
align-items: center; justify-content: center; text-align: center; padding: 1.5rem;
|
||||
background: rgba(0,0,0,0.78);
|
||||
}
|
||||
.cinema-meta {
|
||||
display: flex; align-items: center; justify-content: space-between; gap: 1rem;
|
||||
padding: 0.85rem 1rem; background: rgba(24,24,27,0.8);
|
||||
}
|
||||
|
||||
/* the transcript segment currently matching audio playback */
|
||||
.seg-row { border-left: 2px solid transparent; }
|
||||
.seg-active {
|
||||
|
||||
+12
-2
@@ -1,4 +1,14 @@
|
||||
@echo off
|
||||
chcp 65001 >nul
|
||||
powershell -NoProfile -ExecutionPolicy Bypass -File "%~dp0scripts\start-server.ps1" %*
|
||||
exit /b %errorlevel%
|
||||
powershell -NoProfile -ExecutionPolicy Bypass -File "%~dp0scripts\doctor.ps1" %*
|
||||
set EC=%errorlevel%
|
||||
if not "%EC%"=="0" (
|
||||
echo.
|
||||
echo [!] Algo fallo. Esta ventana queda abierta para que puedas leerlo.
|
||||
echo Si llamas a alguien por ayuda, enviale una captura de pantalla.
|
||||
echo.
|
||||
pause
|
||||
) else (
|
||||
timeout /t 5 /nobreak >nul
|
||||
)
|
||||
exit /b %EC%
|
||||
|
||||
+8
-1
@@ -1,4 +1,11 @@
|
||||
@echo off
|
||||
chcp 65001 >nul
|
||||
powershell -NoProfile -ExecutionPolicy Bypass -File "%~dp0scripts\stop-server.ps1" %*
|
||||
exit /b %errorlevel%
|
||||
set EC=%errorlevel%
|
||||
if not "%EC%"=="0" (
|
||||
echo.
|
||||
echo [!] Algo fallo al apagar. Esta ventana queda abierta.
|
||||
echo.
|
||||
pause
|
||||
)
|
||||
exit /b %EC%
|
||||
|
||||
@@ -0,0 +1,93 @@
|
||||
"""Fechas aproximadas de discovery: bandera upload_date_approx.
|
||||
|
||||
Matriz de precedencia de upsert_videos + upgrade approx->real de
|
||||
set_upload_date + conversion timestamp -> YYYYMMDD en discover.
|
||||
Spec: docs/superpowers/specs/2026-08-22-approximate-upload-dates-design.md
|
||||
"""
|
||||
|
||||
from yt_scraper.store import Store, VideoRef
|
||||
|
||||
|
||||
def _ref(vid, date=None, approx=0, ch="ch1"):
|
||||
return VideoRef(
|
||||
video_id=vid,
|
||||
channel_id=ch,
|
||||
title="t " + vid,
|
||||
url=f"https://youtu.be/{vid}",
|
||||
upload_date=date,
|
||||
date_approx=approx,
|
||||
)
|
||||
|
||||
|
||||
def _store(tmp_path):
|
||||
s = Store(tmp_path / "test.db")
|
||||
s.upsert_channel("ch1", None, "Canal de prueba")
|
||||
return s
|
||||
|
||||
|
||||
def test_approx_into_empty(tmp_path):
|
||||
s = _store(tmp_path)
|
||||
s.upsert_videos([_ref("v1", date="20260820", approx=1)])
|
||||
row = s.get_video("v1")
|
||||
assert row.upload_date == "20260820"
|
||||
assert row.upload_date_approx == 1
|
||||
|
||||
|
||||
def test_real_not_downgraded_by_approx(tmp_path):
|
||||
s = _store(tmp_path)
|
||||
s.upsert_videos([_ref("v1", date="20240101", approx=0)])
|
||||
s.upsert_videos([_ref("v1", date="20260820", approx=1)])
|
||||
row = s.get_video("v1")
|
||||
assert row.upload_date == "20240101"
|
||||
assert row.upload_date_approx == 0
|
||||
|
||||
|
||||
def test_approx_refreshed_by_newer_approx(tmp_path):
|
||||
s = _store(tmp_path)
|
||||
s.upsert_videos([_ref("v1", date="20260101", approx=1)])
|
||||
s.upsert_videos([_ref("v1", date="20260820", approx=1)])
|
||||
row = s.get_video("v1")
|
||||
assert row.upload_date == "20260820"
|
||||
assert row.upload_date_approx == 1
|
||||
|
||||
|
||||
def test_extraction_upgrades_approx_to_real(tmp_path):
|
||||
s = _store(tmp_path)
|
||||
s.upsert_videos([_ref("v1", date="20260820", approx=1)])
|
||||
s.set_upload_date("v1", "20260818")
|
||||
row = s.get_video("v1")
|
||||
assert row.upload_date == "20260818"
|
||||
assert row.upload_date_approx == 0
|
||||
|
||||
|
||||
def test_set_upload_date_still_skips_existing_real(tmp_path):
|
||||
s = _store(tmp_path)
|
||||
s.upsert_videos([_ref("v1", date="20240101", approx=0)])
|
||||
s.set_upload_date("v1", "20260818")
|
||||
assert s.get_video("v1").upload_date == "20240101"
|
||||
|
||||
|
||||
def test_migration_defaults_flag_to_zero(tmp_path):
|
||||
# Filas creadas por la via vieja (INSERT sin la columna) quedan con 0.
|
||||
s = _store(tmp_path)
|
||||
s.upsert_videos([_ref("v2", date=None, approx=0)])
|
||||
row = s.get_video("v2")
|
||||
assert not row.upload_date_approx
|
||||
|
||||
|
||||
# --- conversion en discover -------------------------------------------------
|
||||
|
||||
from yt_scraper.discover import _entry_upload_date
|
||||
|
||||
|
||||
def test_entry_upload_date_prefers_exact():
|
||||
assert _entry_upload_date({"upload_date": "20240101"}) == ("20240101", 0)
|
||||
|
||||
|
||||
def test_entry_upload_date_from_timestamp():
|
||||
ts = 1716115200 # 2024-05-19 12:00 UTC
|
||||
assert _entry_upload_date({"timestamp": ts}) == ("20240519", 1)
|
||||
|
||||
|
||||
def test_entry_upload_date_none():
|
||||
assert _entry_upload_date({}) == (None, 0)
|
||||
@@ -0,0 +1,143 @@
|
||||
"""Members-only / gated videos: identified from discovery, labelled, and kept
|
||||
out of bulk work without becoming unreachable.
|
||||
|
||||
yt-dlp's flat listing reports availability="subscriber_only" for members-only
|
||||
videos, so they are knowable before an extraction attempt is ever spent on them.
|
||||
Measured on a real channel: 120 flat entries in one request, exactly 4 flagged,
|
||||
matching exactly the 4 the DB had learned about the expensive way.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import pytest
|
||||
|
||||
from yt_scraper.store import BLOCKING_AVAILABILITY, Store, VideoRef
|
||||
|
||||
|
||||
def _store(tmp_path) -> Store:
|
||||
store = Store(tmp_path / "state.db")
|
||||
store.upsert_channel("UC1", "@alpha", "Alpha", 0)
|
||||
return store
|
||||
|
||||
|
||||
def _ref(vid: str, availability: str | None = None) -> VideoRef:
|
||||
return VideoRef(vid, "UC1", vid, f"https://www.youtube.com/watch?v={vid}", availability=availability)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- classification
|
||||
|
||||
def test_availability_from_discovery_is_stored_and_classified(tmp_path):
|
||||
store = _store(tmp_path)
|
||||
store.upsert_videos([_ref("gated", "subscriber_only"), _ref("open", "public")])
|
||||
|
||||
assert store.get_video("gated").availability == "subscriber_only"
|
||||
assert store.get_video("gated").block_reason == "members_only"
|
||||
assert store.get_video("open").block_reason is None
|
||||
|
||||
|
||||
@pytest.mark.parametrize("availability,expected", [
|
||||
("subscriber_only", "members_only"),
|
||||
("premium_only", "premium_only"),
|
||||
("private", "private"),
|
||||
("needs_auth", "needs_auth"),
|
||||
])
|
||||
def test_every_blocking_availability_maps_to_a_reason(tmp_path, availability, expected):
|
||||
store = _store(tmp_path)
|
||||
store.upsert_videos([_ref("v", availability)])
|
||||
assert store.get_video("v").block_reason == expected
|
||||
|
||||
|
||||
@pytest.mark.parametrize("availability", ["public", "unlisted", None])
|
||||
def test_fetchable_availability_is_never_blocked(tmp_path, availability):
|
||||
"""`unlisted` downloads perfectly well — mislabelling it would hide videos
|
||||
the user can actually have."""
|
||||
store = _store(tmp_path)
|
||||
store.upsert_videos([_ref("v", availability)])
|
||||
assert store.get_video("v").block_reason is None
|
||||
assert "unlisted" not in BLOCKING_AVAILABILITY
|
||||
|
||||
|
||||
def test_block_reason_falls_back_to_the_recorded_error(tmp_path):
|
||||
"""Rows burned into `error` before availability was captured must still be
|
||||
identifiable without re-fetching them."""
|
||||
store = _store(tmp_path)
|
||||
store.upsert_videos([_ref("old")])
|
||||
store.mark_error("old", "ERROR: [youtube] x: Join this channel to get access to members-only content")
|
||||
|
||||
assert store.get_video("old").availability is None
|
||||
assert store.get_video("old").block_reason == "members_only"
|
||||
|
||||
|
||||
def test_rate_limit_error_is_not_a_block_reason(tmp_path):
|
||||
store = _store(tmp_path)
|
||||
store.upsert_videos([_ref("t")])
|
||||
store.mark_error("t", "ERROR: Video unavailable. The current session has been rate-limited by YouTube")
|
||||
|
||||
assert store.get_video("t").block_reason is None
|
||||
|
||||
|
||||
def test_rediscovery_does_not_wipe_a_known_availability(tmp_path):
|
||||
store = _store(tmp_path)
|
||||
store.upsert_videos([_ref("v", "subscriber_only")])
|
||||
store.upsert_videos([_ref("v", None)]) # a later flat pass omitted the field
|
||||
|
||||
assert store.get_video("v").availability == "subscriber_only"
|
||||
|
||||
|
||||
def test_set_availability_records_what_extraction_learned(tmp_path):
|
||||
store = _store(tmp_path)
|
||||
store.upsert_videos([_ref("v")])
|
||||
store.set_availability("v", "subscriber_only")
|
||||
assert store.get_video("v").block_reason == "members_only"
|
||||
|
||||
store.set_availability("v", None) # must not clear it
|
||||
assert store.get_video("v").availability == "subscriber_only"
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- queue behaviour
|
||||
|
||||
def test_blocked_videos_are_kept_out_of_the_pending_queue(tmp_path):
|
||||
"""Bulk runs must not spend requests on videos that cannot be fetched."""
|
||||
store = _store(tmp_path)
|
||||
store.upsert_videos([_ref("gated", "subscriber_only"), _ref("ok"), _ref("unlisted", "unlisted")])
|
||||
|
||||
ids = {v.video_id for v in store.get_pending("UC1")}
|
||||
assert ids == {"ok", "unlisted"}
|
||||
|
||||
all_ids = {v.video_id for v in store.get_pending("UC1", include_blocked=True)}
|
||||
assert all_ids == {"ok", "unlisted", "gated"}
|
||||
|
||||
|
||||
def test_a_blocked_video_is_still_reachable_by_id(tmp_path):
|
||||
"""Buying the membership must not leave the video permanently stranded."""
|
||||
store = _store(tmp_path)
|
||||
store.upsert_videos([_ref("gated", "subscriber_only")])
|
||||
|
||||
assert store.get_video("gated") is not None, "explicit per-video processing still works"
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- filtering
|
||||
|
||||
def test_blocked_filter_finds_them_across_statuses(tmp_path):
|
||||
store = _store(tmp_path)
|
||||
store.upsert_videos([
|
||||
_ref("gated", "subscriber_only"),
|
||||
_ref("legacy"),
|
||||
_ref("fine"),
|
||||
])
|
||||
store.mark_error("legacy", "ERROR: Join this channel to get access to members-only content")
|
||||
store.mark_status("fine", "no_subtitles")
|
||||
|
||||
rows, total = store.query_videos(blocked=True)
|
||||
assert {r.video_id for r in rows} == {"gated", "legacy"}
|
||||
assert total == 2
|
||||
|
||||
rows, total = store.query_videos(blocked=False)
|
||||
assert {r.video_id for r in rows} == {"fine"}
|
||||
|
||||
|
||||
def test_blocked_filter_absent_means_everything(tmp_path):
|
||||
store = _store(tmp_path)
|
||||
store.upsert_videos([_ref("gated", "subscriber_only"), _ref("fine")])
|
||||
|
||||
_rows, total = store.query_videos()
|
||||
assert total == 2
|
||||
@@ -0,0 +1,75 @@
|
||||
"""Tests for the YAML config loader.
|
||||
|
||||
Focuses on the ``languages`` and ``prefer_manual`` fields, including the
|
||||
legacy/compat behaviour. Network-free, DB-free.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import textwrap
|
||||
|
||||
import pytest
|
||||
|
||||
from yt_scraper.config import load_config, parse_languages
|
||||
|
||||
|
||||
def _write_config(tmp_path, body: str) -> str:
|
||||
p = tmp_path / "config.yaml"
|
||||
p.write_text(textwrap.dedent(body), encoding="utf-8")
|
||||
return str(p)
|
||||
|
||||
|
||||
def test_load_legacy_languages_list(tmp_path) -> None:
|
||||
p = _write_config(tmp_path, """
|
||||
channel_url: "https://example.com/@x/videos"
|
||||
languages: ["es", "en"]
|
||||
prefer_manual: true
|
||||
""")
|
||||
cfg = load_config(p)
|
||||
assert cfg.languages == {"es": "manual", "en": "manual"}
|
||||
|
||||
|
||||
def test_load_legacy_languages_list_with_prefer_manual_false(tmp_path) -> None:
|
||||
p = _write_config(tmp_path, """
|
||||
languages: ["es", "en"]
|
||||
prefer_manual: false
|
||||
""")
|
||||
cfg = load_config(p)
|
||||
assert cfg.languages == {"es": "auto", "en": "auto"}
|
||||
|
||||
|
||||
def test_load_new_dict_languages(tmp_path) -> None:
|
||||
p = _write_config(tmp_path, """
|
||||
languages:
|
||||
en: manual
|
||||
es: auto
|
||||
pt: any
|
||||
prefer_manual: false
|
||||
""")
|
||||
cfg = load_config(p)
|
||||
assert cfg.languages == {"en": "manual", "es": "auto", "pt": "any"}
|
||||
# prefer_manual remains accessible for fallback on `"any"` entries
|
||||
assert cfg.prefer_manual is False
|
||||
|
||||
|
||||
def test_load_unknown_mode_normalises_to_any(tmp_path) -> None:
|
||||
p = _write_config(tmp_path, """
|
||||
languages:
|
||||
en: garbage
|
||||
""")
|
||||
cfg = load_config(p)
|
||||
assert cfg.languages == {"en": "any"}
|
||||
|
||||
|
||||
def test_load_missing_languages_defaults_to_any_not_manual_only(tmp_path) -> None:
|
||||
"""The default must fall back to auto captions.
|
||||
|
||||
A manual-only default silently produces zero transcripts on the many
|
||||
channels that publish only auto-generated captions, and records them as
|
||||
`no_subtitles` — a terminal status that hides a purely configural failure.
|
||||
"""
|
||||
p = _write_config(tmp_path, """
|
||||
channel_url: "https://example.com/@x/videos"
|
||||
""")
|
||||
cfg = load_config(p)
|
||||
assert "es" in cfg.languages and "en" in cfg.languages
|
||||
assert cfg.languages["es"] == "any"
|
||||
@@ -0,0 +1,152 @@
|
||||
from __future__ import annotations
|
||||
|
||||
from pathlib import Path
|
||||
import pytest
|
||||
|
||||
from yt_scraper.cookies import (
|
||||
BrowserCookieLockedError,
|
||||
auto_import_dir,
|
||||
delete,
|
||||
import_from_browser,
|
||||
is_expired,
|
||||
parse_netscape,
|
||||
resolve_active_path,
|
||||
)
|
||||
from yt_scraper.store import CookieRow, Store
|
||||
|
||||
|
||||
SAMPLE_NETSCAPE = """# Netscape HTTP Cookie File
|
||||
.youtube.com\tTRUE\t/\tTRUE\t2000000000\tSID\tsample_sid_token
|
||||
.youtube.com\tTRUE\t/\tTRUE\t2000000000\tLOGIN_INFO\tsample_login_info
|
||||
"""
|
||||
|
||||
|
||||
def test_parse_netscape():
|
||||
ok, info = parse_netscape(SAMPLE_NETSCAPE)
|
||||
assert ok is True
|
||||
assert info["count"] == 2
|
||||
assert info["has_session"] is True
|
||||
assert info["expires_at"] is not None
|
||||
|
||||
|
||||
def test_parse_netscape_httponly_lines_are_data():
|
||||
# SID/HSID are HttpOnly; Netscape exports prefix those lines and a
|
||||
# comment-skipping parser would silently strip the login out.
|
||||
text = (
|
||||
"# Netscape HTTP Cookie File\n"
|
||||
"#HttpOnly_.youtube.com\tTRUE\t/\tTRUE\t2000000000\tSID\ttok\n"
|
||||
"#HttpOnly_.youtube.com\tTRUE\t/\tTRUE\t2000000000\tHSID\ttok\n"
|
||||
"#HttpOnly_.youtube.com\tTRUE\t/\tTRUE\t2000000000\tSSID\ttok\n"
|
||||
"# a real comment\n"
|
||||
)
|
||||
ok, info = parse_netscape(text)
|
||||
assert ok is True
|
||||
assert info["count"] == 3
|
||||
assert info["has_session"] is True
|
||||
|
||||
|
||||
def test_partial_session_export_is_not_a_session():
|
||||
# A lone __Secure-3PSID (partial extension export) is anonymous to
|
||||
# YouTube; it must not be reported as a usable session.
|
||||
text = (
|
||||
"# Netscape HTTP Cookie File\n"
|
||||
".youtube.com\tTRUE\t/\tTRUE\t2000000000\t__Secure-3PSID\ttok\n"
|
||||
".youtube.com\tTRUE\t/\tTRUE\t2000000000\t__Secure-3PAPISID\ttok\n"
|
||||
)
|
||||
ok, info = parse_netscape(text)
|
||||
assert ok is True
|
||||
assert info["has_session"] is False
|
||||
|
||||
|
||||
def _browser_cookie(name, value, *, domain=".youtube.com", httponly=False):
|
||||
import http.cookiejar
|
||||
return http.cookiejar.Cookie(
|
||||
version=0, name=name, value=value, port=None, port_specified=False,
|
||||
domain=domain, domain_specified=True, domain_initial_dot=domain.startswith("."),
|
||||
path="/", path_specified=True, secure=True, expires=2000000000,
|
||||
discard=False, comment=None, comment_url=None,
|
||||
rest={"HttpOnly": None} if httponly else {},
|
||||
)
|
||||
|
||||
|
||||
def test_import_from_browser_writes_vault_file(tmp_path: Path, monkeypatch):
|
||||
store = Store(tmp_path / "state.db")
|
||||
cookie_dir = tmp_path / "cookies"
|
||||
|
||||
def fake_extract(browser, profile=None, logger=None, **kw):
|
||||
return iter([
|
||||
_browser_cookie("SID", "sid_tok", httponly=True),
|
||||
_browser_cookie("HSID", "hsid_tok", httponly=True),
|
||||
_browser_cookie("SSID", "ssid_tok", httponly=True),
|
||||
_browser_cookie("__Secure-3PSID", "psid_tok"),
|
||||
_browser_cookie("PREF", "pref_tok", domain=".google.com"), # not youtube -> dropped
|
||||
])
|
||||
|
||||
import yt_dlp.cookies as ydl_cookies
|
||||
monkeypatch.setattr(ydl_cookies, "extract_cookies_from_browser", fake_extract)
|
||||
|
||||
cid = import_from_browser(store, browser="brave", cookie_dir=cookie_dir)
|
||||
row = store.get_cookie(cid)
|
||||
assert row is not None
|
||||
assert row.cookie_count == 4
|
||||
assert row.has_session is True
|
||||
|
||||
# The written file must round-trip as a valid Netscape file whose
|
||||
# HttpOnly session cookies survive the vault's own parser.
|
||||
from yt_scraper.cookies import parse_netscape_file
|
||||
path = cookie_dir / row.filename
|
||||
ok, info = parse_netscape_file(path)
|
||||
assert ok is True
|
||||
assert info["count"] == 4
|
||||
assert info["has_session"] is True
|
||||
|
||||
|
||||
def test_import_from_browser_maps_locked_db(tmp_path: Path, monkeypatch):
|
||||
store = Store(tmp_path / "state.db")
|
||||
|
||||
def fake_extract(browser, profile=None, logger=None, **kw):
|
||||
raise RuntimeError("Could not copy Chrome cookie database. See https://github.com/yt-dlp/yt-dlp/issues/7271 for more info")
|
||||
|
||||
import yt_dlp.cookies as ydl_cookies
|
||||
monkeypatch.setattr(ydl_cookies, "extract_cookies_from_browser", fake_extract)
|
||||
|
||||
with pytest.raises(BrowserCookieLockedError) as exc:
|
||||
import_from_browser(store, browser="brave", cookie_dir=tmp_path / "cookies")
|
||||
assert "close brave" in str(exc.value).lower()
|
||||
|
||||
|
||||
def test_auto_import_and_prune_dead_cookies(tmp_path: Path):
|
||||
db_path = tmp_path / "state.db"
|
||||
store = Store(db_path)
|
||||
cookie_dir = tmp_path / "cookies"
|
||||
cookie_dir.mkdir()
|
||||
|
||||
# Place two cookie files
|
||||
f1 = cookie_dir / "c1.txt"
|
||||
f1.write_text(SAMPLE_NETSCAPE, encoding="utf-8")
|
||||
f2 = cookie_dir / "c2.txt"
|
||||
f2.write_text(SAMPLE_NETSCAPE, encoding="utf-8")
|
||||
|
||||
# auto_import imports both and activates the first
|
||||
n = auto_import_dir(store, dir_path=cookie_dir)
|
||||
assert n == 2
|
||||
assert len(store.list_cookies()) == 2
|
||||
active = store.get_active_cookie()
|
||||
assert active is not None
|
||||
assert active.filename in ("c1.txt", "c2.txt")
|
||||
|
||||
active_path = resolve_active_path(store, cookie_dir=cookie_dir)
|
||||
assert active_path is not None
|
||||
assert Path(active_path).exists()
|
||||
|
||||
# Now simulate user deleting the active cookie file from disk
|
||||
Path(active_path).unlink()
|
||||
|
||||
# resolve_active_path should fall back to the surviving cookie file
|
||||
fallback_path = resolve_active_path(store, cookie_dir=cookie_dir)
|
||||
assert fallback_path is not None
|
||||
assert Path(fallback_path).exists()
|
||||
|
||||
# auto_import_dir should clean up the deleted cookie row from DB
|
||||
auto_import_dir(store, dir_path=cookie_dir)
|
||||
assert len(store.list_cookies()) == 1
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user