feat: local content-mining platform for YouTube creator scraping
Scraper de canales de YouTube hacia notas Markdown para base de conocimiento (Obsidian-ready), con plataforma web local. Engine + CLI (Workstream A): - Modular pipeline: discover/extract/parse/chapters/render/store + ratelimit - SQLite store con migración idempotente: FTS5 (transcript search), columnas de metadata enriquecida, tablas cookies_meta y scrape_jobs - Módulos: segments, cookies (Netscape vault), export (json/csv/srt/html), analysis (word freq/timeline/wordcloud), monitor (watch loop), pipeline - CLI Click group: search, export, audio, channels, watch, analyze, re-render - Fix del bug de scoping de cookies en cli.py Webapp local (Workstream B): - FastAPI backend: dashboard, channels, videos facetado, transcript, search FTS, analysis, scrape jobs con SSE, cookies drag-and-drop, exports, folders (abrir en OS), tools (re-render, formato) - SPA no-build (Alpine.js + Tailwind + Chart.js por CDN): 9 vistas, tema dark command-center con acento rojo→rosa, cookie vault drag-drop, consola de scrapeo con progreso live vía SSE Launcher + subagentes (Workstream C): - start-server.bat / stop-server.bat con auto port-scan + browser open - .opencode/agent/webapp-builder.md + .opencode/goals/webapp-build.md Tests: 38 pytest verdes. Sin funcionalidad de IA (enfoque data-mining).
This commit is contained in:
+26
@@ -0,0 +1,26 @@
|
|||||||
|
__pycache__/
|
||||||
|
*.py[cod]
|
||||||
|
*.egg-info/
|
||||||
|
build/
|
||||||
|
dist/
|
||||||
|
.eggs/
|
||||||
|
.pytest_cache/
|
||||||
|
.coverage
|
||||||
|
htmlcov/
|
||||||
|
data/state.db
|
||||||
|
data/state.db-journal
|
||||||
|
data/state.db-wal
|
||||||
|
data/markdown/
|
||||||
|
data/exports/
|
||||||
|
data/analysis/
|
||||||
|
data/audio/
|
||||||
|
cookies/
|
||||||
|
.run/
|
||||||
|
.venv/
|
||||||
|
venv/
|
||||||
|
*.log
|
||||||
|
|
||||||
|
# opencode runtime state (keep agent/ + goal definitions, exclude loop/session state)
|
||||||
|
.opencode/goals/state.json
|
||||||
|
.opencode/goals/state.json.ledger.jsonl
|
||||||
|
.opencode/opencode-loop/
|
||||||
@@ -0,0 +1,44 @@
|
|||||||
|
---
|
||||||
|
description: Builds and maintains the yt-channel-scraper local webapp (FastAPI backend + Alpine.js SPA). Use when implementing, fixing, or extending the webapp at src/yt_scraper/webapp/.
|
||||||
|
model: zai-coding-plan/glm-5.2
|
||||||
|
---
|
||||||
|
|
||||||
|
# webapp-builder
|
||||||
|
|
||||||
|
You are the specialist agent for the **yt-channel-scraper local webapp** (Workstream B of the platform).
|
||||||
|
|
||||||
|
## Mission
|
||||||
|
Implement and maintain the local web application that turns the scraper into a "data-mining command center" for content-creator data. You build on top of the already-working engine modules — you do **not** re-implement them.
|
||||||
|
|
||||||
|
## Stack (fixed — do not deviate)
|
||||||
|
- **Backend:** FastAPI. App factory at `src/yt_scraper/webapp/app.py`. Routers under `/api`. SSE via `sse_starlette`.
|
||||||
|
- **Frontend:** Single-page app, **no build step**. `index.html` + `app.js` + `styles.css` served as static files by the same FastAPI process. Alpine.js + Tailwind CSS (both via CDN, no npm). Chart.js (CDN) for visualizations.
|
||||||
|
- **No external services, no auth, no network egress except YouTube via yt-dlp.** This is a 100% local platform.
|
||||||
|
|
||||||
|
## What already exists — REUSE, do not duplicate
|
||||||
|
Read these before writing code. They are your data layer:
|
||||||
|
- `src/yt_scraper/store.py` — `Store(db_path)` with all SQL: `dashboard()`, `query_videos(...)`, `get_video(id)`, `get_segments(id)`, `search_segments(q, channel_id, limit)`, `list_channels()`, `list_cookies()`, `set_active_cookie(id)`, `get_active_cookie()`, `create_job/update_job/get_job/list_jobs`, `stats()`.
|
||||||
|
- `src/yt_scraper/cookies.py` — `import_text(store, text, label)`, `import_file`, `list_vault`, `set_active`, `delete`, `resolve_active_path`, `parse_netscape`, `auto_import_dir`, `is_expired`.
|
||||||
|
- `src/yt_scraper/segments.py` — `search(store, q, channel_id, limit)`, `backfill_from_markdown`, `parse_markdown`.
|
||||||
|
- `src/yt_scraper/analysis.py` — `word_frequency`, `top_words`, `term_timeline`, `render_wordcloud`, `render_top_words_chart`, `render_timeline_chart`.
|
||||||
|
- `src/yt_scraper/export.py` — `export_json/csv/srt/srt_video/html`.
|
||||||
|
- `src/yt_scraper/pipeline.py` — `process_video(row, cfg, store, channel_name, channel_id, channel_url, cookies_file=, cookies_from_browser=, on_log=)` for the scrape job runner.
|
||||||
|
- `src/yt_scraper/config.py` — `Config`, `load_config(path)`.
|
||||||
|
|
||||||
|
The active cookie (`cookies.resolve_active_path(store)`) MUST be passed as `cookies_file` to every scrape job.
|
||||||
|
|
||||||
|
## Design language (non-negotiable visual identity)
|
||||||
|
Dark "data command center": near-black canvas (`zinc-950` / `#0a0a0f`), glassy cards (`backdrop-blur`, `border-zinc-800`), **red→rose accent gradient** (`from-rose-500 to-red-600`, e.g. `#f43f5e`), monospace accents for data/numbers, smooth Alpine `x-transition`. Responsive. Charts via Chart.js with theme-matched colors (grid `#27272a`, text `#a1a1aa`).
|
||||||
|
|
||||||
|
## Required SPA views (hash routes)
|
||||||
|
Dashboard, Channels, Videos (faceted), Video detail (transcript reader + chapters), Search (FTS + snippet + deep-link), Analysis (wordcloud/topwords/timeline), Scrape console (launch job + SSE progress + cancel + watch toggle), Cookie vault (drag-and-drop + activate + test), Exports.
|
||||||
|
|
||||||
|
## Verification you MUST run before declaring done
|
||||||
|
1. `python -c "from yt_scraper.webapp.app import app; print('app ok')"`
|
||||||
|
2. Start server: `python -m uvicorn yt_scraper.webapp.app:app --port 8765` and `curl http://localhost:8765/healthz` → expect `{"status":"ok"}`.
|
||||||
|
3. Smoke each router: `/api/dashboard`, `/api/channels`, `/api/videos?size=5`, `/api/search?q=vaporwave`, `/api/cookies`.
|
||||||
|
4. Load `/` → SPA renders, navigation works.
|
||||||
|
5. `python -m pytest tests/ -q` stays green.
|
||||||
|
6. Confirm the real cookie at `cookies/253506e4-...txt` is auto-imported and shown as active with `has_session=true`.
|
||||||
|
|
||||||
|
Reference the full spec at `docs/superpowers/specs/2026-07-26-platform-design.md` and plan at `docs/superpowers/plans/2026-07-26-platform-build.md`.
|
||||||
@@ -0,0 +1,31 @@
|
|||||||
|
# Sub-goal: Build the local webapp (Workstream B)
|
||||||
|
|
||||||
|
**Parent goal:** Convert yt-channel-scraper into a local content-mining platform.
|
||||||
|
**Owner agent:** `webapp-builder` (`.opencode/agent/webapp-builder.md`)
|
||||||
|
**Spec:** `docs/superpowers/specs/2026-07-26-platform-design.md` (§7 Webapp, §8 Cookies, §7.2 design language)
|
||||||
|
**Status:** Workstream A (engine + CLI) is DONE and verified — `pytest` green, real DB backfilled, FTS search working, cookie vault auto-imports the real cookie.
|
||||||
|
|
||||||
|
## Objective
|
||||||
|
Implement the FastAPI backend + Alpine.js SPA so the platform is usable from a browser, launched by `start-server.bat`. No Node/npm build step.
|
||||||
|
|
||||||
|
## Deliverables
|
||||||
|
1. `src/yt_scraper/webapp/__init__.py`
|
||||||
|
2. `src/yt_scraper/webapp/app.py` — FastAPI factory, `/healthz`, static mount, auto-import cookies on startup.
|
||||||
|
3. `src/yt_scraper/webapp/jobs.py` — single-worker scrape job runner with SSE event bus (queue/progress/log/cancel), persists to `scrape_jobs`.
|
||||||
|
4. `src/yt_scraper/webapp/api.py` — REST routers documented in spec §7.1 (dashboard, channels, videos, search, analysis, scrape+SSE, cookies with drag-and-drop upload, exports).
|
||||||
|
5. `src/yt_scraper/webapp/static/index.html`, `app.js`, `styles.css` — the 9-view SPA, dark command-center theme, drag-and-drop cookie vault, Chart.js viz.
|
||||||
|
|
||||||
|
## API contract (the SPA consumes exactly this)
|
||||||
|
- `GET /healthz` → `{"status":"ok"}`
|
||||||
|
- `GET /api/dashboard` → `{channels[], status_breakdown{}, uploads_over_time[], duration_histogram[], top_tags[]}`
|
||||||
|
- `GET /api/channels` | `POST /api/channels {url}` | `DELETE /api/channels/{id}`
|
||||||
|
- `GET /api/videos?channel=&status=&from=&to=&min_dur=&q=&sort=&page=&size=` → `{items[], total, page, size}`
|
||||||
|
- `GET /api/videos/{id}` | `GET /api/videos/{id}/transcript` → `[{idx,start_sec,end_sec,text}]`
|
||||||
|
- `GET /api/search?q=&channel=&limit=` → `[{video_id,title,start_sec,snippet}]`
|
||||||
|
- `GET /api/analysis?channel=&type=topwords|wordcloud|timeline&term=` → chart-ready JSON
|
||||||
|
- `POST /api/scrape {channel_id,opts}` → `{job_id}` | `GET /api/scrape/{id}/stream` (SSE) | `DELETE /api/scrape/{id}`
|
||||||
|
- `POST /api/cookies/upload` (multipart file(s) or `{text,label}` JSON) | `GET /api/cookies` | `POST /api/cookies/{id}/activate` | `DELETE /api/cookies/{id}` | `POST /api/cookies/{id}/test`
|
||||||
|
- `GET /api/export?format=json|csv|srt|html&channel=` → file stream
|
||||||
|
|
||||||
|
## Done when
|
||||||
|
- `/healthz` returns 200; every router smokes green against the real DB; SPA renders all 9 views; SSE streams a scrape job; the real cookie drag-and-drops and activates; `pytest` green.
|
||||||
@@ -0,0 +1,298 @@
|
|||||||
|
# AUDIT · Oportunidades de extensión
|
||||||
|
|
||||||
|
> Análisis de features, comandos y opciones que se pueden construir sobre la infraestructura actual.
|
||||||
|
> Cada item incluye: valor, esfuerzo estimado, dependencias y comando propuesto.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Infraestructura disponible (lo que ya tenemos)
|
||||||
|
|
||||||
|
| Recurso | Estado | Reutilizable para |
|
||||||
|
|---|---|---|
|
||||||
|
| `yt-dlp` instalado | ✅ v2026.7.4 | Audio/video download, thumbnails, metadata enriquecida |
|
||||||
|
| SQLite 3.49 con FTS5 | ✅ verificado | Búsqueda full-text sobre transcripciones |
|
||||||
|
| SQLite JSON1 | ✅ verificado | Queries estructuradas sobre metadatos |
|
||||||
|
| 33 transcripciones parseadas | ✅ en `data/state.db` | Análisis, estadísticas, búsqueda |
|
||||||
|
| Markdown con frontmatter YAML | ✅ 842 KB en `data/markdown/` | Obsidian, Pandoc, static site generators |
|
||||||
|
| Pipeline modular | ✅ 8 módulos | Insertar nuevos pasos sin romper nada |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## TIER 1 · Alto valor, bajo esfuerzo (1-3 h c/u)
|
||||||
|
|
||||||
|
### 1. `search` — Búsqueda full-text en transcripciones
|
||||||
|
|
||||||
|
```bash
|
||||||
|
yt-scraper search "vaporwave"
|
||||||
|
yt-scraper search "dragon ball" --channel UCmhcYyPg7fsxMzQsY0RJBjw
|
||||||
|
```
|
||||||
|
|
||||||
|
**Qué hace:** busca texto dentro de las transcripciones ya scrapeadas y devuelve vídeo + timestamp exacto.
|
||||||
|
|
||||||
|
**Implementación:**
|
||||||
|
- Nueva tabla `transcript_segments(video_id, start, text)` poblada durante el scrape
|
||||||
|
- Índice `USING fts5(text)` sobre esa tabla
|
||||||
|
- Comando `search` que hace `SELECT ... WHERE transcript_segments MATCH ?`
|
||||||
|
- **Esfuerzo:** 1-2 h. Sin dependencias nuevas.
|
||||||
|
|
||||||
|
### 2. `stats` — Estadísticas del canal
|
||||||
|
|
||||||
|
```bash
|
||||||
|
yt-scraper stats
|
||||||
|
yt-scraper stats --channel UCmhcYyPg7fsxMzQsY0RJBjw
|
||||||
|
```
|
||||||
|
|
||||||
|
**Qué muestra:**
|
||||||
|
```
|
||||||
|
Nostal Vlad (@Nostal-Vlad)
|
||||||
|
Videos: 33 scrapeados / 37 totales
|
||||||
|
Duración total: 11.2 horas
|
||||||
|
Palabras totales: 89,432
|
||||||
|
Fecha rango: 2019-04-28 → 2026-07-26
|
||||||
|
Views promedio: 45,231
|
||||||
|
Likes promedio: 3,201
|
||||||
|
Top tags: videojuegos (12), 2000s (10), nostalgia (8)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Implementación:** queries SQL agregadas + presentación con Rich tables.
|
||||||
|
- **Esfuerzo:** 1 h.
|
||||||
|
|
||||||
|
### 3. `export` — Export multi-formato
|
||||||
|
|
||||||
|
```bash
|
||||||
|
yt-scraper export --format json # un JSON con todo
|
||||||
|
yt-scraper export --format csv # CSV de metadatos
|
||||||
|
yt-scraper export --format srt # subtítulos SRT estándar
|
||||||
|
yt-scraper export --format html # galería navegable
|
||||||
|
```
|
||||||
|
|
||||||
|
**Implementación:**
|
||||||
|
- JSON: serializar `Store.get_all()` + leer transcripciones de los `.md`
|
||||||
|
- CSV: `csv.writer` sobre metadatos
|
||||||
|
- SRT: ya tenemos `{start, end, text}` en `Segment` — conversión trivial
|
||||||
|
- HTML: plantilla Jinja2 con tarjetas por vídeo (thumbnail + título + link)
|
||||||
|
- **Esfuerzo:** 2-3 h los 4 formatos.
|
||||||
|
|
||||||
|
### 4. `clip` — Extracción de segmento por timestamp
|
||||||
|
|
||||||
|
```bash
|
||||||
|
yt-scraper clip gOUyxFwWQqA --from 01:38 --to 03:20
|
||||||
|
```
|
||||||
|
|
||||||
|
**Qué hace:** devuelve el texto de la transcripción entre dos timestamps, listo para citar.
|
||||||
|
|
||||||
|
**Implementación:** cargar el `.md`, parsear segmentos, filtrar por rango.
|
||||||
|
- **Esfuerzo:** 30 min.
|
||||||
|
|
||||||
|
### 5. `audio` — Descarga de audio (podcast)
|
||||||
|
|
||||||
|
```bash
|
||||||
|
yt-scraper --channel "https://www.youtube.com/@Nostal-Vlad/videos" --download-audio
|
||||||
|
```
|
||||||
|
|
||||||
|
**Qué hace:** descarga el audio MP3 de cada vídeo junto a la transcripción.
|
||||||
|
|
||||||
|
**Implementación:** añadir a `extract.py`:
|
||||||
|
```python
|
||||||
|
ydl_opts["format"] = "bestaudio/best"
|
||||||
|
ydl_opts["postprocessors"] = [{"key": "FFmpegExtractAudio", "preferredcodec": "mp3", "preferredquality": "128"}]
|
||||||
|
ydl_opts["outtmpl"] = str(output_dir / "%(title)s.%(ext)s")
|
||||||
|
```
|
||||||
|
Requiere `ffmpeg` instalado.
|
||||||
|
- **Esfuerzo:** 1 h.
|
||||||
|
|
||||||
|
### 6. `re-render` — Regenerar Markdown tras cambiar plantilla
|
||||||
|
|
||||||
|
```bash
|
||||||
|
yt-scraper re-render
|
||||||
|
```
|
||||||
|
|
||||||
|
**Qué hace:** re-procesa todos los `.md` usando la plantilla actual sin re-descargar nada de YouTube. Útil cuando cambias `video.md.j2`.
|
||||||
|
|
||||||
|
**Implementación:** guardar `segments` y `chapters` en SQLite (columnas JSON), o re-parsear los `.md` existentes.
|
||||||
|
- **Esfuerzo:** 1-2 h.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## TIER 2 · Alto valor, esfuerzo medio (3-8 h c/u)
|
||||||
|
|
||||||
|
### 7. Multi-canal — Gestión de varios canales
|
||||||
|
|
||||||
|
```bash
|
||||||
|
yt-scraper channels add "https://www.youtube.com/@otro-canal"
|
||||||
|
yt-scraper channels list
|
||||||
|
yt-scraper --channel @nostal-vlad search "vaporwave"
|
||||||
|
yt-scraper --all-channels stats
|
||||||
|
```
|
||||||
|
|
||||||
|
**Implementación:** tabla `channels` ya existe; ampliar CLI con subcomandos `channels add/list/remove`. El `Store` ya soporta multi-canal (PK por `channel_id`).
|
||||||
|
- **Esfuerzo:** 3-4 h.
|
||||||
|
|
||||||
|
### 8. `watch` — Monitoreo de nuevos vídeos
|
||||||
|
|
||||||
|
```bash
|
||||||
|
yt-scraper watch --interval 6h
|
||||||
|
```
|
||||||
|
|
||||||
|
**Qué hace:** ejecuta discovery periódicamente, y si hay vídeos nuevos los procesa automáticamente.
|
||||||
|
|
||||||
|
**Implementación:** loop con `time.sleep(interval)` + `discover_channel()`. En Windows se puede dejar corriendo o usar Task Scheduler.
|
||||||
|
- **Esfuerzo:** 2-3 h.
|
||||||
|
|
||||||
|
### 9. Análisis de contenido
|
||||||
|
|
||||||
|
```bash
|
||||||
|
yt-scraper analyze --wordcloud
|
||||||
|
yt-scraper analyze --top-words 50
|
||||||
|
yt-scraper analyze --timeline "vaporwave"
|
||||||
|
```
|
||||||
|
|
||||||
|
**Qué hace:**
|
||||||
|
- Nube de palabras (matplotlib + wordcloud)
|
||||||
|
- Frecuencia de términos con gráfico
|
||||||
|
- Timeline: cuándo se mencionó un tema a lo largo del tiempo
|
||||||
|
|
||||||
|
**Dependencias nuevas:** `matplotlib`, `wordcloud`, `nltk` (o regex simple para español).
|
||||||
|
- **Esfuerzo:** 4-5 h.
|
||||||
|
|
||||||
|
### 10. `thumbnails` — Descarga masiva de miniaturas
|
||||||
|
|
||||||
|
```bash
|
||||||
|
yt-scraper thumbnails --size maxres
|
||||||
|
```
|
||||||
|
|
||||||
|
**Implementación:** ya tenemos `thumbnail` URL en el frontmatter. Un `requests.get` por vídeo + guardar en `data/thumbnails/`.
|
||||||
|
- **Esfuerzo:** 30 min.
|
||||||
|
|
||||||
|
### 11. Traducción de transcripciones
|
||||||
|
|
||||||
|
```bash
|
||||||
|
yt-scraper --channel @nostal-vlad --translate-to en
|
||||||
|
```
|
||||||
|
|
||||||
|
**Qué hace:** traduce cada transcripción al idioma especificado antes de renderizar el Markdown.
|
||||||
|
|
||||||
|
**Opciones:**
|
||||||
|
- Google Translate (gratis, vía `deep-translator`): rápido, baja calidad
|
||||||
|
- LLM (OpenAI/Anthropic): calidad alta, requiere API key
|
||||||
|
- **Esfuerzo:** 3-4 h (cualquier ruta).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## TIER 3 · Features avanzadas (8+ h)
|
||||||
|
|
||||||
|
### 12. LLM Summaries — Resúmenes por vídeo
|
||||||
|
|
||||||
|
```bash
|
||||||
|
yt-scraper --channel @nostal-vlad --summarize
|
||||||
|
```
|
||||||
|
|
||||||
|
**Qué hace:** genera un resumen de 3-5 párrafos por vídeo usando un LLM. Se añade al Markdown como campo `summary` en el frontmatter.
|
||||||
|
|
||||||
|
**Implementación:**
|
||||||
|
- Nuevo módulo `summarize.py`
|
||||||
|
- Usa la transcripción completa como contexto
|
||||||
|
- Modelos: OpenAI `gpt-4o-mini` ($0.015 por vídeo de 20 min), Anthropic Claude Haiku, o modelo local con `ollama`
|
||||||
|
- Guardar resumen en SQLite para no re-ejecutar
|
||||||
|
- **Esfuerzo:** 4-6 h.
|
||||||
|
|
||||||
|
**Dependencias:** `openai` o `anthropic` o `ollama` (local, gratis).
|
||||||
|
|
||||||
|
### 13. Q&A / RAG — Preguntas sobre el canal
|
||||||
|
|
||||||
|
```bash
|
||||||
|
yt-scraper ask "¿Qué dijo Vlad sobre los juegos flash?"
|
||||||
|
```
|
||||||
|
|
||||||
|
**Qué hace:** busca en todas las transcripciones y responde con citas + timestamps exactos.
|
||||||
|
|
||||||
|
**Implementación:**
|
||||||
|
- **RAG simple:** FTS5 search → mandar top-K segmentos al LLM como contexto
|
||||||
|
- **RAG con embeddings:** `sentence-transformers` → ChromaDB/FAISS → recuperación semántica
|
||||||
|
- **Esfuerzo:** RAG simple 4-5 h, con embeddings 8-10 h.
|
||||||
|
|
||||||
|
### 14. Extracción de entidades (NER)
|
||||||
|
|
||||||
|
```bash
|
||||||
|
yt-scraper analyze --entities
|
||||||
|
```
|
||||||
|
|
||||||
|
**Qué hace:** detecta personas, marcas, lugares, videojuegos mencionados a lo largo de todos los vídeos.
|
||||||
|
|
||||||
|
**Output:** tabla `entities(video_id, type, name, count)` → "Pokémon mencionado en 8 vídeos", "Sonic en 5 vídeos".
|
||||||
|
|
||||||
|
**Implementación:** `spaCy` (modelo `es_core_news_sm`) o LLM con prompt estructurado.
|
||||||
|
- **Esfuerzo:** 4-6 h.
|
||||||
|
|
||||||
|
### 15. Speaker diarization — Separación de hablantes
|
||||||
|
|
||||||
|
**Qué hace:** distingue "narrador" de "entrevistado" o "invitado" en la transcripción.
|
||||||
|
|
||||||
|
**Implementación:** el audit del Web Clipper ya documenta esto (`groupBySpeaker` en `YoutubeExtractor`). Se puede portar a Python o usar `pyannote-audio`.
|
||||||
|
- **Esfuerzo:** port del algoritmo 2-3 h; con modelo de audio 8+ h.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Matriz de prioridades
|
||||||
|
|
||||||
|
| # | Feature | Valor | Esfuerzo | ROI |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| 1 | `search` (FTS5) | 🔥🔥🔥 | 1-2 h | ⭐⭐⭐ |
|
||||||
|
| 2 | `stats` | 🔥🔥 | 1 h | ⭐⭐⭐ |
|
||||||
|
| 3 | `export json/srt/html` | 🔥🔥 | 2-3 h | ⭐⭐⭐ |
|
||||||
|
| 4 | `clip` | 🔥 | 30 min | ⭐⭐ |
|
||||||
|
| 5 | `audio` download | 🔥🔥 | 1 h | ⭐⭐⭐ |
|
||||||
|
| 6 | `re-render` | 🔥 | 1-2 h | ⭐⭐ |
|
||||||
|
| 7 | Multi-canal | 🔥🔥 | 3-4 h | ⭐⭐ |
|
||||||
|
| 8 | `watch` mode | 🔥🔥 | 2-3 h | ⭐⭐ |
|
||||||
|
| 9 | Análisis (wordcloud) | 🔥 | 4-5 h | ⭐ |
|
||||||
|
| 10 | `thumbnails` | 🔥 | 30 min | ⭐⭐⭐ |
|
||||||
|
| 11 | Traducción | 🔥🔥 | 3-4 h | ⭐⭐ |
|
||||||
|
| 12 | LLM summaries | 🔥🔥🔥 | 4-6 h | ⭐⭐⭐ |
|
||||||
|
| 13 | Q&A / RAG | 🔥🔥🔥 | 4-10 h | ⭐⭐ |
|
||||||
|
| 14 | NER (entidades) | 🔥 | 4-6 h | ⭐ |
|
||||||
|
| 15 | Speaker diarization | 🔥 | 2-8 h | ⭐ |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Dependencias Python adicionales por feature
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Tier 1 — sin dependencias nuevas (solo stdlib)
|
||||||
|
# search, stats, export, clip, re-render usan SQLite + Jinja2 ya instalados
|
||||||
|
|
||||||
|
# Tier 2
|
||||||
|
pip install matplotlib wordcloud # análisis / wordcloud
|
||||||
|
pip install deep-translator # traducción gratuita
|
||||||
|
|
||||||
|
# Tier 3
|
||||||
|
pip install openai # LLM summaries / Q&A
|
||||||
|
pip install sentence-transformers # embeddings para RAG semántico
|
||||||
|
pip install spacy && python -m spacy download es_core_news_sm # NER
|
||||||
|
pip install pyannote.audio # speaker diarization (requiere GPU idealmente)
|
||||||
|
|
||||||
|
# Audio download
|
||||||
|
# Requiere ffmpeg instalado en el sistema:
|
||||||
|
# winget install ffmpeg
|
||||||
|
# o: choco install ffmpeg
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Comandos compuestos propuestos (pipelines)
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Pipeline completo: scrapear + buscar + extraer clip
|
||||||
|
yt-scraper --channel @nostal-vlad scrape
|
||||||
|
yt-scraper search "frutiger aero"
|
||||||
|
yt-scraper clip 8tMQMtIBfTE --from 02:15 --to 04:30 > clip.md
|
||||||
|
|
||||||
|
# Pipeline de análisis
|
||||||
|
yt-scraper stats > stats.txt
|
||||||
|
yt-scraper export --format html > index.html
|
||||||
|
yt-scraper analyze --top-words 30 > palabras.csv
|
||||||
|
|
||||||
|
# Pipeline LLM
|
||||||
|
yt-scraper --channel @nostal-vlad --summarize
|
||||||
|
yt-scraper ask "¿Qué opina Vlad sobre la cultura otaku?"
|
||||||
|
```
|
||||||
@@ -0,0 +1,163 @@
|
|||||||
|
# yt-channel-scraper
|
||||||
|
|
||||||
|
Scraper de canales de YouTube que extrae transcripciones, capítulos y metadatos completos, generando notas Markdown listas para Obsidian.
|
||||||
|
|
||||||
|
Basado en la ingeniería inversa de [Obsidian Web Clipper](https://obsidian.md/) — usa `yt-dlp` internamente, que implementa el mismo mecanismo InnerTube (`youtubei/v1/player` con clientes ANDROID/IOS/WEB) que la extensión audita.
|
||||||
|
|
||||||
|
## Instalación
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd yt-channel-scraper
|
||||||
|
pip install -e ".[dev]"
|
||||||
|
```
|
||||||
|
|
||||||
|
Requiere Python ≥ 3.10.
|
||||||
|
|
||||||
|
## Uso
|
||||||
|
|
||||||
|
### Scrape completo de un canal
|
||||||
|
|
||||||
|
```bash
|
||||||
|
yt-scraper --channel "https://www.youtube.com/@Nostal-Vlad/videos"
|
||||||
|
```
|
||||||
|
|
||||||
|
### Dry run (ver qué descubriría sin descargar)
|
||||||
|
|
||||||
|
```bash
|
||||||
|
yt-scraper --channel "https://www.youtube.com/@Nostal-Vlad/videos" --dry-run --limit 10
|
||||||
|
```
|
||||||
|
|
||||||
|
### Solo vídeos recientes
|
||||||
|
|
||||||
|
```bash
|
||||||
|
yt-scraper --channel "https://www.youtube.com/@Nostal-Vlad/videos" --since 2025-01-01
|
||||||
|
```
|
||||||
|
|
||||||
|
### Reanudar tras interrupción
|
||||||
|
|
||||||
|
```bash
|
||||||
|
yt-scraper --channel "https://www.youtube.com/@Nostal-Vlad/videos"
|
||||||
|
```
|
||||||
|
|
||||||
|
El estado se guarda en `data/state.db` (SQLite). Los vídeos ya procesados se saltan automáticamente.
|
||||||
|
|
||||||
|
### Reintentar vídeos con error
|
||||||
|
|
||||||
|
```bash
|
||||||
|
yt-scraper --channel "https://www.youtube.com/@Nostal-Vlad/videos" --reset-errors
|
||||||
|
```
|
||||||
|
|
||||||
|
## Opciones
|
||||||
|
|
||||||
|
| Flag | Descripción | Default |
|
||||||
|
|---|---|---|
|
||||||
|
| `--channel, -c` | URL del canal | de config.yaml |
|
||||||
|
| `--config` | Ruta al YAML de configuración | `config.yaml` |
|
||||||
|
| `--limit N` | Procesar solo N vídeos | sin límite |
|
||||||
|
| `--since DATE` | Solo vídeos desde YYYY-MM-DD | sin filtro |
|
||||||
|
| `--languages, -l` | Idiomas preferidos (coma-sep) | `es,en` |
|
||||||
|
| `--no-auto` | Ignorar subtítulos auto-generados | false |
|
||||||
|
| `--no-shorts` | Excluir Shorts | de config |
|
||||||
|
| `--include-shorts` | Incluir Shorts | de config |
|
||||||
|
| `--resume/--no-resume` | Saltar procesados | `--resume` |
|
||||||
|
| `--dry-run` | Solo discovery, no descargar | false |
|
||||||
|
| `--reset-errors` | Reintentar vídeos con error | false |
|
||||||
|
| `--verbose, -v` | Logging DEBUG | false |
|
||||||
|
|
||||||
|
## Configuración
|
||||||
|
|
||||||
|
Copia `config.example.yaml` a `config.yaml` y edita:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
channel_url: "https://www.youtube.com/@Nostal-Vlad/videos"
|
||||||
|
languages: ["es", "es-419", "en"]
|
||||||
|
prefer_manual: true
|
||||||
|
include_shorts: false
|
||||||
|
min_duration_sec: 30
|
||||||
|
|
||||||
|
delay:
|
||||||
|
min_seconds: 1.5
|
||||||
|
max_seconds: 3.5
|
||||||
|
```
|
||||||
|
|
||||||
|
## Output
|
||||||
|
|
||||||
|
Cada vídeo genera un archivo Markdown en `data/markdown/@canal/`:
|
||||||
|
|
||||||
|
```
|
||||||
|
data/markdown/Nostal Vlad/
|
||||||
|
├── 2026-07-26_asi-era-ser-una-adolescente-edgy-en-los-2000.md
|
||||||
|
├── 2026-07-12_la-estetica-que-romantiza-ser-un-perdedor-losercore.md
|
||||||
|
└── ...
|
||||||
|
```
|
||||||
|
|
||||||
|
Formato de cada nota:
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
---
|
||||||
|
video_id: gOUyxFwWQqA
|
||||||
|
title: "La Estética Que ROMANTIZA ser un \"PERDEDOR\" | Losercore"
|
||||||
|
channel: Nostal Vlad
|
||||||
|
upload_date: 2026-07-12
|
||||||
|
duration: 1069
|
||||||
|
url: https://www.youtube.com/watch?v=gOUyxFwWQqA
|
||||||
|
transcript_lang: es-orig
|
||||||
|
transcript_src: auto
|
||||||
|
views: 100523
|
||||||
|
likes: 8196
|
||||||
|
---
|
||||||
|
|
||||||
|
# La Estética Que ROMANTIZA ser un "PERDEDOR" | Losercore
|
||||||
|
|
||||||
|
> [Ver en YouTube](https://www.youtube.com/watch?v=gOUyxFwWQqA)
|
||||||
|
|
||||||
|
## Transcripcion
|
||||||
|
|
||||||
|
### Intro (00:15)
|
||||||
|
|
||||||
|
**00:15** · Una de las estéticas que ha cobrado más relevancia últimamente...
|
||||||
|
**00:28** · que hacer esto. Y es que el loser core como tal es muy difuso...
|
||||||
|
|
||||||
|
### Losercore (01:38)
|
||||||
|
|
||||||
|
**01:38** · ...
|
||||||
|
```
|
||||||
|
|
||||||
|
## Estados en SQLite
|
||||||
|
|
||||||
|
| Status | Significado |
|
||||||
|
|---|---|
|
||||||
|
| `pending` | Descubierto, sin procesar |
|
||||||
|
| `done` | Transcripción extraída y Markdown generado |
|
||||||
|
| `no_subtitles` | El vídeo no tiene subtítulos (ni manuales ni auto) |
|
||||||
|
| `error` | Error al procesar (miembros-only, privado, bloqueo, etc.) |
|
||||||
|
|
||||||
|
## Arquitectura
|
||||||
|
|
||||||
|
```
|
||||||
|
cli.py Orquestador (Click + Rich progress)
|
||||||
|
├── discover.py yt-dlp --flat-playlist → lista de videoIds
|
||||||
|
├── store.py SQLite: estado, resume, dedup
|
||||||
|
├── extract.py yt-dlp.extract_info → metadata + subtítulos
|
||||||
|
├── parse.py JSON3/VTT → segmentos {start, end, text}
|
||||||
|
├── chapters.py align_chapters: capítulos ↔ segmentos
|
||||||
|
├── render.py Jinja2 → Markdown con frontmatter YAML
|
||||||
|
└── ratelimit.py delays aleatorios + backoff exponencial
|
||||||
|
```
|
||||||
|
|
||||||
|
## Tests
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python -m pytest tests/ -v
|
||||||
|
```
|
||||||
|
|
||||||
|
## Mantenimiento
|
||||||
|
|
||||||
|
Si la extracción falla tras una actualización de YouTube:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
yt-dlp -U # actualizar yt-dlp
|
||||||
|
pip install -U yt-dlp
|
||||||
|
```
|
||||||
|
|
||||||
|
Las versiones de cliente InnerTube (ANDROID `20.10.x`, IOS `20.10.x`, WEB `2.2024xxxx`) rotan mensualmente. `yt-dlp` las mantiene actualizadas.
|
||||||
@@ -0,0 +1,23 @@
|
|||||||
|
channel_url: "https://www.youtube.com/@Nostal-Vlad/videos"
|
||||||
|
|
||||||
|
languages: ["es", "es-419", "en"]
|
||||||
|
prefer_manual: true
|
||||||
|
include_shorts: false
|
||||||
|
include_live: true
|
||||||
|
min_duration_sec: 30
|
||||||
|
|
||||||
|
delay:
|
||||||
|
min_seconds: 1.5
|
||||||
|
max_seconds: 3.5
|
||||||
|
backoff_base: 2.0
|
||||||
|
backoff_cap: 60.0
|
||||||
|
|
||||||
|
yt_dlp:
|
||||||
|
retries: 10
|
||||||
|
sleep_subrequests: 2
|
||||||
|
|
||||||
|
database_path: "data/state.db"
|
||||||
|
output_dir: "data/markdown"
|
||||||
|
template_path: "templates/video.md.j2"
|
||||||
|
|
||||||
|
filename_template: "{upload_date}_{slug}"
|
||||||
@@ -0,0 +1,122 @@
|
|||||||
|
# yt-scraper Platform Implementation Plan
|
||||||
|
|
||||||
|
> **Execution mode:** Subagent-driven + inline, autonomous, self-iterating. User mandated "completa todo el proceso de manera agentica y subagentica creando y resolviendo todos los problemas y dudas en un loop auto iterativo." Proceed without further checkpoints; verify with tests and fix in a loop until green.
|
||||||
|
|
||||||
|
**Goal:** Turn the CLI scraper into a local content-mining platform: CLI features (no AI) + FastAPI/Alpine webapp + .bat launcher + opencode subagent.
|
||||||
|
|
||||||
|
**Architecture:** Monorepo. `Store` extends with non-destructive migration (segments FTS5 + metadata + cookies + jobs). New modules `segments/cookies/export/analysis/monitor`. FastAPI app at `src/yt_scraper/webapp/` serves a no-build Alpine SPA. `.bat` launcher auto-scans ports.
|
||||||
|
|
||||||
|
**Tech:** Python ≥3.10, FastAPI, uvicorn, Alpine.js + Tailwind + Chart.js (CDN), SQLite FTS5, yt-dlp, matplotlib, Jinja2.
|
||||||
|
|
||||||
|
**Spec:** `docs/superpowers/specs/2026-07-26-platform-design.md`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## File map & responsibilities
|
||||||
|
|
||||||
|
| File | Responsibility | Status |
|
||||||
|
|---|---|---|
|
||||||
|
| `src/yt_scraper/store.py` | Add migration (new cols + 3 new tables), segment/cookie/job queries | modify |
|
||||||
|
| `src/yt_scraper/segments.py` | `store_segments`, `populate_fts`, `search`, `backfill_from_markdown` | new |
|
||||||
|
| `src/yt_scraper/cookies.py` | Netscape parse/validate, vault CRUD, `resolve_active_path` | new |
|
||||||
|
| `src/yt_scraper/export.py` | json/csv/srt/html exporters | new |
|
||||||
|
| `src/yt_scraper/analysis.py` | word freq / top-N / timeline (data + optional matplotlib) | new |
|
||||||
|
| `src/yt_scraper/monitor.py` | `watch_loop(config, interval, once)` | new |
|
||||||
|
| `src/yt_scraper/cli.py` | Click group + subcommands; fix cookies bug | modify |
|
||||||
|
| `src/yt_scraper/extract.py` | unchanged (already accepts cookies) | - |
|
||||||
|
| `src/yt_scraper/webapp/app.py` | FastAPI factory + static + healthz | new |
|
||||||
|
| `src/yt_scraper/webapp/api.py` | REST routers | new |
|
||||||
|
| `src/yt_scraper/webapp/jobs.py` | scrape job runner + SSE bus | new |
|
||||||
|
| `src/yt_scraper/webapp/static/{index.html,app.js,styles.css}` | SPA | new |
|
||||||
|
| `templates/gallery.html.j2` | HTML export gallery | new |
|
||||||
|
| `start-server.bat`, `stop-server.bat` | launcher | new |
|
||||||
|
| `.opencode/agent/webapp-builder.md`, `.opencode/goals/webapp-build.md` | subagent | new |
|
||||||
|
| `tests/` | cover migration, segments, cookies, export, analysis, cli | new/extend |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Key interfaces (contract between modules)
|
||||||
|
|
||||||
|
```python
|
||||||
|
# segments.py
|
||||||
|
@dataclass
|
||||||
|
class Hit: video_id: str; channel_id: str; title: str; start_sec: float; end_sec: float; snippet: str; rank: float
|
||||||
|
def store_segments(store, video_id, segments: list[Segment]) -> None
|
||||||
|
def search(store, query: str, channel_id: str|None=None, limit: int=50) -> list[Hit]
|
||||||
|
def backfill_from_markdown(store, md_root: Path, log=print) -> int # returns n videos
|
||||||
|
|
||||||
|
# cookies.py
|
||||||
|
@dataclass
|
||||||
|
class CookieMeta: id:str; filename:str; label:str|None; added_at:str; expires_at:str|None; is_active:bool; has_session:bool; cookie_count:int
|
||||||
|
def parse_netscape(path) -> tuple[bool, dict] # (ok, {expires_at, has_session, count})
|
||||||
|
def import_file(store, path) -> str # id
|
||||||
|
def import_text(store, text, label) -> str
|
||||||
|
def set_active(store, cookie_id) -> None
|
||||||
|
def list_vault(store) -> list[CookieMeta]
|
||||||
|
def delete(store, cookie_id) -> None
|
||||||
|
def resolve_active_path(store) -> str|None
|
||||||
|
def auto_import_dir(store, dir_path) -> None # import loose .txt in cookies/
|
||||||
|
|
||||||
|
# export.py
|
||||||
|
def export_json(store, segments_store_fn, channel_id, out_path) -> Path
|
||||||
|
def export_csv(store, channel_id, out_path) -> Path
|
||||||
|
def export_srt(store, video_id, out_path) -> Path
|
||||||
|
def export_html(store, channel_id, out_path, template_path) -> Path
|
||||||
|
|
||||||
|
# analysis.py
|
||||||
|
def word_frequency(store, channel_id=None) -> dict[str,int]
|
||||||
|
def top_words(store, n=50, channel_id=None) -> list[tuple[str,int]]
|
||||||
|
def term_timeline(store, term, channel_id=None) -> list[tuple[str,int]] # (YYYY-MM, count)
|
||||||
|
def render_wordcloud(freq, out_path) -> None # matplotlib
|
||||||
|
|
||||||
|
# monitor.py
|
||||||
|
def watch_loop(cfg, interval_seconds, channel_id=None, once=False) -> None
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Tricky bits (resolved decisions)
|
||||||
|
|
||||||
|
- **FTS5 sync:** manual `populate_fts` (delete+insert) after `store_segments`. No triggers (simpler, idempotent). FTS uses external-content `transcript_fts` with `content='transcript_segments'`; queries via `MATCH` + join.
|
||||||
|
- **Migration:** `ALTER TABLE ADD COLUMN` is guarded by `PRAGMA table_info` checks (idempotent). New tables via `CREATE TABLE IF NOT EXISTS`.
|
||||||
|
- **Backfill parse:** re-use existing `.md` format: lines `**MM:SS** · text` under `### Title (MM:SS)` chapter headers; frontmatter between `---` fences has YAML metadata. Reuse `parse.Segment` + `chapters.Section`.
|
||||||
|
- **Cookie active:** single row `is_active=1`; `set_active` flips all others to 0 in one tx.
|
||||||
|
- **Click structure:** convert top-level `@click.command` to `@click.group` with the original scrape flow under `scrape` command (default invocation preserved via a wrapper), plus `search/export/audio/channels/watch/analyze/re-render`.
|
||||||
|
- **Port scan:** PowerShell `System.Net.Sockets.TcpListener` bind-test 8000..8100.
|
||||||
|
- **SSE:** `EventSourceResponse`-style via `sse-starlette` OR a hand-rolled streaming `Response` with `text/event-stream`. Use `sse-starlette` (add to deps) — cleaner.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task list (build order)
|
||||||
|
|
||||||
|
### A1 — Foundation
|
||||||
|
- T1: Store migration (new cols + `transcript_segments` + `transcript_fts` + `cookies_meta` + `scrape_jobs`) + queries. Test: migration idempotent, columns exist, tables exist.
|
||||||
|
- T2: `segments.py` (`store_segments`, `populate_fts`, `search`). Test: store + search returns Hit.
|
||||||
|
- T3: `cookies.py`. Test: parse real cookie file, validate, active toggling, resolve path.
|
||||||
|
- T4: Fix cli.py cookies bug (thread cookies through `_process_one`; fallback to active vault cookie).
|
||||||
|
- T5: `backfill_from_markdown`. Test: parse one real `.md`, assert segments inserted.
|
||||||
|
|
||||||
|
### A2 — Feature modules
|
||||||
|
- T6: `export.py` (json/csv/srt/html) + `gallery.html.j2`. Test: each format round-trips.
|
||||||
|
- T7: `analysis.py` (freq, top, timeline, wordcloud). Test: counts on synthetic segments.
|
||||||
|
- T8: `monitor.py` (`watch_loop`, `--once`). Test: single iteration with mocked discover.
|
||||||
|
|
||||||
|
### A3 — CLI wiring
|
||||||
|
- T9: Convert `cli.py` to Click group; add `search`, `export`, `audio`, `channels`, `watch`, `analyze`, `re-render`, `--all-channels`. Smoke-test each via `CliRunner`.
|
||||||
|
|
||||||
|
### A4 — Tests green
|
||||||
|
- T10: Full `pytest` green; integration tests marked.
|
||||||
|
|
||||||
|
### B — Webapp (subagent: `webapp-builder`)
|
||||||
|
- T11: `.opencode/agent/webapp-builder.md` + `.opencode/goals/webapp-build.md`.
|
||||||
|
- T12: webapp backend (`app.py`, `api.py`, `jobs.py`) — all routers + SSE.
|
||||||
|
- T13: SPA (`index.html`, `app.js`, `styles.css`) — 9 views, dark command-center, drag-drop cookies, Chart.js, SSE scrape console.
|
||||||
|
- T14: Verify `/healthz` + smoke each endpoint; auto-import real cookie on startup.
|
||||||
|
|
||||||
|
### C — Launcher + deps
|
||||||
|
- T15: `pyproject.toml` optional deps (`web`, `analysis`); `.gitignore` (`cookies/`, `.run/`, `data/analysis/`).
|
||||||
|
- T16: `start-server.bat` + `stop-server.bat` (PowerShell port scan).
|
||||||
|
- T17: End-to-end: `start-server.bat` → browser → `/healthz` 200 → dashboard loads.
|
||||||
|
|
||||||
|
### Verification loop
|
||||||
|
- Run `pytest`, fix until green. Run `start-server.bat`, confirm health + SPA. Iterate.
|
||||||
@@ -0,0 +1,296 @@
|
|||||||
|
# Design: yt-channel-scraper → Local Content-Mining Platform
|
||||||
|
|
||||||
|
**Date:** 2026-07-26
|
||||||
|
**Status:** Approved (design review)
|
||||||
|
**Stack decision:** FastAPI + Alpine.js + Tailwind (CDN), no Node/npm build step
|
||||||
|
**Scope rule:** Data-mining + content focus. **No AI** (excluded: LLM summaries, RAG, NER, speaker diarization).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Goal
|
||||||
|
|
||||||
|
Turn the existing yt-dlp-based CLI scraper into a **local "intelligent platform"** for scraping and mining content-creator data. Add the high-value CLI features from `OPPORTUNITIES.md` (minus AI), a visually attractive local webapp with drag-and-drop cookie session login, `.bat` launchers with automatic port handling, and an opencode subagent + sub-goal that owns the webapp implementation.
|
||||||
|
|
||||||
|
Everything runs **100% locally** — no remote deploy, no auth, no rate limits beyond YouTube's own.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Context (current state)
|
||||||
|
|
||||||
|
- Modular pipeline: `cli → discover → extract → parse → chapters → render → store`, plus `ratelimit` and `config`.
|
||||||
|
- SQLite (`data/state.db`): tables `channels` (2 rows), `videos` (883 rows). Statuses: `pending/done/no_subtitles/error`.
|
||||||
|
- Rich metadata (views, likes, tags, thumbnail, description) currently lives **only in Markdown frontmatter**, not in the DB.
|
||||||
|
- 33 transcripts already rendered as `.md` under `data/markdown/<channel>/`.
|
||||||
|
- A valid Netscape cookie file exists at `cookies/253506e4-387c-4cda-9f11-e95548f8ca0b.txt`.
|
||||||
|
- **Known bug:** `cli.py:228-229` — `_process_one()` references `cookies` and `cookies_from_browser` that are local to `main()` and therefore out of scope. Cookies are currently non-functional. This design fixes it.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Architecture decision
|
||||||
|
|
||||||
|
**Approach A — Monorepo, in-package webapp (chosen).**
|
||||||
|
|
||||||
|
The FastAPI app lives inside the existing package at `src/yt_scraper/webapp/` and reuses `Store` + pipeline modules directly. The same FastAPI process serves the SPA static files, the JSON API, and SSE. One repo, one process, zero query duplication.
|
||||||
|
|
||||||
|
Rejected alternatives:
|
||||||
|
- B (separate top-level package importing `yt_scraper`): doubles plumbing for no local benefit.
|
||||||
|
- C (separate processes sharing only DB): duplicates query logic, breaks the "one platform" feel.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. New repo layout
|
||||||
|
|
||||||
|
```
|
||||||
|
src/yt_scraper/
|
||||||
|
cli.py +subcommands: search, export, audio, channels, watch, analyze, re-render
|
||||||
|
store.py +schema migration, +segments/metadata/cookies/jobs queries
|
||||||
|
segments.py NEW populate transcript_segments + FTS5; backfill from .md
|
||||||
|
export.py NEW json/csv/srt/html exporters
|
||||||
|
analysis.py NEW word freq, top-N terms, term-over-time timeline (matplotlib)
|
||||||
|
monitor.py NEW watch-mode loop (interval discovery + process)
|
||||||
|
cookies.py NEW Netscape parse/validate/vault; resolve active cookie path
|
||||||
|
webapp/
|
||||||
|
__init__.py
|
||||||
|
app.py FastAPI factory: static mount, CORS off (local), /healthz, SSE mount
|
||||||
|
api.py REST routers (channels, videos, search, transcript, analysis, scrape, cookies, exports)
|
||||||
|
jobs.py in-process scrape job runner + SSE event bus (queue/stream/cancel)
|
||||||
|
static/
|
||||||
|
index.html SPA shell (Tailwind CDN, Alpine.js, Chart.js CDN)
|
||||||
|
app.js Alpine app: views, state, API client, SSE consumer
|
||||||
|
styles.css custom dark "command-center" theme
|
||||||
|
pipeline glue: extract.py / cli.py now accept the active cookie path
|
||||||
|
templates/
|
||||||
|
video.md.j2 unchanged (re-render reuses it)
|
||||||
|
gallery.html.j2 NEW template for `export --format html`
|
||||||
|
docs/superpowers/specs/2026-07-26-platform-design.md this file
|
||||||
|
.opencode/
|
||||||
|
agent/webapp-builder.md NEW subagent owning Workstream B
|
||||||
|
goals/webapp-build.md NEW sub-goal the subagent executes
|
||||||
|
start-server.bat NEW auto port-scan + launch + open browser (repo root)
|
||||||
|
stop-server.bat NEW graceful shutdown + port cleanup (repo root)
|
||||||
|
.run/ gitignored runtime: server.info (port/PID)
|
||||||
|
data/analysis/ gitignored output of `analyze` (PNG/CSV)
|
||||||
|
```
|
||||||
|
|
||||||
|
`pyproject.toml` gains optional deps: `fastapi`, `uvicorn[standard]`, `matplotlib`, `wordcloud`, and `ffmpeg` is a system requirement for `audio`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Data model (non-destructive, idempotent migration)
|
||||||
|
|
||||||
|
Run automatically on every `Store` instantiation via `_init_schema` + `ALTER TABLE ADD COLUMN` guards.
|
||||||
|
|
||||||
|
### 5.1 `videos` — added columns
|
||||||
|
```
|
||||||
|
view_count INTEGER
|
||||||
|
like_count INTEGER
|
||||||
|
tags TEXT -- JSON array
|
||||||
|
thumbnail TEXT
|
||||||
|
description TEXT
|
||||||
|
chapters_json TEXT -- JSON of chapters (for re-render)
|
||||||
|
segments_json TEXT -- JSON of segments (for re-render, avoids re-download)
|
||||||
|
```
|
||||||
|
|
||||||
|
### 5.2 New `transcript_segments` + FTS5
|
||||||
|
```
|
||||||
|
CREATE TABLE transcript_segments (
|
||||||
|
video_id TEXT NOT NULL,
|
||||||
|
idx INTEGER NOT NULL,
|
||||||
|
start_sec REAL NOT NULL,
|
||||||
|
end_sec REAL NOT NULL,
|
||||||
|
text TEXT NOT NULL,
|
||||||
|
PRIMARY KEY (video_id, idx),
|
||||||
|
FOREIGN KEY (video_id) REFERENCES videos(video_id)
|
||||||
|
);
|
||||||
|
CREATE INDEX idx_segments_video ON transcript_segments(video_id);
|
||||||
|
|
||||||
|
-- FTS5 external-content table over transcript_segments(text)
|
||||||
|
CREATE VIRTUAL TABLE transcript_fts USING fts5(
|
||||||
|
text, content='transcript_segments', content_rowid='rowid'
|
||||||
|
);
|
||||||
|
```
|
||||||
|
Populated during scrape (after parse) **and** by a one-time backfill that parses the 33 existing `.md` files. Drives `search` (CLI) and webapp search (snippet + timestamp).
|
||||||
|
|
||||||
|
### 5.3 New `cookies_meta`
|
||||||
|
```
|
||||||
|
CREATE TABLE cookies_meta (
|
||||||
|
id TEXT PRIMARY KEY, -- uuid (matches filename stem)
|
||||||
|
filename TEXT NOT NULL,
|
||||||
|
label TEXT,
|
||||||
|
added_at TEXT NOT NULL,
|
||||||
|
expires_at TEXT, -- earliest cookie expiry, ISO
|
||||||
|
is_active INTEGER DEFAULT 0, -- exactly one row may be 1
|
||||||
|
has_session INTEGER DEFAULT 0, -- 1 if SID/SAPISID/LOGIN_INFO present
|
||||||
|
cookie_count INTEGER DEFAULT 0
|
||||||
|
);
|
||||||
|
```
|
||||||
|
Cookie files are stored in `cookies/`. Exactly **one** cookie may be active globally. `cookies.py` resolves the active path for yt-dlp's `cookiefile`.
|
||||||
|
|
||||||
|
### 5.4 New `scrape_jobs`
|
||||||
|
```
|
||||||
|
CREATE TABLE scrape_jobs (
|
||||||
|
id TEXT PRIMARY KEY,
|
||||||
|
channel_id TEXT,
|
||||||
|
opts_json TEXT,
|
||||||
|
status TEXT NOT NULL, -- queued/running/done/cancelled/error
|
||||||
|
started_at TEXT,
|
||||||
|
finished_at TEXT,
|
||||||
|
total INTEGER DEFAULT 0,
|
||||||
|
completed INTEGER DEFAULT 0,
|
||||||
|
last_error TEXT
|
||||||
|
);
|
||||||
|
```
|
||||||
|
The scrape console + SSE consumer read/write this table.
|
||||||
|
|
||||||
|
### 5.5 Backfill
|
||||||
|
One-time `segments.backfill_from_markdown()` invoked on first run (or via `re-render --backfill`): parse each done `.md`, extract segments + chapters + frontmatter metadata, write into `transcript_segments`/`transcript_fts` and the new `videos` columns. Idempotent (skip videos already having segments).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Workstream A — Core engine + CLI
|
||||||
|
|
||||||
|
### 6.1 CLI subcommands (new)
|
||||||
|
|
||||||
|
| Command | Behavior |
|
||||||
|
|---|---|
|
||||||
|
| `search "query" [-c CHANNEL] [-l N]` | FTS5 MATCH → prints video_id, start timestamp, snippet. |
|
||||||
|
| `export --format {json,csv,srt,html} [-c CHANNEL] [-o DIR]` | JSON = full dump; CSV = metadata table; SRT = subtitles from segments; HTML = Jinja2 gallery (`gallery.html.j2`). |
|
||||||
|
| `audio [-c CHANNEL] [--limit N] [--from DATE]` | yt-dlp `bestaudio` → mp3 via FFmpegExtractAudio. Skips videos already having audio unless `--force`. Records nothing new in DB beyond existing `done`. |
|
||||||
|
| `re-render [-c CHANNEL]` | Rebuild all `.md` from `segments_json`+`chapters_json`+metadata using the current `video.md.j2`. Errors if segments missing (hint: run backfill). |
|
||||||
|
| `channels add <URL>` / `channels list` / `channels remove <id|handle>` | Multi-channel management. `list` prints a Rich table. |
|
||||||
|
| `--all-channels` (global flag) | When set, `search`/`export`/`analyze`/`re-render` operate across every channel instead of one. Mutually exclusive with `-c`. |
|
||||||
|
| `watch --interval 6h [-c CHANNEL] [--once]` | Loop: `discover_channel` → upsert → process pending → sleep. `--once` runs a single iteration (useful for Task Scheduler). |
|
||||||
|
| `analyze --wordcloud --top-words N --timeline TERM [-c CHANNEL] [-o DIR]` | Saves `wordcloud.png`, `top_words.csv`, `timeline.png` to `data/analysis/`. Pure frequency mining (regex tokenization, Spanish stopwords). |
|
||||||
|
|
||||||
|
### 6.2 Bug fix
|
||||||
|
`_process_one()` and `extract_video()` already accept `cookies_file`/`cookies_from_browser` — the bug is that `main()` doesn't pass its local `cookies`/`cookies_from_browser` into `_process_one()`. Fix: thread the two values through `_process_one()`'s signature. Additionally, when neither flag is given, fall back to the **active cookie from the vault** (`cookies.resolve_active_path()`), so the webapp's drag-and-drop selection works for CLI too.
|
||||||
|
|
||||||
|
### 6.3 Module responsibilities
|
||||||
|
- `segments.py`: `store_segments(video_id, segments)`, `populate_fts(video_id)`, `search(query, channel_id, limit) -> list[Hit]`, `backfill_from_markdown(store, md_dir)`.
|
||||||
|
- `export.py`: `export_json/csv/srt/html(store, segments_store, channel_id, out_dir)`.
|
||||||
|
- `analysis.py`: `word_frequency(channel_id)`, `top_words(n)`, `term_timeline(term)` — returns data structures; matplotlib rendering optional behind a `render=True` flag so the webapp can call the same functions for Chart.js data without PNG.
|
||||||
|
- `monitor.py`: `watch_loop(config, interval, once)`.
|
||||||
|
- `cookies.py`: `parse_netscape(path) -> CookieFile`, `validate(meta) -> {valid, has_session, expires_at}`, `import_upload(text_or_path) -> id`, `set_active(id)`, `resolve_active_path() -> str|None`, `list_vault() -> list[CookieMeta]`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Workstream B — Webapp (owned by the `webapp-builder` subagent)
|
||||||
|
|
||||||
|
### 7.1 Backend (FastAPI)
|
||||||
|
- `app.py`: factory mounts `/static` (SPA), includes API routers under `/api`, exposes `/healthz`, and an SSE endpoint `/api/scrape/{job_id}/stream`.
|
||||||
|
- `api.py` routers:
|
||||||
|
- `GET /api/dashboard` → aggregate stats (per-channel counts, total duration, date range, uploads-over-time buckets, top tags, duration histogram, status breakdown).
|
||||||
|
- `GET /api/channels`, `POST /api/channels {url}`, `DELETE /api/channels/{id}`.
|
||||||
|
- `GET /api/videos?channel=&status=&from=&to=&min_dur=&q=&sort=&page=&size=`.
|
||||||
|
- `GET /api/videos/{id}` → full metadata + chapters.
|
||||||
|
- `GET /api/videos/{id}/transcript` → segments list (for the reader).
|
||||||
|
- `GET /api/search?q=&channel=&limit=` → FTS hits with snippet + timestamp.
|
||||||
|
- `GET /api/analysis?channel=&type=wordcloud|topwords|timeline&term=` → JSON for Chart.js.
|
||||||
|
- `POST /api/scrape {channel_id, opts}` → creates a `scrape_jobs` row, enqueues in `jobs.py`, returns `{job_id}`.
|
||||||
|
- `GET /api/scrape/{job_id}/stream` (SSE) → progress + log lines.
|
||||||
|
- `DELETE /api/scrape/{job_id}` → cancel.
|
||||||
|
- `POST /api/cookies/upload` (multipart/text) → import; `GET /api/cookies`; `POST /api/cookies/{id}/activate`; `DELETE /api/cookies/{id}`; `POST /api/cookies/{id}/test`.
|
||||||
|
- `GET /api/export?format=&channel=` → streams the file.
|
||||||
|
- `jobs.py`: a single-threaded `queue.Queue` worker runs scrape jobs sequentially (avoids concurrent yt-dlp hammering). Emits SSE events (`progress`, `log`, `done`, `error`). Persisted in `scrape_jobs`.
|
||||||
|
- The active cookie from the vault is passed to every scrape job automatically.
|
||||||
|
|
||||||
|
### 7.2 Frontend (Alpine.js SPA)
|
||||||
|
Design language: **dark "data command center."** Near-black canvas (`zinc-950`), glassy cards (`backdrop-blur`, subtle borders), red→rose accent gradient (creator-content nod), monospace data accents, Chart.js viz, smooth Alpine `x-transition`s. Responsive (works on the laptop you launch it on). No auth.
|
||||||
|
|
||||||
|
Views (hash routes):
|
||||||
|
1. **Dashboard** — channel cards + global charts (uploads-over-time line, top-tags bar, duration histogram, status donut).
|
||||||
|
2. **Channels** — list / add (paste URL) / remove / drill-in to a channel detail.
|
||||||
|
3. **Videos** — faceted table/grid: filter by channel, status, date range, min duration, free-text; thumbnail grid or list; sort; click → detail.
|
||||||
|
4. **Video detail** — metadata header, chapters sidebar, full transcript reader with per-segment timestamps, in-page term highlight, "open on YouTube at t", per-video export (srt/json), audio-download trigger.
|
||||||
|
5. **Search** — FTS across all transcripts → result cards with highlighted snippet + timestamp deep-link into the video detail at that segment.
|
||||||
|
6. **Analysis** — word cloud (canvas), top-N terms bar, term-over-time timeline.
|
||||||
|
7. **Scrape console** — launch a job (pick channel + options: limit, since, languages, shorts, which cookie), live progress via SSE, streaming log, cancel button, watch-mode toggle.
|
||||||
|
8. **Cookie vault** — drag-and-drop zone (see §8), vault list, test-session.
|
||||||
|
9. **Exports** — pick format + scope → download.
|
||||||
|
|
||||||
|
### 7.3 "Intelligent" without AI
|
||||||
|
Smart faceted filters, cross-channel term mining, term-over-time timelines, saved searches (localStorage), live monitoring + auto-processing, auto-suggestions from top terms. Feels like a command center for content-creator data mining.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. Cookie drag-and-drop UX
|
||||||
|
|
||||||
|
- Drop zone accepts **multiple `.txt` files** *or* **pasted Netscape cookie text**.
|
||||||
|
- On upload (`POST /api/cookies/upload`): `cookies.parse_netscape()` validates the format, extracts the earliest expiry, detects session cookies (`SID`, `SAPISID`, `__Secure-3PSID`, `LOGIN_INFO`), and returns a card: `{label, expires_at, has_session, cookie_count, size}`.
|
||||||
|
- Vault list with **radio = active** (exactly one active globally), rename, delete.
|
||||||
|
- The active cookie path is passed to yt-dlp as `cookiefile` for every scrape (members-only, age-gate, etc.).
|
||||||
|
- **"Test session"** → lightweight yt-dlp call (fetch the logged-in user's identity / a known members-only-free check) → reports `OK / expired / weak`.
|
||||||
|
- The existing `cookies/253506e4-...txt` is auto-imported into the vault on first server start.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. Launcher `.bat` (automatic port system)
|
||||||
|
|
||||||
|
PowerShell-assisted `.bat` files at the repo root.
|
||||||
|
|
||||||
|
- **`start-server.bat`**
|
||||||
|
1. Resolve project root (script dir).
|
||||||
|
2. Find a free TCP port: scan from `8000` upward until `bind` succeeds (PowerShell `TcpListener` probe).
|
||||||
|
3. Write `port`, `PID`, and ISO timestamp to `.run/server.info`.
|
||||||
|
4. Launch `python -m uvicorn yt_scraper.webapp.app:app --port <port>` via `start` (minimized).
|
||||||
|
5. Poll `http://localhost:<port>/healthz` until 200 (or timeout ~30s).
|
||||||
|
6. Open `http://localhost:<port>` in the default browser.
|
||||||
|
7. Echo a clear success banner (port, URL) or an error.
|
||||||
|
8. Idempotent: if `.run/server.info` already describes a live server, just re-open the browser.
|
||||||
|
- **`stop-server.bat`**
|
||||||
|
1. Read `.run/server.info`.
|
||||||
|
2. Graceful terminate the uvicorn process (taskkill by PID, then `/T` for children).
|
||||||
|
3. Confirm the port is freed (re-probe).
|
||||||
|
4. Remove `.run/server.info`.
|
||||||
|
5. Echo a clear banner.
|
||||||
|
|
||||||
|
`.run/` is added to `.gitignore`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 10. Subagent + sub-goal
|
||||||
|
|
||||||
|
- **`.opencode/agent/webapp-builder.md`** — an opencode agent specialized for this webapp. It knows: the FastAPI + Alpine + Tailwind (CDN, no-build) stack; the `Store` API surface; the design language (§7.2); and its verification steps (start server, `curl /healthz`, smoke-test each SPA view, run `pytest`). It receives Workstream B scope and is expected to reuse `Store`, `segments`, `analysis`, `cookies` rather than duplicating logic.
|
||||||
|
- **`.opencode/goals/webapp-build.md`** — the sub-goal: "Implement the local webapp (Workstream B) per `docs/superpowers/specs/2026-07-26-platform-design.md`. Build backend routers, SSE job runner, and the 9-view Alpine SPA. Verify with `/healthz`, smoke tests, and `pytest`."
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 11. Dependencies
|
||||||
|
|
||||||
|
Add to `pyproject.toml` `[project.optional-dependencies]`:
|
||||||
|
```
|
||||||
|
web = ["fastapi>=0.110", "uvicorn[standard]>=0.27"]
|
||||||
|
analysis = ["matplotlib>=3.8", "wordcloud>=1.9"]
|
||||||
|
```
|
||||||
|
System requirement for `audio`: `ffmpeg` (`winget install ffmpeg` or `choco install ffmpeg`). The CLI will detect ffmpeg and print a clear hint if missing.
|
||||||
|
|
||||||
|
`cookies/` is added to `.gitignore` (real session secrets must never be committed).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 12. Build order
|
||||||
|
|
||||||
|
A → B → C:
|
||||||
|
1. **A1** Schema migration + backfill + fix cookies bug.
|
||||||
|
2. **A2** `segments`/`export`/`analysis`/`monitor`/`cookies` modules.
|
||||||
|
3. **A3** CLI subcommands (`search`, `export`, `audio`, `channels`, `watch`, `analyze`, `re-render`).
|
||||||
|
4. **A4** Tests; `pytest` green.
|
||||||
|
5. **B** Webapp (subagent) — backend routers, job runner, SPA, drag-drop cookies. Verify `/healthz` + smoke.
|
||||||
|
6. **C** `.bat` launcher + `.opencode` agent/goal files.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 13. Testing & verification
|
||||||
|
|
||||||
|
- `pytest tests/ -v` covering: schema migration idempotency, FTS search, segment backfill from `.md`, cookie Netscape parse/validate, exporters (json/csv/srt/html), analysis frequency/timeline, multi-channel store ops.
|
||||||
|
- Webapp: `/healthz` 200; representative endpoint per router; SPA loads; SSE streams a fake job.
|
||||||
|
- Manual: run `start-server.bat`, confirm browser opens, drag-drop the real cookie, trigger a small scrape (`--limit 1`) against the existing Nostal Vlad channel, verify a new `.md` + FTS row appear and search returns the new content.
|
||||||
|
- `python -m yt_scraper.cli search "..."` returns hits after backfill.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 14. Out of scope (explicitly excluded)
|
||||||
|
|
||||||
|
- AI features: LLM summaries, Q&A/RAG, NER (spaCy), speaker diarization (pyannote).
|
||||||
|
- CLI `stats`, `clip`, `thumbnails`, `translation` (dropped per user scope decision — aggregate *queries* still exist internally to feed the dashboard).
|
||||||
|
- Remote deployment, auth, multi-user, cloud sync.
|
||||||
@@ -0,0 +1,48 @@
|
|||||||
|
[build-system]
|
||||||
|
requires = ["setuptools>=68.0", "wheel"]
|
||||||
|
build-backend = "setuptools.build_meta"
|
||||||
|
|
||||||
|
[project]
|
||||||
|
name = "yt-channel-scraper"
|
||||||
|
version = "0.1.0"
|
||||||
|
description = "Scraper de canales de YouTube hacia notas Markdown para Obsidian"
|
||||||
|
requires-python = ">=3.10"
|
||||||
|
dependencies = [
|
||||||
|
"yt-dlp>=2024.10.7",
|
||||||
|
"click>=8.1",
|
||||||
|
"rich>=13.7",
|
||||||
|
"jinja2>=3.1",
|
||||||
|
"pyyaml>=6.0",
|
||||||
|
"python-slugify>=8.0",
|
||||||
|
"requests>=2.31",
|
||||||
|
]
|
||||||
|
|
||||||
|
[project.optional-dependencies]
|
||||||
|
dev = [
|
||||||
|
"pytest>=8.0",
|
||||||
|
"pytest-cov>=4.1",
|
||||||
|
]
|
||||||
|
web = [
|
||||||
|
"fastapi>=0.110",
|
||||||
|
"uvicorn[standard]>=0.27",
|
||||||
|
"sse-starlette>=2.0",
|
||||||
|
]
|
||||||
|
analysis = [
|
||||||
|
"matplotlib>=3.8",
|
||||||
|
"wordcloud>=1.9",
|
||||||
|
]
|
||||||
|
|
||||||
|
[project.scripts]
|
||||||
|
yt-scraper = "yt_scraper.cli:main"
|
||||||
|
|
||||||
|
[tool.setuptools.packages.find]
|
||||||
|
where = ["src"]
|
||||||
|
|
||||||
|
[tool.setuptools.package-data]
|
||||||
|
yt_scraper = ["../templates/*.j2"]
|
||||||
|
|
||||||
|
[tool.pytest.ini_options]
|
||||||
|
testpaths = ["tests"]
|
||||||
|
markers = [
|
||||||
|
"integration: tests que requieren red (deselect with -m 'not integration')",
|
||||||
|
]
|
||||||
@@ -0,0 +1,84 @@
|
|||||||
|
$ErrorActionPreference = 'Stop'
|
||||||
|
$root = Split-Path -Parent $PSScriptRoot
|
||||||
|
Set-Location -LiteralPath $root
|
||||||
|
|
||||||
|
Write-Host ""
|
||||||
|
Write-Host " [yt-scraper] iniciando servidor local..." -ForegroundColor Cyan
|
||||||
|
|
||||||
|
if (-not (Test-Path '.run')) { New-Item -ItemType Directory -Path '.run' | Out-Null }
|
||||||
|
|
||||||
|
# Already running?
|
||||||
|
if (Test-Path '.run\server.info') {
|
||||||
|
$firstLine = (Get-Content '.run\server.info' -TotalCount 1)
|
||||||
|
$existingPort = ($firstLine -split '\s+')[0]
|
||||||
|
try {
|
||||||
|
$r = Invoke-WebRequest -Uri "http://localhost:$existingPort/healthz" -UseBasicParsing -TimeoutSec 2
|
||||||
|
if ($r.Content -match 'ok') {
|
||||||
|
Write-Host " [ok] servidor ya activo en el puerto $existingPort" -ForegroundColor Green
|
||||||
|
Start-Process "http://localhost:$existingPort"
|
||||||
|
exit 0
|
||||||
|
}
|
||||||
|
} catch {}
|
||||||
|
}
|
||||||
|
|
||||||
|
# Collect candidate free ports scanning from 8000.
|
||||||
|
$candidates = @()
|
||||||
|
foreach ($p in 8000..8100) {
|
||||||
|
try {
|
||||||
|
$listener = [System.Net.Sockets.TcpListener]::new([System.Net.IPAddress]::Loopback, $p)
|
||||||
|
$listener.Start(); $listener.Stop()
|
||||||
|
$candidates += $p
|
||||||
|
if ($candidates.Count -ge 8) { break }
|
||||||
|
} catch {}
|
||||||
|
}
|
||||||
|
if ($candidates.Count -eq 0) {
|
||||||
|
Write-Host " [error] no se encontro ningun puerto libre (8000-8100)" -ForegroundColor Red
|
||||||
|
exit 2
|
||||||
|
}
|
||||||
|
|
||||||
|
# Launch uvicorn DETACHED in its own persistent minimized console window, logging to .run/server.log.
|
||||||
|
# The `cmd /c` window is an independent console: it survives this launcher exiting.
|
||||||
|
$launched = $false
|
||||||
|
$proc = $null
|
||||||
|
$port = 0
|
||||||
|
foreach ($candidate in $candidates) {
|
||||||
|
# clear previous log
|
||||||
|
Remove-Item '.run\server.log' -Force -ErrorAction SilentlyContinue
|
||||||
|
$cmdLine = '/c python -m uvicorn yt_scraper.webapp.app:app --host 127.0.0.1 --port ' + $candidate + ' --log-level info > .run\server.log 2>&1'
|
||||||
|
$proc = Start-Process -FilePath 'cmd.exe' -ArgumentList $cmdLine -WorkingDirectory $root -WindowStyle Minimized -PassThru
|
||||||
|
Set-Content -Path '.run\server.info' -Value "$candidate $($proc.Id) $(Get-Date -Format o)" -Encoding ASCII
|
||||||
|
|
||||||
|
$ok = $false
|
||||||
|
for ($i = 0; $i -lt 40; $i++) {
|
||||||
|
Start-Sleep -Milliseconds 500
|
||||||
|
try {
|
||||||
|
$r = Invoke-WebRequest -Uri "http://localhost:$candidate/healthz" -UseBasicParsing -TimeoutSec 2
|
||||||
|
if ($r.Content -match 'ok') { $ok = $true; break }
|
||||||
|
} catch {}
|
||||||
|
}
|
||||||
|
if ($ok) { $port = $candidate; $launched = $true; break }
|
||||||
|
|
||||||
|
# failed on this port — show why from the log, then try the next candidate
|
||||||
|
Stop-Process -Id $proc.Id -Force -ErrorAction SilentlyContinue
|
||||||
|
Remove-Item '.run\server.info' -Force -ErrorAction SilentlyContinue
|
||||||
|
}
|
||||||
|
|
||||||
|
if (-not $launched) {
|
||||||
|
Write-Host " [error] no se pudo iniciar el servidor." -ForegroundColor Red
|
||||||
|
if (Test-Path '.run\server.log') {
|
||||||
|
Write-Host " --- ultimo log del servidor ---" -ForegroundColor Yellow
|
||||||
|
Get-Content '.run\server.log' -Tail 25 | ForEach-Object { Write-Host " $_" }
|
||||||
|
} else {
|
||||||
|
Write-Host " instala dependencias: pip install -e `".[web]`"" -ForegroundColor Yellow
|
||||||
|
}
|
||||||
|
exit 3
|
||||||
|
}
|
||||||
|
|
||||||
|
Start-Process "http://localhost:$port"
|
||||||
|
Write-Host ""
|
||||||
|
Write-Host " [ok] servidor activo" -ForegroundColor Green
|
||||||
|
Write-Host " URL: http://localhost:$port"
|
||||||
|
Write-Host " PID: $($proc.Id) (ventana minimizada 'cmd' independiente)"
|
||||||
|
Write-Host " log: .run\server.log"
|
||||||
|
Write-Host " stop: stop-server.bat"
|
||||||
|
Write-Host ""
|
||||||
@@ -0,0 +1,46 @@
|
|||||||
|
$ErrorActionPreference = 'Stop'
|
||||||
|
$root = Split-Path -Parent $PSScriptRoot
|
||||||
|
Set-Location -LiteralPath $root
|
||||||
|
|
||||||
|
if (-not (Test-Path '.run\server.info')) {
|
||||||
|
Write-Host ""
|
||||||
|
Write-Host " [yt-scraper] no hay servidor registrado." -ForegroundColor Yellow
|
||||||
|
Write-Host ""
|
||||||
|
exit 0
|
||||||
|
}
|
||||||
|
|
||||||
|
$firstLine = (Get-Content '.run\server.info' -TotalCount 1)
|
||||||
|
$parts = $firstLine -split '\s+'
|
||||||
|
$port = $parts[0]
|
||||||
|
$pidv = $parts[1]
|
||||||
|
|
||||||
|
Write-Host ""
|
||||||
|
Write-Host " [yt-scraper] deteniendo servidor (PID $pidv, puerto $port)..." -ForegroundColor Cyan
|
||||||
|
|
||||||
|
if ($pidv -and $pidv -ne 'pending') {
|
||||||
|
# Kill the process tree (uvicorn + any workers).
|
||||||
|
try {
|
||||||
|
Stop-Process -Id ([int]$pidv) -Force -ErrorAction Stop
|
||||||
|
} catch {
|
||||||
|
# fallback: taskkill
|
||||||
|
& taskkill /PID $pidv /T /F 2>$null | Out-Null
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
# Confirm the port is freed.
|
||||||
|
$freed = $false
|
||||||
|
for ($i = 0; $i -lt 10; $i++) {
|
||||||
|
try {
|
||||||
|
Invoke-WebRequest -Uri "http://localhost:$port/healthz" -UseBasicParsing -TimeoutSec 1 | Out-Null
|
||||||
|
} catch { $freed = $true; break }
|
||||||
|
Start-Sleep -Milliseconds 400
|
||||||
|
}
|
||||||
|
|
||||||
|
Remove-Item '.run\server.info' -Force -ErrorAction SilentlyContinue
|
||||||
|
|
||||||
|
if ($freed) {
|
||||||
|
Write-Host " [ok] puerto liberado, servidor detenido." -ForegroundColor Green
|
||||||
|
} else {
|
||||||
|
Write-Host " [warn] el puerto $port sigue ocupado; revisa procesos python." -ForegroundColor Yellow
|
||||||
|
}
|
||||||
|
Write-Host ""
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
__version__ = "0.1.0"
|
||||||
@@ -0,0 +1,163 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import logging
|
||||||
|
import re
|
||||||
|
from collections import Counter
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Iterable
|
||||||
|
|
||||||
|
from .store import Store
|
||||||
|
|
||||||
|
|
||||||
|
log = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
# Minimal Spanish + English stopword list (no extra deps).
|
||||||
|
_STOPWORDS = {
|
||||||
|
# es
|
||||||
|
"el", "la", "los", "las", "un", "una", "unos", "unas", "de", "del", "al", "a", "y", "o", "u",
|
||||||
|
"que", "en", "como", "por", "para", "su", "sus", "se", "si", "no", "con", "es", "son", "fue",
|
||||||
|
"era", "este", "esta", "estos", "estas", "eso", "esa", "esos", "esas", "lo", "le", "les", "te",
|
||||||
|
"me", "se", "nos", "os", "pero", "mas", "muy", "ya", "cuando", "donde", "quien", "como", "todo",
|
||||||
|
"todos", "toda", "todas", "nada", "algo", "tambien", "asi", "hay", "habia", "tiene", "tener",
|
||||||
|
"puede", "pueden", "esto", "esta", "sin", "sobre", "entre", "hasta", "desde", "mi", "tu", "yo",
|
||||||
|
"el", "ella", "ellos", "ellas", "nosotros", "vosotros", "ustedes", "porque", "pues", "entonces",
|
||||||
|
# en
|
||||||
|
"the", "a", "an", "and", "or", "but", "of", "to", "in", "on", "at", "by", "for", "with", "from",
|
||||||
|
"is", "are", "was", "were", "be", "been", "being", "this", "that", "these", "those", "it", "its",
|
||||||
|
"as", "if", "not", "no", "so", "do", "does", "did", "have", "has", "had", "i", "you", "he", "she",
|
||||||
|
"we", "they", "my", "your", "his", "her", "our", "their", "me", "him", "us", "them", "can", "could",
|
||||||
|
"would", "should", "will", "shall", "may", "might", "just", "like", "what", "which", "who", "when",
|
||||||
|
"where", "why", "how", "all", "any", "both", "each", "more", "most", "other", "some", "such",
|
||||||
|
}
|
||||||
|
|
||||||
|
_WORD = re.compile(r"[A-Za-zÁÉÍÓÚÜÑáéíóúüñ]{3,}")
|
||||||
|
|
||||||
|
|
||||||
|
def _iter_texts(store: Store, channel_id: str | None) -> Iterable[str]:
|
||||||
|
sql = "SELECT text FROM transcript_segments"
|
||||||
|
params: list = []
|
||||||
|
if channel_id:
|
||||||
|
sql += " WHERE video_id IN (SELECT video_id FROM videos WHERE channel_id = ?)"
|
||||||
|
params.append(channel_id)
|
||||||
|
conn = store._connect()
|
||||||
|
try:
|
||||||
|
cur = conn.execute(sql, params)
|
||||||
|
for row in cur:
|
||||||
|
yield row["text"]
|
||||||
|
finally:
|
||||||
|
conn.close()
|
||||||
|
|
||||||
|
|
||||||
|
def _tokenize(text: str) -> Iterable[str]:
|
||||||
|
for m in _WORD.finditer(text.lower()):
|
||||||
|
w = m.group(0)
|
||||||
|
if w not in _STOPWORDS:
|
||||||
|
yield w
|
||||||
|
|
||||||
|
|
||||||
|
def word_frequency(store: Store, channel_id: str | None = None, top: int | None = None) -> dict[str, int]:
|
||||||
|
counter: Counter[str] = Counter()
|
||||||
|
for text in _iter_texts(store, channel_id):
|
||||||
|
counter.update(_tokenize(text))
|
||||||
|
items = counter.most_common(top) if top else counter.most_common()
|
||||||
|
return dict(items)
|
||||||
|
|
||||||
|
|
||||||
|
def top_words(store: Store, n: int = 50, channel_id: str | None = None) -> list[tuple[str, int]]:
|
||||||
|
return list(word_frequency(store, channel_id, top=n).items())
|
||||||
|
|
||||||
|
|
||||||
|
def term_timeline(store: Store, term: str, channel_id: str | None = None) -> list[tuple[str, int]]:
|
||||||
|
"""Return [(YYYY-MM, count)] of months where `term` appears in transcripts."""
|
||||||
|
term_l = term.lower().strip()
|
||||||
|
if not term_l:
|
||||||
|
return []
|
||||||
|
sql = (
|
||||||
|
"SELECT substr(v.upload_date,1,6) AS month, COUNT(*) AS n "
|
||||||
|
"FROM transcript_fts f JOIN videos v ON v.video_id = f.video_id "
|
||||||
|
"WHERE transcript_fts MATCH ?"
|
||||||
|
)
|
||||||
|
params: list = [f'"{term_l}"']
|
||||||
|
if channel_id:
|
||||||
|
sql += " AND v.channel_id = ?"
|
||||||
|
params.append(channel_id)
|
||||||
|
sql += " AND v.upload_date IS NOT NULL GROUP BY month ORDER BY month"
|
||||||
|
conn = store._connect()
|
||||||
|
try:
|
||||||
|
rows = conn.execute(sql, params).fetchall()
|
||||||
|
finally:
|
||||||
|
conn.close()
|
||||||
|
out: list[tuple[str, int]] = []
|
||||||
|
for r in rows:
|
||||||
|
m = r["month"]
|
||||||
|
if m and len(m) == 6:
|
||||||
|
out.append((f"{m[:4]}-{m[4:6]}", r["n"]))
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def render_wordcloud(freq: dict[str, int], out_path: str | Path) -> Path:
|
||||||
|
from wordcloud import WordCloud
|
||||||
|
import matplotlib
|
||||||
|
matplotlib.use("Agg")
|
||||||
|
import matplotlib.pyplot as plt
|
||||||
|
|
||||||
|
out = Path(out_path)
|
||||||
|
out.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
wc = WordCloud(
|
||||||
|
width=1600, height=900, background_color="#0a0a0f",
|
||||||
|
colormap="Reds", max_words=200,
|
||||||
|
).generate_from_frequencies(freq)
|
||||||
|
fig, ax = plt.subplots(figsize=(16, 9), dpi=100)
|
||||||
|
ax.imshow(wc, interpolation="bilinear")
|
||||||
|
ax.set_axis_off()
|
||||||
|
fig.tight_layout(pad=0)
|
||||||
|
fig.savefig(str(out), facecolor="#0a0a0f")
|
||||||
|
plt.close(fig)
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def render_top_words_chart(top: list[tuple[str, int]], out_path: str | Path) -> Path:
|
||||||
|
import matplotlib
|
||||||
|
matplotlib.use("Agg")
|
||||||
|
import matplotlib.pyplot as plt
|
||||||
|
|
||||||
|
out = Path(out_path)
|
||||||
|
out.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
items = top[:25][::-1]
|
||||||
|
labels = [t[0] for t in items]
|
||||||
|
values = [t[1] for t in items]
|
||||||
|
fig, ax = plt.subplots(figsize=(10, 8), dpi=100)
|
||||||
|
ax.barh(labels, values, color="#f43f5e")
|
||||||
|
ax.set_facecolor("#0a0a0f")
|
||||||
|
fig.patch.set_facecolor("#0a0a0f")
|
||||||
|
ax.tick_params(colors="#e5e7eb")
|
||||||
|
for spine in ax.spines.values():
|
||||||
|
spine.set_color("#27272a")
|
||||||
|
ax.set_title("Top terms", color="#e5e7eb")
|
||||||
|
fig.tight_layout()
|
||||||
|
fig.savefig(str(out), facecolor="#0a0a0f")
|
||||||
|
plt.close(fig)
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def render_timeline_chart(timeline: list[tuple[str, int]], term: str, out_path: str | Path) -> Path:
|
||||||
|
import matplotlib
|
||||||
|
matplotlib.use("Agg")
|
||||||
|
import matplotlib.pyplot as plt
|
||||||
|
|
||||||
|
out = Path(out_path)
|
||||||
|
out.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
labels = [t[0] for t in timeline]
|
||||||
|
values = [t[1] for t in timeline]
|
||||||
|
fig, ax = plt.subplots(figsize=(12, 5), dpi=100)
|
||||||
|
ax.plot(labels, values, marker="o", color="#f43f5e")
|
||||||
|
ax.set_facecolor("#0a0a0f")
|
||||||
|
fig.patch.set_facecolor("#0a0a0f")
|
||||||
|
ax.tick_params(colors="#e5e7eb", axisxlabelrotation=45)
|
||||||
|
for spine in ax.spines.values():
|
||||||
|
spine.set_color("#27272a")
|
||||||
|
ax.set_title(f"Mentions over time: {term}", color="#e5e7eb")
|
||||||
|
fig.tight_layout()
|
||||||
|
fig.savefig(str(out), facecolor="#0a0a0f")
|
||||||
|
plt.close(fig)
|
||||||
|
return out
|
||||||
@@ -0,0 +1,49 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from dataclasses import dataclass, field
|
||||||
|
|
||||||
|
from .parse import Segment
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class Chapter:
|
||||||
|
title: str
|
||||||
|
start_time: float
|
||||||
|
end_time: float
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class Section:
|
||||||
|
title: str
|
||||||
|
start: float
|
||||||
|
segments: list[Segment] = field(default_factory=list)
|
||||||
|
|
||||||
|
|
||||||
|
def align_chapters(segments: list[Segment], chapters: list[Chapter]) -> list[Section]:
|
||||||
|
if not chapters:
|
||||||
|
return [Section(title="Transcripcion", start=0.0, segments=list(segments))]
|
||||||
|
|
||||||
|
ordered = sorted(chapters, key=lambda c: c.start_time)
|
||||||
|
sections: list[Section] = []
|
||||||
|
for i, ch in enumerate(ordered):
|
||||||
|
end = ordered[i + 1].start_time if i + 1 < len(ordered) else float("inf")
|
||||||
|
segs = [s for s in segments if ch.start_time <= s.start < end]
|
||||||
|
if i == 0:
|
||||||
|
pre = [s for s in segments if s.start < ch.start_time]
|
||||||
|
if pre:
|
||||||
|
sections.append(Section(title="Inicio", start=0.0, segments=pre))
|
||||||
|
sections.append(Section(title=ch.title, start=ch.start_time, segments=segs))
|
||||||
|
return sections
|
||||||
|
|
||||||
|
|
||||||
|
def chapters_from_info(info: dict) -> list[Chapter]:
|
||||||
|
raw = info.get("chapters") or []
|
||||||
|
out: list[Chapter] = []
|
||||||
|
for ch in raw:
|
||||||
|
title = (ch.get("title") or "").strip()
|
||||||
|
if not title:
|
||||||
|
continue
|
||||||
|
start = float(ch.get("start_time", 0))
|
||||||
|
end = float(ch.get("end_time", start))
|
||||||
|
out.append(Chapter(title=title, start_time=start, end_time=end))
|
||||||
|
return out
|
||||||
@@ -0,0 +1,542 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import logging
|
||||||
|
import shutil
|
||||||
|
import sys
|
||||||
|
from pathlib import Path
|
||||||
|
from types import SimpleNamespace
|
||||||
|
|
||||||
|
import click
|
||||||
|
from rich.console import Console
|
||||||
|
from rich.progress import Progress, SpinnerColumn, TextColumn, BarColumn, TaskProgressColumn, TimeRemainingColumn
|
||||||
|
from rich.table import Table
|
||||||
|
|
||||||
|
from .config import Config, load_config
|
||||||
|
from .store import Store, VideoRef, VideoRow
|
||||||
|
from .discover import discover_channel
|
||||||
|
from .chapters import align_chapters, chapters_from_info, Chapter, Section
|
||||||
|
from .parse import Segment
|
||||||
|
from .render import build_filename_stem, render_markdown
|
||||||
|
from .ratelimit import polite_sleep
|
||||||
|
from .pipeline import process_video
|
||||||
|
from .cookies import auto_import_dir, resolve_active_path
|
||||||
|
|
||||||
|
console = Console()
|
||||||
|
log = logging.getLogger("yt_scraper")
|
||||||
|
|
||||||
|
|
||||||
|
def _setup_logging(verbose: bool) -> None:
|
||||||
|
level = logging.DEBUG if verbose else logging.INFO
|
||||||
|
logging.basicConfig(
|
||||||
|
level=level,
|
||||||
|
format="%(asctime)s [%(levelname)s] %(name)s: %(message)s",
|
||||||
|
datefmt="%H:%M:%S",
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
# --------------------------------------------------------------------------- group
|
||||||
|
|
||||||
|
@click.group(invoke_without_command=True, context_settings={"help_option_names": ["-h", "--help"]})
|
||||||
|
@click.option("--config", "config_path", default="config.yaml", help="Ruta al config YAML")
|
||||||
|
@click.option("--channel", "-c", default=None, help="URL o id del canal")
|
||||||
|
@click.option("--all-channels", is_flag=True, default=False, help="Aplicar subcomando a todos los canales")
|
||||||
|
@click.option("--cookies", default=None, help="Ruta a cookies.txt (Netscape)")
|
||||||
|
@click.option("--cookies-from-browser", default=None, help="chrome|firefox|edge|brave")
|
||||||
|
@click.option("--verbose", "-v", is_flag=True, default=False, help="Logging DEBUG")
|
||||||
|
@click.pass_context
|
||||||
|
def cli(ctx, config_path, channel, all_channels, cookies, cookies_from_browser, verbose):
|
||||||
|
"""yt-channel-scraper: plataforma local de minería de contenido de creadores."""
|
||||||
|
_setup_logging(verbose)
|
||||||
|
cfg_path = Path(config_path)
|
||||||
|
cfg = load_config(cfg_path) if cfg_path.exists() else Config()
|
||||||
|
if channel:
|
||||||
|
cfg.channel_url = channel
|
||||||
|
store = Store(cfg.database_path_resolved)
|
||||||
|
# import any loose cookie files into the vault (idempotent)
|
||||||
|
try:
|
||||||
|
auto_import_dir(store)
|
||||||
|
except Exception as exc:
|
||||||
|
log.debug("cookie auto-import skipped: %s", exc)
|
||||||
|
ctx.obj = SimpleNamespace(
|
||||||
|
cfg=cfg, store=store, channel=channel, all_channels=all_channels,
|
||||||
|
cookies=cookies, cookies_from_browser=cookies_from_browser, config_path=config_path,
|
||||||
|
)
|
||||||
|
if ctx.invoked_subcommand is None:
|
||||||
|
_run_scrape(ctx.obj, limit=None, since=None, languages=None, no_auto=False,
|
||||||
|
no_shorts=None, include_shorts=None, no_live=None, resume=True,
|
||||||
|
dry_run=False, reset_errors=False)
|
||||||
|
|
||||||
|
|
||||||
|
# --------------------------------------------------------------------------- scrape
|
||||||
|
|
||||||
|
def _apply_filters(refs, since, no_shorts, no_live, min_duration, limit):
|
||||||
|
filtered = list(refs)
|
||||||
|
if since:
|
||||||
|
filtered = [r for r in filtered if (r.upload_date or "") >= since.replace("-", "")]
|
||||||
|
if no_shorts:
|
||||||
|
filtered = [r for r in filtered if "/shorts/" not in (r.url or "")]
|
||||||
|
if no_live:
|
||||||
|
filtered = [r for r in filtered if not (r.url or "").startswith("https://www.youtube.com/live/")]
|
||||||
|
if min_duration > 0:
|
||||||
|
filtered = [r for r in filtered if (r.duration or 0) >= min_duration]
|
||||||
|
if limit:
|
||||||
|
filtered = filtered[:limit]
|
||||||
|
return filtered
|
||||||
|
|
||||||
|
|
||||||
|
@cli.command("scrape")
|
||||||
|
@click.option("--limit", type=int, default=None)
|
||||||
|
@click.option("--since", default=None, help="Solo desde YYYY-MM-DD")
|
||||||
|
@click.option("--languages", "-l", default=None)
|
||||||
|
@click.option("--no-auto", is_flag=True, default=False)
|
||||||
|
@click.option("--no-shorts", is_flag=True, default=None)
|
||||||
|
@click.option("--include-shorts", is_flag=True, default=None)
|
||||||
|
@click.option("--no-live", is_flag=True, default=None)
|
||||||
|
@click.option("--resume/--no-resume", default=True)
|
||||||
|
@click.option("--dry-run", is_flag=True, default=False)
|
||||||
|
@click.option("--reset-errors", is_flag=True, default=False)
|
||||||
|
@click.pass_obj
|
||||||
|
def scrape(obj, limit, since, languages, no_auto, no_shorts, include_shorts, no_live, resume, dry_run, reset_errors):
|
||||||
|
"""Scrape completo: discovery + extraccion + markdown."""
|
||||||
|
_run_scrape(obj, limit, since, languages, no_auto, no_shorts, include_shorts, no_live, resume, dry_run, reset_errors)
|
||||||
|
|
||||||
|
|
||||||
|
def _run_scrape(obj, limit, since, languages, no_auto, no_shorts, include_shorts, no_live, resume, dry_run, reset_errors):
|
||||||
|
cfg: Config = obj.cfg
|
||||||
|
store: Store = obj.store
|
||||||
|
if not cfg.channel_url:
|
||||||
|
console.print("[red]Error:[/red] falta la URL del canal. Usa --channel o config.yaml")
|
||||||
|
sys.exit(1)
|
||||||
|
if languages:
|
||||||
|
cfg.languages = [l.strip() for l in languages.split(",") if l.strip()]
|
||||||
|
if no_auto:
|
||||||
|
cfg.prefer_manual = True
|
||||||
|
if include_shorts:
|
||||||
|
cfg.include_shorts = True
|
||||||
|
if no_shorts is True:
|
||||||
|
cfg.include_shorts = False
|
||||||
|
if no_live is True:
|
||||||
|
cfg.include_live = False
|
||||||
|
|
||||||
|
console.print(f"[cyan]Canal:[/cyan] {cfg.channel_url}\n[cyan]DB:[/cyan] {cfg.database_path_resolved}")
|
||||||
|
console.print("\n[bold blue]Paso 1:[/bold blue] Discovery")
|
||||||
|
try:
|
||||||
|
with Progress(SpinnerColumn(), TextColumn("[progress.description]{task.description}"), transient=True) as prog:
|
||||||
|
task = prog.add_task("Descubriendo videos...", total=None)
|
||||||
|
channel_id, channel_name, refs = discover_channel(cfg.channel_url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests)
|
||||||
|
prog.update(task, completed=1, total=1)
|
||||||
|
except Exception as exc:
|
||||||
|
console.print(f"[red]Error en discovery:[/red] {exc}")
|
||||||
|
sys.exit(1)
|
||||||
|
|
||||||
|
store.upsert_channel(channel_id, _extract_handle(cfg.channel_url), channel_name, len(refs))
|
||||||
|
console.print(f"[green]Canal:[/green] {channel_name} ({channel_id}) — {len(refs)} videos")
|
||||||
|
|
||||||
|
refs = _apply_filters(refs, since, not cfg.include_shorts, not cfg.include_live, cfg.min_duration_sec, limit)
|
||||||
|
console.print(f"[yellow]Tras filtros:[/yellow] {len(refs)} videos")
|
||||||
|
store.upsert_videos(refs)
|
||||||
|
|
||||||
|
if reset_errors:
|
||||||
|
n = store.reset_errors(channel_id)
|
||||||
|
if n:
|
||||||
|
console.print(f"[yellow]Resetados {n} videos con error.[/yellow]")
|
||||||
|
if dry_run:
|
||||||
|
_print_dry_run(refs)
|
||||||
|
return
|
||||||
|
|
||||||
|
pending = store.get_pending(channel_id) if resume else store.get_all(channel_id)
|
||||||
|
ref_ids = {r.video_id for r in refs}
|
||||||
|
pending = [p for p in pending if p.video_id in ref_ids] if ref_ids else pending
|
||||||
|
if limit:
|
||||||
|
pending = pending[:limit]
|
||||||
|
if not pending:
|
||||||
|
console.print("[green]No hay videos pendientes.[/green]")
|
||||||
|
_print_stats(store, channel_id)
|
||||||
|
return
|
||||||
|
|
||||||
|
console.print(f"\n[bold blue]Paso 2:[/bold blue] Extraccion ({len(pending)} videos)")
|
||||||
|
cookie_path = obj.cookies or resolve_active_path(store)
|
||||||
|
if cookie_path and not obj.cookies:
|
||||||
|
console.print(f"[dim]Usando cookie del vault: {Path(cookie_path).name}[/dim]")
|
||||||
|
|
||||||
|
with Progress(SpinnerColumn(), TextColumn("[progress.description]{task.description}"),
|
||||||
|
BarColumn(), TaskProgressColumn(), TimeRemainingColumn()) as progress:
|
||||||
|
task = progress.add_task("Procesando", total=len(pending))
|
||||||
|
for i, row in enumerate(pending):
|
||||||
|
progress.update(task, description=f"{row.video_id} {(row.title or '')[:30]}", completed=i)
|
||||||
|
process_video(row, cfg, store, channel_name, channel_id, cfg.channel_url,
|
||||||
|
cookies_file=cookie_path, cookies_from_browser=obj.cookies_from_browser)
|
||||||
|
progress.advance(task)
|
||||||
|
if i < len(pending) - 1:
|
||||||
|
polite_sleep(cfg.delay.min_seconds, cfg.delay.max_seconds)
|
||||||
|
console.print()
|
||||||
|
_print_stats(store, channel_id)
|
||||||
|
|
||||||
|
|
||||||
|
def _print_dry_run(refs):
|
||||||
|
table = Table(show_lines=False)
|
||||||
|
table.add_column("Fecha", style="dim")
|
||||||
|
table.add_column("Video ID", style="cyan")
|
||||||
|
table.add_column("Duracion")
|
||||||
|
table.add_column("Titulo")
|
||||||
|
for r in refs[:50]:
|
||||||
|
table.add_row(_fmt_date(r.upload_date), r.video_id, _fmt_duration(r.duration), (r.title or "")[:60])
|
||||||
|
console.print(table)
|
||||||
|
if len(refs) > 50:
|
||||||
|
console.print(f"... y {len(refs) - 50} mas")
|
||||||
|
|
||||||
|
|
||||||
|
# --------------------------------------------------------------------------- search
|
||||||
|
|
||||||
|
@cli.command("search")
|
||||||
|
@click.argument("query")
|
||||||
|
@click.option("-l", "--limit", type=int, default=20)
|
||||||
|
@click.pass_obj
|
||||||
|
def search_cmd(obj, query, limit):
|
||||||
|
"""Busqueda full-text en transcripciones."""
|
||||||
|
store: Store = obj.store
|
||||||
|
channels = _channel_targets(obj)
|
||||||
|
hits = store.search_segments(query, channel_id=channels[0] if len(channels) == 1 else None, limit=limit)
|
||||||
|
if not hits:
|
||||||
|
console.print("[yellow]Sin resultados.[/yellow]")
|
||||||
|
return
|
||||||
|
table = Table(show_lines=True)
|
||||||
|
table.add_column("Video", style="cyan")
|
||||||
|
table.add_column("Tiempo")
|
||||||
|
table.add_column("Snippet")
|
||||||
|
for h in hits:
|
||||||
|
table.add_row(f"{h.video_id} {(_title_for(store, h.video_id) or h.title or '')[:40]}",
|
||||||
|
seconds_to_ts(h.start_sec), h.snippet[:120])
|
||||||
|
console.print(table)
|
||||||
|
|
||||||
|
|
||||||
|
# --------------------------------------------------------------------------- export
|
||||||
|
|
||||||
|
@cli.command("export")
|
||||||
|
@click.option("--format", "fmt", type=click.Choice(["json", "csv", "srt", "html"]), default="json")
|
||||||
|
@click.option("-o", "--out", "out_dir", default="data/exports")
|
||||||
|
@click.pass_obj
|
||||||
|
def export_cmd(obj, fmt, out_dir):
|
||||||
|
"""Export multi-formato."""
|
||||||
|
from . import export as export_mod
|
||||||
|
store: Store = obj.store
|
||||||
|
targets = _channel_targets(obj)
|
||||||
|
out_dir = Path(out_dir)
|
||||||
|
out_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
for cid in targets:
|
||||||
|
label = cid or "all"
|
||||||
|
if fmt == "json":
|
||||||
|
p = export_mod.export_json(store, cid, out_dir / f"{label}.json")
|
||||||
|
elif fmt == "csv":
|
||||||
|
p = export_mod.export_csv(store, cid, out_dir / f"{label}.csv")
|
||||||
|
elif fmt == "srt":
|
||||||
|
files = export_mod.export_srt(store, cid, out_dir / "srt")
|
||||||
|
console.print(f"[green]SRT:[/green] {len(files)} archivos en {out_dir / 'srt'}")
|
||||||
|
continue
|
||||||
|
else:
|
||||||
|
p = export_mod.export_html(store, cid, out_dir / f"{label}.html")
|
||||||
|
console.print(f"[green]{fmt.upper()}:[/green] {p}")
|
||||||
|
|
||||||
|
|
||||||
|
# --------------------------------------------------------------------------- audio
|
||||||
|
|
||||||
|
@cli.command("audio")
|
||||||
|
@click.option("--limit", type=int, default=None)
|
||||||
|
@click.option("--from", "since", default=None, help="Solo desde YYYY-MM-DD")
|
||||||
|
@click.option("--force", is_flag=True, default=False, help="Re-descargar aunque exista")
|
||||||
|
@click.pass_obj
|
||||||
|
def audio_cmd(obj, limit, since, force):
|
||||||
|
"""Descarga audio MP3 (requiere ffmpeg)."""
|
||||||
|
if not shutil.which("ffmpeg"):
|
||||||
|
console.print("[red]ffmpeg no encontrado.[/red] Instala: winget install ffmpeg")
|
||||||
|
sys.exit(1)
|
||||||
|
store: Store = obj.store
|
||||||
|
cfg: Config = obj.cfg
|
||||||
|
cookie_path = obj.cookies or resolve_active_path(store)
|
||||||
|
out_dir = Path(cfg.output_dir_resolved).parent / "audio"
|
||||||
|
out_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
import yt_dlp
|
||||||
|
targets = _channel_targets(obj)
|
||||||
|
videos = [v for cid in targets for v in store.get_all(cid) if v.status == "done"]
|
||||||
|
if since:
|
||||||
|
videos = [v for v in videos if (v.upload_date or "") >= since.replace("-", "")]
|
||||||
|
if limit:
|
||||||
|
videos = videos[:limit]
|
||||||
|
console.print(f"[cyan]Descargando {len(videos)} audios -> {out_dir}[/cyan]")
|
||||||
|
ydl_opts = {
|
||||||
|
"format": "bestaudio/best",
|
||||||
|
"outtmpl": str(out_dir / "%(title)s.%(ext)s"),
|
||||||
|
"postprocessors": [{"key": "FFmpegExtractAudio", "preferredcodec": "mp3", "preferredquality": "128"}],
|
||||||
|
"quiet": True, "no_warnings": True, "noprogress": True,
|
||||||
|
}
|
||||||
|
if cookie_path:
|
||||||
|
ydl_opts["cookiefile"] = cookie_path
|
||||||
|
if obj.cookies_from_browser:
|
||||||
|
ydl_opts["cookiesfrombrowser"] = (obj.cookies_from_browser,)
|
||||||
|
with yt_dlp.YoutubeDL(ydl_opts) as ydl:
|
||||||
|
for v in videos:
|
||||||
|
target = out_dir / f"{_safe_filename(v.title or v.video_id)}.mp3"
|
||||||
|
if target.exists() and not force:
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
ydl.download([v.url])
|
||||||
|
console.print(f"[green]OK[/green] {v.video_id}")
|
||||||
|
except Exception as exc:
|
||||||
|
console.print(f"[red]FAIL[/red] {v.video_id}: {exc}")
|
||||||
|
|
||||||
|
|
||||||
|
# --------------------------------------------------------------------------- channels
|
||||||
|
|
||||||
|
@cli.group("channels")
|
||||||
|
def channels_grp():
|
||||||
|
"""Gestion de canales."""
|
||||||
|
|
||||||
|
|
||||||
|
@channels_grp.command("list")
|
||||||
|
@click.pass_obj
|
||||||
|
def channels_list(obj):
|
||||||
|
store: Store = obj.store
|
||||||
|
table = Table(title="Canales")
|
||||||
|
table.add_column("ID", style="cyan")
|
||||||
|
table.add_column("Handle")
|
||||||
|
table.add_column("Nombre")
|
||||||
|
table.add_column("Videos", justify="right")
|
||||||
|
for ch in store.list_channels():
|
||||||
|
table.add_row(ch["channel_id"], ch.get("handle") or "", ch.get("name") or "", str(ch.get("video_count") or 0))
|
||||||
|
console.print(table)
|
||||||
|
|
||||||
|
|
||||||
|
@channels_grp.command("add")
|
||||||
|
@click.argument("url")
|
||||||
|
@click.pass_obj
|
||||||
|
def channels_add(obj, url):
|
||||||
|
cfg: Config = obj.cfg
|
||||||
|
cfg.channel_url = url
|
||||||
|
console.print("[cyan]Resolviendo canal...[/cyan]")
|
||||||
|
channel_id, name, refs = discover_channel(url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests)
|
||||||
|
obj.store.upsert_channel(channel_id, _extract_handle(url), name, len(refs))
|
||||||
|
obj.store.upsert_videos(refs)
|
||||||
|
console.print(f"[green]Added:[/green] {name} ({channel_id}) — {len(refs)} videos")
|
||||||
|
|
||||||
|
|
||||||
|
@channels_grp.command("remove")
|
||||||
|
@click.argument("ident")
|
||||||
|
@click.pass_obj
|
||||||
|
def channels_remove(obj, ident):
|
||||||
|
store: Store = obj.store
|
||||||
|
match = None
|
||||||
|
for ch in store.list_channels():
|
||||||
|
if ident in (ch["channel_id"], (ch.get("handle") or "").lstrip("@"), ch.get("name")):
|
||||||
|
match = ch["channel_id"]
|
||||||
|
break
|
||||||
|
if not match:
|
||||||
|
console.print(f"[red]No encontrado:[/red] {ident}")
|
||||||
|
sys.exit(1)
|
||||||
|
store.delete_channel(match)
|
||||||
|
console.print(f"[green]Removed:[/green] {match}")
|
||||||
|
|
||||||
|
|
||||||
|
# --------------------------------------------------------------------------- watch
|
||||||
|
|
||||||
|
@cli.command("watch")
|
||||||
|
@click.option("--interval", default="6h", help="Ej: 30m, 6h, 1d")
|
||||||
|
@click.option("--once", is_flag=True, default=False)
|
||||||
|
@click.pass_obj
|
||||||
|
def watch_cmd(obj, interval, once):
|
||||||
|
"""Monitoreo de nuevos videos."""
|
||||||
|
from .monitor import watch_loop
|
||||||
|
sec = _parse_interval(interval)
|
||||||
|
console.print(f"[cyan]Watch mode:[/cyan] intervalo {sec}s {'(once)' if once else '(loop)'}")
|
||||||
|
watch_loop(obj.cfg, obj.store, sec, channel_id=None, once=once,
|
||||||
|
cookies_file=obj.cookies, cookies_from_browser=obj.cookies_from_browser,
|
||||||
|
on_log=lambda m: console.print(f"[dim]{m}[/dim]"))
|
||||||
|
|
||||||
|
|
||||||
|
# --------------------------------------------------------------------------- analyze
|
||||||
|
|
||||||
|
@cli.command("analyze")
|
||||||
|
@click.option("--wordcloud", is_flag=True, default=False)
|
||||||
|
@click.option("--top-words", type=int, default=None)
|
||||||
|
@click.option("--timeline", default=None, help="Termino a rastrear en el tiempo")
|
||||||
|
@click.option("-o", "--out", "out_dir", default="data/analysis")
|
||||||
|
@click.pass_obj
|
||||||
|
def analyze_cmd(obj, wordcloud, top_words, timeline, out_dir):
|
||||||
|
"""Analisis estadistico de contenido (frecuencia, timeline)."""
|
||||||
|
from . import analysis as A
|
||||||
|
store: Store = obj.store
|
||||||
|
out_dir = Path(out_dir)
|
||||||
|
out_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
targets = _channel_targets(obj)
|
||||||
|
cid = targets[0] if len(targets) == 1 else None
|
||||||
|
if wordcloud or (top_words is None and not timeline):
|
||||||
|
top_words = 50
|
||||||
|
if top_words:
|
||||||
|
tw = A.top_words(store, top_words, cid)
|
||||||
|
import csv
|
||||||
|
with open(out_dir / "top_words.csv", "w", encoding="utf-8", newline="") as f:
|
||||||
|
w = csv.writer(f)
|
||||||
|
w.writerow(["word", "count"])
|
||||||
|
w.writerows(tw)
|
||||||
|
console.print(f"[green]top_words.csv[/green] ({len(tw)} terms)")
|
||||||
|
try:
|
||||||
|
A.render_top_words_chart(tw, out_dir / "top_words.png")
|
||||||
|
console.print(f"[green]top_words.png[/green]")
|
||||||
|
except Exception as exc:
|
||||||
|
console.print(f"[yellow]chart skip:[/yellow] {exc}")
|
||||||
|
if wordcloud:
|
||||||
|
freq = dict(tw)
|
||||||
|
if freq:
|
||||||
|
A.render_wordcloud(freq, out_dir / "wordcloud.png")
|
||||||
|
console.print(f"[green]wordcloud.png[/green]")
|
||||||
|
if timeline:
|
||||||
|
tl = A.term_timeline(store, timeline, cid)
|
||||||
|
console.print(f"[cyan]Timeline '{timeline}':[/cyan] {len(tl)} meses")
|
||||||
|
for month, n in tl:
|
||||||
|
console.print(f" {month}: {n}")
|
||||||
|
try:
|
||||||
|
A.render_timeline_chart(tl, timeline, out_dir / "timeline.png")
|
||||||
|
console.print(f"[green]timeline.png[/green]")
|
||||||
|
except Exception as exc:
|
||||||
|
console.print(f"[yellow]chart skip:[/yellow] {exc}")
|
||||||
|
|
||||||
|
|
||||||
|
# --------------------------------------------------------------------------- re-render
|
||||||
|
|
||||||
|
@cli.command("re-render")
|
||||||
|
@click.option("--backfill", is_flag=True, default=False, help="Importar segmentos desde .md antes")
|
||||||
|
@click.pass_obj
|
||||||
|
def re_render_cmd(obj, backfill):
|
||||||
|
"""Regenerar Markdown desde segmentos almacenados."""
|
||||||
|
store: Store = obj.store
|
||||||
|
cfg: Config = obj.cfg
|
||||||
|
from .segments import backfill_from_markdown
|
||||||
|
if backfill:
|
||||||
|
md_root = Path(cfg.output_dir_resolved)
|
||||||
|
n = backfill_from_markdown(store, md_root, log=lambda m: console.print(f"[dim]{m}[/dim]"))
|
||||||
|
console.print(f"[green]Backfill:[/green] {n} videos")
|
||||||
|
targets = _channel_targets(obj)
|
||||||
|
videos = [v for cid in targets for v in store.get_all(cid) if v.status == "done" and v.segments_json]
|
||||||
|
if not videos:
|
||||||
|
console.print("[yellow]No hay videos con segments_json. Usa --backfill.[/yellow]")
|
||||||
|
return
|
||||||
|
import json
|
||||||
|
from .render import render_markdown
|
||||||
|
console.print(f"[cyan]Re-renderizando {len(videos)} videos...[/cyan]")
|
||||||
|
for v in videos:
|
||||||
|
segs = [Segment(start=s["start"], end=s["end"], text=s["text"]) for s in json.loads(v.segments_json)]
|
||||||
|
chapters = [Chapter(title=c["title"], start_time=c["start"], end_time=c.get("end", c["start"])) for c in json.loads(v.chapters_json or "[]")]
|
||||||
|
sections = align_chapters(segs, chapters)
|
||||||
|
context = {
|
||||||
|
"video_id": v.video_id, "title": v.title or v.video_id, "channel_name": "",
|
||||||
|
"channel_id": v.channel_id, "channel_url": "", "upload_date": v.upload_date or "",
|
||||||
|
"duration": v.duration or 0, "url": v.url, "transcript_lang": v.transcript_lang or "",
|
||||||
|
"transcript_src": v.transcript_src or "", "view_count": v.view_count, "like_count": v.like_count,
|
||||||
|
"tags": _parse_tags(v.tags), "thumbnail": v.thumbnail or "", "description": v.description or "",
|
||||||
|
"sections": sections,
|
||||||
|
}
|
||||||
|
stem = build_filename_stem(v.upload_date, v.title or v.video_id, cfg.filename_template)
|
||||||
|
ch = store.get_channel(v.channel_id)
|
||||||
|
out_subdir = Path(cfg.output_dir_resolved) / _safe_dirname((ch or {}).get("name") or "unknown")
|
||||||
|
render_markdown(cfg.template_path_resolved, out_subdir, stem, context)
|
||||||
|
console.print(f"[green]Re-render completo.[/green]")
|
||||||
|
|
||||||
|
|
||||||
|
# --------------------------------------------------------------------------- helpers
|
||||||
|
|
||||||
|
def _channel_targets(obj) -> list[str | None]:
|
||||||
|
if obj.all_channels:
|
||||||
|
return [ch["channel_id"] for ch in obj.store.list_channels()]
|
||||||
|
if obj.channel:
|
||||||
|
# try resolve a channel id from store by handle/name
|
||||||
|
for ch in obj.store.list_channels():
|
||||||
|
if obj.channel in (ch["channel_id"], (ch.get("handle") or "").lstrip("@"), ch.get("name")):
|
||||||
|
return [ch["channel_id"]]
|
||||||
|
return [None]
|
||||||
|
return [None]
|
||||||
|
|
||||||
|
|
||||||
|
def _print_stats(store: Store, channel_id: str) -> None:
|
||||||
|
stats = store.stats(channel_id)
|
||||||
|
table = Table(title="Resumen")
|
||||||
|
table.add_column("Estado", style="bold")
|
||||||
|
table.add_column("Cantidad", justify="right")
|
||||||
|
for status in ("done", "no_subtitles", "error", "pending"):
|
||||||
|
if status in stats:
|
||||||
|
color = {"done": "green", "no_subtitles": "yellow", "error": "red", "pending": "cyan"}.get(status, "white")
|
||||||
|
table.add_row(f"[{color}]{status}[/{color}]", str(stats[status]))
|
||||||
|
console.print(table)
|
||||||
|
|
||||||
|
|
||||||
|
def _title_for(store: Store, video_id: str) -> str | None:
|
||||||
|
v = store.get_video(video_id)
|
||||||
|
return v.title if v else None
|
||||||
|
|
||||||
|
|
||||||
|
def _extract_handle(url: str) -> str:
|
||||||
|
if "@" in url:
|
||||||
|
return "@" + url.split("@", 1)[1].split("/", 1)[0]
|
||||||
|
return ""
|
||||||
|
|
||||||
|
|
||||||
|
def _fmt_date(d: str | None) -> str:
|
||||||
|
if not d:
|
||||||
|
return ""
|
||||||
|
if len(d) == 8:
|
||||||
|
return f"{d[:4]}-{d[4:6]}-{d[6:8]}"
|
||||||
|
return d
|
||||||
|
|
||||||
|
|
||||||
|
def _fmt_duration(seconds: int | None) -> str:
|
||||||
|
if not seconds:
|
||||||
|
return ""
|
||||||
|
m, s = divmod(int(seconds), 60)
|
||||||
|
h, m = divmod(m, 60)
|
||||||
|
return f"{h}:{m:02d}:{s:02d}" if h else f"{m}:{s:02d}"
|
||||||
|
|
||||||
|
|
||||||
|
def seconds_to_ts(sec: float) -> str:
|
||||||
|
total = int(sec)
|
||||||
|
h, rem = divmod(total, 3600)
|
||||||
|
m, s = divmod(rem, 60)
|
||||||
|
return f"{h:d}:{m:02d}:{s:02d}" if h else f"{m:d}:{s:02d}"
|
||||||
|
|
||||||
|
|
||||||
|
def _safe_dirname(name: str) -> str:
|
||||||
|
safe = "".join(c for c in name if c not in r'\/:*?"<>|')
|
||||||
|
return safe.strip().strip(".") or "unknown"
|
||||||
|
|
||||||
|
|
||||||
|
def _safe_filename(name: str) -> str:
|
||||||
|
safe = "".join(c for c in name if c not in r'\/:*?"<>|')
|
||||||
|
return safe.strip().strip(".") or "untitled"
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_tags(tags_json: str | None) -> list[str]:
|
||||||
|
if not tags_json:
|
||||||
|
return []
|
||||||
|
import json
|
||||||
|
try:
|
||||||
|
return json.loads(tags_json)
|
||||||
|
except (json.JSONDecodeError, TypeError):
|
||||||
|
return []
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_interval(s: str) -> float:
|
||||||
|
s = s.strip().lower()
|
||||||
|
if s.endswith("s"):
|
||||||
|
return float(s[:-1])
|
||||||
|
if s.endswith("m"):
|
||||||
|
return float(s[:-1]) * 60
|
||||||
|
if s.endswith("h"):
|
||||||
|
return float(s[:-1]) * 3600
|
||||||
|
if s.endswith("d"):
|
||||||
|
return float(s[:-1]) * 86400
|
||||||
|
return float(s)
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
cli()
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -0,0 +1,85 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from dataclasses import dataclass, field
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
import yaml
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class DelayConfig:
|
||||||
|
min_seconds: float = 1.5
|
||||||
|
max_seconds: float = 3.5
|
||||||
|
backoff_base: float = 2.0
|
||||||
|
backoff_cap: float = 60.0
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class YtDlpConfig:
|
||||||
|
retries: int = 10
|
||||||
|
sleep_subrequests: float = 2.0
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class Config:
|
||||||
|
channel_url: str = ""
|
||||||
|
languages: list[str] = field(default_factory=lambda: ["es", "en"])
|
||||||
|
prefer_manual: bool = True
|
||||||
|
include_shorts: bool = False
|
||||||
|
include_live: bool = True
|
||||||
|
min_duration_sec: int = 0
|
||||||
|
delay: DelayConfig = field(default_factory=DelayConfig)
|
||||||
|
yt_dlp: YtDlpConfig = field(default_factory=YtDlpConfig)
|
||||||
|
database_path: str = "data/state.db"
|
||||||
|
output_dir: str = "data/markdown"
|
||||||
|
template_path: str = "templates/video.md.j2"
|
||||||
|
filename_template: str = "{upload_date}_{slug}"
|
||||||
|
|
||||||
|
@property
|
||||||
|
def database_path_resolved(self) -> Path:
|
||||||
|
return Path(self.database_path).resolve()
|
||||||
|
|
||||||
|
@property
|
||||||
|
def output_dir_resolved(self) -> Path:
|
||||||
|
return Path(self.output_dir).resolve()
|
||||||
|
|
||||||
|
@property
|
||||||
|
def template_path_resolved(self) -> Path:
|
||||||
|
return Path(self.template_path).resolve()
|
||||||
|
|
||||||
|
|
||||||
|
def load_config(path: str | Path) -> Config:
|
||||||
|
p = Path(path)
|
||||||
|
if not p.exists():
|
||||||
|
raise FileNotFoundError(f"Config file not found: {p}")
|
||||||
|
with open(p, encoding="utf-8") as f:
|
||||||
|
raw: dict[str, Any] = yaml.safe_load(f) or {}
|
||||||
|
return _build_config(raw)
|
||||||
|
|
||||||
|
|
||||||
|
def _build_config(raw: dict[str, Any]) -> Config:
|
||||||
|
delay_raw = raw.get("delay") or {}
|
||||||
|
ydl_raw = raw.get("yt_dlp") or {}
|
||||||
|
return Config(
|
||||||
|
channel_url=raw.get("channel_url", ""),
|
||||||
|
languages=list(raw.get("languages", ["es", "en"])),
|
||||||
|
prefer_manual=bool(raw.get("prefer_manual", True)),
|
||||||
|
include_shorts=bool(raw.get("include_shorts", False)),
|
||||||
|
include_live=bool(raw.get("include_live", True)),
|
||||||
|
min_duration_sec=int(raw.get("min_duration_sec", 0)),
|
||||||
|
delay=DelayConfig(
|
||||||
|
min_seconds=float(delay_raw.get("min_seconds", 1.5)),
|
||||||
|
max_seconds=float(delay_raw.get("max_seconds", 3.5)),
|
||||||
|
backoff_base=float(delay_raw.get("backoff_base", 2.0)),
|
||||||
|
backoff_cap=float(delay_raw.get("backoff_cap", 60.0)),
|
||||||
|
),
|
||||||
|
yt_dlp=YtDlpConfig(
|
||||||
|
retries=int(ydl_raw.get("retries", 10)),
|
||||||
|
sleep_subrequests=float(ydl_raw.get("sleep_subrequests", 2.0)),
|
||||||
|
),
|
||||||
|
database_path=raw.get("database_path", "data/state.db"),
|
||||||
|
output_dir=raw.get("output_dir", "data/markdown"),
|
||||||
|
template_path=raw.get("template_path", "templates/video.md.j2"),
|
||||||
|
filename_template=raw.get("filename_template", "{upload_date}_{slug}"),
|
||||||
|
)
|
||||||
@@ -0,0 +1,161 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import logging
|
||||||
|
import uuid
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from .store import CookieRow, Store
|
||||||
|
|
||||||
|
|
||||||
|
log = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
SESSION_COOKIE_NAMES = {"SID", "SAPISID", "__Secure-3PSID", "SSID", "LOGIN_INFO", "HSID", "APISID"}
|
||||||
|
|
||||||
|
_DEFAULT_DIR = Path("cookies")
|
||||||
|
|
||||||
|
|
||||||
|
def cookies_dir(override: str | Path | None = None) -> Path:
|
||||||
|
d = Path(override) if override else _DEFAULT_DIR
|
||||||
|
d.mkdir(parents=True, exist_ok=True)
|
||||||
|
return d.resolve()
|
||||||
|
|
||||||
|
|
||||||
|
def parse_netscape(text: str) -> tuple[bool, dict]:
|
||||||
|
"""Parse Netscape cookie text. Returns (ok, info).
|
||||||
|
|
||||||
|
info = {count, has_session, expires_at (ISO or None), names: set}.
|
||||||
|
"""
|
||||||
|
lines = text.splitlines()
|
||||||
|
expiries: list[int] = []
|
||||||
|
names: set[str] = set()
|
||||||
|
count = 0
|
||||||
|
for line in lines:
|
||||||
|
line = line.rstrip("\n")
|
||||||
|
if not line.strip() or line.startswith("#"):
|
||||||
|
continue
|
||||||
|
parts = line.split("\t")
|
||||||
|
if len(parts) < 7:
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
expiry = int(parts[4])
|
||||||
|
except ValueError:
|
||||||
|
expiry = 0
|
||||||
|
name = parts[5].strip()
|
||||||
|
if not name:
|
||||||
|
continue
|
||||||
|
names.add(name)
|
||||||
|
count += 1
|
||||||
|
if expiry:
|
||||||
|
expiries.append(expiry)
|
||||||
|
if count == 0:
|
||||||
|
return False, {"count": 0, "has_session": False, "expires_at": None, "names": set()}
|
||||||
|
has_session = bool(names & SESSION_COOKIE_NAMES)
|
||||||
|
expires_at = None
|
||||||
|
if expiries:
|
||||||
|
earliest = min(expiries)
|
||||||
|
if earliest > 0:
|
||||||
|
expires_at = datetime.fromtimestamp(earliest, tz=timezone.utc).isoformat(timespec="seconds")
|
||||||
|
return True, {"count": count, "has_session": has_session, "expires_at": expires_at, "names": names}
|
||||||
|
|
||||||
|
|
||||||
|
def parse_netscape_file(path: str | Path) -> tuple[bool, dict]:
|
||||||
|
p = Path(path)
|
||||||
|
if not p.exists():
|
||||||
|
return False, {"count": 0, "has_session": False, "expires_at": None, "names": set()}
|
||||||
|
try:
|
||||||
|
text = p.read_text(encoding="utf-8")
|
||||||
|
except OSError as exc:
|
||||||
|
log.warning("cannot read cookie file %s: %s", path, exc)
|
||||||
|
return False, {"count": 0, "has_session": False, "expires_at": None, "names": set()}
|
||||||
|
return parse_netscape(text)
|
||||||
|
|
||||||
|
|
||||||
|
def import_file(store: Store, path: str | Path, label: str | None = None, cookie_dir: str | Path | None = None) -> str:
|
||||||
|
src = Path(path)
|
||||||
|
text = src.read_text(encoding="utf-8")
|
||||||
|
return import_text(store, text, label=label or src.stem, cookie_dir=cookie_dir)
|
||||||
|
|
||||||
|
|
||||||
|
def import_text(store: Store, text: str, label: str, cookie_dir: str | Path | None = None) -> str:
|
||||||
|
ok, info = parse_netscape(text)
|
||||||
|
if not ok:
|
||||||
|
raise ValueError("Invalid Netscape cookie file: no parseable cookie lines")
|
||||||
|
cookie_id = str(uuid.uuid4())
|
||||||
|
cdir = cookies_dir(cookie_dir)
|
||||||
|
out_path = cdir / f"{cookie_id}.txt"
|
||||||
|
out_path.write_text(text, encoding="utf-8")
|
||||||
|
store.upsert_cookie(
|
||||||
|
cookie_id=cookie_id,
|
||||||
|
filename=out_path.name,
|
||||||
|
label=label,
|
||||||
|
expires_at=info["expires_at"],
|
||||||
|
has_session=info["has_session"],
|
||||||
|
cookie_count=info["count"],
|
||||||
|
)
|
||||||
|
return cookie_id
|
||||||
|
|
||||||
|
|
||||||
|
def auto_import_dir(store: Store, dir_path: str | Path | None = None) -> int:
|
||||||
|
"""Import any loose .txt Netscape files in dir that aren't tracked yet. Returns count."""
|
||||||
|
cdir = cookies_dir(dir_path)
|
||||||
|
tracked = {c.filename for c in store.list_cookies()}
|
||||||
|
n = 0
|
||||||
|
for f in sorted(cdir.glob("*.txt")):
|
||||||
|
if f.name in tracked:
|
||||||
|
continue
|
||||||
|
ok, info = parse_netscape_file(f)
|
||||||
|
if not ok:
|
||||||
|
continue
|
||||||
|
cookie_id = f.stem
|
||||||
|
# store the tracked id == filename stem; keep file in place
|
||||||
|
store.upsert_cookie(
|
||||||
|
cookie_id=cookie_id,
|
||||||
|
filename=f.name,
|
||||||
|
label=f.stem,
|
||||||
|
expires_at=info["expires_at"],
|
||||||
|
has_session=info["has_session"],
|
||||||
|
cookie_count=info["count"],
|
||||||
|
)
|
||||||
|
n += 1
|
||||||
|
# activate first cookie if none active
|
||||||
|
if not store.get_active_cookie():
|
||||||
|
cookies = store.list_cookies()
|
||||||
|
if cookies:
|
||||||
|
store.set_active_cookie(cookies[0].id)
|
||||||
|
return n
|
||||||
|
|
||||||
|
|
||||||
|
def set_active(store: Store, cookie_id: str) -> None:
|
||||||
|
store.set_active_cookie(cookie_id)
|
||||||
|
|
||||||
|
|
||||||
|
def list_vault(store: Store) -> list[CookieRow]:
|
||||||
|
return store.list_cookies()
|
||||||
|
|
||||||
|
|
||||||
|
def delete(store: Store, cookie_id: str, cookie_dir: str | Path | None = None) -> None:
|
||||||
|
row = store.get_cookie(cookie_id)
|
||||||
|
store.delete_cookie(cookie_id)
|
||||||
|
if row:
|
||||||
|
cdir = cookies_dir(cookie_dir)
|
||||||
|
(cdir / row.filename).unlink(missing_ok=True)
|
||||||
|
|
||||||
|
|
||||||
|
def resolve_active_path(store: Store, cookie_dir: str | Path | None = None) -> str | None:
|
||||||
|
row = store.get_active_cookie()
|
||||||
|
if not row:
|
||||||
|
return None
|
||||||
|
cdir = cookies_dir(cookie_dir)
|
||||||
|
path = cdir / row.filename
|
||||||
|
return str(path) if path.exists() else None
|
||||||
|
|
||||||
|
|
||||||
|
def is_expired(row: CookieRow) -> bool:
|
||||||
|
if not row.expires_at:
|
||||||
|
return False
|
||||||
|
try:
|
||||||
|
exp = datetime.fromisoformat(row.expires_at)
|
||||||
|
except ValueError:
|
||||||
|
return False
|
||||||
|
return datetime.now(timezone.utc) > exp
|
||||||
@@ -0,0 +1,68 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import logging
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
import yt_dlp
|
||||||
|
|
||||||
|
from .store import VideoRef
|
||||||
|
|
||||||
|
log = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
|
||||||
|
def discover_channel(channel_url: str, sleep_subrequests: float = 2.0) -> tuple[str, str, list[VideoRef]]:
|
||||||
|
"""Returns (channel_id, channel_name, video_refs)."""
|
||||||
|
ydl_opts: dict[str, Any] = {
|
||||||
|
"extract_flat": "in_playlist",
|
||||||
|
"quiet": True,
|
||||||
|
"no_warnings": True,
|
||||||
|
"skip_download": True,
|
||||||
|
"extract_flat_args": None,
|
||||||
|
"sleep_subrequests": sleep_subrequests,
|
||||||
|
}
|
||||||
|
|
||||||
|
with yt_dlp.YoutubeDL(ydl_opts) as ydl:
|
||||||
|
info = ydl.extract_info(channel_url, download=False)
|
||||||
|
|
||||||
|
channel_id = info.get("channel_id") or info.get("uploader_id") or info.get("id") or "unknown"
|
||||||
|
channel_name = info.get("channel") or info.get("title") or info.get("uploader") or "Unknown"
|
||||||
|
|
||||||
|
entries = _flatten_entries(info)
|
||||||
|
refs: list[VideoRef] = []
|
||||||
|
for entry in entries:
|
||||||
|
video_id = entry.get("id")
|
||||||
|
if not video_id:
|
||||||
|
continue
|
||||||
|
if entry.get("_type") == "playlist":
|
||||||
|
continue
|
||||||
|
url = entry.get("url") or f"https://www.youtube.com/watch?v={video_id}"
|
||||||
|
upload_date = entry.get("upload_date")
|
||||||
|
duration = entry.get("duration")
|
||||||
|
title = entry.get("title") or video_id
|
||||||
|
refs.append(
|
||||||
|
VideoRef(
|
||||||
|
video_id=str(video_id),
|
||||||
|
channel_id=str(channel_id),
|
||||||
|
title=str(title),
|
||||||
|
url=str(url),
|
||||||
|
upload_date=str(upload_date) if upload_date else None,
|
||||||
|
duration=int(duration) if duration else None,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
log.info("Discovered %d videos on channel %s (%s)", len(refs), channel_name, channel_id)
|
||||||
|
return str(channel_id), str(channel_name), refs
|
||||||
|
|
||||||
|
|
||||||
|
def _flatten_entries(info: dict[str, Any]) -> list[dict[str, Any]]:
|
||||||
|
if "entries" not in info:
|
||||||
|
return []
|
||||||
|
out: list[dict[str, Any]] = []
|
||||||
|
for entry in info["entries"]:
|
||||||
|
if entry is None:
|
||||||
|
continue
|
||||||
|
if isinstance(entry, dict) and entry.get("_type") == "playlist":
|
||||||
|
out.extend(_flatten_entries(entry))
|
||||||
|
else:
|
||||||
|
out.append(entry)
|
||||||
|
return out
|
||||||
@@ -0,0 +1,118 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import csv
|
||||||
|
import json
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from jinja2 import Environment, FileSystemLoader, select_autoescape
|
||||||
|
|
||||||
|
from .store import SegmentRow, Store, VideoRow
|
||||||
|
|
||||||
|
|
||||||
|
def export_json(store: Store, channel_id: str | None, out_path: str | Path) -> Path:
|
||||||
|
out = Path(out_path)
|
||||||
|
out.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
videos = store.get_all(channel_id)
|
||||||
|
payload = {
|
||||||
|
"channels": store.list_channels(),
|
||||||
|
"videos": [_video_to_dict(store, v) for v in videos],
|
||||||
|
}
|
||||||
|
out.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def export_csv(store: Store, channel_id: str | None, out_path: str | Path) -> Path:
|
||||||
|
out = Path(out_path)
|
||||||
|
out.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
videos = store.get_all(channel_id)
|
||||||
|
fields = [
|
||||||
|
"video_id", "channel_id", "title", "url", "upload_date", "duration",
|
||||||
|
"status", "transcript_lang", "transcript_src", "has_chapters",
|
||||||
|
"view_count", "like_count", "tags",
|
||||||
|
]
|
||||||
|
with open(out, "w", encoding="utf-8", newline="") as f:
|
||||||
|
writer = csv.DictWriter(f, fieldnames=fields, extrasaction="ignore")
|
||||||
|
writer.writeheader()
|
||||||
|
for v in videos:
|
||||||
|
d = _video_to_dict(store, v)
|
||||||
|
writer.writerow(d)
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def export_srt_video(store: Store, video_id: str, out_path: str | Path) -> Path | None:
|
||||||
|
segs = store.get_segments(video_id)
|
||||||
|
if not segs:
|
||||||
|
return None
|
||||||
|
out = Path(out_path)
|
||||||
|
out.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
out.write_text(_segments_to_srt(segs), encoding="utf-8")
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def export_srt(store: Store, channel_id: str | None, out_dir: str | Path) -> list[Path]:
|
||||||
|
out_dir = Path(out_dir)
|
||||||
|
out_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
written: list[Path] = []
|
||||||
|
for v in store.get_all(channel_id):
|
||||||
|
if v.status != "done":
|
||||||
|
continue
|
||||||
|
target = out_dir / f"{v.video_id}.srt"
|
||||||
|
res = export_srt_video(store, v.video_id, target)
|
||||||
|
if res:
|
||||||
|
written.append(res)
|
||||||
|
return written
|
||||||
|
|
||||||
|
|
||||||
|
def export_html(store: Store, channel_id: str | None, out_path: str | Path, template_path: str | Path | None = None) -> Path:
|
||||||
|
out = Path(out_path)
|
||||||
|
out.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
tpl_path = Path(template_path).resolve() if template_path else (Path(__file__).resolve().parent.parent.parent / "templates" / "gallery.html.j2")
|
||||||
|
env = Environment(loader=FileSystemLoader(str(tpl_path.parent)), autoescape=select_autoescape(["html", "j2"]))
|
||||||
|
template = env.get_template(tpl_path.name)
|
||||||
|
videos = [_video_to_dict(store, v) for v in store.get_all(channel_id)]
|
||||||
|
channels = store.list_channels()
|
||||||
|
html = template.render(videos=videos, channels=channels, channel_id=channel_id)
|
||||||
|
out.write_text(html, encoding="utf-8")
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def _video_to_dict(store: Store, v: VideoRow) -> dict:
|
||||||
|
tags = []
|
||||||
|
if v.tags:
|
||||||
|
try:
|
||||||
|
tags = json.loads(v.tags)
|
||||||
|
except (json.JSONDecodeError, TypeError):
|
||||||
|
tags = []
|
||||||
|
return {
|
||||||
|
"video_id": v.video_id,
|
||||||
|
"channel_id": v.channel_id,
|
||||||
|
"title": v.title,
|
||||||
|
"url": v.url,
|
||||||
|
"upload_date": v.upload_date,
|
||||||
|
"duration": v.duration,
|
||||||
|
"status": v.status,
|
||||||
|
"transcript_lang": v.transcript_lang,
|
||||||
|
"transcript_src": v.transcript_src,
|
||||||
|
"has_chapters": bool(v.has_chapters),
|
||||||
|
"view_count": v.view_count,
|
||||||
|
"like_count": v.like_count,
|
||||||
|
"tags": tags,
|
||||||
|
"thumbnail": v.thumbnail,
|
||||||
|
"markdown_path": v.markdown_path,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _segments_to_srt(segs: list[SegmentRow]) -> str:
|
||||||
|
blocks: list[str] = []
|
||||||
|
for i, s in enumerate(segs, start=1):
|
||||||
|
end = s.end_sec if s.end_sec and s.end_sec > s.start_sec else s.start_sec + 2.0
|
||||||
|
blocks.append(f"{i}\n{_srt_ts(s.start_sec)} --> {_srt_ts(end)}\n{s.text}\n")
|
||||||
|
return "\n".join(blocks)
|
||||||
|
|
||||||
|
|
||||||
|
def _srt_ts(sec: float) -> str:
|
||||||
|
ms = int(round((sec - int(sec)) * 1000))
|
||||||
|
total = int(sec)
|
||||||
|
h, rem = divmod(total, 3600)
|
||||||
|
m, s = divmod(rem, 60)
|
||||||
|
return f"{h:02d}:{m:02d}:{s:02d},{ms:03d}"
|
||||||
@@ -0,0 +1,125 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import logging
|
||||||
|
from dataclasses import dataclass
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
import requests
|
||||||
|
import yt_dlp
|
||||||
|
|
||||||
|
from .parse import Segment, parse_auto_dump
|
||||||
|
|
||||||
|
log = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class SubtitlePick:
|
||||||
|
url: str
|
||||||
|
ext: str
|
||||||
|
lang: str
|
||||||
|
source: str # "manual" | "auto"
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class VideoData:
|
||||||
|
info: dict[str, Any]
|
||||||
|
segments: list[Segment]
|
||||||
|
subtitle: SubtitlePick | None
|
||||||
|
has_chapters: bool
|
||||||
|
|
||||||
|
|
||||||
|
def extract_video(
|
||||||
|
video_url: str,
|
||||||
|
languages: list[str],
|
||||||
|
retries: int = 10,
|
||||||
|
sleep_subrequests: float = 2.0,
|
||||||
|
prefer_manual: bool = True,
|
||||||
|
cookies_file: str | None = None,
|
||||||
|
cookies_from_browser: str | None = None,
|
||||||
|
) -> VideoData:
|
||||||
|
ydl_opts: dict[str, Any] = {
|
||||||
|
"writesubtitles": True,
|
||||||
|
"writeautomaticsub": True,
|
||||||
|
"subtitleslangs": languages,
|
||||||
|
"skip_download": True,
|
||||||
|
"quiet": True,
|
||||||
|
"no_warnings": True,
|
||||||
|
"retries": retries,
|
||||||
|
"sleep_subrequests": sleep_subrequests,
|
||||||
|
"noprogress": True,
|
||||||
|
}
|
||||||
|
if cookies_file:
|
||||||
|
ydl_opts["cookiefile"] = cookies_file
|
||||||
|
if cookies_from_browser:
|
||||||
|
ydl_opts["cookiesfrombrowser"] = (cookies_from_browser,)
|
||||||
|
|
||||||
|
with yt_dlp.YoutubeDL(ydl_opts) as ydl:
|
||||||
|
info = ydl.extract_info(video_url, download=False)
|
||||||
|
|
||||||
|
pick = pick_subtitle(info, languages, prefer_manual)
|
||||||
|
segments: list[Segment] = []
|
||||||
|
if pick:
|
||||||
|
raw = _download_subtitle(pick.url)
|
||||||
|
if raw:
|
||||||
|
segments = parse_auto_dump(raw)
|
||||||
|
if not segments:
|
||||||
|
log.warning("Could not parse subtitle for %s (format=%s)", video_url, pick.ext)
|
||||||
|
|
||||||
|
has_chapters = bool(info.get("chapters"))
|
||||||
|
return VideoData(info=info, segments=segments, subtitle=pick, has_chapters=has_chapters)
|
||||||
|
|
||||||
|
|
||||||
|
def pick_subtitle(info: dict[str, Any], languages: list[str], prefer_manual: bool = True) -> SubtitlePick | None:
|
||||||
|
manual = info.get("subtitles") or {}
|
||||||
|
auto = info.get("automatic_captions") or {}
|
||||||
|
|
||||||
|
ordered_sources: list[tuple[str, dict[str, Any]]]
|
||||||
|
if prefer_manual:
|
||||||
|
ordered_sources = [("manual", manual), ("auto", auto)]
|
||||||
|
else:
|
||||||
|
ordered_sources = [("auto", auto), ("manual", manual)]
|
||||||
|
|
||||||
|
for source_label, tracks in ordered_sources:
|
||||||
|
for lang in languages:
|
||||||
|
normalized = _normalize_lang(lang)
|
||||||
|
for track_lang, formats in tracks.items():
|
||||||
|
if _normalize_lang(track_lang) != normalized:
|
||||||
|
continue
|
||||||
|
if not formats:
|
||||||
|
continue
|
||||||
|
pick = _pick_best_format(formats)
|
||||||
|
if pick:
|
||||||
|
return SubtitlePick(
|
||||||
|
url=pick["url"],
|
||||||
|
ext=pick["ext"],
|
||||||
|
lang=track_lang,
|
||||||
|
source="manual" if source_label == "manual" else "auto",
|
||||||
|
)
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def _pick_best_format(formats: list[dict[str, Any]]) -> dict[str, Any] | None:
|
||||||
|
priority = ["json3", "srv1", "srv3", "vtt", "ttml"]
|
||||||
|
for ext in priority:
|
||||||
|
for fmt in formats:
|
||||||
|
if fmt.get("ext") == ext and fmt.get("url"):
|
||||||
|
return fmt
|
||||||
|
for fmt in formats:
|
||||||
|
if fmt.get("url"):
|
||||||
|
return fmt
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def _normalize_lang(code: str) -> str:
|
||||||
|
base = code.replace("_", "-").split("-")[0].lower()
|
||||||
|
return base
|
||||||
|
|
||||||
|
|
||||||
|
def _download_subtitle(url: str) -> str | None:
|
||||||
|
try:
|
||||||
|
resp = requests.get(url, timeout=15, headers={"User-Agent": "Mozilla/5.0"})
|
||||||
|
resp.raise_for_status()
|
||||||
|
return resp.text
|
||||||
|
except requests.RequestException as exc:
|
||||||
|
log.error("Failed to download subtitle from %s: %s", url, exc)
|
||||||
|
return None
|
||||||
@@ -0,0 +1,102 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import logging
|
||||||
|
import time
|
||||||
|
from typing import Callable
|
||||||
|
|
||||||
|
from .config import Config
|
||||||
|
from .cookies import resolve_active_path
|
||||||
|
from .discover import discover_channel
|
||||||
|
from .pipeline import process_video
|
||||||
|
from .ratelimit import polite_sleep
|
||||||
|
from .store import Store
|
||||||
|
|
||||||
|
|
||||||
|
log = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
|
||||||
|
def watch_loop(
|
||||||
|
cfg: Config,
|
||||||
|
store: Store,
|
||||||
|
interval_seconds: float,
|
||||||
|
channel_id: str | None = None,
|
||||||
|
once: bool = False,
|
||||||
|
cookies_file: str | None = None,
|
||||||
|
cookies_from_browser: str | None = None,
|
||||||
|
on_log: Callable[[str], None] | None = None,
|
||||||
|
) -> None:
|
||||||
|
"""Discover + process pending videos for a channel, optionally on an interval."""
|
||||||
|
|
||||||
|
def _emit(msg: str) -> None:
|
||||||
|
log.info(msg)
|
||||||
|
if on_log:
|
||||||
|
on_log(msg)
|
||||||
|
|
||||||
|
while True:
|
||||||
|
if not cfg.channel_url and not channel_id:
|
||||||
|
_emit("watch: no channel_url configured; sleeping")
|
||||||
|
if once:
|
||||||
|
return
|
||||||
|
time.sleep(interval_seconds)
|
||||||
|
continue
|
||||||
|
|
||||||
|
try:
|
||||||
|
_run_once(cfg, store, channel_id, cookies_file, cookies_from_browser, on_log)
|
||||||
|
except Exception as exc:
|
||||||
|
_emit(f"watch: iteration error: {exc}")
|
||||||
|
log.exception("watch iteration failed")
|
||||||
|
|
||||||
|
if once:
|
||||||
|
return
|
||||||
|
_emit(f"watch: sleeping {interval_seconds}s")
|
||||||
|
time.sleep(interval_seconds)
|
||||||
|
|
||||||
|
|
||||||
|
def _run_once(
|
||||||
|
cfg: Config,
|
||||||
|
store: Store,
|
||||||
|
channel_id: str | None,
|
||||||
|
cookies_file: str | None,
|
||||||
|
cookies_from_browser: str | None,
|
||||||
|
on_log: Callable[[str], None] | None,
|
||||||
|
) -> None:
|
||||||
|
def _emit(msg: str) -> None:
|
||||||
|
log.info(msg)
|
||||||
|
if on_log:
|
||||||
|
on_log(msg)
|
||||||
|
|
||||||
|
channel_id_found, channel_name, refs = discover_channel(
|
||||||
|
cfg.channel_url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests
|
||||||
|
)
|
||||||
|
target_channel = channel_id or channel_id_found
|
||||||
|
store.upsert_channel(target_channel, _extract_handle(cfg.channel_url), channel_name, len(refs))
|
||||||
|
store.upsert_videos(refs)
|
||||||
|
_emit(f"watch: discovered {len(refs)} videos on {channel_name}")
|
||||||
|
|
||||||
|
pending = store.get_pending(target_channel)
|
||||||
|
if not pending:
|
||||||
|
_emit("watch: nothing pending")
|
||||||
|
return
|
||||||
|
|
||||||
|
# fallback to active vault cookie if no explicit cookie given
|
||||||
|
cookie_path = cookies_file
|
||||||
|
if not cookie_path and not cookies_from_browser:
|
||||||
|
vault = resolve_active_path(store)
|
||||||
|
if vault:
|
||||||
|
cookie_path = vault
|
||||||
|
_emit(f"watch: using active vault cookie {vault}")
|
||||||
|
|
||||||
|
for row in pending:
|
||||||
|
process_video(
|
||||||
|
row, cfg, store, channel_name, target_channel, cfg.channel_url,
|
||||||
|
cookies_file=cookie_path,
|
||||||
|
cookies_from_browser=cookies_from_browser,
|
||||||
|
on_log=on_log,
|
||||||
|
)
|
||||||
|
polite_sleep(cfg.delay.min_seconds, cfg.delay.max_seconds)
|
||||||
|
|
||||||
|
|
||||||
|
def _extract_handle(url: str) -> str:
|
||||||
|
if "@" in url:
|
||||||
|
return "@" + url.split("@", 1)[1].split("/", 1)[0]
|
||||||
|
return ""
|
||||||
@@ -0,0 +1,95 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import re
|
||||||
|
from dataclasses import dataclass
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class Segment:
|
||||||
|
start: float
|
||||||
|
end: float
|
||||||
|
text: str
|
||||||
|
|
||||||
|
|
||||||
|
_VTT_TIMESTAMP = re.compile(
|
||||||
|
r"(\d{1,2}):(\d{2}):(\d{2})[.,](\d{3})\s*-->\s*(\d{1,2}):(\d{2}):(\d{2})[.,](\d{3})"
|
||||||
|
)
|
||||||
|
_VTT_TAG = re.compile(r"<[^>]+>")
|
||||||
|
|
||||||
|
|
||||||
|
def parse_json3(text: str) -> list[Segment]:
|
||||||
|
try:
|
||||||
|
data = json.loads(text)
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
return []
|
||||||
|
events = data.get("events") or []
|
||||||
|
segments: list[Segment] = []
|
||||||
|
for ev in events:
|
||||||
|
if "segs" not in ev:
|
||||||
|
continue
|
||||||
|
start_ms = ev.get("tStartMs", 0)
|
||||||
|
dur_ms = ev.get("dDurationMs", 0)
|
||||||
|
parts = []
|
||||||
|
for seg in ev["segs"]:
|
||||||
|
piece = seg.get("utf8", "")
|
||||||
|
if piece and piece != "\n":
|
||||||
|
parts.append(piece)
|
||||||
|
body = "".join(parts).strip()
|
||||||
|
if not body:
|
||||||
|
continue
|
||||||
|
segments.append(
|
||||||
|
Segment(
|
||||||
|
start=start_ms / 1000.0,
|
||||||
|
end=(start_ms + dur_ms) / 1000.0 if dur_ms else start_ms / 1000.0,
|
||||||
|
text=body,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
return merge_adjacent(segments)
|
||||||
|
|
||||||
|
|
||||||
|
def parse_vtt(text: str) -> list[Segment]:
|
||||||
|
segments: list[Segment] = []
|
||||||
|
lines = text.splitlines()
|
||||||
|
i = 0
|
||||||
|
while i < len(lines):
|
||||||
|
line = lines[i].strip()
|
||||||
|
match = _VTT_TIMESTAMP.search(line)
|
||||||
|
if match:
|
||||||
|
start = _vtt_ts_to_seconds(match.group(1), match.group(2), match.group(3), match.group(4))
|
||||||
|
end = _vtt_ts_to_seconds(match.group(5), match.group(6), match.group(7), match.group(8))
|
||||||
|
i += 1
|
||||||
|
cue_lines: list[str] = []
|
||||||
|
while i < len(lines) and lines[i].strip():
|
||||||
|
cue_lines.append(lines[i].strip())
|
||||||
|
i += 1
|
||||||
|
body = " ".join(cue_lines)
|
||||||
|
body = _VTT_TAG.sub("", body).strip()
|
||||||
|
if body:
|
||||||
|
segments.append(Segment(start=start, end=end, text=body))
|
||||||
|
i += 1
|
||||||
|
return merge_adjacent(segments)
|
||||||
|
|
||||||
|
|
||||||
|
def merge_adjacent(segments: list[Segment], max_gap: float = 0.0) -> list[Segment]:
|
||||||
|
if not segments:
|
||||||
|
return []
|
||||||
|
merged: list[Segment] = [segments[0]]
|
||||||
|
for seg in segments[1:]:
|
||||||
|
last = merged[-1]
|
||||||
|
if seg.start - last.end <= max_gap and len(last.text) < 200:
|
||||||
|
last.end = seg.end
|
||||||
|
last.text = f"{last.text} {seg.text}".strip()
|
||||||
|
else:
|
||||||
|
merged.append(seg)
|
||||||
|
return merged
|
||||||
|
|
||||||
|
|
||||||
|
def parse_auto_dump(raw: str) -> list[Segment]:
|
||||||
|
if raw.lstrip().startswith("{"):
|
||||||
|
return parse_json3(raw)
|
||||||
|
return parse_vtt(raw)
|
||||||
|
|
||||||
|
|
||||||
|
def _vtt_ts_to_seconds(h: str, m: str, s: str, ms: str) -> float:
|
||||||
|
return int(h) * 3600 + int(m) * 60 + int(s) + int(ms) / 1000.0
|
||||||
@@ -0,0 +1,181 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import logging
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Callable
|
||||||
|
|
||||||
|
from .chapters import align_chapters, chapters_from_info
|
||||||
|
from .config import Config
|
||||||
|
from .extract import extract_video
|
||||||
|
from .render import build_filename_stem, render_markdown
|
||||||
|
from .store import Store, VideoRow
|
||||||
|
|
||||||
|
|
||||||
|
log = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
|
||||||
|
def process_video(
|
||||||
|
row: VideoRow,
|
||||||
|
cfg: Config,
|
||||||
|
store: Store,
|
||||||
|
channel_name: str,
|
||||||
|
channel_id: str,
|
||||||
|
channel_url: str,
|
||||||
|
cookies_file: str | None = None,
|
||||||
|
cookies_from_browser: str | None = None,
|
||||||
|
on_log: Callable[[str], None] | None = None,
|
||||||
|
) -> str:
|
||||||
|
"""Extract + parse + store + render a single video. Returns final status string."""
|
||||||
|
|
||||||
|
def _emit(msg: str) -> None:
|
||||||
|
log.info(msg)
|
||||||
|
if on_log:
|
||||||
|
on_log(msg)
|
||||||
|
|
||||||
|
try:
|
||||||
|
data = extract_video(
|
||||||
|
row.url,
|
||||||
|
cfg.languages,
|
||||||
|
retries=cfg.yt_dlp.retries,
|
||||||
|
sleep_subrequests=cfg.yt_dlp.sleep_subrequests,
|
||||||
|
prefer_manual=cfg.prefer_manual,
|
||||||
|
cookies_file=cookies_file,
|
||||||
|
cookies_from_browser=cookies_from_browser,
|
||||||
|
)
|
||||||
|
except Exception as exc:
|
||||||
|
log.error("Error extrayendo %s: %s", row.video_id, exc)
|
||||||
|
store.mark_error(row.video_id, str(exc))
|
||||||
|
return "error"
|
||||||
|
|
||||||
|
if not data.segments:
|
||||||
|
log.warning("Sin transcripcion para %s", row.video_id)
|
||||||
|
store.mark_status(row.video_id, "no_subtitles")
|
||||||
|
return "no_subtitles"
|
||||||
|
|
||||||
|
chapters = chapters_from_info(data.info)
|
||||||
|
sections = align_chapters(data.segments, chapters)
|
||||||
|
|
||||||
|
info = data.info
|
||||||
|
|
||||||
|
# persist segments + rich metadata to DB (for search, stats, webapp)
|
||||||
|
store.store_segments(row.video_id, data.segments)
|
||||||
|
seg_json = json.dumps(
|
||||||
|
[{"start": s.start, "end": s.end, "text": s.text} for s in data.segments],
|
||||||
|
ensure_ascii=False,
|
||||||
|
)
|
||||||
|
ch_json = json.dumps(
|
||||||
|
[{"title": c.title, "start": c.start_time, "end": c.end_time} for c in chapters],
|
||||||
|
ensure_ascii=False,
|
||||||
|
)
|
||||||
|
store.update_video_metadata(
|
||||||
|
row.video_id,
|
||||||
|
view_count=info.get("view_count"),
|
||||||
|
like_count=info.get("like_count"),
|
||||||
|
tags=info.get("tags") or None,
|
||||||
|
thumbnail=info.get("thumbnail"),
|
||||||
|
description=(info.get("description") or "").strip() or None,
|
||||||
|
chapters_json=ch_json,
|
||||||
|
segments_json=seg_json,
|
||||||
|
)
|
||||||
|
|
||||||
|
# ensure upload_date is populated (discovery sometimes lacks it)
|
||||||
|
ud = info.get("upload_date") or row.upload_date
|
||||||
|
if ud:
|
||||||
|
store.set_upload_date(row.video_id, ud.replace("-", "") if "-" in str(ud) else str(ud))
|
||||||
|
|
||||||
|
context = {
|
||||||
|
"video_id": row.video_id,
|
||||||
|
"title": info.get("title") or row.title or row.video_id,
|
||||||
|
"channel_name": info.get("channel") or channel_name,
|
||||||
|
"channel_id": channel_id,
|
||||||
|
"channel_url": channel_url,
|
||||||
|
"upload_date": _normalize_date(info.get("upload_date") or row.upload_date),
|
||||||
|
"duration": info.get("duration") or row.duration or 0,
|
||||||
|
"url": row.url,
|
||||||
|
"transcript_lang": data.subtitle.lang if data.subtitle else "",
|
||||||
|
"transcript_src": data.subtitle.source if data.subtitle else "",
|
||||||
|
"view_count": info.get("view_count"),
|
||||||
|
"like_count": info.get("like_count"),
|
||||||
|
"tags": info.get("tags") or [],
|
||||||
|
"thumbnail": info.get("thumbnail") or "",
|
||||||
|
"description": (info.get("description") or "").strip(),
|
||||||
|
"sections": sections,
|
||||||
|
}
|
||||||
|
|
||||||
|
stem = build_filename_stem(
|
||||||
|
upload_date=context["upload_date"],
|
||||||
|
title=context["title"],
|
||||||
|
template=cfg.filename_template,
|
||||||
|
)
|
||||||
|
out_subdir = Path(cfg.output_dir_resolved) / _safe_dirname(channel_name)
|
||||||
|
md_path = render_markdown(cfg.template_path_resolved, out_subdir, stem, context)
|
||||||
|
|
||||||
|
out_root = Path(cfg.output_dir_resolved).parent
|
||||||
|
try:
|
||||||
|
rel = md_path.relative_to(out_root) if md_path.is_relative_to(out_root) else md_path
|
||||||
|
except ValueError:
|
||||||
|
rel = md_path
|
||||||
|
store.mark_done(
|
||||||
|
row.video_id,
|
||||||
|
str(rel),
|
||||||
|
data.subtitle.lang if data.subtitle else None,
|
||||||
|
data.subtitle.source if data.subtitle else None,
|
||||||
|
data.has_chapters,
|
||||||
|
)
|
||||||
|
_emit(f"OK {row.video_id} -> {md_path.name}")
|
||||||
|
return "done"
|
||||||
|
|
||||||
|
|
||||||
|
def _normalize_date(d: str | None) -> str:
|
||||||
|
if not d:
|
||||||
|
return ""
|
||||||
|
if len(d) == 8:
|
||||||
|
return f"{d[:4]}-{d[4:6]}-{d[6:8]}"
|
||||||
|
return d
|
||||||
|
|
||||||
|
|
||||||
|
def _safe_dirname(name: str) -> str:
|
||||||
|
safe = "".join(c for c in name if c not in r'\/:*?"<>|')
|
||||||
|
return safe.strip().strip(".") or "unknown"
|
||||||
|
|
||||||
|
|
||||||
|
def re_render_videos(store: Store, cfg: Config, channel_id: str | None = None) -> int:
|
||||||
|
"""Regenerate .md for done videos that have stored segments. Returns count re-rendered."""
|
||||||
|
import json
|
||||||
|
from .chapters import align_chapters, Chapter
|
||||||
|
from .parse import Segment
|
||||||
|
from .render import build_filename_stem, render_markdown
|
||||||
|
|
||||||
|
videos = [v for v in store.get_all(channel_id) if v.status == "done" and v.segments_json]
|
||||||
|
n = 0
|
||||||
|
for v in videos:
|
||||||
|
try:
|
||||||
|
segs = [Segment(start=s["start"], end=s["end"], text=s["text"]) for s in json.loads(v.segments_json)]
|
||||||
|
chapters = [
|
||||||
|
Chapter(title=c["title"], start_time=c["start"], end_time=c.get("end", c["start"]))
|
||||||
|
for c in json.loads(v.chapters_json or "[]")
|
||||||
|
]
|
||||||
|
sections = align_chapters(segs, chapters)
|
||||||
|
ch = store.get_channel(v.channel_id) or {}
|
||||||
|
tags = []
|
||||||
|
if v.tags:
|
||||||
|
try:
|
||||||
|
tags = json.loads(v.tags)
|
||||||
|
except Exception:
|
||||||
|
tags = []
|
||||||
|
context = {
|
||||||
|
"video_id": v.video_id, "title": v.title or v.video_id,
|
||||||
|
"channel_name": ch.get("name") or "", "channel_id": v.channel_id, "channel_url": "",
|
||||||
|
"upload_date": v.upload_date or "", "duration": v.duration or 0, "url": v.url,
|
||||||
|
"transcript_lang": v.transcript_lang or "", "transcript_src": v.transcript_src or "",
|
||||||
|
"view_count": v.view_count, "like_count": v.like_count, "tags": tags,
|
||||||
|
"thumbnail": v.thumbnail or "", "description": v.description or "", "sections": sections,
|
||||||
|
}
|
||||||
|
stem = build_filename_stem(v.upload_date, v.title or v.video_id, cfg.filename_template)
|
||||||
|
out_subdir = Path(cfg.output_dir_resolved) / _safe_dirname(ch.get("name") or "unknown")
|
||||||
|
render_markdown(cfg.template_path_resolved, out_subdir, stem, context)
|
||||||
|
n += 1
|
||||||
|
except Exception as exc:
|
||||||
|
log.warning("re-render failed for %s: %s", v.video_id, exc)
|
||||||
|
return n
|
||||||
@@ -0,0 +1,14 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import random
|
||||||
|
import time
|
||||||
|
|
||||||
|
|
||||||
|
def polite_sleep(min_s: float = 1.5, max_s: float = 3.5) -> None:
|
||||||
|
delay = random.uniform(min_s, max_s)
|
||||||
|
time.sleep(delay)
|
||||||
|
|
||||||
|
|
||||||
|
def backoff_sleep(attempt: int, base: float = 2.0, cap: float = 60.0) -> None:
|
||||||
|
delay = min(cap, base * (2 ** attempt) + random.uniform(0, 1))
|
||||||
|
time.sleep(delay)
|
||||||
@@ -0,0 +1,72 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from jinja2 import Environment, FileSystemLoader, select_autoescape
|
||||||
|
|
||||||
|
from .chapters import Section
|
||||||
|
|
||||||
|
|
||||||
|
def format_timestamp(seconds: float) -> str:
|
||||||
|
total = int(seconds)
|
||||||
|
hours, remainder = divmod(total, 3600)
|
||||||
|
minutes, secs = divmod(remainder, 60)
|
||||||
|
if hours > 0:
|
||||||
|
return f"{hours:02d}:{minutes:02d}:{secs:02d}"
|
||||||
|
return f"{minutes:02d}:{secs:02d}"
|
||||||
|
|
||||||
|
|
||||||
|
def quote_yaml(value: str | None) -> str:
|
||||||
|
if value is None:
|
||||||
|
return '""'
|
||||||
|
needs_quote = any(c in value for c in [':', '#', '{', '}', '[', ']', ',', '&', '*', '!', '|', '>', "'", '"', '%', '@', '`', '\n'])
|
||||||
|
if needs_quote:
|
||||||
|
escaped = value.replace('"', '\\"')
|
||||||
|
return f'"{escaped}"'
|
||||||
|
return value
|
||||||
|
|
||||||
|
|
||||||
|
def to_json(value) -> str:
|
||||||
|
import json
|
||||||
|
return json.dumps(value, ensure_ascii=False)
|
||||||
|
|
||||||
|
|
||||||
|
def _make_env(template_dir: Path) -> Environment:
|
||||||
|
env = Environment(
|
||||||
|
loader=FileSystemLoader(str(template_dir)),
|
||||||
|
autoescape=select_autoescape(disabled_extensions=("j2", "txt")),
|
||||||
|
trim_blocks=True,
|
||||||
|
lstrip_blocks=True,
|
||||||
|
)
|
||||||
|
env.filters["format_timestamp"] = format_timestamp
|
||||||
|
env.filters["quote_yaml"] = quote_yaml
|
||||||
|
env.filters["to_json"] = to_json
|
||||||
|
return env
|
||||||
|
|
||||||
|
|
||||||
|
def render_markdown(
|
||||||
|
template_path: str | Path,
|
||||||
|
output_dir: str | Path,
|
||||||
|
filename_stem: str,
|
||||||
|
context: dict,
|
||||||
|
) -> Path:
|
||||||
|
tpl_path = Path(template_path).resolve()
|
||||||
|
env = _make_env(tpl_path.parent)
|
||||||
|
template = env.get_template(tpl_path.name)
|
||||||
|
rendered = template.render(**context)
|
||||||
|
|
||||||
|
out_dir = Path(output_dir)
|
||||||
|
out_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
out_file = out_dir / f"{filename_stem}.md"
|
||||||
|
out_file.write_text(rendered, encoding="utf-8")
|
||||||
|
return out_file
|
||||||
|
|
||||||
|
|
||||||
|
def build_filename_stem(upload_date: str | None, title: str, template: str = "{upload_date}_{slug}") -> str:
|
||||||
|
from slugify import slugify
|
||||||
|
date_part = upload_date or "unknown-date"
|
||||||
|
slug = slugify(title, max_length=60) or "untitled"
|
||||||
|
stem = template.format(upload_date=date_part, slug=slug, title=title)
|
||||||
|
safe = "".join(c for c in stem if c not in r'\/:*?"<>|')
|
||||||
|
return safe.strip().strip(".")
|
||||||
@@ -0,0 +1,185 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import logging
|
||||||
|
import re
|
||||||
|
from dataclasses import dataclass
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Callable
|
||||||
|
|
||||||
|
from .parse import Segment
|
||||||
|
from .store import SearchHit, Store
|
||||||
|
|
||||||
|
|
||||||
|
log = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
_TS = re.compile(r"^(\d{1,2}):(\d{2})(?::(\d{2}))?$")
|
||||||
|
_SEG_LINE = re.compile(r"^\*\*(?P<ts>\d{1,2}:\d{2}(?::\d{2})?)\*\*\s*[·\-]\s*(?P<text>.+?)\s*$")
|
||||||
|
_CHAPTER = re.compile(r"^###\s+(?P<title>.+?)\s*\((?P<ts>\d{1,2}:\d{2}(?::\d{2})?)\)\s*$")
|
||||||
|
_FRONTMATTER_KEY = re.compile(r"^([A-Za-z0-9_]+)\s*:\s*(.*)$")
|
||||||
|
|
||||||
|
|
||||||
|
def ts_to_seconds(ts: str) -> float:
|
||||||
|
m = _TS.match(ts.strip())
|
||||||
|
if not m:
|
||||||
|
return 0.0
|
||||||
|
if m.group(3): # h:m:s
|
||||||
|
return int(m.group(1)) * 3600 + int(m.group(2)) * 60 + int(m.group(3))
|
||||||
|
return int(m.group(1)) * 60 + int(m.group(2))
|
||||||
|
|
||||||
|
|
||||||
|
def seconds_to_ts(sec: float) -> str:
|
||||||
|
total = int(sec)
|
||||||
|
h, rem = divmod(total, 3600)
|
||||||
|
m, s = divmod(rem, 60)
|
||||||
|
return f"{h:d}:{m:02d}:{s:02d}" if h else f"{m:d}:{s:02d}"
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class ParsedMarkdown:
|
||||||
|
metadata: dict
|
||||||
|
segments: list[Segment]
|
||||||
|
chapters: list[dict] # {"title": str, "start": float}
|
||||||
|
|
||||||
|
|
||||||
|
def parse_markdown(text: str) -> ParsedMarkdown:
|
||||||
|
lines = text.splitlines()
|
||||||
|
metadata: dict[str, str] = {}
|
||||||
|
segments: list[Segment] = []
|
||||||
|
chapters: list[dict] = []
|
||||||
|
in_front = False
|
||||||
|
front_done = False
|
||||||
|
i = 0
|
||||||
|
# frontmatter
|
||||||
|
if lines and lines[0].strip() == "---":
|
||||||
|
i = 1
|
||||||
|
in_front = True
|
||||||
|
while i < len(lines):
|
||||||
|
if lines[i].strip() == "---":
|
||||||
|
front_done = True
|
||||||
|
i += 1
|
||||||
|
break
|
||||||
|
m = _FRONTMATTER_KEY.match(lines[i])
|
||||||
|
if m:
|
||||||
|
metadata[m.group(1).lower()] = _strip_quotes(m.group(2).strip())
|
||||||
|
i += 1
|
||||||
|
if not front_done:
|
||||||
|
i = 0
|
||||||
|
# body
|
||||||
|
current_start = 0.0
|
||||||
|
while i < len(lines):
|
||||||
|
line = lines[i]
|
||||||
|
s = line.strip()
|
||||||
|
cm = _CHAPTER.match(s)
|
||||||
|
if cm:
|
||||||
|
current_start = ts_to_seconds(cm.group("ts"))
|
||||||
|
chapters.append({"title": cm.group("title").strip(), "start": current_start})
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
|
sm = _SEG_LINE.match(s)
|
||||||
|
if sm:
|
||||||
|
start = ts_to_seconds(sm.group("ts"))
|
||||||
|
segments.append(Segment(start=start, end=start, text=sm.group("text").strip()))
|
||||||
|
i += 1
|
||||||
|
return ParsedMarkdown(metadata=metadata, segments=segments, chapters=chapters)
|
||||||
|
|
||||||
|
|
||||||
|
def _strip_quotes(value: str) -> str:
|
||||||
|
v = value.strip()
|
||||||
|
if len(v) >= 2 and v[0] == v[-1] and v[0] in ("'", '"'):
|
||||||
|
return v[1:-1]
|
||||||
|
return v
|
||||||
|
|
||||||
|
|
||||||
|
def store_segments(store: Store, video_id: str, segments: list[Segment]) -> None:
|
||||||
|
store.store_segments(video_id, segments)
|
||||||
|
|
||||||
|
|
||||||
|
def search(store: Store, query: str, channel_id: str | None = None, limit: int = 50) -> list[SearchHit]:
|
||||||
|
return store.search_segments(query, channel_id=channel_id, limit=limit)
|
||||||
|
|
||||||
|
|
||||||
|
def backfill_from_markdown(
|
||||||
|
store: Store,
|
||||||
|
md_root: Path,
|
||||||
|
log: Callable[[str], None] | None = None,
|
||||||
|
) -> int:
|
||||||
|
"""Parse all done .md files under md_root, populate segments/FTS/metadata. Idempotent."""
|
||||||
|
def _log(msg: str) -> None:
|
||||||
|
if log:
|
||||||
|
log(msg)
|
||||||
|
else:
|
||||||
|
logmod = logging.getLogger(__name__)
|
||||||
|
logmod.info(msg)
|
||||||
|
|
||||||
|
md_root = Path(md_root)
|
||||||
|
if not md_root.exists():
|
||||||
|
_log(f"backfill: markdown root not found: {md_root}")
|
||||||
|
return 0
|
||||||
|
md_files = sorted(md_root.rglob("*.md"))
|
||||||
|
_log(f"backfill: scanning {len(md_files)} markdown files")
|
||||||
|
n = 0
|
||||||
|
for md_path in md_files:
|
||||||
|
try:
|
||||||
|
text = md_path.read_text(encoding="utf-8")
|
||||||
|
except OSError as exc:
|
||||||
|
_log(f"backfill: skip unreadable {md_path}: {exc}")
|
||||||
|
continue
|
||||||
|
parsed = parse_markdown(text)
|
||||||
|
video_id = parsed.metadata.get("video_id")
|
||||||
|
if not video_id:
|
||||||
|
continue
|
||||||
|
existing = store.get_video(video_id)
|
||||||
|
if not existing:
|
||||||
|
_log(f"backfill: video {video_id} not in DB (orphan md), skipping")
|
||||||
|
continue
|
||||||
|
# always (re)populate upload_date if missing — cheap, idempotent, gated by NULL
|
||||||
|
ud = parsed.metadata.get("upload_date")
|
||||||
|
if ud:
|
||||||
|
compact = ud.replace("-", "")
|
||||||
|
if compact.isdigit():
|
||||||
|
store.set_upload_date(video_id, compact)
|
||||||
|
if store.has_segments(video_id) and existing.segments_json:
|
||||||
|
continue # segments + rich metadata already populated
|
||||||
|
# metadata from frontmatter
|
||||||
|
tags_raw = parsed.metadata.get("tags", "")
|
||||||
|
tags: list[str] = []
|
||||||
|
if tags_raw:
|
||||||
|
try:
|
||||||
|
tags = json.loads(tags_raw) if tags_raw.startswith("[") else [t.strip() for t in tags_raw.split(",")]
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
tags = []
|
||||||
|
seg_json = json.dumps(
|
||||||
|
[{"start": s.start, "end": s.end, "text": s.text} for s in parsed.segments],
|
||||||
|
ensure_ascii=False,
|
||||||
|
)
|
||||||
|
ch_json = json.dumps(parsed.chapters, ensure_ascii=False)
|
||||||
|
store.update_video_metadata(
|
||||||
|
video_id,
|
||||||
|
view_count=_to_int(parsed.metadata.get("views")),
|
||||||
|
like_count=_to_int(parsed.metadata.get("likes")),
|
||||||
|
tags=tags or None,
|
||||||
|
thumbnail=parsed.metadata.get("thumbnail") or None,
|
||||||
|
description=None,
|
||||||
|
chapters_json=ch_json,
|
||||||
|
segments_json=seg_json,
|
||||||
|
)
|
||||||
|
store.store_segments(video_id, parsed.segments)
|
||||||
|
n += 1
|
||||||
|
_log(f"backfill: populated {n} videos")
|
||||||
|
return n
|
||||||
|
|
||||||
|
|
||||||
|
def _to_int(value: str | None) -> int | None:
|
||||||
|
if value is None:
|
||||||
|
return None
|
||||||
|
value = str(value).strip().strip('"').strip("'")
|
||||||
|
if not value:
|
||||||
|
return None
|
||||||
|
try:
|
||||||
|
return int(value)
|
||||||
|
except ValueError:
|
||||||
|
try:
|
||||||
|
return int(float(value))
|
||||||
|
except ValueError:
|
||||||
|
return None
|
||||||
@@ -0,0 +1,731 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import sqlite3
|
||||||
|
from contextlib import contextmanager
|
||||||
|
from dataclasses import dataclass
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any, Iterator
|
||||||
|
|
||||||
|
|
||||||
|
SCHEMA = """
|
||||||
|
CREATE TABLE IF NOT EXISTS channels (
|
||||||
|
channel_id TEXT PRIMARY KEY,
|
||||||
|
handle TEXT,
|
||||||
|
name TEXT,
|
||||||
|
last_scraped TEXT,
|
||||||
|
video_count INTEGER DEFAULT 0
|
||||||
|
);
|
||||||
|
|
||||||
|
CREATE TABLE IF NOT EXISTS videos (
|
||||||
|
video_id TEXT PRIMARY KEY,
|
||||||
|
channel_id TEXT NOT NULL,
|
||||||
|
title TEXT,
|
||||||
|
url TEXT NOT NULL,
|
||||||
|
upload_date TEXT,
|
||||||
|
duration INTEGER,
|
||||||
|
status TEXT NOT NULL DEFAULT 'pending',
|
||||||
|
error_msg TEXT,
|
||||||
|
transcript_lang TEXT,
|
||||||
|
transcript_src TEXT,
|
||||||
|
has_chapters INTEGER DEFAULT 0,
|
||||||
|
markdown_path TEXT,
|
||||||
|
discovered_at TEXT NOT NULL,
|
||||||
|
processed_at TEXT,
|
||||||
|
FOREIGN KEY (channel_id) REFERENCES channels(channel_id)
|
||||||
|
);
|
||||||
|
|
||||||
|
CREATE INDEX IF NOT EXISTS idx_videos_status ON videos(status);
|
||||||
|
CREATE INDEX IF NOT EXISTS idx_videos_channel ON videos(channel_id);
|
||||||
|
"""
|
||||||
|
|
||||||
|
# New columns added by the platform migration (idempotent ALTERs)
|
||||||
|
_VIDEO_COLUMNS: dict[str, str] = {
|
||||||
|
"view_count": "INTEGER",
|
||||||
|
"like_count": "INTEGER",
|
||||||
|
"tags": "TEXT",
|
||||||
|
"thumbnail": "TEXT",
|
||||||
|
"description": "TEXT",
|
||||||
|
"chapters_json": "TEXT",
|
||||||
|
"segments_json": "TEXT",
|
||||||
|
}
|
||||||
|
|
||||||
|
_EXTRA_SCHEMA = """
|
||||||
|
CREATE TABLE IF NOT EXISTS transcript_segments (
|
||||||
|
video_id TEXT NOT NULL,
|
||||||
|
idx INTEGER NOT NULL,
|
||||||
|
start_sec REAL NOT NULL,
|
||||||
|
end_sec REAL NOT NULL,
|
||||||
|
text TEXT NOT NULL,
|
||||||
|
PRIMARY KEY (video_id, idx),
|
||||||
|
FOREIGN KEY (video_id) REFERENCES videos(video_id)
|
||||||
|
);
|
||||||
|
CREATE INDEX IF NOT EXISTS idx_segments_video ON transcript_segments(video_id);
|
||||||
|
|
||||||
|
CREATE VIRTUAL TABLE IF NOT EXISTS transcript_fts USING fts5(
|
||||||
|
video_id UNINDEXED, idx UNINDEXED, start_sec UNINDEXED, end_sec UNINDEXED, text
|
||||||
|
);
|
||||||
|
|
||||||
|
CREATE TABLE IF NOT EXISTS cookies_meta (
|
||||||
|
id TEXT PRIMARY KEY,
|
||||||
|
filename TEXT NOT NULL,
|
||||||
|
label TEXT,
|
||||||
|
added_at TEXT NOT NULL,
|
||||||
|
expires_at TEXT,
|
||||||
|
is_active INTEGER DEFAULT 0,
|
||||||
|
has_session INTEGER DEFAULT 0,
|
||||||
|
cookie_count INTEGER DEFAULT 0
|
||||||
|
);
|
||||||
|
|
||||||
|
CREATE TABLE IF NOT EXISTS scrape_jobs (
|
||||||
|
id TEXT PRIMARY KEY,
|
||||||
|
channel_id TEXT,
|
||||||
|
opts_json TEXT,
|
||||||
|
status TEXT NOT NULL,
|
||||||
|
started_at TEXT,
|
||||||
|
finished_at TEXT,
|
||||||
|
total INTEGER DEFAULT 0,
|
||||||
|
completed INTEGER DEFAULT 0,
|
||||||
|
last_error TEXT
|
||||||
|
);
|
||||||
|
"""
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class VideoRef:
|
||||||
|
video_id: str
|
||||||
|
channel_id: str
|
||||||
|
title: str
|
||||||
|
url: str
|
||||||
|
upload_date: str | None = None
|
||||||
|
duration: int | None = None
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class VideoRow:
|
||||||
|
video_id: str
|
||||||
|
channel_id: str
|
||||||
|
title: str | None
|
||||||
|
url: str
|
||||||
|
upload_date: str | None
|
||||||
|
duration: int | None
|
||||||
|
status: str
|
||||||
|
error_msg: str | None
|
||||||
|
transcript_lang: str | None
|
||||||
|
transcript_src: str | None
|
||||||
|
has_chapters: int
|
||||||
|
markdown_path: str | None
|
||||||
|
view_count: int | None = None
|
||||||
|
like_count: int | None = None
|
||||||
|
tags: str | None = None
|
||||||
|
thumbnail: str | None = None
|
||||||
|
description: str | None = None
|
||||||
|
chapters_json: str | None = None
|
||||||
|
segments_json: str | None = None
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class SegmentRow:
|
||||||
|
video_id: str
|
||||||
|
idx: int
|
||||||
|
start_sec: float
|
||||||
|
end_sec: float
|
||||||
|
text: str
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class SearchHit:
|
||||||
|
video_id: str
|
||||||
|
channel_id: str
|
||||||
|
title: str | None
|
||||||
|
idx: int
|
||||||
|
start_sec: float
|
||||||
|
end_sec: float
|
||||||
|
snippet: str
|
||||||
|
rank: float
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class CookieRow:
|
||||||
|
id: str
|
||||||
|
filename: str
|
||||||
|
label: str | None
|
||||||
|
added_at: str
|
||||||
|
expires_at: str | None
|
||||||
|
is_active: bool
|
||||||
|
has_session: bool
|
||||||
|
cookie_count: int
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class JobRow:
|
||||||
|
id: str
|
||||||
|
channel_id: str | None
|
||||||
|
opts_json: str | None
|
||||||
|
status: str
|
||||||
|
started_at: str | None
|
||||||
|
finished_at: str | None
|
||||||
|
total: int
|
||||||
|
completed: int
|
||||||
|
last_error: str | None
|
||||||
|
|
||||||
|
|
||||||
|
class Store:
|
||||||
|
def __init__(self, db_path: str | Path):
|
||||||
|
self.db_path = Path(db_path)
|
||||||
|
self.db_path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
self._init_schema()
|
||||||
|
|
||||||
|
def _connect(self) -> sqlite3.Connection:
|
||||||
|
conn = sqlite3.connect(str(self.db_path))
|
||||||
|
conn.row_factory = sqlite3.Row
|
||||||
|
conn.execute("PRAGMA journal_mode=WAL")
|
||||||
|
conn.execute("PRAGMA foreign_keys=ON")
|
||||||
|
return conn
|
||||||
|
|
||||||
|
def _init_schema(self) -> None:
|
||||||
|
with self._connect() as conn:
|
||||||
|
conn.executescript(SCHEMA)
|
||||||
|
# idempotent column adds
|
||||||
|
existing = {row["name"] for row in conn.execute("PRAGMA table_info(videos)")}
|
||||||
|
for col, coltype in _VIDEO_COLUMNS.items():
|
||||||
|
if col not in existing:
|
||||||
|
conn.execute(f"ALTER TABLE videos ADD COLUMN {col} {coltype}")
|
||||||
|
conn.executescript(_EXTRA_SCHEMA)
|
||||||
|
|
||||||
|
@contextmanager
|
||||||
|
def _cursor(self) -> Iterator[sqlite3.Cursor]:
|
||||||
|
conn = self._connect()
|
||||||
|
try:
|
||||||
|
yield conn.cursor()
|
||||||
|
conn.commit()
|
||||||
|
finally:
|
||||||
|
conn.close()
|
||||||
|
|
||||||
|
# ------------------------------------------------------------------ channels
|
||||||
|
|
||||||
|
def upsert_channel(self, channel_id: str, handle: str | None, name: str | None, video_count: int = 0) -> None:
|
||||||
|
now = _now_iso()
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute(
|
||||||
|
"""INSERT INTO channels (channel_id, handle, name, last_scraped, video_count)
|
||||||
|
VALUES (?, ?, ?, ?, ?)
|
||||||
|
ON CONFLICT(channel_id) DO UPDATE SET
|
||||||
|
handle = excluded.handle,
|
||||||
|
name = excluded.name,
|
||||||
|
last_scraped = excluded.last_scraped,
|
||||||
|
video_count = excluded.video_count""",
|
||||||
|
(channel_id, handle, name, now, video_count),
|
||||||
|
)
|
||||||
|
|
||||||
|
def list_channels(self) -> list[dict]:
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute("SELECT * FROM channels ORDER BY name")
|
||||||
|
return [dict(r) for r in cur.fetchall()]
|
||||||
|
|
||||||
|
def get_channel(self, channel_id: str) -> dict | None:
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute("SELECT * FROM channels WHERE channel_id = ?", (channel_id,))
|
||||||
|
row = cur.fetchone()
|
||||||
|
return dict(row) if row else None
|
||||||
|
|
||||||
|
def delete_channel(self, channel_id: str) -> None:
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute("DELETE FROM transcript_segments WHERE video_id IN (SELECT video_id FROM videos WHERE channel_id = ?)", (channel_id,))
|
||||||
|
cur.execute("DELETE FROM transcript_fts WHERE video_id IN (SELECT video_id FROM videos WHERE channel_id = ?)", (channel_id,))
|
||||||
|
cur.execute("DELETE FROM videos WHERE channel_id = ?", (channel_id,))
|
||||||
|
cur.execute("DELETE FROM channels WHERE channel_id = ?", (channel_id,))
|
||||||
|
|
||||||
|
# ------------------------------------------------------------------ videos
|
||||||
|
|
||||||
|
def upsert_videos(self, refs: list[VideoRef]) -> int:
|
||||||
|
now = _now_iso()
|
||||||
|
inserted = 0
|
||||||
|
with self._cursor() as cur:
|
||||||
|
for r in refs:
|
||||||
|
cur.execute(
|
||||||
|
"""INSERT INTO videos
|
||||||
|
(video_id, channel_id, title, url, upload_date, duration, status, discovered_at)
|
||||||
|
VALUES (?, ?, ?, ?, ?, ?, 'pending', ?)
|
||||||
|
ON CONFLICT(video_id) DO UPDATE SET
|
||||||
|
title = excluded.title,
|
||||||
|
upload_date = excluded.upload_date,
|
||||||
|
duration = excluded.duration""",
|
||||||
|
(r.video_id, r.channel_id, r.title, r.url, r.upload_date, r.duration, now),
|
||||||
|
)
|
||||||
|
if cur.rowcount > 0:
|
||||||
|
inserted += 1
|
||||||
|
return inserted
|
||||||
|
|
||||||
|
def get_pending(self, channel_id: str | None = None, limit: int | None = None) -> list[VideoRow]:
|
||||||
|
sql = "SELECT * FROM videos WHERE status = 'pending'"
|
||||||
|
params: list[Any] = []
|
||||||
|
if channel_id:
|
||||||
|
sql += " AND channel_id = ?"
|
||||||
|
params.append(channel_id)
|
||||||
|
sql += " ORDER BY discovered_at ASC"
|
||||||
|
if limit:
|
||||||
|
sql += " LIMIT ?"
|
||||||
|
params.append(limit)
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute(sql, params)
|
||||||
|
return [_row_to_videorow(row) for row in cur.fetchall()]
|
||||||
|
|
||||||
|
def get_all(self, channel_id: str | None = None) -> list[VideoRow]:
|
||||||
|
sql = "SELECT * FROM videos"
|
||||||
|
params: list[Any] = []
|
||||||
|
if channel_id:
|
||||||
|
sql += " WHERE channel_id = ?"
|
||||||
|
params.append(channel_id)
|
||||||
|
sql += " ORDER BY discovered_at ASC"
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute(sql, params)
|
||||||
|
return [_row_to_videorow(row) for row in cur.fetchall()]
|
||||||
|
|
||||||
|
def get_video(self, video_id: str) -> VideoRow | None:
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute("SELECT * FROM videos WHERE video_id = ?", (video_id,))
|
||||||
|
row = cur.fetchone()
|
||||||
|
return _row_to_videorow(row) if row else None
|
||||||
|
|
||||||
|
def query_videos(
|
||||||
|
self,
|
||||||
|
channel_id: str | None = None,
|
||||||
|
status: str | None = None,
|
||||||
|
date_from: str | None = None,
|
||||||
|
date_to: str | None = None,
|
||||||
|
min_duration: int | None = None,
|
||||||
|
q: str | None = None,
|
||||||
|
sort: str = "upload_date_desc",
|
||||||
|
page: int = 1,
|
||||||
|
size: int = 50,
|
||||||
|
) -> tuple[list[VideoRow], int]:
|
||||||
|
where: list[str] = []
|
||||||
|
params: list[Any] = []
|
||||||
|
if channel_id:
|
||||||
|
where.append("channel_id = ?"); params.append(channel_id)
|
||||||
|
if status:
|
||||||
|
where.append("status = ?"); params.append(status)
|
||||||
|
if date_from:
|
||||||
|
where.append("upload_date >= ?"); params.append(date_from.replace("-", ""))
|
||||||
|
if date_to:
|
||||||
|
where.append("upload_date <= ?"); params.append(date_to.replace("-", ""))
|
||||||
|
if min_duration is not None:
|
||||||
|
where.append("duration >= ?"); params.append(min_duration)
|
||||||
|
if q:
|
||||||
|
where.append("(title LIKE ? OR description LIKE ?)")
|
||||||
|
params.extend([f"%{q}%", f"%{q}%"])
|
||||||
|
clause = ("WHERE " + " AND ".join(where)) if where else ""
|
||||||
|
order = {
|
||||||
|
"upload_date_desc": "upload_date DESC",
|
||||||
|
"upload_date_asc": "upload_date ASC",
|
||||||
|
"duration_desc": "duration DESC",
|
||||||
|
"views_desc": "view_count DESC",
|
||||||
|
"title_asc": "title ASC",
|
||||||
|
}.get(sort, "upload_date DESC")
|
||||||
|
offset = max(0, (page - 1) * size)
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute(f"SELECT COUNT(*) AS n FROM videos {clause}", params)
|
||||||
|
total = cur.fetchone()["n"]
|
||||||
|
cur.execute(
|
||||||
|
f"SELECT * FROM videos {clause} ORDER BY {order} LIMIT ? OFFSET ?",
|
||||||
|
[*params, size, offset],
|
||||||
|
)
|
||||||
|
rows = [_row_to_videorow(r) for r in cur.fetchall()]
|
||||||
|
return rows, total
|
||||||
|
|
||||||
|
def update_video_metadata(
|
||||||
|
self,
|
||||||
|
video_id: str,
|
||||||
|
*,
|
||||||
|
view_count: int | None = None,
|
||||||
|
like_count: int | None = None,
|
||||||
|
tags: list[str] | None = None,
|
||||||
|
thumbnail: str | None = None,
|
||||||
|
description: str | None = None,
|
||||||
|
chapters_json: str | None = None,
|
||||||
|
segments_json: str | None = None,
|
||||||
|
) -> None:
|
||||||
|
sets: list[str] = []
|
||||||
|
params: list[Any] = []
|
||||||
|
if view_count is not None:
|
||||||
|
sets.append("view_count = ?"); params.append(view_count)
|
||||||
|
if like_count is not None:
|
||||||
|
sets.append("like_count = ?"); params.append(like_count)
|
||||||
|
if tags is not None:
|
||||||
|
sets.append("tags = ?"); params.append(json.dumps(tags, ensure_ascii=False))
|
||||||
|
if thumbnail is not None:
|
||||||
|
sets.append("thumbnail = ?"); params.append(thumbnail)
|
||||||
|
if description is not None:
|
||||||
|
sets.append("description = ?"); params.append(description[:4000])
|
||||||
|
if chapters_json is not None:
|
||||||
|
sets.append("chapters_json = ?"); params.append(chapters_json)
|
||||||
|
if segments_json is not None:
|
||||||
|
sets.append("segments_json = ?"); params.append(segments_json)
|
||||||
|
if not sets:
|
||||||
|
return
|
||||||
|
params.append(video_id)
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute(f"UPDATE videos SET {', '.join(sets)} WHERE video_id = ?", params)
|
||||||
|
|
||||||
|
def mark_done(
|
||||||
|
self,
|
||||||
|
video_id: str,
|
||||||
|
markdown_path: str,
|
||||||
|
transcript_lang: str | None,
|
||||||
|
transcript_src: str | None,
|
||||||
|
has_chapters: bool,
|
||||||
|
) -> None:
|
||||||
|
now = _now_iso()
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute(
|
||||||
|
"""UPDATE videos SET
|
||||||
|
status = 'done',
|
||||||
|
markdown_path = ?,
|
||||||
|
transcript_lang = ?,
|
||||||
|
transcript_src = ?,
|
||||||
|
has_chapters = ?,
|
||||||
|
processed_at = ?,
|
||||||
|
error_msg = NULL
|
||||||
|
WHERE video_id = ?""",
|
||||||
|
(markdown_path, transcript_lang, transcript_src, int(has_chapters), now, video_id),
|
||||||
|
)
|
||||||
|
|
||||||
|
def mark_status(self, video_id: str, status: str) -> None:
|
||||||
|
now = _now_iso()
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute(
|
||||||
|
"UPDATE videos SET status = ?, processed_at = ? WHERE video_id = ?",
|
||||||
|
(status, now, video_id),
|
||||||
|
)
|
||||||
|
|
||||||
|
def set_upload_date(self, video_id: str, upload_date: str | None) -> None:
|
||||||
|
if not upload_date:
|
||||||
|
return
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute(
|
||||||
|
"UPDATE videos SET upload_date = ? WHERE video_id = ? AND upload_date IS NULL",
|
||||||
|
(upload_date, video_id),
|
||||||
|
)
|
||||||
|
|
||||||
|
def mark_error(self, video_id: str, error_msg: str) -> None:
|
||||||
|
now = _now_iso()
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute(
|
||||||
|
"""UPDATE videos SET status = 'error', error_msg = ?, processed_at = ?
|
||||||
|
WHERE video_id = ?""",
|
||||||
|
(error_msg[:500], now, video_id),
|
||||||
|
)
|
||||||
|
|
||||||
|
# ------------------------------------------------------------------ segments / FTS
|
||||||
|
|
||||||
|
def store_segments(self, video_id: str, segments: list) -> None:
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute("DELETE FROM transcript_segments WHERE video_id = ?", (video_id,))
|
||||||
|
cur.execute("DELETE FROM transcript_fts WHERE video_id = ?", (video_id,))
|
||||||
|
rows = []
|
||||||
|
fts_rows = []
|
||||||
|
for i, seg in enumerate(segments):
|
||||||
|
start = float(getattr(seg, "start", 0.0))
|
||||||
|
end = float(getattr(seg, "end", start))
|
||||||
|
text = (getattr(seg, "text", "") or "").strip()
|
||||||
|
if not text:
|
||||||
|
continue
|
||||||
|
rows.append((video_id, i, start, end, text))
|
||||||
|
fts_rows.append((video_id, i, start, end, text))
|
||||||
|
if rows:
|
||||||
|
cur.executemany(
|
||||||
|
"INSERT INTO transcript_segments (video_id, idx, start_sec, end_sec, text) VALUES (?, ?, ?, ?, ?)",
|
||||||
|
rows,
|
||||||
|
)
|
||||||
|
cur.executemany(
|
||||||
|
"INSERT INTO transcript_fts (video_id, idx, start_sec, end_sec, text) VALUES (?, ?, ?, ?, ?)",
|
||||||
|
fts_rows,
|
||||||
|
)
|
||||||
|
|
||||||
|
def get_segments(self, video_id: str) -> list[SegmentRow]:
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute(
|
||||||
|
"SELECT * FROM transcript_segments WHERE video_id = ? ORDER BY idx ASC",
|
||||||
|
(video_id,),
|
||||||
|
)
|
||||||
|
return [
|
||||||
|
SegmentRow(
|
||||||
|
video_id=r["video_id"], idx=r["idx"], start_sec=r["start_sec"],
|
||||||
|
end_sec=r["end_sec"], text=r["text"],
|
||||||
|
)
|
||||||
|
for r in cur.fetchall()
|
||||||
|
]
|
||||||
|
|
||||||
|
def has_segments(self, video_id: str) -> bool:
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute("SELECT 1 FROM transcript_segments WHERE video_id = ? LIMIT 1", (video_id,))
|
||||||
|
return cur.fetchone() is not None
|
||||||
|
|
||||||
|
def search_segments(self, query: str, channel_id: str | None = None, limit: int = 50) -> list[SearchHit]:
|
||||||
|
# FTS5 MATCH; join videos for title + optional channel filter.
|
||||||
|
fts_query = _sanitize_fts(query)
|
||||||
|
if not fts_query:
|
||||||
|
return []
|
||||||
|
sql = (
|
||||||
|
"SELECT f.video_id, f.idx, f.start_sec, f.end_sec, f.text, v.channel_id, v.title, bm25(transcript_fts) AS rank "
|
||||||
|
"FROM transcript_fts f JOIN videos v ON v.video_id = f.video_id "
|
||||||
|
"WHERE transcript_fts MATCH ?"
|
||||||
|
)
|
||||||
|
params: list[Any] = [fts_query]
|
||||||
|
if channel_id:
|
||||||
|
sql += " AND v.channel_id = ?"
|
||||||
|
params.append(channel_id)
|
||||||
|
sql += " ORDER BY rank LIMIT ?"
|
||||||
|
params.append(limit)
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute(sql, params)
|
||||||
|
return [
|
||||||
|
SearchHit(
|
||||||
|
video_id=r["video_id"], channel_id=r["channel_id"], title=r["title"],
|
||||||
|
idx=r["idx"], start_sec=r["start_sec"], end_sec=r["end_sec"],
|
||||||
|
snippet=r["text"], rank=r["rank"],
|
||||||
|
)
|
||||||
|
for r in cur.fetchall()
|
||||||
|
]
|
||||||
|
|
||||||
|
# ------------------------------------------------------------------ cookies
|
||||||
|
|
||||||
|
def upsert_cookie(self, cookie_id: str, filename: str, label: str | None, expires_at: str | None, has_session: bool, cookie_count: int) -> None:
|
||||||
|
now = _now_iso()
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute(
|
||||||
|
"""INSERT INTO cookies_meta (id, filename, label, added_at, expires_at, is_active, has_session, cookie_count)
|
||||||
|
VALUES (?, ?, ?, ?, ?, 0, ?, ?)
|
||||||
|
ON CONFLICT(id) DO UPDATE SET
|
||||||
|
filename = excluded.filename,
|
||||||
|
label = excluded.label,
|
||||||
|
expires_at = excluded.expires_at,
|
||||||
|
has_session = excluded.has_session,
|
||||||
|
cookie_count = excluded.cookie_count""",
|
||||||
|
(cookie_id, filename, label, now, expires_at, int(has_session), cookie_count),
|
||||||
|
)
|
||||||
|
|
||||||
|
def list_cookies(self) -> list[CookieRow]:
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute("SELECT * FROM cookies_meta ORDER BY added_at DESC")
|
||||||
|
return [_row_to_cookierow(r) for r in cur.fetchall()]
|
||||||
|
|
||||||
|
def get_cookie(self, cookie_id: str) -> CookieRow | None:
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute("SELECT * FROM cookies_meta WHERE id = ?", (cookie_id,))
|
||||||
|
r = cur.fetchone()
|
||||||
|
return _row_to_cookierow(r) if r else None
|
||||||
|
|
||||||
|
def set_active_cookie(self, cookie_id: str | None) -> None:
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute("UPDATE cookies_meta SET is_active = 0")
|
||||||
|
if cookie_id:
|
||||||
|
cur.execute("UPDATE cookies_meta SET is_active = 1 WHERE id = ?", (cookie_id,))
|
||||||
|
|
||||||
|
def get_active_cookie(self) -> CookieRow | None:
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute("SELECT * FROM cookies_meta WHERE is_active = 1 LIMIT 1")
|
||||||
|
r = cur.fetchone()
|
||||||
|
return _row_to_cookierow(r) if r else None
|
||||||
|
|
||||||
|
def delete_cookie(self, cookie_id: str) -> None:
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute("DELETE FROM cookies_meta WHERE id = ?", (cookie_id,))
|
||||||
|
|
||||||
|
# ------------------------------------------------------------------ jobs
|
||||||
|
|
||||||
|
def create_job(self, job_id: str, channel_id: str | None, opts: dict) -> None:
|
||||||
|
now = _now_iso()
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute(
|
||||||
|
"""INSERT INTO scrape_jobs (id, channel_id, opts_json, status, started_at, total, completed)
|
||||||
|
VALUES (?, ?, ?, 'queued', ?, 0, 0)""",
|
||||||
|
(job_id, channel_id, json.dumps(opts, ensure_ascii=False), now),
|
||||||
|
)
|
||||||
|
|
||||||
|
def update_job(self, job_id: str, *, status: str | None = None, total: int | None = None, completed: int | None = None, last_error: str | None = None, finished: bool = False) -> None:
|
||||||
|
sets: list[str] = []
|
||||||
|
params: list[Any] = []
|
||||||
|
if status:
|
||||||
|
sets.append("status = ?"); params.append(status)
|
||||||
|
if total is not None:
|
||||||
|
sets.append("total = ?"); params.append(total)
|
||||||
|
if completed is not None:
|
||||||
|
sets.append("completed = ?"); params.append(completed)
|
||||||
|
if last_error is not None:
|
||||||
|
sets.append("last_error = ?"); params.append(last_error[:500])
|
||||||
|
if finished:
|
||||||
|
sets.append("finished_at = ?"); params.append(_now_iso())
|
||||||
|
if not sets:
|
||||||
|
return
|
||||||
|
params.append(job_id)
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute(f"UPDATE scrape_jobs SET {', '.join(sets)} WHERE id = ?", params)
|
||||||
|
|
||||||
|
def get_job(self, job_id: str) -> JobRow | None:
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute("SELECT * FROM scrape_jobs WHERE id = ?", (job_id,))
|
||||||
|
r = cur.fetchone()
|
||||||
|
return _row_to_jobrow(r) if r else None
|
||||||
|
|
||||||
|
def list_jobs(self, limit: int = 50) -> list[JobRow]:
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute("SELECT * FROM scrape_jobs ORDER BY started_at DESC LIMIT ?", (limit,))
|
||||||
|
return [_row_to_jobrow(r) for r in cur.fetchall()]
|
||||||
|
|
||||||
|
# ------------------------------------------------------------------ aggregates
|
||||||
|
|
||||||
|
def stats(self, channel_id: str | None = None) -> dict[str, int]:
|
||||||
|
sql = "SELECT status, COUNT(*) as n FROM videos"
|
||||||
|
params: list[Any] = []
|
||||||
|
if channel_id:
|
||||||
|
sql += " WHERE channel_id = ?"
|
||||||
|
params.append(channel_id)
|
||||||
|
sql += " GROUP BY status"
|
||||||
|
result: dict[str, int] = {}
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute(sql, params)
|
||||||
|
for row in cur.fetchall():
|
||||||
|
result[row["status"]] = row["n"]
|
||||||
|
return result
|
||||||
|
|
||||||
|
def dashboard(self) -> dict:
|
||||||
|
with self._cursor() as cur:
|
||||||
|
channels = [dict(r) for r in cur.execute("SELECT * FROM channels ORDER BY name").fetchall()]
|
||||||
|
for ch in channels:
|
||||||
|
cid = ch["channel_id"]
|
||||||
|
row = cur.execute(
|
||||||
|
"SELECT COUNT(*) AS n, COALESCE(SUM(duration),0) AS dur, MIN(upload_date) AS mind, MAX(upload_date) AS maxd, COALESCE(SUM(view_count),0) AS views, COALESCE(SUM(like_count),0) AS likes FROM videos WHERE channel_id = ?",
|
||||||
|
(cid,),
|
||||||
|
).fetchone()
|
||||||
|
ch["video_count_db"] = row["n"]
|
||||||
|
ch["total_duration"] = row["dur"]
|
||||||
|
ch["date_min"] = row["mind"]
|
||||||
|
ch["date_max"] = row["maxd"]
|
||||||
|
ch["total_views"] = row["views"]
|
||||||
|
ch["total_likes"] = row["likes"]
|
||||||
|
status_breakdown = {
|
||||||
|
r["status"]: r["n"]
|
||||||
|
for r in cur.execute("SELECT status, COUNT(*) AS n FROM videos GROUP BY status").fetchall()
|
||||||
|
}
|
||||||
|
# uploads over time (by month)
|
||||||
|
uploads = [
|
||||||
|
{"month": r["m"], "count": r["n"]}
|
||||||
|
for r in cur.execute(
|
||||||
|
"SELECT substr(upload_date,1,6) AS m, COUNT(*) AS n FROM videos WHERE upload_date IS NOT NULL GROUP BY m ORDER BY m"
|
||||||
|
).fetchall()
|
||||||
|
]
|
||||||
|
# duration histogram (buckets)
|
||||||
|
hist = [
|
||||||
|
{"bucket": r["b"], "count": r["n"]}
|
||||||
|
for r in cur.execute(
|
||||||
|
"""SELECT
|
||||||
|
CASE WHEN duration < 300 THEN '<5m'
|
||||||
|
WHEN duration < 600 THEN '5-10m'
|
||||||
|
WHEN duration < 1200 THEN '10-20m'
|
||||||
|
WHEN duration < 2400 THEN '20-40m'
|
||||||
|
ELSE '40m+' END AS b,
|
||||||
|
COUNT(*) AS n
|
||||||
|
FROM videos WHERE duration IS NOT NULL GROUP BY b"""
|
||||||
|
).fetchall()
|
||||||
|
]
|
||||||
|
# top tags (tags is JSON array text)
|
||||||
|
tag_rows = cur.execute("SELECT tags FROM videos WHERE tags IS NOT NULL AND tags != '[]'").fetchall()
|
||||||
|
tag_counts: dict[str, int] = {}
|
||||||
|
for tr in tag_rows:
|
||||||
|
try:
|
||||||
|
for t in json.loads(tr["tags"]):
|
||||||
|
t = (t or "").strip().lower()
|
||||||
|
if t:
|
||||||
|
tag_counts[t] = tag_counts.get(t, 0) + 1
|
||||||
|
except (json.JSONDecodeError, TypeError):
|
||||||
|
continue
|
||||||
|
top_tags = sorted(tag_counts.items(), key=lambda x: x[1], reverse=True)[:20]
|
||||||
|
return {
|
||||||
|
"channels": channels,
|
||||||
|
"status_breakdown": status_breakdown,
|
||||||
|
"uploads_over_time": uploads,
|
||||||
|
"duration_histogram": hist,
|
||||||
|
"top_tags": [{"tag": t, "count": c} for t, c in top_tags],
|
||||||
|
}
|
||||||
|
|
||||||
|
def reset_errors(self, channel_id: str | None = None) -> int:
|
||||||
|
sql = "UPDATE videos SET status = 'pending', error_msg = NULL WHERE status = 'error'"
|
||||||
|
params: list[Any] = []
|
||||||
|
if channel_id:
|
||||||
|
sql += " AND channel_id = ?"
|
||||||
|
params.append(channel_id)
|
||||||
|
with self._cursor() as cur:
|
||||||
|
cur.execute(sql, params)
|
||||||
|
return cur.rowcount
|
||||||
|
|
||||||
|
|
||||||
|
def _sanitize_fts(query: str) -> str:
|
||||||
|
# Build a safe AND FTS5 query from whitespace-separated terms.
|
||||||
|
cleaned = (query or "").strip()
|
||||||
|
if not cleaned:
|
||||||
|
return ""
|
||||||
|
terms = []
|
||||||
|
for tok in cleaned.split():
|
||||||
|
tok = tok.strip('"')
|
||||||
|
if tok:
|
||||||
|
terms.append(f'"{tok.replace("\"", "")}"')
|
||||||
|
return " AND ".join(terms)
|
||||||
|
|
||||||
|
|
||||||
|
def _now_iso() -> str:
|
||||||
|
return datetime.now(timezone.utc).isoformat(timespec="seconds")
|
||||||
|
|
||||||
|
|
||||||
|
def _row_to_videorow(row: sqlite3.Row) -> VideoRow:
|
||||||
|
keys = row.keys()
|
||||||
|
return VideoRow(
|
||||||
|
video_id=row["video_id"],
|
||||||
|
channel_id=row["channel_id"],
|
||||||
|
title=row["title"],
|
||||||
|
url=row["url"],
|
||||||
|
upload_date=row["upload_date"],
|
||||||
|
duration=row["duration"],
|
||||||
|
status=row["status"],
|
||||||
|
error_msg=row["error_msg"],
|
||||||
|
transcript_lang=row["transcript_lang"],
|
||||||
|
transcript_src=row["transcript_src"],
|
||||||
|
has_chapters=row["has_chapters"],
|
||||||
|
markdown_path=row["markdown_path"],
|
||||||
|
view_count=row["view_count"] if "view_count" in keys else None,
|
||||||
|
like_count=row["like_count"] if "like_count" in keys else None,
|
||||||
|
tags=row["tags"] if "tags" in keys else None,
|
||||||
|
thumbnail=row["thumbnail"] if "thumbnail" in keys else None,
|
||||||
|
description=row["description"] if "description" in keys else None,
|
||||||
|
chapters_json=row["chapters_json"] if "chapters_json" in keys else None,
|
||||||
|
segments_json=row["segments_json"] if "segments_json" in keys else None,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _row_to_cookierow(row: sqlite3.Row) -> CookieRow:
|
||||||
|
return CookieRow(
|
||||||
|
id=row["id"],
|
||||||
|
filename=row["filename"],
|
||||||
|
label=row["label"],
|
||||||
|
added_at=row["added_at"],
|
||||||
|
expires_at=row["expires_at"],
|
||||||
|
is_active=bool(row["is_active"]),
|
||||||
|
has_session=bool(row["has_session"]),
|
||||||
|
cookie_count=row["cookie_count"],
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _row_to_jobrow(row: sqlite3.Row) -> JobRow:
|
||||||
|
return JobRow(
|
||||||
|
id=row["id"],
|
||||||
|
channel_id=row["channel_id"],
|
||||||
|
opts_json=row["opts_json"],
|
||||||
|
status=row["status"],
|
||||||
|
started_at=row["started_at"],
|
||||||
|
finished_at=row["finished_at"],
|
||||||
|
total=row["total"],
|
||||||
|
completed=row["completed"],
|
||||||
|
last_error=row["last_error"],
|
||||||
|
)
|
||||||
@@ -0,0 +1,371 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from fastapi import APIRouter, HTTPException, Query, Request, UploadFile, File
|
||||||
|
from fastapi.responses import FileResponse, JSONResponse, StreamingResponse
|
||||||
|
from sse_starlette.sse import EventSourceResponse
|
||||||
|
|
||||||
|
from .. import analysis as analysis_mod
|
||||||
|
from .. import cookies as cookies_mod
|
||||||
|
from .. import export as export_mod
|
||||||
|
from ..config import Config
|
||||||
|
from ..store import Store
|
||||||
|
|
||||||
|
|
||||||
|
def build_router(store: Store, cfg: Config, jobs) -> APIRouter:
|
||||||
|
r = APIRouter(prefix="/api")
|
||||||
|
|
||||||
|
# -------------------------------------------------- dashboard
|
||||||
|
@r.get("/dashboard")
|
||||||
|
def dashboard():
|
||||||
|
return store.dashboard()
|
||||||
|
|
||||||
|
# -------------------------------------------------- channels
|
||||||
|
@r.get("/channels")
|
||||||
|
def channels_list():
|
||||||
|
return {"items": store.list_channels()}
|
||||||
|
|
||||||
|
@r.post("/channels")
|
||||||
|
async def add_channel(payload: dict):
|
||||||
|
from ..discover import discover_channel
|
||||||
|
url = (payload or {}).get("url")
|
||||||
|
if not url:
|
||||||
|
raise HTTPException(400, "url required")
|
||||||
|
cid, name, refs = discover_channel(url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests)
|
||||||
|
store.upsert_channel(cid, _handle(url), name, len(refs))
|
||||||
|
store.upsert_videos(refs)
|
||||||
|
return {"channel_id": cid, "name": name, "video_count": len(refs)}
|
||||||
|
|
||||||
|
@r.delete("/channels/{channel_id}")
|
||||||
|
def del_channel(channel_id: str):
|
||||||
|
store.delete_channel(channel_id)
|
||||||
|
return {"deleted": channel_id}
|
||||||
|
|
||||||
|
# -------------------------------------------------- videos
|
||||||
|
@r.get("/videos")
|
||||||
|
def videos(
|
||||||
|
channel: str | None = None, status: str | None = None,
|
||||||
|
date_from: str | None = Query(None, alias="from"),
|
||||||
|
date_to: str | None = Query(None, alias="to"),
|
||||||
|
min_dur: int | None = None, q: str | None = None,
|
||||||
|
sort: str = "upload_date_desc", page: int = 1, size: int = 50,
|
||||||
|
):
|
||||||
|
rows, total = store.query_videos(channel, status, date_from, date_to, min_dur, q, sort, page, size)
|
||||||
|
return {"items": [_video_dict(v) for v in rows], "total": total, "page": page, "size": size}
|
||||||
|
|
||||||
|
@r.get("/videos/{video_id}")
|
||||||
|
def video_detail(video_id: str):
|
||||||
|
v = store.get_video(video_id)
|
||||||
|
if not v:
|
||||||
|
raise HTTPException(404, "video not found")
|
||||||
|
return _video_dict(v)
|
||||||
|
|
||||||
|
@r.get("/videos/{video_id}/transcript")
|
||||||
|
def video_transcript(video_id: str):
|
||||||
|
segs = store.get_segments(video_id)
|
||||||
|
return {"items": [
|
||||||
|
{"idx": s.idx, "start_sec": s.start_sec, "end_sec": s.end_sec, "text": s.text}
|
||||||
|
for s in segs
|
||||||
|
]}
|
||||||
|
|
||||||
|
# -------------------------------------------------- search
|
||||||
|
@r.get("/search")
|
||||||
|
def search(q: str, channel: str | None = None, limit: int = 30):
|
||||||
|
hits = store.search_segments(q, channel_id=channel, limit=limit)
|
||||||
|
return {"items": [
|
||||||
|
{"video_id": h.video_id, "title": h.title, "channel_id": h.channel_id,
|
||||||
|
"start_sec": h.start_sec, "end_sec": h.end_sec, "snippet": h.snippet}
|
||||||
|
for h in hits
|
||||||
|
]}
|
||||||
|
|
||||||
|
# -------------------------------------------------- analysis
|
||||||
|
@r.get("/analysis")
|
||||||
|
def analysis(channel: str | None = None, type: str = "topwords", term: str | None = None, n: int = 50):
|
||||||
|
if type == "topwords":
|
||||||
|
tw = analysis_mod.top_words(store, n, channel)
|
||||||
|
return {"items": [{"word": w, "count": c} for w, c in tw]}
|
||||||
|
if type == "timeline":
|
||||||
|
tl = analysis_mod.term_timeline(store, term or "", channel)
|
||||||
|
return {"term": term, "items": [{"month": m, "count": c} for m, c in tl]}
|
||||||
|
if type == "wordcloud":
|
||||||
|
freq = analysis_mod.word_frequency(store, channel, top=300)
|
||||||
|
return {"items": [{"word": w, "count": c} for w, c in freq.items()]}
|
||||||
|
raise HTTPException(400, "unknown analysis type")
|
||||||
|
|
||||||
|
# -------------------------------------------------- scrape jobs
|
||||||
|
@r.post("/scrape")
|
||||||
|
async def start_scrape(payload: dict):
|
||||||
|
channel_id = (payload or {}).get("channel_id")
|
||||||
|
opts = (payload or {}).get("opts", {})
|
||||||
|
job_id = jobs.enqueue(channel_id, opts)
|
||||||
|
return {"job_id": job_id}
|
||||||
|
|
||||||
|
@r.get("/scrape")
|
||||||
|
def list_jobs(limit: int = 30):
|
||||||
|
return {"items": [_job_dict(j) for j in store.list_jobs(limit)]}
|
||||||
|
|
||||||
|
@r.get("/scrape/{job_id}")
|
||||||
|
def get_job(job_id: str):
|
||||||
|
j = store.get_job(job_id)
|
||||||
|
if not j:
|
||||||
|
raise HTTPException(404, "job not found")
|
||||||
|
return _job_dict(j)
|
||||||
|
|
||||||
|
@r.get("/scrape/{job_id}/stream")
|
||||||
|
async def stream_job(job_id: str, request: Request):
|
||||||
|
import asyncio
|
||||||
|
async def gen():
|
||||||
|
cursor = 0
|
||||||
|
while True:
|
||||||
|
if await request.is_disconnected():
|
||||||
|
return
|
||||||
|
for ev in jobs.events_since(job_id, cursor):
|
||||||
|
cursor += 1
|
||||||
|
yield {"event": ev["event"], "data": json.dumps(ev["data"])}
|
||||||
|
if ev["event"] in ("done", "error", "cancelled"):
|
||||||
|
return
|
||||||
|
terminal = jobs.is_terminal(job_id)
|
||||||
|
if terminal and not jobs.events_since(job_id, cursor):
|
||||||
|
yield {"event": terminal, "data": json.dumps({"status": terminal})}
|
||||||
|
return
|
||||||
|
await asyncio.sleep(0.25)
|
||||||
|
return EventSourceResponse(gen())
|
||||||
|
|
||||||
|
@r.delete("/scrape/{job_id}")
|
||||||
|
def cancel_job(job_id: str):
|
||||||
|
ok = jobs.cancel(job_id)
|
||||||
|
return {"cancelled": ok}
|
||||||
|
|
||||||
|
# -------------------------------------------------- cookies
|
||||||
|
@r.get("/cookies")
|
||||||
|
def cookies_list():
|
||||||
|
items = []
|
||||||
|
for c in cookies_mod.list_vault(store):
|
||||||
|
items.append({
|
||||||
|
"id": c.id, "filename": c.filename, "label": c.label,
|
||||||
|
"added_at": c.added_at, "expires_at": c.expires_at,
|
||||||
|
"is_active": c.is_active, "has_session": c.has_session,
|
||||||
|
"cookie_count": c.cookie_count, "expired": cookies_mod.is_expired(c),
|
||||||
|
})
|
||||||
|
return {"items": items}
|
||||||
|
|
||||||
|
@r.post("/cookies/upload")
|
||||||
|
async def cookies_upload(request: Request, files: list[UploadFile] = File(default=[])):
|
||||||
|
# accept multipart files OR a JSON body {text, label}
|
||||||
|
content_type = request.headers.get("content-type", "")
|
||||||
|
imported = []
|
||||||
|
if "application/json" in content_type:
|
||||||
|
payload = await request.json()
|
||||||
|
text = payload.get("text", "")
|
||||||
|
label = payload.get("label") or "pasted"
|
||||||
|
if not text:
|
||||||
|
raise HTTPException(400, "no text")
|
||||||
|
cid = cookies_mod.import_text(store, text, label)
|
||||||
|
imported.append(cid)
|
||||||
|
else:
|
||||||
|
if not files:
|
||||||
|
raise HTTPException(400, "no files")
|
||||||
|
for f in files:
|
||||||
|
raw = await f.read()
|
||||||
|
text = raw.decode("utf-8", errors="replace")
|
||||||
|
try:
|
||||||
|
cid = cookies_mod.import_text(store, text, label=Path(f.filename or "upload").stem)
|
||||||
|
imported.append(cid)
|
||||||
|
except ValueError:
|
||||||
|
continue
|
||||||
|
if not imported:
|
||||||
|
raise HTTPException(400, "no valid Netscape cookie files")
|
||||||
|
# activate the first if nothing active
|
||||||
|
if not store.get_active_cookie():
|
||||||
|
cookies_mod.set_active(store, imported[0])
|
||||||
|
return {"imported": imported}
|
||||||
|
|
||||||
|
@r.post("/cookies/{cookie_id}/activate")
|
||||||
|
def cookies_activate(cookie_id: str):
|
||||||
|
cookies_mod.set_active(store, cookie_id)
|
||||||
|
return {"active": cookie_id}
|
||||||
|
|
||||||
|
@r.delete("/cookies/{cookie_id}")
|
||||||
|
def cookies_delete(cookie_id: str):
|
||||||
|
cookies_mod.delete(store, cookie_id)
|
||||||
|
return {"deleted": cookie_id}
|
||||||
|
|
||||||
|
@r.post("/cookies/{cookie_id}/test")
|
||||||
|
def cookies_test(cookie_id: str):
|
||||||
|
c = store.get_cookie(cookie_id)
|
||||||
|
if not c:
|
||||||
|
raise HTTPException(404, "cookie not found")
|
||||||
|
path = cookies_mod.resolve_active_path(store) if c.is_active else str(cookies_mod.cookies_dir() / c.filename)
|
||||||
|
if not Path(path).exists():
|
||||||
|
return {"ok": False, "message": "cookie file missing"}
|
||||||
|
# lightweight validation via parse
|
||||||
|
ok, info = cookies_mod.parse_netscape_file(path)
|
||||||
|
expired = cookies_mod.is_expired(c)
|
||||||
|
return {"ok": ok and not expired, "has_session": info["has_session"], "expired": expired,
|
||||||
|
"expires_at": c.expires_at, "cookie_count": info["count"]}
|
||||||
|
|
||||||
|
# -------------------------------------------------- export
|
||||||
|
@r.get("/export")
|
||||||
|
def export(format: str = "json", channel: str | None = None):
|
||||||
|
import tempfile, os
|
||||||
|
tmp = Path(tempfile.gettempdir()) / "yt_scraper_export"
|
||||||
|
tmp.mkdir(parents=True, exist_ok=True)
|
||||||
|
label = channel or "all"
|
||||||
|
try:
|
||||||
|
if format == "json":
|
||||||
|
p = export_mod.export_json(store, channel, tmp / f"{label}.json")
|
||||||
|
elif format == "csv":
|
||||||
|
p = export_mod.export_csv(store, channel, tmp / f"{label}.csv")
|
||||||
|
elif format == "html":
|
||||||
|
p = export_mod.export_html(store, channel, tmp / f"{label}.html")
|
||||||
|
elif format == "srt":
|
||||||
|
# stream a zip-less approach: return count + note (SRT is multi-file)
|
||||||
|
files = export_mod.export_srt(store, channel, tmp / "srt")
|
||||||
|
return {"files": [str(f.name) for f in files], "dir": str(tmp / "srt")}
|
||||||
|
else:
|
||||||
|
raise HTTPException(400, "unknown format")
|
||||||
|
except Exception as exc:
|
||||||
|
raise HTTPException(500, str(exc))
|
||||||
|
return FileResponse(str(p), filename=p.name)
|
||||||
|
|
||||||
|
# -------------------------------------------------- folders (open in OS)
|
||||||
|
@r.get("/folders")
|
||||||
|
def folders():
|
||||||
|
data_root = Path(cfg.output_dir_resolved).parent
|
||||||
|
return {"items": {
|
||||||
|
"markdown": _folder_info(Path(cfg.output_dir_resolved)),
|
||||||
|
"exports": _folder_info(data_root / "exports"),
|
||||||
|
"audio": _folder_info(data_root / "audio"),
|
||||||
|
"analysis": _folder_info(data_root / "analysis"),
|
||||||
|
"database": _folder_info(Path(cfg.database_path_resolved).parent),
|
||||||
|
}}
|
||||||
|
|
||||||
|
@r.post("/folders/open")
|
||||||
|
async def open_folder(payload: dict):
|
||||||
|
kind = (payload or {}).get("kind")
|
||||||
|
channel_id = (payload or {}).get("channel_id")
|
||||||
|
data_root = Path(cfg.output_dir_resolved).parent
|
||||||
|
if kind == "markdown":
|
||||||
|
target = Path(cfg.output_dir_resolved)
|
||||||
|
if channel_id:
|
||||||
|
ch = store.get_channel(channel_id) or {}
|
||||||
|
if ch.get("name"):
|
||||||
|
target = Path(cfg.output_dir_resolved) / _safe_dir(ch["name"])
|
||||||
|
elif kind == "exports":
|
||||||
|
target = data_root / "exports"
|
||||||
|
elif kind == "audio":
|
||||||
|
target = data_root / "audio"
|
||||||
|
elif kind == "analysis":
|
||||||
|
target = data_root / "analysis"
|
||||||
|
elif kind == "database":
|
||||||
|
target = Path(cfg.database_path_resolved).parent
|
||||||
|
elif kind == "path" and (payload or {}).get("path"):
|
||||||
|
req = Path(payload["path"]).resolve()
|
||||||
|
proj = Path.cwd().resolve()
|
||||||
|
try:
|
||||||
|
req.relative_to(proj)
|
||||||
|
except ValueError:
|
||||||
|
raise HTTPException(403, "path outside project root")
|
||||||
|
target = req
|
||||||
|
else:
|
||||||
|
raise HTTPException(400, "unknown kind")
|
||||||
|
target.mkdir(parents=True, exist_ok=True)
|
||||||
|
_open_in_os(target)
|
||||||
|
return {"opened": str(target)}
|
||||||
|
|
||||||
|
# -------------------------------------------------- tools (knowledge base)
|
||||||
|
@r.post("/tools/re-render")
|
||||||
|
def tools_rerender(payload: dict | None = None):
|
||||||
|
from ..pipeline import re_render_videos
|
||||||
|
from ..segments import backfill_from_markdown
|
||||||
|
# ensure segments are populated first (idempotent)
|
||||||
|
backfill_from_markdown(store, Path(cfg.output_dir_resolved))
|
||||||
|
channel_id = (payload or {}).get("channel_id") if payload else None
|
||||||
|
n = re_render_videos(store, cfg, channel_id)
|
||||||
|
return {"re_rendered": n}
|
||||||
|
|
||||||
|
@r.get("/tools/format")
|
||||||
|
def tools_format():
|
||||||
|
# describe the established .md knowledge-base format
|
||||||
|
return {
|
||||||
|
"template": "templates/video.md.j2",
|
||||||
|
"frontmatter": [
|
||||||
|
"video_id", "title", "channel", "channel_id", "channel_url",
|
||||||
|
"upload_date", "duration", "url", "transcript_lang", "transcript_src",
|
||||||
|
"views", "likes", "tags", "thumbnail",
|
||||||
|
],
|
||||||
|
"body": "H1 title -> YouTube link -> (optional description) -> ## Transcripcion -> ### Chapter (MM:SS) -> **MM:SS** · segment lines",
|
||||||
|
"output_dir": str(Path(cfg.output_dir_resolved)),
|
||||||
|
}
|
||||||
|
|
||||||
|
return r
|
||||||
|
|
||||||
|
|
||||||
|
def _video_dict(v) -> dict:
|
||||||
|
import json as _json
|
||||||
|
tags = []
|
||||||
|
if v.tags:
|
||||||
|
try:
|
||||||
|
tags = _json.loads(v.tags)
|
||||||
|
except Exception:
|
||||||
|
tags = []
|
||||||
|
return {
|
||||||
|
"video_id": v.video_id, "channel_id": v.channel_id, "title": v.title, "url": v.url,
|
||||||
|
"upload_date": v.upload_date, "duration": v.duration, "status": v.status,
|
||||||
|
"transcript_lang": v.transcript_lang, "has_chapters": bool(v.has_chapters),
|
||||||
|
"view_count": v.view_count, "like_count": v.like_count, "tags": tags,
|
||||||
|
"thumbnail": v.thumbnail, "markdown_path": v.markdown_path,
|
||||||
|
"error_msg": v.error_msg,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _job_dict(j) -> dict:
|
||||||
|
return {
|
||||||
|
"id": j.id, "channel_id": j.channel_id, "opts": _loads(j.opts_json),
|
||||||
|
"status": j.status, "started_at": j.started_at, "finished_at": j.finished_at,
|
||||||
|
"total": j.total, "completed": j.completed, "last_error": j.last_error,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _loads(s):
|
||||||
|
if not s:
|
||||||
|
return {}
|
||||||
|
try:
|
||||||
|
return json.loads(s)
|
||||||
|
except Exception:
|
||||||
|
return {}
|
||||||
|
|
||||||
|
|
||||||
|
def _handle(url: str) -> str:
|
||||||
|
if "@" in url:
|
||||||
|
return "@" + url.split("@", 1)[1].split("/", 1)[0]
|
||||||
|
return ""
|
||||||
|
|
||||||
|
|
||||||
|
def _folder_info(path: Path) -> dict:
|
||||||
|
path = Path(path)
|
||||||
|
return {"path": str(path.resolve()), "exists": path.exists()}
|
||||||
|
|
||||||
|
|
||||||
|
def _safe_dir(name: str) -> str:
|
||||||
|
safe = "".join(c for c in (name or "") if c not in r'\/:*?"<>|')
|
||||||
|
return (safe.strip().strip(".") or "unknown")
|
||||||
|
|
||||||
|
|
||||||
|
def _open_in_os(path: Path) -> None:
|
||||||
|
import os
|
||||||
|
import sys
|
||||||
|
import subprocess
|
||||||
|
p = str(path)
|
||||||
|
try:
|
||||||
|
if sys.platform.startswith("win"):
|
||||||
|
os.startfile(p) # type: ignore[attr-defined]
|
||||||
|
elif sys.platform == "darwin":
|
||||||
|
subprocess.Popen(["open", p])
|
||||||
|
else:
|
||||||
|
subprocess.Popen(["xdg-open", p])
|
||||||
|
except Exception as exc:
|
||||||
|
raise HTTPException(500, f"cannot open folder: {exc}")
|
||||||
@@ -0,0 +1,84 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from fastapi import FastAPI
|
||||||
|
from fastapi.responses import FileResponse, HTMLResponse
|
||||||
|
from fastapi.staticfiles import StaticFiles
|
||||||
|
|
||||||
|
from ..config import Config, load_config
|
||||||
|
from ..cookies import auto_import_dir
|
||||||
|
from ..segments import backfill_from_markdown
|
||||||
|
from ..store import Store
|
||||||
|
from .api import build_router
|
||||||
|
from .jobs import JobManager
|
||||||
|
|
||||||
|
|
||||||
|
_STATIC_DIR = Path(__file__).resolve().parent / "static"
|
||||||
|
|
||||||
|
|
||||||
|
def create_app(cfg: Config | None = None, db_path: str | Path | None = None) -> FastAPI:
|
||||||
|
cfg = cfg or _load_cfg()
|
||||||
|
if db_path:
|
||||||
|
cfg.database_path = str(db_path)
|
||||||
|
store = Store(cfg.database_path_resolved)
|
||||||
|
|
||||||
|
# auto-import any loose cookies into the vault
|
||||||
|
try:
|
||||||
|
auto_import_dir(store)
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
# auto-backfill segments from existing markdown (idempotent)
|
||||||
|
try:
|
||||||
|
md_root = cfg.output_dir_resolved
|
||||||
|
if Path(md_root).exists():
|
||||||
|
backfill_from_markdown(store, Path(md_root))
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
|
||||||
|
app = FastAPI(title="yt-scraper platform", version="1.0.0")
|
||||||
|
|
||||||
|
@app.get("/healthz")
|
||||||
|
def healthz():
|
||||||
|
return {"status": "ok"}
|
||||||
|
|
||||||
|
jobs = JobManager(store, cfg)
|
||||||
|
app.state.jobs = jobs
|
||||||
|
app.state.store = store
|
||||||
|
app.state.cfg = cfg
|
||||||
|
|
||||||
|
app.include_router(build_router(store, cfg, jobs))
|
||||||
|
|
||||||
|
if _STATIC_DIR.exists():
|
||||||
|
app.mount("/static", StaticFiles(directory=str(_STATIC_DIR)), name="static")
|
||||||
|
|
||||||
|
@app.get("/", response_class=HTMLResponse)
|
||||||
|
def index():
|
||||||
|
return FileResponse(str(_STATIC_DIR / "index.html"))
|
||||||
|
|
||||||
|
@app.get("/{path:path}", response_class=HTMLResponse)
|
||||||
|
def spa_fallback(path: str):
|
||||||
|
# SPA hash routes — serve index for any non-API, non-file path
|
||||||
|
if path.startswith("api/") or path.startswith("static/") or "." in path.split("/")[-1]:
|
||||||
|
from fastapi import HTTPException
|
||||||
|
raise HTTPException(404)
|
||||||
|
return FileResponse(str(_STATIC_DIR / "index.html"))
|
||||||
|
|
||||||
|
@app.on_event("startup")
|
||||||
|
async def _startup():
|
||||||
|
jobs.start()
|
||||||
|
|
||||||
|
return app
|
||||||
|
|
||||||
|
|
||||||
|
def _load_cfg() -> Config:
|
||||||
|
for candidate in ("config.yaml", "config.example.yaml"):
|
||||||
|
if Path(candidate).exists():
|
||||||
|
try:
|
||||||
|
return load_config(candidate)
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
return Config()
|
||||||
|
|
||||||
|
|
||||||
|
app = create_app()
|
||||||
@@ -0,0 +1,179 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import asyncio
|
||||||
|
import json
|
||||||
|
import threading
|
||||||
|
import time
|
||||||
|
import uuid
|
||||||
|
from collections import deque
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from ..config import Config, load_config
|
||||||
|
from ..cookies import resolve_active_path
|
||||||
|
from ..discover import discover_channel
|
||||||
|
from ..pipeline import process_video
|
||||||
|
from ..ratelimit import polite_sleep
|
||||||
|
from ..store import Store
|
||||||
|
|
||||||
|
|
||||||
|
class JobManager:
|
||||||
|
"""Single-worker scrape job runner with an in-memory event log per job (SSE-polled)."""
|
||||||
|
|
||||||
|
def __init__(self, store: Store, cfg: Config):
|
||||||
|
self.store = store
|
||||||
|
self.cfg = cfg
|
||||||
|
self._lock = threading.Lock()
|
||||||
|
self._events: dict[str, list[dict]] = {}
|
||||||
|
self._cancel: set[str] = set()
|
||||||
|
self._queue: deque[str] = deque()
|
||||||
|
self._thread = threading.Thread(target=self._worker, daemon=True, name="scrape-worker")
|
||||||
|
self._running = False
|
||||||
|
|
||||||
|
def start(self) -> None:
|
||||||
|
if self._running:
|
||||||
|
return
|
||||||
|
self._running = True
|
||||||
|
self._thread.start()
|
||||||
|
|
||||||
|
def enqueue(self, channel_id: str | None, opts: dict[str, Any]) -> str:
|
||||||
|
job_id = uuid.uuid4().hex[:12]
|
||||||
|
self.store.create_job(job_id, channel_id, opts)
|
||||||
|
with self._lock:
|
||||||
|
self._events[job_id] = []
|
||||||
|
self._emit(job_id, "queued", {"job_id": job_id, "channel_id": channel_id, "opts": opts})
|
||||||
|
self._queue.append(job_id)
|
||||||
|
return job_id
|
||||||
|
|
||||||
|
def cancel(self, job_id: str) -> bool:
|
||||||
|
job = self.store.get_job(job_id)
|
||||||
|
if not job or job.status not in ("queued", "running"):
|
||||||
|
return False
|
||||||
|
self._cancel.add(job_id)
|
||||||
|
self._emit(job_id, "log", {"msg": "cancel requested"})
|
||||||
|
return True
|
||||||
|
|
||||||
|
def events_since(self, job_id: str, cursor: int) -> list[dict]:
|
||||||
|
with self._lock:
|
||||||
|
evs = self._events.get(job_id, [])
|
||||||
|
return evs[cursor:]
|
||||||
|
|
||||||
|
def is_terminal(self, job_id: str) -> str | None:
|
||||||
|
job = self.store.get_job(job_id)
|
||||||
|
if not job:
|
||||||
|
return None
|
||||||
|
if job.status in ("done", "error", "cancelled"):
|
||||||
|
return job.status
|
||||||
|
return None
|
||||||
|
|
||||||
|
# ---------------------------------------------------------- internals
|
||||||
|
|
||||||
|
def _emit(self, job_id: str, event: str, data: dict) -> None:
|
||||||
|
with self._lock:
|
||||||
|
self._events.setdefault(job_id, []).append({"event": event, "data": data})
|
||||||
|
|
||||||
|
def _worker(self) -> None:
|
||||||
|
while True:
|
||||||
|
try:
|
||||||
|
job_id = self._queue.popleft()
|
||||||
|
except IndexError:
|
||||||
|
time.sleep(0.2)
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
self._run_job(job_id)
|
||||||
|
except Exception as exc:
|
||||||
|
self.store.update_job(job_id, status="error", last_error=str(exc), finished=True)
|
||||||
|
self._emit(job_id, "error", {"message": str(exc)})
|
||||||
|
|
||||||
|
def _run_job(self, job_id: str) -> None:
|
||||||
|
job = self.store.get_job(job_id)
|
||||||
|
if not job:
|
||||||
|
return
|
||||||
|
opts = json.loads(job.opts_json) if job.opts_json else {}
|
||||||
|
channel_id = job.channel_id
|
||||||
|
cfg = _clone_config(self.cfg)
|
||||||
|
channel_url = _resolve_channel_url(self.store, cfg, channel_id)
|
||||||
|
if not channel_url:
|
||||||
|
self.store.update_job(job_id, status="error", last_error="no channel url", finished=True)
|
||||||
|
self._emit(job_id, "error", {"message": "no channel url"})
|
||||||
|
return
|
||||||
|
cfg.channel_url = channel_url
|
||||||
|
|
||||||
|
self.store.update_job(job_id, status="running")
|
||||||
|
self._emit(job_id, "log", {"msg": f"discovering {channel_url}"})
|
||||||
|
try:
|
||||||
|
_channel_id, channel_name, refs = discover_channel(channel_url, sleep_subrequests=cfg.yt_dlp.sleep_subrequests)
|
||||||
|
except Exception as exc:
|
||||||
|
self.store.update_job(job_id, status="error", last_error=str(exc), finished=True)
|
||||||
|
self._emit(job_id, "error", {"message": f"discovery failed: {exc}"})
|
||||||
|
return
|
||||||
|
|
||||||
|
limit = opts.get("limit")
|
||||||
|
since = opts.get("since")
|
||||||
|
languages = opts.get("languages")
|
||||||
|
include_shorts = opts.get("include_shorts", cfg.include_shorts)
|
||||||
|
no_live = opts.get("no_live", not cfg.include_live)
|
||||||
|
cookie_override = opts.get("cookies_file")
|
||||||
|
|
||||||
|
if languages:
|
||||||
|
cfg.languages = languages
|
||||||
|
if since:
|
||||||
|
refs = [r for r in refs if (r.upload_date or "") >= since.replace("-", "")]
|
||||||
|
if not include_shorts:
|
||||||
|
refs = [r for r in refs if "/shorts/" not in (r.url or "")]
|
||||||
|
if no_live:
|
||||||
|
refs = [r for r in refs if not (r.url or "").startswith("https://www.youtube.com/live/")]
|
||||||
|
if limit:
|
||||||
|
refs = refs[: int(limit)]
|
||||||
|
|
||||||
|
self.store.upsert_channel(_channel_id, _handle(channel_url), channel_name, len(refs))
|
||||||
|
self.store.upsert_videos(refs)
|
||||||
|
|
||||||
|
pending = [r for r in self.store.get_pending(_channel_id) if r.video_id in {x.video_id for x in refs}]
|
||||||
|
total = len(pending)
|
||||||
|
self.store.update_job(job_id, total=total)
|
||||||
|
self._emit(job_id, "progress", {"completed": 0, "total": total})
|
||||||
|
self._emit(job_id, "log", {"msg": f"{total} videos to process"})
|
||||||
|
|
||||||
|
cookie_path = cookie_override or resolve_active_path(self.store)
|
||||||
|
completed = 0
|
||||||
|
for row in pending:
|
||||||
|
if job_id in self._cancel:
|
||||||
|
self.store.update_job(job_id, status="cancelled", completed=completed, finished=True)
|
||||||
|
self._emit(job_id, "cancelled", {"completed": completed, "total": total})
|
||||||
|
return
|
||||||
|
status = process_video(
|
||||||
|
row, cfg, self.store, channel_name, _channel_id, channel_url,
|
||||||
|
cookies_file=cookie_path,
|
||||||
|
on_log=lambda m: self._emit(job_id, "log", {"msg": m}),
|
||||||
|
)
|
||||||
|
completed += 1
|
||||||
|
self.store.update_job(job_id, completed=completed)
|
||||||
|
self._emit(job_id, "progress", {"completed": completed, "total": total, "video_id": row.video_id, "status": status})
|
||||||
|
if completed < total:
|
||||||
|
polite_sleep(cfg.delay.min_seconds, cfg.delay.max_seconds)
|
||||||
|
|
||||||
|
self.store.update_job(job_id, status="done", completed=completed, finished=True)
|
||||||
|
self._emit(job_id, "done", {"completed": completed, "total": total})
|
||||||
|
|
||||||
|
|
||||||
|
def _clone_config(cfg: Config) -> Config:
|
||||||
|
import copy
|
||||||
|
return copy.deepcopy(cfg)
|
||||||
|
|
||||||
|
|
||||||
|
def _resolve_channel_url(store: Store, cfg: Config, channel_id: str | None) -> str | None:
|
||||||
|
if not channel_id:
|
||||||
|
return cfg.channel_url or None
|
||||||
|
ch = store.get_channel(channel_id)
|
||||||
|
if not ch:
|
||||||
|
return None
|
||||||
|
handle = (ch.get("handle") or "").lstrip("@")
|
||||||
|
if handle:
|
||||||
|
return f"https://www.youtube.com/@{handle}/videos"
|
||||||
|
return f"https://www.youtube.com/channel/{channel_id}"
|
||||||
|
|
||||||
|
|
||||||
|
def _handle(url: str) -> str:
|
||||||
|
if "@" in url:
|
||||||
|
return "@" + url.split("@", 1)[1].split("/", 1)[0]
|
||||||
|
return ""
|
||||||
@@ -0,0 +1,596 @@
|
|||||||
|
/* yt-scraper platform — Alpine SPA controller.
|
||||||
|
* Defines window.platform BEFORE Alpine initializes (both scripts are deferred;
|
||||||
|
* app.js is listed before alpinejs so this runs first).
|
||||||
|
*/
|
||||||
|
(function () {
|
||||||
|
"use strict";
|
||||||
|
|
||||||
|
// ---- Chart.js dark defaults ----
|
||||||
|
function setupChartDefaults() {
|
||||||
|
if (!window.Chart) return;
|
||||||
|
Chart.defaults.color = "#a1a1aa";
|
||||||
|
Chart.defaults.borderColor = "rgba(39,39,42,0.7)";
|
||||||
|
Chart.defaults.font.family =
|
||||||
|
'ui-sans-serif, system-ui, "Segoe UI", Roboto, sans-serif';
|
||||||
|
}
|
||||||
|
|
||||||
|
const NAV = [
|
||||||
|
{ id: "dashboard", label: "Dashboard", icon: '<svg viewBox="0 0 24 24" fill="currentColor" class="w-4 h-4"><path d="M3 13h8V3H3v10zm0 8h8v-6H3v6zm10 0h8V11h-8v10zm0-18v6h8V3h-8z"/></svg>' },
|
||||||
|
{ id: "channels", label: "Channels", icon: '<svg viewBox="0 0 24 24" fill="currentColor" class="w-4 h-4"><path d="M4 4h16v3H4zm0 6h16v3H4zm0 6h10v3H4z"/></svg>' },
|
||||||
|
{ id: "videos", label: "Videos", icon: '<svg viewBox="0 0 24 24" fill="currentColor" class="w-4 h-4"><path d="M4 6h16v12H4zM10 9l5 3-5 3z"/></svg>' },
|
||||||
|
{ id: "search", label: "Search", icon: '<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><circle cx="11" cy="11" r="7"/><path d="M21 21l-4.3-4.3" stroke-linecap="round"/></svg>' },
|
||||||
|
{ id: "analysis", label: "Analysis", icon: '<svg viewBox="0 0 24 24" fill="currentColor" class="w-4 h-4"><path d="M4 20h16v2H4zM6 16V8h3v8zm4 0V4h3v12zm4 0v-6h3v6z"/></svg>' },
|
||||||
|
{ id: "scrape", label: "Scrape", icon: '<svg viewBox="0 0 24 24" fill="currentColor" class="w-4 h-4"><path d="M12 4V1L8 5l4 4V6a6 6 0 010 12 6 6 0 01-6-6H4a8 8 0 108-8z"/></svg>' },
|
||||||
|
{ id: "cookies", label: "Cookies", icon: '<svg viewBox="0 0 24 24" fill="currentColor" class="w-4 h-4"><path d="M21 11.5a8.5 8.5 0 11-9-8.46 2.5 2.5 0 003 3 2.5 2.5 0 003 3 2.5 2.5 0 003 2.46zM9 14a1 1 0 100-2 1 1 0 000 2zm3 3a1 1 0 100-2 1 1 0 000 2zm4-4a1 1 0 100-2 1 1 0 000 2z"/></svg>' },
|
||||||
|
{ id: "export", label: "Export", icon: '<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><path stroke-linecap="round" stroke-linejoin="round" d="M12 3v12m0 0l-4-4m4 4l4-4M4 17v2a2 2 0 002 2h12a2 2 0 002-2v-2"/></svg>' },
|
||||||
|
{ id: "tools", label: "Tools", icon: '<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><path stroke-linecap="round" stroke-linejoin="round" d="M11.42 15.17L17.25 21A2.652 2.652 0 0021 17.25l-5.877-5.877M11.42 15.17l2.496-3.03c.317-.384.74-.626 1.208-.766M11.42 15.17l-4.655 5.653a2.548 2.548 0 11-3.586-3.586l6.837-5.63m5.34-.11l-3.65-3.65a2.21 2.21 0 00-3.125 0l-3.65 3.65m7.425 3.65L19 11.83"/></svg>' },
|
||||||
|
];
|
||||||
|
|
||||||
|
window.platform = function platform() {
|
||||||
|
return {
|
||||||
|
// ----- nav + global -----
|
||||||
|
nav: NAV,
|
||||||
|
view: "dashboard",
|
||||||
|
health: false,
|
||||||
|
|
||||||
|
// ----- data stores -----
|
||||||
|
dash: { channels: [], status_breakdown: {}, uploads_over_time: [], duration_histogram: [], top_tags: [] },
|
||||||
|
channels: { items: [] },
|
||||||
|
videos: { items: [], total: 0, page: 1, size: 25 },
|
||||||
|
detail: { video: null, transcript: [] },
|
||||||
|
search: { q: "", channel: "", items: [], ran: false },
|
||||||
|
analysis: { tab: "wordcloud", channel: "", term: "", words: [], timeline: [] },
|
||||||
|
|
||||||
|
scrape: {
|
||||||
|
form: { channel_id: "", limit: "", since: "", languages: "", include_shorts: false, no_live: false },
|
||||||
|
jobId: null, progress: { completed: 0, total: 0 }, log: [], jobs: [], es: null, done: false, error: null,
|
||||||
|
},
|
||||||
|
cookies: { items: [], drag: false, testResult: {}, uploadMsg: "" },
|
||||||
|
|
||||||
|
// ----- tools / folders -----
|
||||||
|
folders: { markdown: { path: "", exists: false }, exports: { path: "", exists: false }, audio: { path: "", exists: false }, analysis: { path: "", exists: false }, database: { path: "", exists: false } },
|
||||||
|
tools: { reRendering: false, msg: "", fmt: null },
|
||||||
|
|
||||||
|
// ----- forms -----
|
||||||
|
forms: { channelUrl: "", channelError: "", exportFormat: "json", exportChannel: "" },
|
||||||
|
filters: { channel: "", status: "", from: "", to: "", min_dur: "", q: "", sort: "upload_date" },
|
||||||
|
|
||||||
|
// ----- loading flags -----
|
||||||
|
loading: { dashboard: false, channels: false, videos: false, detail: false, transcript: false, search: false, analysis: false, jobs: false, cookies: false, addChannel: false, startScrape: false },
|
||||||
|
|
||||||
|
// ----- chart instances -----
|
||||||
|
charts: {},
|
||||||
|
|
||||||
|
// =================================================================
|
||||||
|
// lifecycle
|
||||||
|
// =================================================================
|
||||||
|
init() {
|
||||||
|
setupChartDefaults();
|
||||||
|
this.checkHealth();
|
||||||
|
this.loadDashboard();
|
||||||
|
this.loadChannels();
|
||||||
|
},
|
||||||
|
|
||||||
|
async api(url, opts) {
|
||||||
|
const res = await fetch(url, opts);
|
||||||
|
if (!res.ok) {
|
||||||
|
let msg = res.status + " " + res.statusText;
|
||||||
|
try { const j = await res.json(); if (j && j.detail) msg = typeof j.detail === "string" ? j.detail : JSON.stringify(j.detail); } catch (_) {}
|
||||||
|
throw new Error(msg);
|
||||||
|
}
|
||||||
|
const ct = res.headers.get("content-type") || "";
|
||||||
|
if (ct.includes("application/json")) return res.json();
|
||||||
|
return res.text();
|
||||||
|
},
|
||||||
|
|
||||||
|
async checkHealth() {
|
||||||
|
try { const d = await this.api("/healthz"); this.health = d && d.status === "ok"; }
|
||||||
|
catch (_) { this.health = false; }
|
||||||
|
},
|
||||||
|
|
||||||
|
setView(v) {
|
||||||
|
this.view = v;
|
||||||
|
// destroy stray canvases when leaving chart-bearing views
|
||||||
|
if (v !== "dashboard") this.destroyCharts(["uploads", "duration", "tags"]);
|
||||||
|
if (v !== "analysis") this.destroyCharts(["topwords", "timeline"]);
|
||||||
|
if (v === "dashboard") this.$nextTick(() => this.renderDashboardCharts());
|
||||||
|
if (v === "videos" && this.videos.items.length === 0) this.loadVideos(1);
|
||||||
|
if (v === "scrape") this.loadJobs();
|
||||||
|
if (v === "cookies") this.loadCookies();
|
||||||
|
if (v === "analysis") this.loadAnalysis();
|
||||||
|
if (v === "tools") { this.loadFolders(); this.loadFormat(); }
|
||||||
|
},
|
||||||
|
|
||||||
|
// =================================================================
|
||||||
|
// dashboard
|
||||||
|
// =================================================================
|
||||||
|
async loadDashboard() {
|
||||||
|
this.loading.dashboard = true;
|
||||||
|
try {
|
||||||
|
const d = await this.api("/api/dashboard");
|
||||||
|
this.dash = d;
|
||||||
|
this.$nextTick(() => this.renderDashboardCharts());
|
||||||
|
} catch (e) { this.toast("Failed to load dashboard: " + e.message, "error"); }
|
||||||
|
finally { this.loading.dashboard = false; }
|
||||||
|
},
|
||||||
|
|
||||||
|
renderDashboardCharts() {
|
||||||
|
if (!window.Chart) return;
|
||||||
|
const up = this.dash.uploads_over_time || [];
|
||||||
|
this.makeChart("uploads", document.getElementById("chart-uploads"), {
|
||||||
|
type: "line",
|
||||||
|
data: { labels: up.map(u => this.fmtMonth(u.month)), datasets: [{ label: "uploads", data: up.map(u => u.count), borderColor: "#f43f5e", backgroundColor: "rgba(244,63,94,0.18)", fill: true, tension: 0.35, pointRadius: 2 }] },
|
||||||
|
options: this.lineOpts(),
|
||||||
|
});
|
||||||
|
const dh = this.dash.duration_histogram || [];
|
||||||
|
this.makeChart("duration", document.getElementById("chart-duration"), {
|
||||||
|
type: "bar",
|
||||||
|
data: { labels: dh.map(d => d.bucket), datasets: [{ label: "videos", data: dh.map(d => d.count), backgroundColor: "rgba(244,63,94,0.6)", borderRadius: 4 }] },
|
||||||
|
options: this.barOpts(),
|
||||||
|
});
|
||||||
|
const tt = (this.dash.top_tags || []).slice(0, 12);
|
||||||
|
this.makeChart("tags", document.getElementById("chart-tags"), {
|
||||||
|
type: "bar",
|
||||||
|
data: { labels: tt.map(t => t.tag), datasets: [{ label: "count", data: tt.map(t => t.count), backgroundColor: "rgba(251,113,133,0.7)", borderRadius: 4 }] },
|
||||||
|
options: this.barOpts(true),
|
||||||
|
});
|
||||||
|
},
|
||||||
|
|
||||||
|
// =================================================================
|
||||||
|
// channels
|
||||||
|
// =================================================================
|
||||||
|
async loadChannels() {
|
||||||
|
this.loading.channels = true;
|
||||||
|
try { const d = await this.api("/api/channels"); this.channels = d; }
|
||||||
|
catch (e) { this.toast("Failed to load channels: " + e.message, "error"); }
|
||||||
|
finally { this.loading.channels = false; }
|
||||||
|
},
|
||||||
|
|
||||||
|
async addChannel() {
|
||||||
|
const url = (this.forms.channelUrl || "").trim();
|
||||||
|
if (!url) return;
|
||||||
|
this.forms.channelError = "";
|
||||||
|
this.loading.addChannel = true;
|
||||||
|
try {
|
||||||
|
const d = await this.api("/api/channels", { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify({ url }) });
|
||||||
|
this.forms.channelUrl = "";
|
||||||
|
await this.loadChannels();
|
||||||
|
this.toast("Added " + (d.name || "channel") + " (" + (d.video_count || 0) + " videos)");
|
||||||
|
} catch (e) { this.forms.channelError = e.message; }
|
||||||
|
finally { this.loading.addChannel = false; }
|
||||||
|
},
|
||||||
|
|
||||||
|
async removeChannel(id) {
|
||||||
|
if (!confirm("Remove this channel and its videos?")) return;
|
||||||
|
try { await this.api("/api/channels/" + encodeURIComponent(id), { method: "DELETE" }); await this.loadChannels(); this.loadDashboard(); }
|
||||||
|
catch (e) { this.toast("Remove failed: " + e.message, "error"); }
|
||||||
|
},
|
||||||
|
|
||||||
|
// =================================================================
|
||||||
|
// videos
|
||||||
|
// =================================================================
|
||||||
|
async loadVideos(page) {
|
||||||
|
if (page) this.videos.page = page;
|
||||||
|
this.loading.videos = true;
|
||||||
|
const p = new URLSearchParams();
|
||||||
|
const f = this.filters;
|
||||||
|
if (f.channel) p.set("channel", f.channel);
|
||||||
|
if (f.status) p.set("status", f.status);
|
||||||
|
if (f.from) p.set("from", f.from);
|
||||||
|
if (f.to) p.set("to", f.to);
|
||||||
|
if (f.min_dur) p.set("min_dur", f.min_dur);
|
||||||
|
if (f.q) p.set("q", f.q);
|
||||||
|
if (f.sort) p.set("sort", f.sort);
|
||||||
|
p.set("page", this.videos.page);
|
||||||
|
p.set("size", this.videos.size);
|
||||||
|
try {
|
||||||
|
const d = await this.api("/api/videos?" + p.toString());
|
||||||
|
this.videos = Object.assign({}, this.videos, { items: d.items, total: d.total, page: d.page, size: d.size });
|
||||||
|
} catch (e) { this.toast("Failed to load videos: " + e.message, "error"); }
|
||||||
|
finally { this.loading.videos = false; }
|
||||||
|
},
|
||||||
|
|
||||||
|
async openVideo(id) {
|
||||||
|
this.view = "detail";
|
||||||
|
this.loading.detail = true;
|
||||||
|
this.loading.transcript = true;
|
||||||
|
this.detail = { video: null, transcript: [] };
|
||||||
|
try {
|
||||||
|
const v = await this.api("/api/videos/" + encodeURIComponent(id));
|
||||||
|
// chapters may come embedded or be absent
|
||||||
|
if (v && !v.chapters) v.chapters = [];
|
||||||
|
this.detail.video = v;
|
||||||
|
} catch (e) { this.toast("Failed to load video: " + e.message, "error"); }
|
||||||
|
finally { this.loading.detail = false; }
|
||||||
|
try {
|
||||||
|
const t = await this.api("/api/videos/" + encodeURIComponent(id) + "/transcript");
|
||||||
|
this.detail.transcript = (t && t.items) || [];
|
||||||
|
} catch (_) { this.detail.transcript = []; }
|
||||||
|
finally { this.loading.transcript = false; }
|
||||||
|
},
|
||||||
|
|
||||||
|
seekTranscript(sec) {
|
||||||
|
// best-effort scroll to nearest transcript segment
|
||||||
|
const segs = this.detail.transcript || [];
|
||||||
|
let best = null, bestDelta = Infinity;
|
||||||
|
for (const s of segs) { const d = Math.abs((s.start_sec || 0) - sec); if (d < bestDelta) { bestDelta = d; best = s; } }
|
||||||
|
if (best) { const el = document.getElementById("seg-" + best.idx); if (el) el.scrollIntoView({ behavior: "smooth", block: "center" }); }
|
||||||
|
},
|
||||||
|
|
||||||
|
// =================================================================
|
||||||
|
// search
|
||||||
|
// =================================================================
|
||||||
|
async loadSearch() {
|
||||||
|
const q = (this.search.q || "").trim();
|
||||||
|
if (!q) return;
|
||||||
|
this.loading.search = true;
|
||||||
|
this.search.ran = true;
|
||||||
|
const p = new URLSearchParams({ q, limit: 50 });
|
||||||
|
if (this.search.channel) p.set("channel", this.search.channel);
|
||||||
|
try {
|
||||||
|
const d = await this.api("/api/search?" + p.toString());
|
||||||
|
this.search.items = (d && d.items) || [];
|
||||||
|
} catch (e) { this.toast("Search failed: " + e.message, "error"); this.search.items = []; }
|
||||||
|
finally { this.loading.search = false; }
|
||||||
|
},
|
||||||
|
|
||||||
|
highlight(snippet, q) {
|
||||||
|
if (!snippet) return "";
|
||||||
|
let s = String(snippet).replace(/&/g, "&").replace(/</g, "<").replace(/>/g, ">");
|
||||||
|
if (!q) return s;
|
||||||
|
const terms = q.trim().split(/\s+/).filter(Boolean).map(t => t.replace(/[.*+?^${}()|[\]\\]/g, "\\$&"));
|
||||||
|
if (!terms.length) return s;
|
||||||
|
const re = new RegExp("(" + terms.join("|") + ")", "gi");
|
||||||
|
return s.replace(re, '<mark class="bg-rose-500/30 text-rose-200 rounded px-0.5">$1</mark>');
|
||||||
|
},
|
||||||
|
|
||||||
|
// =================================================================
|
||||||
|
// analysis
|
||||||
|
// =================================================================
|
||||||
|
async loadAnalysis() {
|
||||||
|
this.loading.analysis = true;
|
||||||
|
const base = "/api/analysis";
|
||||||
|
try {
|
||||||
|
if (this.analysis.tab === "timeline") {
|
||||||
|
const p = new URLSearchParams({ type: "timeline", n: 200 });
|
||||||
|
if (this.analysis.term.trim()) p.set("term", this.analysis.term.trim());
|
||||||
|
if (this.analysis.channel) p.set("channel", this.analysis.channel);
|
||||||
|
const d = await this.api(base + "?" + p.toString());
|
||||||
|
this.analysis.timeline = (d && d.items) || [];
|
||||||
|
this.$nextTick(() => this.renderTimelineChart());
|
||||||
|
} else {
|
||||||
|
const p = new URLSearchParams({ type: "topwords", n: this.analysis.tab === "wordcloud" ? 300 : 25 });
|
||||||
|
if (this.analysis.channel) p.set("channel", this.analysis.channel);
|
||||||
|
const d = await this.api(base + "?" + p.toString());
|
||||||
|
this.analysis.words = (d && d.items) || [];
|
||||||
|
if (this.analysis.tab === "topwords") this.$nextTick(() => this.renderTopWordsChart());
|
||||||
|
}
|
||||||
|
} catch (e) { this.toast("Analysis failed: " + e.message, "error"); }
|
||||||
|
finally { this.loading.analysis = false; }
|
||||||
|
},
|
||||||
|
|
||||||
|
renderTopWordsChart() {
|
||||||
|
if (!window.Chart) return;
|
||||||
|
const w = this.analysis.words.slice(0, 25);
|
||||||
|
this.makeChart("topwords", document.getElementById("chart-topwords"), {
|
||||||
|
type: "bar",
|
||||||
|
data: { labels: w.map(x => x.word), datasets: [{ label: "count", data: w.map(x => x.count), backgroundColor: "rgba(244,63,94,0.65)", borderRadius: 4 }] },
|
||||||
|
options: this.barOpts(true),
|
||||||
|
});
|
||||||
|
},
|
||||||
|
|
||||||
|
renderTimelineChart() {
|
||||||
|
if (!window.Chart) return;
|
||||||
|
const t = this.analysis.timeline;
|
||||||
|
this.makeChart("timeline", document.getElementById("chart-timeline"), {
|
||||||
|
type: "line",
|
||||||
|
data: { labels: t.map(x => this.fmtMonth(String(x.month))), datasets: [{ label: this.analysis.term || "count", data: t.map(x => x.count), borderColor: "#f43f5e", backgroundColor: "rgba(244,63,94,0.18)", fill: true, tension: 0.3, pointRadius: 2 }] },
|
||||||
|
options: this.lineOpts(),
|
||||||
|
});
|
||||||
|
},
|
||||||
|
|
||||||
|
cloudStyle(w) {
|
||||||
|
const counts = this.analysis.words.map(x => x.count);
|
||||||
|
const max = Math.max.apply(null, counts.concat([1]));
|
||||||
|
const min = Math.min.apply(null, counts.concat([1]));
|
||||||
|
const span = Math.max(max - min, 1);
|
||||||
|
const r = (w.count - min) / span;
|
||||||
|
const size = 12 + r * 34; // 12px..46px
|
||||||
|
const weights = ["#71717a", "#a1a1aa", "#d4d4d8", "#fb7185", "#f43f5e", "#dc2626"];
|
||||||
|
const ci = Math.min(weights.length - 1, Math.floor(r * weights.length));
|
||||||
|
const op = 0.55 + r * 0.45;
|
||||||
|
return `font-size:${size.toFixed(1)}px;color:${weights[ci]};opacity:${op.toFixed(2)}`;
|
||||||
|
},
|
||||||
|
|
||||||
|
// =================================================================
|
||||||
|
// scrape
|
||||||
|
// =================================================================
|
||||||
|
async startScrape() {
|
||||||
|
const f = this.scrape.form;
|
||||||
|
if (!f.channel_id) { this.toast("Pick a channel first", "error"); return; }
|
||||||
|
this.loading.startScrape = true;
|
||||||
|
const body = {
|
||||||
|
channel_id: f.channel_id,
|
||||||
|
opts: {
|
||||||
|
limit: f.limit ? parseInt(f.limit, 10) : null,
|
||||||
|
since: f.since || null,
|
||||||
|
languages: f.languages ? f.languages.split(",").map(s => s.trim()).filter(Boolean) : null,
|
||||||
|
include_shorts: !!f.include_shorts,
|
||||||
|
no_live: !!f.no_live,
|
||||||
|
},
|
||||||
|
};
|
||||||
|
try {
|
||||||
|
const d = await this.api("/api/scrape", { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify(body) });
|
||||||
|
this.scrape.done = false; this.scrape.error = null;
|
||||||
|
this.subscribeJob(d.job_id);
|
||||||
|
} catch (e) { this.toast("Scrape failed: " + e.message, "error"); }
|
||||||
|
finally { this.loading.startScrape = false; }
|
||||||
|
},
|
||||||
|
|
||||||
|
subscribeJob(jobId) {
|
||||||
|
this.closeStream();
|
||||||
|
this.scrape.jobId = jobId;
|
||||||
|
this.scrape.log = ["[stream] connecting…"];
|
||||||
|
this.scrape.progress = { completed: 0, total: 0 };
|
||||||
|
const es = new EventSource("/api/scrape/" + encodeURIComponent(jobId) + "/stream");
|
||||||
|
this.scrape.es = es;
|
||||||
|
es.addEventListener("log", e => { try { const d = JSON.parse(e.data); this.scrape.log.push(d.msg || ""); } catch (_) {} });
|
||||||
|
es.addEventListener("progress", e => { try { const d = JSON.parse(e.data); this.scrape.progress = d; } catch (_) {} });
|
||||||
|
es.addEventListener("done", () => { this.scrape.done = true; this.scrape.log.push("[done] job finished"); this.closeStream(); this.loadJobs(); });
|
||||||
|
es.addEventListener("cancelled", () => { this.scrape.log.push("[cancelled]"); this.closeStream(); this.loadJobs(); });
|
||||||
|
es.addEventListener("error", e => {
|
||||||
|
if (this.scrape.es === null) return; // already closed by done/cancelled
|
||||||
|
let msg = "connection error";
|
||||||
|
try { if (e.data) { const d = JSON.parse(e.data); msg = d.msg || msg; this.scrape.error = msg; this.scrape.log.push("[error] " + msg); } } catch (_) { this.scrape.error = msg; }
|
||||||
|
this.closeStream(); this.loadJobs();
|
||||||
|
});
|
||||||
|
},
|
||||||
|
|
||||||
|
closeStream() { if (this.scrape.es) { try { this.scrape.es.close(); } catch (_) {} this.scrape.es = null; } },
|
||||||
|
|
||||||
|
cancelScrape() {
|
||||||
|
if (!this.scrape.jobId) return;
|
||||||
|
fetch("/api/scrape/" + encodeURIComponent(this.scrape.jobId), { method: "DELETE" }).then(() => { this.scrape.log.push("[cancel] requested"); this.closeStream(); this.loadJobs(); }).catch(() => {});
|
||||||
|
},
|
||||||
|
|
||||||
|
async deleteJob(id) {
|
||||||
|
try { await this.api("/api/scrape/" + encodeURIComponent(id), { method: "DELETE" }); this.loadJobs(); }
|
||||||
|
catch (e) { this.toast("Delete failed: " + e.message, "error"); }
|
||||||
|
},
|
||||||
|
|
||||||
|
scrapePct() {
|
||||||
|
const p = this.scrape.progress || {};
|
||||||
|
if (!p.total) return this.scrape.done ? 100 : 0;
|
||||||
|
return Math.min(100, Math.round((p.completed / p.total) * 100));
|
||||||
|
},
|
||||||
|
|
||||||
|
async loadJobs() {
|
||||||
|
this.loading.jobs = true;
|
||||||
|
try { const d = await this.api("/api/scrape"); this.scrape.jobs = (d && d.items) || []; }
|
||||||
|
catch (_) {}
|
||||||
|
finally { this.loading.jobs = false; }
|
||||||
|
},
|
||||||
|
|
||||||
|
// =================================================================
|
||||||
|
// cookies
|
||||||
|
// =================================================================
|
||||||
|
async loadCookies() {
|
||||||
|
this.loading.cookies = true;
|
||||||
|
try { const d = await this.api("/api/cookies"); this.cookies.items = (d && d.items) || []; }
|
||||||
|
catch (e) { this.toast("Cookie load failed: " + e.message, "error"); }
|
||||||
|
finally { this.loading.cookies = false; }
|
||||||
|
},
|
||||||
|
|
||||||
|
handleDrop(e) {
|
||||||
|
this.cookies.drag = false;
|
||||||
|
const files = e.dataTransfer && e.dataTransfer.files;
|
||||||
|
if (files && files.length) this.handleFiles(files);
|
||||||
|
},
|
||||||
|
|
||||||
|
handleFiles(fileList) {
|
||||||
|
const files = Array.from(fileList || []).filter(f => /\.txt$/i.test(f.name) || f.type === "text/plain" || f.type === "");
|
||||||
|
if (!files.length) { this.toast("Drop .txt cookie files only", "error"); return; }
|
||||||
|
this.uploadCookies(files);
|
||||||
|
// reset the file input so the same file can be picked again
|
||||||
|
try { if (this.$refs.cookieFile) this.$refs.cookieFile.value = ""; } catch (_) {}
|
||||||
|
},
|
||||||
|
|
||||||
|
async uploadCookies(files) {
|
||||||
|
const fd = new FormData();
|
||||||
|
for (const f of files) fd.append("files", f, f.name);
|
||||||
|
try {
|
||||||
|
const d = await fetch("/api/cookies/upload", { method: "POST", body: fd });
|
||||||
|
if (!d.ok) throw new Error(d.status + " " + d.statusText);
|
||||||
|
const j = await d.json();
|
||||||
|
this.cookies.uploadMsg = "Imported " + ((j.imported || []).length) + " file(s).";
|
||||||
|
await this.loadCookies();
|
||||||
|
setTimeout(() => { this.cookies.uploadMsg = ""; }, 4000);
|
||||||
|
} catch (e) { this.toast("Upload failed: " + e.message, "error"); }
|
||||||
|
},
|
||||||
|
|
||||||
|
async activateCookie(id) {
|
||||||
|
try { await this.api("/api/cookies/" + encodeURIComponent(id) + "/activate", { method: "POST" }); this.loadCookies(); }
|
||||||
|
catch (e) { this.toast("Activate failed: " + e.message, "error"); }
|
||||||
|
},
|
||||||
|
|
||||||
|
async testCookie(id) {
|
||||||
|
this.cookies.testResult[id] = null;
|
||||||
|
try {
|
||||||
|
const d = await this.api("/api/cookies/" + encodeURIComponent(id) + "/test", { method: "POST" });
|
||||||
|
this.cookies.testResult[id] = d;
|
||||||
|
} catch (e) { this.cookies.testResult[id] = { ok: false, msg: e.message }; }
|
||||||
|
},
|
||||||
|
|
||||||
|
testLabel(id) {
|
||||||
|
const r = this.cookies.testResult[id];
|
||||||
|
if (r === null || r === undefined) return "—";
|
||||||
|
if (r === false || r.ok === false) return "✕ fail";
|
||||||
|
return "✓ ok (" + (r.cookie_count || 0) + ")";
|
||||||
|
},
|
||||||
|
|
||||||
|
async removeCookie(id) {
|
||||||
|
if (!confirm("Delete this cookie file?")) return;
|
||||||
|
try { await this.api("/api/cookies/" + encodeURIComponent(id), { method: "DELETE" }); this.loadCookies(); }
|
||||||
|
catch (e) { this.toast("Delete failed: " + e.message, "error"); }
|
||||||
|
},
|
||||||
|
|
||||||
|
// =================================================================
|
||||||
|
// export
|
||||||
|
// =================================================================
|
||||||
|
doExport() {
|
||||||
|
const p = new URLSearchParams({ format: this.forms.exportFormat });
|
||||||
|
if (this.forms.exportChannel) p.set("channel", this.forms.exportChannel);
|
||||||
|
window.location.href = "/api/export?" + p.toString();
|
||||||
|
},
|
||||||
|
|
||||||
|
// =================================================================
|
||||||
|
// folders + tools (knowledge base)
|
||||||
|
// =================================================================
|
||||||
|
openChannel(id) {
|
||||||
|
this.filters.channel = id;
|
||||||
|
this.setView("videos");
|
||||||
|
this.loadVideos(1);
|
||||||
|
},
|
||||||
|
|
||||||
|
async loadFolders() {
|
||||||
|
try { const d = await this.api("/api/folders"); if (d && d.items) this.folders = d.items; }
|
||||||
|
catch (_) {}
|
||||||
|
},
|
||||||
|
|
||||||
|
async loadFormat() {
|
||||||
|
try { const d = await this.api("/api/tools/format"); this.tools.fmt = d; }
|
||||||
|
catch (_) {}
|
||||||
|
},
|
||||||
|
|
||||||
|
async openFolder(kind, channelId) {
|
||||||
|
const body = { kind };
|
||||||
|
if (channelId) body.channel_id = channelId;
|
||||||
|
try {
|
||||||
|
await this.api("/api/folders/open", { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify(body) });
|
||||||
|
this.toast("Folder opened in file manager");
|
||||||
|
} catch (e) { this.toast("Open folder failed: " + e.message, "error"); }
|
||||||
|
},
|
||||||
|
|
||||||
|
async reRender() {
|
||||||
|
this.tools.reRendering = true; this.tools.msg = "Working…";
|
||||||
|
try {
|
||||||
|
const d = await this.api("/api/tools/re-render", { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify({}) });
|
||||||
|
this.tools.msg = "Re-rendered " + (d.re_rendered || 0) + " markdown files.";
|
||||||
|
this.toast(this.tools.msg);
|
||||||
|
} catch (e) { this.tools.msg = "Failed: " + e.message; this.toast("Re-render failed: " + e.message, "error"); }
|
||||||
|
finally { this.tools.reRendering = false; }
|
||||||
|
},
|
||||||
|
|
||||||
|
// =================================================================
|
||||||
|
// chart helpers
|
||||||
|
// =================================================================
|
||||||
|
makeChart(key, ctx, config) {
|
||||||
|
if (!ctx) return;
|
||||||
|
if (this.charts[key]) { try { this.charts[key].destroy(); } catch (_) {} }
|
||||||
|
const merged = Object.assign({}, config, {
|
||||||
|
options: Object.assign({
|
||||||
|
responsive: true, maintainAspectRatio: false,
|
||||||
|
plugins: { legend: { display: false }, tooltip: { backgroundColor: "#18181b", borderColor: "#27272a", borderWidth: 1, titleColor: "#fafafa", bodyColor: "#d4d4d8" } },
|
||||||
|
}, config.options || {}),
|
||||||
|
});
|
||||||
|
this.charts[key] = new Chart(ctx, merged);
|
||||||
|
},
|
||||||
|
|
||||||
|
destroyCharts(keys) {
|
||||||
|
for (const k of keys) { if (this.charts[k]) { try { this.charts[k].destroy(); } catch (_) {} delete this.charts[k]; } }
|
||||||
|
},
|
||||||
|
|
||||||
|
lineOpts() {
|
||||||
|
return { responsive: true, maintainAspectRatio: false, scales: this.axisOpts(), plugins: { legend: { display: false } } };
|
||||||
|
},
|
||||||
|
barOpts(horizontal) {
|
||||||
|
return {
|
||||||
|
responsive: true, maintainAspectRatio: false,
|
||||||
|
indexAxis: horizontal ? "y" : "x",
|
||||||
|
scales: horizontal ? { x: this.axisOpts().x, y: { ticks: { color: "#d4d4d8" }, grid: { display: false } } } : this.axisOpts(),
|
||||||
|
plugins: { legend: { display: false } },
|
||||||
|
};
|
||||||
|
},
|
||||||
|
axisOpts() {
|
||||||
|
return {
|
||||||
|
x: { ticks: { color: "#71717a", maxRotation: 0, autoSkip: true }, grid: { color: "rgba(39,39,42,0.5)" } },
|
||||||
|
y: { ticks: { color: "#71717a" }, grid: { color: "rgba(39,39,42,0.5)" }, beginAtZero: true },
|
||||||
|
};
|
||||||
|
},
|
||||||
|
|
||||||
|
// =================================================================
|
||||||
|
// formatting helpers
|
||||||
|
// =================================================================
|
||||||
|
ts(sec) {
|
||||||
|
sec = Math.max(0, Math.floor(Number(sec) || 0));
|
||||||
|
const h = Math.floor(sec / 3600), m = Math.floor((sec % 3600) / 60), s = sec % 60;
|
||||||
|
const pad = n => String(n).padStart(2, "0");
|
||||||
|
return h > 0 ? h + ":" + pad(m) + ":" + pad(s) : m + ":" + pad(s);
|
||||||
|
},
|
||||||
|
|
||||||
|
humanDur(sec) {
|
||||||
|
sec = Number(sec) || 0;
|
||||||
|
if (sec < 60) return sec + "s";
|
||||||
|
const h = sec / 3600, m = (sec % 3600) / 60;
|
||||||
|
if (h >= 1) return h.toFixed(0) + "h " + m.toFixed(0) + "m";
|
||||||
|
return m.toFixed(0) + "m";
|
||||||
|
},
|
||||||
|
|
||||||
|
fmtNum(n) {
|
||||||
|
n = Number(n || 0);
|
||||||
|
if (!isFinite(n)) return "0";
|
||||||
|
if (n >= 1e9) return (n / 1e9).toFixed(1) + "B";
|
||||||
|
if (n >= 1e6) return (n / 1e6).toFixed(1) + "M";
|
||||||
|
if (n >= 1e3) return (n / 1e3).toFixed(1) + "K";
|
||||||
|
return n.toLocaleString();
|
||||||
|
},
|
||||||
|
|
||||||
|
fmtDate(d) {
|
||||||
|
if (!d) return "—";
|
||||||
|
const s = String(d);
|
||||||
|
// accept YYYYMMDD, YYYY-MM-DD, or ISO
|
||||||
|
const m = s.match(/^(\d{4})(\d{2})(\d{2})/);
|
||||||
|
if (m) return m[1] + "-" + m[2] + "-" + m[3];
|
||||||
|
if (s.length >= 10) return s.slice(0, 10);
|
||||||
|
return s;
|
||||||
|
},
|
||||||
|
|
||||||
|
fmtDateTime(d) {
|
||||||
|
if (!d) return "—";
|
||||||
|
const s = String(d).replace("T", " ");
|
||||||
|
return s.length > 16 ? s.slice(0, 16) : s;
|
||||||
|
},
|
||||||
|
|
||||||
|
fmtMonth(yyyymm) {
|
||||||
|
const s = String(yyyymm);
|
||||||
|
const m = s.match(/^(\d{4})(\d{2})/);
|
||||||
|
if (!m) return s;
|
||||||
|
return m[1] + "-" + m[2];
|
||||||
|
},
|
||||||
|
|
||||||
|
chanName(id) {
|
||||||
|
const c = (this.channels.items || []).find(c => c.channel_id === id);
|
||||||
|
return c ? c.name : (id || "—");
|
||||||
|
},
|
||||||
|
|
||||||
|
statusClass(status) {
|
||||||
|
status = String(status || "").toLowerCase();
|
||||||
|
const map = { done: "st-done", ok: "st-done", completed: "st-done", success: "st-done", error: "st-error", failed: "st-error", fail: "st-error", pending: "st-pending", queued: "st-pending", running: "st-running", active: "st-running", skipped: "st-skipped", skip: "st-skipped", new: "st-new" };
|
||||||
|
return map[status] || "st-skipped";
|
||||||
|
},
|
||||||
|
|
||||||
|
toast(msg, kind) {
|
||||||
|
// simple visible status — log + alert fallback
|
||||||
|
if (kind === "error") console.error("[platform]", msg);
|
||||||
|
else console.log("[platform]", msg);
|
||||||
|
// transient on-page toast
|
||||||
|
const el = document.createElement("div");
|
||||||
|
el.textContent = msg;
|
||||||
|
el.style.cssText = "position:fixed;bottom:1.25rem;right:1.25rem;z-index:9999;padding:0.7rem 1rem;border-radius:0.6rem;font-size:0.85rem;font-weight:600;color:#fff;background:" + (kind === "error" ? "linear-gradient(to right,#dc2626,#b91c1c)" : "linear-gradient(to right,#f43f5e,#dc2626)") + ";box-shadow:0 10px 30px rgba(0,0,0,0.5);opacity:0;transition:opacity .2s ease;";
|
||||||
|
document.body.appendChild(el);
|
||||||
|
requestAnimationFrame(() => { el.style.opacity = "1"; });
|
||||||
|
setTimeout(() => { el.style.opacity = "0"; setTimeout(() => el.remove(), 250); }, 3800);
|
||||||
|
},
|
||||||
|
};
|
||||||
|
};
|
||||||
|
})();
|
||||||
@@ -0,0 +1,627 @@
|
|||||||
|
<!DOCTYPE html>
|
||||||
|
<html lang="en" class="dark" style="background:#0a0a0f;color:#e5e7eb">
|
||||||
|
<head>
|
||||||
|
<meta charset="UTF-8" />
|
||||||
|
<meta name="viewport" content="width=device-width, initial-scale=1.0" />
|
||||||
|
<title>yt-scraper · command center</title>
|
||||||
|
<!-- Tailwind -->
|
||||||
|
<script src="https://cdn.tailwindcss.com"></script>
|
||||||
|
<script>
|
||||||
|
tailwind.config = {
|
||||||
|
darkMode: 'class',
|
||||||
|
theme: { extend: { colors: { accent: '#f43f5e' }, fontFamily: { mono: ['ui-monospace', 'JetBrains Mono', 'Cascadia Code', 'monospace'] } } }
|
||||||
|
};
|
||||||
|
</script>
|
||||||
|
<!-- Chart.js -->
|
||||||
|
<script src="https://cdn.jsdelivr.net/npm/chart.js@4"></script>
|
||||||
|
<!-- app (defines window.platform BEFORE Alpine runs; both deferred, app first) -->
|
||||||
|
<script defer src="/static/app.js"></script>
|
||||||
|
<!-- Alpine (loaded after app.js so window.platform exists at init) -->
|
||||||
|
<script defer src="https://unpkg.com/[email protected]/dist/cdn.min.js"></script>
|
||||||
|
<link rel="stylesheet" href="/static/styles.css" />
|
||||||
|
</head>
|
||||||
|
<body class="min-h-screen relative">
|
||||||
|
<div id="root" x-cloak x-data="platform()" x-init="init()" class="relative z-10 flex min-h-screen">
|
||||||
|
|
||||||
|
<!-- ============ SIDEBAR ============ -->
|
||||||
|
<aside class="w-64 shrink-0 border-r border-zinc-800/80 bg-zinc-950/70 backdrop-blur-xl flex flex-col fixed inset-y-0 left-0 z-30">
|
||||||
|
<div class="px-5 py-5 flex items-center gap-3 border-b border-zinc-800/70">
|
||||||
|
<div class="w-9 h-9 rounded-xl accent-grad flex items-center justify-center font-bold text-white shadow-lg shadow-rose-500/20">
|
||||||
|
<svg viewBox="0 0 24 24" class="w-5 h-5" fill="currentColor"><path d="M10 15.5v-7l6 3.5-6 3.5zM21 6.5c0-1.1-.9-2-2-2H5c-1.1 0-2 .9-2 2v11c0 1.1.9 2 2 2h14c1.1 0 2-.9 2-2v-11z"/></svg>
|
||||||
|
</div>
|
||||||
|
<div>
|
||||||
|
<div class="text-sm font-bold text-white leading-tight">yt-scraper</div>
|
||||||
|
<div class="text-[0.65rem] text-zinc-500 uppercase tracking-widest">command center</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<nav class="flex-1 overflow-y-auto px-3 py-4 space-y-1">
|
||||||
|
<template x-for="n in nav" :key="n.id">
|
||||||
|
<button class="nav-item" :class="view===n.id ? 'active' : ''" @click="setView(n.id)">
|
||||||
|
<span class="nav-ico" x-html="n.icon"></span>
|
||||||
|
<span x-text="n.label"></span>
|
||||||
|
</button>
|
||||||
|
</template>
|
||||||
|
</nav>
|
||||||
|
|
||||||
|
<div class="px-4 py-4 border-t border-zinc-800/70 text-[0.7rem] text-zinc-500">
|
||||||
|
<div class="flex items-center gap-2">
|
||||||
|
<span class="w-2 h-2 rounded-full" :class="health ? 'bg-emerald-500' : 'bg-zinc-600'"></span>
|
||||||
|
<span x-text="health ? 'backend online' : 'checking…'"></span>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</aside>
|
||||||
|
|
||||||
|
<!-- ============ MAIN ============ -->
|
||||||
|
<main class="flex-1 ml-64 min-w-0">
|
||||||
|
<div class="max-w-7xl mx-auto px-8 py-8">
|
||||||
|
|
||||||
|
<!-- ===== Dashboard ===== -->
|
||||||
|
<template x-if="view==='dashboard'">
|
||||||
|
<section class="space-y-6" x-transition.opacity>
|
||||||
|
<header class="flex items-end justify-between">
|
||||||
|
<div>
|
||||||
|
<h1 class="text-2xl font-bold text-white">Dashboard</h1>
|
||||||
|
<p class="text-sm text-zinc-500">Overview across all tracked channels</p>
|
||||||
|
</div>
|
||||||
|
<button class="btn btn-ghost" @click="loadDashboard()" :disabled="loading.dashboard">
|
||||||
|
<svg x-show="loading.dashboard" class="spin w-4 h-4" viewBox="0 0 24 24" fill="none"><circle cx="12" cy="12" r="9" stroke="currentColor" stroke-width="3" stroke-dasharray="40 20"/></svg>
|
||||||
|
Refresh
|
||||||
|
</button>
|
||||||
|
</header>
|
||||||
|
|
||||||
|
<!-- stat row -->
|
||||||
|
<div class="grid grid-cols-2 md:grid-cols-4 gap-4">
|
||||||
|
<div class="glass p-4">
|
||||||
|
<div class="text-xs text-zinc-500 uppercase tracking-wider">Channels</div>
|
||||||
|
<div class="mt-1 text-2xl font-bold text-white font-mono" x-text="dash.channels.length"></div>
|
||||||
|
</div>
|
||||||
|
<div class="glass p-4">
|
||||||
|
<div class="text-xs text-zinc-500 uppercase tracking-wider">Videos</div>
|
||||||
|
<div class="mt-1 text-2xl font-bold text-white font-mono" x-text="fmtNum(dash.channels.reduce((s,c)=>s+(c.video_count||0),0))"></div>
|
||||||
|
</div>
|
||||||
|
<div class="glass p-4">
|
||||||
|
<div class="text-xs text-zinc-500 uppercase tracking-wider">Total Views</div>
|
||||||
|
<div class="mt-1 text-2xl font-bold text-white font-mono" x-text="fmtNum(dash.channels.reduce((s,c)=>s+(c.total_views||0),0))"></div>
|
||||||
|
</div>
|
||||||
|
<div class="glass p-4">
|
||||||
|
<div class="text-xs text-zinc-500 uppercase tracking-wider">Total Duration</div>
|
||||||
|
<div class="mt-1 text-2xl font-bold text-white font-mono" x-text="humanDur(dash.channels.reduce((s,c)=>s+(c.total_duration||0),0))"></div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<!-- status pills -->
|
||||||
|
<div class="glass p-4 flex flex-wrap items-center gap-2">
|
||||||
|
<span class="text-xs text-zinc-500 uppercase tracking-wider mr-2">Status</span>
|
||||||
|
<template x-if="loading.dashboard"><span class="text-sm text-zinc-500">loading…</span></template>
|
||||||
|
<template x-for="(n, k) in dash.status_breakdown" :key="k">
|
||||||
|
<span class="pill" :class="statusClass(k)"><span x-text="k"></span><span class="font-mono" x-text="n"></span></span>
|
||||||
|
</template>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<!-- channel cards -->
|
||||||
|
<div class="space-y-3">
|
||||||
|
<h2 class="text-sm font-semibold text-zinc-300 uppercase tracking-wider">Channels</h2>
|
||||||
|
<template x-if="!loading.dashboard && dash.channels.length===0">
|
||||||
|
<div class="glass p-8 text-center text-zinc-500 text-sm">No channels yet. Add one in the <button class="accent-text underline" @click="setView('channels')">Channels</button> view.</div>
|
||||||
|
</template>
|
||||||
|
<div class="grid md:grid-cols-2 gap-4">
|
||||||
|
<template x-for="c in dash.channels" :key="c.channel_id">
|
||||||
|
<div class="glass p-5 hover:border-rose-500/40 transition-colors">
|
||||||
|
<div class="flex items-start justify-between gap-3">
|
||||||
|
<div class="min-w-0">
|
||||||
|
<div class="text-white font-semibold truncate" x-text="c.name"></div>
|
||||||
|
<div class="text-xs text-zinc-500 truncate" x-text="c.handle || c.channel_id"></div>
|
||||||
|
</div>
|
||||||
|
<div class="text-right shrink-0">
|
||||||
|
<div class="text-lg font-bold text-white font-mono" x-text="fmtNum(c.video_count)"></div>
|
||||||
|
<div class="text-[0.65rem] text-zinc-500 uppercase">videos</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
<div class="grid grid-cols-3 gap-2 mt-4 pt-4 border-t border-zinc-800/70 text-xs">
|
||||||
|
<div><div class="text-zinc-500">Views</div><div class="text-zinc-200 font-mono" x-text="fmtNum(c.total_views)"></div></div>
|
||||||
|
<div><div class="text-zinc-500">Length</div><div class="text-zinc-200 font-mono" x-text="humanDur(c.total_duration)"></div></div>
|
||||||
|
<div><div class="text-zinc-500">Range</div><div class="text-zinc-200 font-mono text-[0.7rem]" x-text="fmtDate(c.date_min)+' → '+fmtDate(c.date_max)"></div></div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</template>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<!-- charts -->
|
||||||
|
<div class="grid lg:grid-cols-2 gap-4">
|
||||||
|
<div class="glass p-5">
|
||||||
|
<h3 class="text-sm font-semibold text-zinc-200 mb-3">Uploads over time</h3>
|
||||||
|
<div class="h-64"><canvas id="chart-uploads"></canvas></div>
|
||||||
|
</div>
|
||||||
|
<div class="glass p-5">
|
||||||
|
<h3 class="text-sm font-semibold text-zinc-200 mb-3">Duration distribution</h3>
|
||||||
|
<div class="h-64"><canvas id="chart-duration"></canvas></div>
|
||||||
|
</div>
|
||||||
|
<div class="glass p-5 lg:col-span-2">
|
||||||
|
<h3 class="text-sm font-semibold text-zinc-200 mb-3">Top tags</h3>
|
||||||
|
<div class="h-64"><canvas id="chart-tags"></canvas></div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
</template>
|
||||||
|
|
||||||
|
<!-- ===== Channels ===== -->
|
||||||
|
<template x-if="view==='channels'">
|
||||||
|
<section class="space-y-6" x-transition.opacity>
|
||||||
|
<header><h1 class="text-2xl font-bold text-white">Channels</h1><p class="text-sm text-zinc-500">Add a channel URL to start tracking</p></header>
|
||||||
|
|
||||||
|
<form class="glass p-4 flex gap-2" @submit.prevent="addChannel()">
|
||||||
|
<input class="field flex-1" type="text" placeholder="https://www.youtube.com/@handle" x-model="forms.channelUrl" />
|
||||||
|
<button class="btn accent-grad btn-primary" :disabled="loading.addChannel">
|
||||||
|
<svg x-show="loading.addChannel" class="spin w-4 h-4" viewBox="0 0 24 24" fill="none"><circle cx="12" cy="12" r="9" stroke="currentColor" stroke-width="3" stroke-dasharray="40 20"/></svg>
|
||||||
|
Add channel
|
||||||
|
</button>
|
||||||
|
</form>
|
||||||
|
<p x-show="forms.channelError" class="text-sm text-rose-400 -mt-3" x-text="forms.channelError"></p>
|
||||||
|
|
||||||
|
<div class="glass overflow-hidden">
|
||||||
|
<table class="tbl">
|
||||||
|
<thead><tr><th>Name</th><th>Handle</th><th>Videos</th><th>Last scraped</th><th class="text-right">Actions</th></tr></thead>
|
||||||
|
<tbody>
|
||||||
|
<template x-if="loading.channels"><tr><td colspan="5" class="text-center text-zinc-500 py-8">loading…</td></tr></template>
|
||||||
|
<template x-if="!loading.channels && channels.items.length===0"><tr><td colspan="5" class="text-center text-zinc-500 py-8">No channels tracked yet.</td></tr></template>
|
||||||
|
<template x-for="c in channels.items" :key="c.channel_id">
|
||||||
|
<tr class="row-hover">
|
||||||
|
<td><button class="text-white font-medium hover:text-rose-400 transition-colors text-left" @click="openChannel(c.channel_id)" x-text="c.name"></button></td>
|
||||||
|
<td class="text-zinc-400" x-text="c.handle || '—'"></td>
|
||||||
|
<td><button class="font-mono text-zinc-300 hover:text-rose-400 transition-colors" @click="openChannel(c.channel_id)" x-text="fmtNum(c.video_count)"></button></td>
|
||||||
|
<td class="text-zinc-400 text-xs" x-text="c.last_scraped ? fmtDate(c.last_scraped) : '—'"></td>
|
||||||
|
<td class="text-right whitespace-nowrap">
|
||||||
|
<button class="btn btn-ghost !py-1 !px-2 text-xs" @click="openChannel(c.channel_id)">Videos</button>
|
||||||
|
<button class="btn btn-ghost !py-1 !px-2 text-xs" @click="openFolder('markdown', c.channel_id)" title="Open .md folder">.md</button>
|
||||||
|
<button class="btn btn-danger !py-1 !px-2 text-xs" @click="removeChannel(c.channel_id)">Remove</button>
|
||||||
|
</td>
|
||||||
|
</tr>
|
||||||
|
</template>
|
||||||
|
</tbody>
|
||||||
|
</table>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
</template>
|
||||||
|
|
||||||
|
<!-- ===== Videos ===== -->
|
||||||
|
<template x-if="view==='videos'">
|
||||||
|
<section class="space-y-5" x-transition.opacity>
|
||||||
|
<header><h1 class="text-2xl font-bold text-white">Videos</h1><p class="text-sm text-zinc-500">Browse and filter the catalog</p></header>
|
||||||
|
|
||||||
|
<div class="glass p-4 grid md:grid-cols-6 gap-3">
|
||||||
|
<select class="field" x-model="filters.channel" @change="loadVideos(1)"><option value="">All channels</option><template x-for="c in channels.items" :key="c.channel_id"><option :value="c.channel_id" x-text="c.name"></option></template></select>
|
||||||
|
<select class="field" x-model="filters.status" @change="loadVideos(1)">
|
||||||
|
<option value="">Any status</option>
|
||||||
|
<option value="done">Done</option><option value="pending">Pending</option><option value="error">Error</option><option value="skipped">Skipped</option><option value="new">New</option>
|
||||||
|
</select>
|
||||||
|
<input class="field" type="date" x-model="filters.from" @change="loadVideos(1)" />
|
||||||
|
<input class="field" type="date" x-model="filters.to" @change="loadVideos(1)" />
|
||||||
|
<input class="field" type="number" min="0" step="60" placeholder="min sec" x-model="filters.min_dur" @change="loadVideos(1)" />
|
||||||
|
<div class="flex gap-2">
|
||||||
|
<input class="field flex-1" type="text" placeholder="search title…" x-model="filters.q" @keydown.enter="loadVideos(1)" />
|
||||||
|
<button class="btn btn-ghost" @click="loadVideos(1)">Go</button>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="glass overflow-hidden">
|
||||||
|
<table class="tbl">
|
||||||
|
<thead><tr>
|
||||||
|
<th class="w-24">Thumb</th><th>Title</th><th>Channel</th>
|
||||||
|
<th><select class="bg-transparent text-zinc-500 uppercase text-[0.7rem] tracking-wider font-semibold" x-model="filters.sort" @change="loadVideos(videos.page)">
|
||||||
|
<option value="upload_date">Uploaded</option><option value="duration">Duration</option><option value="view_count">Views</option><option value="like_count">Likes</option>
|
||||||
|
</select></th>
|
||||||
|
<th>Duration</th><th>Views</th><th>Status</th>
|
||||||
|
</tr></thead>
|
||||||
|
<tbody>
|
||||||
|
<template x-if="loading.videos"><tr><td colspan="7" class="text-center text-zinc-500 py-10">loading…</td></tr></template>
|
||||||
|
<template x-if="!loading.videos && videos.items.length===0"><tr><td colspan="7" class="text-center text-zinc-500 py-10">No videos match these filters.</td></tr></template>
|
||||||
|
<template x-for="v in videos.items" :key="v.video_id">
|
||||||
|
<tr class="row-hover" @click="openVideo(v.video_id)">
|
||||||
|
<td><img class="thumb w-20 h-12" :src="v.thumbnail || ''" :alt="v.title" onerror="this.style.visibility='hidden'" /></td>
|
||||||
|
<td class="text-zinc-100 max-w-md truncate" x-text="v.title"></td>
|
||||||
|
<td class="text-zinc-400 text-xs" x-text="chanName(v.channel_id)"></td>
|
||||||
|
<td class="text-zinc-400 font-mono text-xs" x-text="fmtDate(v.upload_date)"></td>
|
||||||
|
<td class="font-mono text-zinc-300 text-xs" x-text="ts(v.duration)"></td>
|
||||||
|
<td class="font-mono text-zinc-300 text-xs" x-text="fmtNum(v.view_count)"></td>
|
||||||
|
<td><span class="pill" :class="statusClass(v.status)" x-text="v.status"></span></td>
|
||||||
|
</tr>
|
||||||
|
</template>
|
||||||
|
</tbody>
|
||||||
|
</table>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<!-- pagination -->
|
||||||
|
<div class="flex items-center justify-between text-sm">
|
||||||
|
<div class="text-zinc-500" x-text="videos.total ? (videos.size*(videos.page-1)+1)+'–'+Math.min(videos.size*videos.page, videos.total)+' of '+videos.total : '0 results'"></div>
|
||||||
|
<div class="flex gap-2">
|
||||||
|
<button class="btn btn-ghost" :disabled="videos.page<=1" @click="loadVideos(videos.page-1)">Prev</button>
|
||||||
|
<button class="btn btn-ghost" :disabled="videos.size*videos.page >= videos.total" @click="loadVideos(videos.page+1)">Next</button>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
</template>
|
||||||
|
|
||||||
|
<!-- ===== Detail ===== -->
|
||||||
|
<template x-if="view==='detail'">
|
||||||
|
<section class="space-y-5" x-transition.opacity>
|
||||||
|
<button class="btn btn-ghost" @click="setView('videos')">← Back to videos</button>
|
||||||
|
<template x-if="loading.detail"><div class="glass p-10 text-center text-zinc-500">loading…</div></template>
|
||||||
|
<template x-if="!loading.detail && detail.video">
|
||||||
|
<div class="space-y-5">
|
||||||
|
<div class="glass p-5 flex flex-col md:flex-row gap-5">
|
||||||
|
<img class="thumb w-full md:w-80 h-44 md:h-48" :src="detail.video.thumbnail || ''" :alt="detail.video.title" onerror="this.style.visibility='hidden'" />
|
||||||
|
<div class="flex-1 min-w-0">
|
||||||
|
<h1 class="text-xl font-bold text-white" x-text="detail.video.title"></h1>
|
||||||
|
<div class="text-sm text-zinc-500 mt-1" x-text="chanName(detail.video.channel_id)"></div>
|
||||||
|
<div class="flex flex-wrap gap-4 mt-4 text-sm">
|
||||||
|
<div><span class="text-zinc-500">Uploaded </span><span class="text-zinc-200 font-mono" x-text="fmtDate(detail.video.upload_date)"></span></div>
|
||||||
|
<div><span class="text-zinc-500">Duration </span><span class="text-zinc-200 font-mono" x-text="ts(detail.video.duration)"></span></div>
|
||||||
|
<div><span class="text-zinc-500">Views </span><span class="text-zinc-200 font-mono" x-text="fmtNum(detail.video.view_count)"></span></div>
|
||||||
|
<div><span class="text-zinc-500">Likes </span><span class="text-zinc-200 font-mono" x-text="fmtNum(detail.video.like_count)"></span></div>
|
||||||
|
<div><span class="pill" :class="statusClass(detail.video.status)" x-text="detail.video.status"></span></div>
|
||||||
|
</div>
|
||||||
|
<div class="flex flex-wrap gap-2 mt-4">
|
||||||
|
<template x-if="detail.video.url"><a class="btn accent-grad btn-primary" :href="detail.video.url" target="_blank" rel="noopener">Open on YouTube ↗</a></template>
|
||||||
|
<template x-for="t in (detail.video.tags||[]).slice(0,8)" :key="t"><span class="pill st-skipped" x-text="t"></span></template>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="grid lg:grid-cols-3 gap-5">
|
||||||
|
<!-- chapters -->
|
||||||
|
<div class="glass p-5 lg:col-span-1" x-show="detail.video.has_chapters && (detail.video.chapters||[]).length">
|
||||||
|
<h3 class="text-sm font-semibold text-zinc-200 mb-3">Chapters</h3>
|
||||||
|
<div class="space-y-1 max-h-[28rem] overflow-y-auto">
|
||||||
|
<template x-for="(ch, i) in (detail.video.chapters||[])" :key="i">
|
||||||
|
<button class="w-full flex items-center gap-2 text-left text-xs px-2 py-1.5 rounded hover:bg-zinc-800/60" @click="seekTranscript(ch.start_sec)">
|
||||||
|
<span class="font-mono accent-text shrink-0" x-text="ts(ch.start_sec)"></span>
|
||||||
|
<span class="text-zinc-300 truncate" x-text="ch.title"></span>
|
||||||
|
</button>
|
||||||
|
</template>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<!-- transcript -->
|
||||||
|
<div class="glass p-5" :class="(detail.video.has_chapters && (detail.video.chapters||[]).length) ? 'lg:col-span-2' : 'lg:col-span-3'">
|
||||||
|
<h3 class="text-sm font-semibold text-zinc-200 mb-3">Transcript</h3>
|
||||||
|
<template x-if="loading.transcript"><div class="text-zinc-500 text-sm">loading…</div></template>
|
||||||
|
<template x-if="!loading.transcript && detail.transcript.length===0"><div class="text-zinc-500 text-sm">No transcript available.</div></template>
|
||||||
|
<div class="space-y-2 max-h-[34rem] overflow-y-auto pr-1">
|
||||||
|
<template x-for="seg in detail.transcript" :key="seg.idx">
|
||||||
|
<div class="flex gap-3 text-sm leading-relaxed hover:bg-zinc-800/40 rounded px-2 py-1" :id="'seg-'+seg.idx">
|
||||||
|
<a class="font-mono accent-text shrink-0 mt-0.5 cursor-pointer" :href="detail.video.url ? detail.video.url+'&t='+Math.floor(seg.start_sec)+'s' : null" target="_blank" rel="noopener" x-text="ts(seg.start_sec)"></a>
|
||||||
|
<span class="text-zinc-300" x-text="seg.text"></span>
|
||||||
|
</div>
|
||||||
|
</template>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</template>
|
||||||
|
</section>
|
||||||
|
</template>
|
||||||
|
|
||||||
|
<!-- ===== Search ===== -->
|
||||||
|
<template x-if="view==='search'">
|
||||||
|
<section class="space-y-5" x-transition.opacity>
|
||||||
|
<header><h1 class="text-2xl font-bold text-white">Search transcripts</h1><p class="text-sm text-zinc-500">Find moments across all videos</p></header>
|
||||||
|
<form class="glass p-4 flex flex-col md:flex-row gap-2" @submit.prevent="loadSearch()">
|
||||||
|
<input class="field flex-1" type="text" placeholder="search query…" x-model="search.q" />
|
||||||
|
<select class="field md:w-56" x-model="search.channel"><option value="">All channels</option><template x-for="c in channels.items" :key="c.channel_id"><option :value="c.channel_id" x-text="c.name"></option></template></select>
|
||||||
|
<button class="btn accent-grad btn-primary" :disabled="loading.search">Search</button>
|
||||||
|
</form>
|
||||||
|
<template x-if="loading.search"><div class="glass p-10 text-center text-zinc-500">searching…</div></template>
|
||||||
|
<template x-if="!loading.search && search.ran && search.items.length===0"><div class="glass p-10 text-center text-zinc-500">No matches found.</div></template>
|
||||||
|
<div class="space-y-3">
|
||||||
|
<template x-for="r in search.items" :key="r.video_id+r.start_sec">
|
||||||
|
<div class="glass p-4 hover:border-rose-500/40 transition-colors cursor-pointer" @click="openVideo(r.video_id)">
|
||||||
|
<div class="flex items-center justify-between gap-3">
|
||||||
|
<div class="text-white font-medium truncate" x-text="r.title"></div>
|
||||||
|
<div class="font-mono accent-text text-sm shrink-0" x-text="ts(r.start_sec)"></div>
|
||||||
|
</div>
|
||||||
|
<div class="text-xs text-zinc-500 mt-0.5" x-text="chanName(r.channel_id)"></div>
|
||||||
|
<div class="text-sm text-zinc-300 mt-2 line-clamp-2" x-html="highlight(r.snippet, search.q)"></div>
|
||||||
|
</div>
|
||||||
|
</template>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
</template>
|
||||||
|
|
||||||
|
<!-- ===== Analysis ===== -->
|
||||||
|
<template x-if="view==='analysis'">
|
||||||
|
<section class="space-y-5" x-transition.opacity>
|
||||||
|
<header class="flex items-end justify-between">
|
||||||
|
<div><h1 class="text-2xl font-bold text-white">Analysis</h1><p class="text-sm text-zinc-500">Language patterns across the corpus</p></div>
|
||||||
|
<div class="flex gap-1 glass-soft p-1">
|
||||||
|
<button class="tab" :class="analysis.tab==='wordcloud'?'active':''" @click="analysis.tab='wordcloud'; loadAnalysis()">Word cloud</button>
|
||||||
|
<button class="tab" :class="analysis.tab==='topwords'?'active':''" @click="analysis.tab='topwords'; loadAnalysis()">Top words</button>
|
||||||
|
<button class="tab" :class="analysis.tab==='timeline'?'active':''" @click="analysis.tab='timeline'; loadAnalysis()">Timeline</button>
|
||||||
|
</div>
|
||||||
|
</header>
|
||||||
|
|
||||||
|
<div class="glass p-4 flex flex-wrap gap-3 items-center">
|
||||||
|
<select class="field md:w-56" x-model="analysis.channel" @change="loadAnalysis()"><option value="">All channels</option><template x-for="c in channels.items" :key="c.channel_id"><option :value="c.channel_id" x-text="c.name"></option></template></select>
|
||||||
|
<template x-if="analysis.tab==='timeline'">
|
||||||
|
<input class="field md:w-56" type="text" placeholder="term e.g. ai" x-model="analysis.term" @keydown.enter="loadAnalysis()" />
|
||||||
|
</template>
|
||||||
|
<button class="btn btn-ghost" @click="loadAnalysis()" :disabled="loading.analysis">Run</button>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<template x-if="loading.analysis"><div class="glass p-16 text-center text-zinc-500">crunching…</div></template>
|
||||||
|
|
||||||
|
<template x-if="!loading.analysis && analysis.tab==='wordcloud'">
|
||||||
|
<div class="glass p-8 min-h-[24rem] flex flex-wrap items-center justify-center">
|
||||||
|
<template x-if="analysis.words.length===0"><div class="text-zinc-500 text-sm">No data.</div></template>
|
||||||
|
<template x-for="(w,i) in analysis.words" :key="w.word+i">
|
||||||
|
<span class="cloud-word" :style="cloudStyle(w)" x-text="w.word"></span>
|
||||||
|
</template>
|
||||||
|
</div>
|
||||||
|
</template>
|
||||||
|
|
||||||
|
<template x-if="!loading.analysis && analysis.tab==='topwords'">
|
||||||
|
<div class="glass p-5"><div class="h-96"><canvas id="chart-topwords"></canvas></div></div>
|
||||||
|
</template>
|
||||||
|
|
||||||
|
<template x-if="!loading.analysis && analysis.tab==='timeline'">
|
||||||
|
<div class="glass p-5"><div class="h-96"><canvas id="chart-timeline"></canvas></div></div>
|
||||||
|
</template>
|
||||||
|
</section>
|
||||||
|
</template>
|
||||||
|
|
||||||
|
<!-- ===== Scrape ===== -->
|
||||||
|
<template x-if="view==='scrape'">
|
||||||
|
<section class="space-y-5" x-transition.opacity>
|
||||||
|
<header><h1 class="text-2xl font-bold text-white">Scrape</h1><p class="text-sm text-zinc-500">Launch a channel harvest job</p></header>
|
||||||
|
|
||||||
|
<div class="grid lg:grid-cols-2 gap-5">
|
||||||
|
<form class="glass p-5 space-y-3" @submit.prevent="startScrape()">
|
||||||
|
<label class="block text-xs text-zinc-500 uppercase tracking-wider">Channel</label>
|
||||||
|
<select class="field" x-model="scrape.form.channel_id" required>
|
||||||
|
<option value="">Select a channel…</option>
|
||||||
|
<template x-for="c in channels.items" :key="c.channel_id"><option :value="c.channel_id" x-text="c.name"></option></template>
|
||||||
|
</select>
|
||||||
|
<div class="grid grid-cols-2 gap-3">
|
||||||
|
<div>
|
||||||
|
<label class="block text-xs text-zinc-500 uppercase tracking-wider mb-1">Limit</label>
|
||||||
|
<input class="field" type="number" min="1" placeholder="∞" x-model="scrape.form.limit" />
|
||||||
|
</div>
|
||||||
|
<div>
|
||||||
|
<label class="block text-xs text-zinc-500 uppercase tracking-wider mb-1">Since</label>
|
||||||
|
<input class="field" type="date" x-model="scrape.form.since" />
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
<div>
|
||||||
|
<label class="block text-xs text-zinc-500 uppercase tracking-wider mb-1">Languages <span class="text-zinc-600 normal-case">(comma-sep)</span></label>
|
||||||
|
<input class="field" type="text" placeholder="en, es" x-model="scrape.form.languages" />
|
||||||
|
</div>
|
||||||
|
<div class="flex gap-6 pt-1">
|
||||||
|
<label class="flex items-center gap-2 text-sm text-zinc-300 cursor-pointer"><input type="checkbox" class="accent-rose-500" x-model="scrape.form.include_shorts" /> Include shorts</label>
|
||||||
|
<label class="flex items-center gap-2 text-sm text-zinc-300 cursor-pointer"><input type="checkbox" class="accent-rose-500" x-model="scrape.form.no_live" /> Skip live</label>
|
||||||
|
</div>
|
||||||
|
<button class="btn accent-grad btn-primary w-full" :disabled="loading.startScrape || !scrape.form.channel_id">
|
||||||
|
<svg x-show="loading.startScrape" class="spin w-4 h-4" viewBox="0 0 24 24" fill="none"><circle cx="12" cy="12" r="9" stroke="currentColor" stroke-width="3" stroke-dasharray="40 20"/></svg>
|
||||||
|
Start scrape
|
||||||
|
</button>
|
||||||
|
</form>
|
||||||
|
|
||||||
|
<!-- live progress -->
|
||||||
|
<div class="glass p-5 flex flex-col">
|
||||||
|
<div class="flex items-center justify-between">
|
||||||
|
<h3 class="text-sm font-semibold text-zinc-200">Live job</h3>
|
||||||
|
<template x-if="scrape.jobId">
|
||||||
|
<span class="pill" :class="scrape.error ? 'st-error' : (scrape.done ? 'st-done' : 'st-running')" x-text="scrape.error ? 'error' : (scrape.done ? 'done' : 'running')"></span>
|
||||||
|
</template>
|
||||||
|
</div>
|
||||||
|
<template x-if="!scrape.jobId"><div class="flex-1 flex items-center justify-center text-zinc-600 text-sm">No active job. Start one to see live progress.</div></template>
|
||||||
|
<template x-if="scrape.jobId">
|
||||||
|
<div class="mt-4 space-y-3 flex-1 flex flex-col">
|
||||||
|
<div class="flex justify-between text-xs text-zinc-500"><span>Progress</span><span class="font-mono" x-text="(scrape.progress.completed||0)+' / '+(scrape.progress.total||0)"></span></div>
|
||||||
|
<div class="progress-track"><div class="progress-fill" :style="'width:' + scrapePct() + '%'"></div></div>
|
||||||
|
<div class="log-panel flex-1 overflow-y-auto p-3 space-y-0.5 max-h-72">
|
||||||
|
<template x-for="(l,i) in scrape.log" :key="i"><div :class="l.startsWith('[') ? 'accent-text' : 'text-zinc-400'" x-text="l"></div></template>
|
||||||
|
</div>
|
||||||
|
<div class="flex gap-2">
|
||||||
|
<button class="btn btn-danger flex-1" @click="cancelScrape()" :disabled="scrape.done || scrape.error">Cancel</button>
|
||||||
|
<button class="btn btn-ghost flex-1" @click="closeStream()">Close stream</button>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</template>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<!-- recent jobs -->
|
||||||
|
<div class="glass overflow-hidden">
|
||||||
|
<div class="px-4 py-3 border-b border-zinc-800/70 flex items-center justify-between">
|
||||||
|
<h3 class="text-sm font-semibold text-zinc-200">Recent jobs</h3>
|
||||||
|
<button class="btn btn-ghost !py-1 !px-2 text-xs" @click="loadJobs()">Refresh</button>
|
||||||
|
</div>
|
||||||
|
<table class="tbl">
|
||||||
|
<thead><tr><th>Channel</th><th>Status</th><th>Progress</th><th>Started</th><th>Finished</th><th></th></tr></thead>
|
||||||
|
<tbody>
|
||||||
|
<template x-if="loading.jobs"><tr><td colspan="6" class="text-center text-zinc-500 py-6">loading…</td></tr></template>
|
||||||
|
<template x-if="!loading.jobs && scrape.jobs.length===0"><tr><td colspan="6" class="text-center text-zinc-500 py-6">No jobs yet.</td></tr></template>
|
||||||
|
<template x-for="j in scrape.jobs" :key="j.id">
|
||||||
|
<tr>
|
||||||
|
<td class="text-zinc-200 font-mono text-xs" x-text="j.channel_id"></td>
|
||||||
|
<td><span class="pill" :class="statusClass(j.status)" x-text="j.status"></span></td>
|
||||||
|
<td class="font-mono text-xs text-zinc-400" x-text="(j.completed||0)+' / '+(j.total||0)"></td>
|
||||||
|
<td class="font-mono text-xs text-zinc-500" x-text="j.started_at ? fmtDateTime(j.started_at) : '—'"></td>
|
||||||
|
<td class="font-mono text-xs text-zinc-500" x-text="j.finished_at ? fmtDateTime(j.finished_at) : '—'"></td>
|
||||||
|
<td class="text-right"><button class="btn btn-danger !py-1 !px-2 text-xs" @click="deleteJob(j.id)">Delete</button></td>
|
||||||
|
</tr>
|
||||||
|
</template>
|
||||||
|
</tbody>
|
||||||
|
</table>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
</template>
|
||||||
|
|
||||||
|
<!-- ===== Cookies ===== -->
|
||||||
|
<template x-if="view==='cookies'">
|
||||||
|
<section class="space-y-5" x-transition.opacity>
|
||||||
|
<header><h1 class="text-2xl font-bold text-white">Cookie vault</h1><p class="text-sm text-zinc-500">Upload YouTube cookie files to authenticate</p></header>
|
||||||
|
|
||||||
|
<div class="dropzone p-8 text-center transition-all" :class="cookies.drag ? 'drag' : ''"
|
||||||
|
@click="$refs.cookieFile.click()"
|
||||||
|
@dragover.prevent="cookies.drag=true"
|
||||||
|
@dragleave.prevent="cookies.drag=false"
|
||||||
|
@drop.prevent="handleDrop($event)">
|
||||||
|
<svg class="w-10 h-10 mx-auto text-zinc-600 mb-2" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="1.5"><path stroke-linecap="round" stroke-linejoin="round" d="M3 16.5v2.25A2.25 2.25 0 005.25 21h13.5A2.25 2.25 0 0021 18.75V16.5M16.5 12L12 7.5 7.5 12M12 7.5V18"/></svg>
|
||||||
|
<div class="text-sm text-zinc-300 font-medium">Drop <span class="font-mono accent-text">.txt</span> cookie files here</div>
|
||||||
|
<div class="text-xs text-zinc-500 mt-1">or click to browse</div>
|
||||||
|
<input x-ref="cookieFile" type="file" multiple accept=".txt,text/plain" class="hidden" @change="handleFiles($event.target.files)" />
|
||||||
|
</div>
|
||||||
|
<p x-show="cookies.uploadMsg" class="text-sm text-emerald-400" x-text="cookies.uploadMsg"></p>
|
||||||
|
|
||||||
|
<div class="glass overflow-hidden">
|
||||||
|
<div class="px-4 py-3 border-b border-zinc-800/70 flex items-center justify-between">
|
||||||
|
<h3 class="text-sm font-semibold text-zinc-200">Vault</h3>
|
||||||
|
<button class="btn btn-ghost !py-1 !px-2 text-xs" @click="loadCookies()">Refresh</button>
|
||||||
|
</div>
|
||||||
|
<table class="tbl">
|
||||||
|
<thead><tr><th>Label</th><th>File</th><th>Expires</th><th>Session</th><th>Active</th><th>Test</th><th></th></tr></thead>
|
||||||
|
<tbody>
|
||||||
|
<template x-if="loading.cookies"><tr><td colspan="7" class="text-center text-zinc-500 py-6">loading…</td></tr></template>
|
||||||
|
<template x-if="!loading.cookies && cookies.items.length===0"><tr><td colspan="7" class="text-center text-zinc-500 py-6">Vault is empty.</td></tr></template>
|
||||||
|
<template x-for="c in cookies.items" :key="c.id">
|
||||||
|
<tr>
|
||||||
|
<td class="text-white font-medium" x-text="c.label || c.filename"></td>
|
||||||
|
<td class="text-zinc-400 text-xs font-mono truncate max-w-[12rem]" x-text="c.filename"></td>
|
||||||
|
<td class="text-xs"><span class="font-mono" :class="c.expired ? 'text-rose-400' : 'text-zinc-300'" x-text="c.expires_at ? fmtDateTime(c.expires_at) : '—'"></span></td>
|
||||||
|
<td><span class="pill" :class="c.has_session ? 'st-done' : 'st-skipped'" x-text="c.has_session ? 'session' : 'none'"></span></td>
|
||||||
|
<td><input type="radio" :name="'cookie-active'" :checked="c.is_active" @change="activateCookie(c.id)" class="accent-rose-500" /></td>
|
||||||
|
<td>
|
||||||
|
<button class="btn btn-ghost !py-1 !px-2 text-xs" @click="testCookie(c.id)">Test</button>
|
||||||
|
<span class="text-xs ml-2" :class="cookies.testResult[c.id]===null ? 'text-zinc-500' : (cookies.testResult[c.id]?.ok ? 'text-emerald-400' : 'text-rose-400')" x-text="testLabel(c.id)"></span>
|
||||||
|
</td>
|
||||||
|
<td class="text-right"><button class="btn btn-danger !py-1 !px-2 text-xs" @click="removeCookie(c.id)">Delete</button></td>
|
||||||
|
</tr>
|
||||||
|
</template>
|
||||||
|
</tbody>
|
||||||
|
</table>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
</template>
|
||||||
|
|
||||||
|
<!-- ===== Tools (knowledge base) ===== -->
|
||||||
|
<template x-if="view==='tools'">
|
||||||
|
<section class="space-y-6" x-transition.opacity>
|
||||||
|
<header><h1 class="text-2xl font-bold text-white">Tools · Knowledge Base</h1><p class="text-sm text-zinc-500">Download channels as <span class="font-mono accent-text">.md</span> notes in the established Obsidian-ready format</p></header>
|
||||||
|
|
||||||
|
<!-- hero tool: scrape to markdown -->
|
||||||
|
<div class="glass p-6 flex flex-col md:flex-row md:items-center gap-4 border-rose-500/20">
|
||||||
|
<div class="w-12 h-12 rounded-xl accent-grad flex items-center justify-center shrink-0">
|
||||||
|
<svg viewBox="0 0 24 24" class="w-6 h-6 text-white" fill="none" stroke="currentColor" stroke-width="2"><path stroke-linecap="round" stroke-linejoin="round" d="M12 4v12m0 0l-4-4m4 4l4-4M4 18v2h16v-2"/></svg>
|
||||||
|
</div>
|
||||||
|
<div class="flex-1 min-w-0">
|
||||||
|
<div class="text-white font-semibold">Download as Markdown</div>
|
||||||
|
<div class="text-sm text-zinc-500 mt-0.5">Scrape a channel → one <span class="font-mono">.md</span> per video (frontmatter + transcript + chapters). Resumes automatically.</div>
|
||||||
|
</div>
|
||||||
|
<button class="btn accent-grad btn-primary" @click="setView('scrape')">Open scraper</button>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<!-- tool grid -->
|
||||||
|
<div class="grid md:grid-cols-2 lg:grid-cols-3 gap-4">
|
||||||
|
<div class="glass p-5">
|
||||||
|
<h3 class="text-sm font-semibold text-zinc-200">Re-render Markdown</h3>
|
||||||
|
<p class="text-xs text-zinc-500 mt-1 mb-3">Rebuild every <span class="font-mono">.md</span> from stored segments using the current template.</p>
|
||||||
|
<button class="btn btn-ghost w-full" :disabled="tools.reRendering" @click="reRender()">
|
||||||
|
<svg x-show="tools.reRendering" class="spin w-4 h-4" viewBox="0 0 24 24" fill="none"><circle cx="12" cy="12" r="9" stroke="currentColor" stroke-width="3" stroke-dasharray="40 20"/></svg>
|
||||||
|
Re-render all
|
||||||
|
</button>
|
||||||
|
<p x-show="tools.msg" class="text-xs mt-2" :class="tools.msg.startsWith('Failed') ? 'text-rose-400' : 'text-emerald-400'" x-text="tools.msg"></p>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="glass p-5">
|
||||||
|
<h3 class="text-sm font-semibold text-zinc-200">Export catalog</h3>
|
||||||
|
<p class="text-xs text-zinc-500 mt-1 mb-3">JSON / CSV / browsable HTML gallery of all videos.</p>
|
||||||
|
<button class="btn btn-ghost w-full" @click="setView('export')">Export tools</button>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="glass p-5">
|
||||||
|
<h3 class="text-sm font-semibold text-zinc-200">Cookie vault</h3>
|
||||||
|
<p class="text-xs text-zinc-500 mt-1 mb-3">Drag-and-drop a <span class="font-mono">.txt</span> session for members-only scraping.</p>
|
||||||
|
<button class="btn btn-ghost w-full" @click="setView('cookies')">Manage cookies</button>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<!-- open folders -->
|
||||||
|
<div class="space-y-3">
|
||||||
|
<h2 class="text-sm font-semibold text-zinc-300 uppercase tracking-wider">Open folders</h2>
|
||||||
|
<div class="grid md:grid-cols-2 lg:grid-cols-3 gap-4">
|
||||||
|
<button class="glass p-4 text-left hover:border-rose-500/40 transition-colors" @click="openFolder('markdown')">
|
||||||
|
<div class="flex items-center gap-2">
|
||||||
|
<svg viewBox="0 0 24 24" class="w-5 h-5 accent-text" fill="currentColor"><path d="M10 4H4a2 2 0 00-2 2v12a2 2 0 002 2h16a2 2 0 002-2V8a2 2 0 00-2-2h-8l-2-2z"/></svg>
|
||||||
|
<span class="text-white font-medium text-sm">Knowledge base (.md)</span>
|
||||||
|
</div>
|
||||||
|
<div class="text-[0.7rem] text-zinc-500 mt-1 font-mono truncate" x-text="folders.markdown?.path || '—'"></div>
|
||||||
|
</button>
|
||||||
|
<button class="glass p-4 text-left hover:border-rose-500/40 transition-colors" @click="openFolder('exports')">
|
||||||
|
<div class="flex items-center gap-2">
|
||||||
|
<svg viewBox="0 0 24 24" class="w-5 h-5 accent-text" fill="currentColor"><path d="M10 4H4a2 2 0 00-2 2v12a2 2 0 002 2h16a2 2 0 002-2V8a2 2 0 00-2-2h-8l-2-2z"/></svg>
|
||||||
|
<span class="text-white font-medium text-sm">Exports</span>
|
||||||
|
</div>
|
||||||
|
<div class="text-[0.7rem] text-zinc-500 mt-1 font-mono truncate" x-text="folders.exports?.path || '—'"></div>
|
||||||
|
</button>
|
||||||
|
<button class="glass p-4 text-left hover:border-rose-500/40 transition-colors" @click="openFolder('database')">
|
||||||
|
<div class="flex items-center gap-2">
|
||||||
|
<svg viewBox="0 0 24 24" class="w-5 h-5 accent-text" fill="currentColor"><path d="M12 3C7 3 3 4.8 3 7s4 4 9 4 9-1.8 9-4-4-4-9-4zm0 6c-3.3 0-6-.9-6-2s2.7-2 6-2 6 .9 6 2-2.7 2-6 2z"/></svg>
|
||||||
|
<span class="text-white font-medium text-sm">Database + data</span>
|
||||||
|
</div>
|
||||||
|
<div class="text-[0.7rem] text-zinc-500 mt-1 font-mono truncate" x-text="folders.database?.path || '—'"></div>
|
||||||
|
</button>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<!-- format reference -->
|
||||||
|
<div class="glass p-5" x-show="tools.fmt">
|
||||||
|
<h3 class="text-sm font-semibold text-zinc-200 mb-2">Established Markdown format</h3>
|
||||||
|
<p class="text-xs text-zinc-500 mb-3">Template: <span class="font-mono accent-text" x-text="tools.fmt?.template"></span></p>
|
||||||
|
<div class="text-xs text-zinc-500 mb-1 uppercase tracking-wider">Frontmatter</div>
|
||||||
|
<div class="flex flex-wrap gap-1 mb-3">
|
||||||
|
<template x-for="f in (tools.fmt?.frontmatter||[])" :key="f"><span class="pill st-skipped font-mono" x-text="f"></span></template>
|
||||||
|
</div>
|
||||||
|
<div class="text-xs text-zinc-500 mb-1 uppercase tracking-wider">Body</div>
|
||||||
|
<pre class="text-[0.72rem] text-zinc-400 font-mono bg-black/40 rounded-lg p-3 overflow-x-auto" x-text="tools.fmt?.body"></pre>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
</template>
|
||||||
|
|
||||||
|
<!-- ===== Export ===== -->
|
||||||
|
<template x-if="view==='export'">
|
||||||
|
<section class="space-y-5" x-transition.opacity>
|
||||||
|
<header class="flex items-end justify-between">
|
||||||
|
<div><h1 class="text-2xl font-bold text-white">Export</h1><p class="text-sm text-zinc-500">Download the catalog</p></div>
|
||||||
|
<button class="btn btn-ghost" @click="openFolder('exports')">Open exports folder ↗</button>
|
||||||
|
</header>
|
||||||
|
<div class="glass p-6 space-y-5 max-w-xl">
|
||||||
|
<div>
|
||||||
|
<label class="block text-xs text-zinc-500 uppercase tracking-wider mb-2">Format</label>
|
||||||
|
<div class="flex gap-2">
|
||||||
|
<template x-for="f in ['json','csv','html']" :key="f">
|
||||||
|
<button class="tab flex-1" :class="forms.exportFormat===f?'active':''" @click="forms.exportFormat=f" x-text="f.toUpperCase()"></button>
|
||||||
|
</template>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
<div>
|
||||||
|
<label class="block text-xs text-zinc-500 uppercase tracking-wider mb-2">Channel</label>
|
||||||
|
<select class="field" x-model="forms.exportChannel">
|
||||||
|
<option value="">All channels</option>
|
||||||
|
<template x-for="c in channels.items" :key="c.channel_id"><option :value="c.channel_id" x-text="c.name"></option></template>
|
||||||
|
</select>
|
||||||
|
</div>
|
||||||
|
<button class="btn accent-grad btn-primary w-full" @click="doExport()">
|
||||||
|
<svg class="w-4 h-4" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><path stroke-linecap="round" stroke-linejoin="round" d="M3 16.5v2.25A2.25 2.25 0 005.25 21h13.5A2.25 2.25 0 0021 18.75V16.5M7.5 12l4.5 4.5m0 0l4.5-4.5M12 3v13.5"/></svg>
|
||||||
|
Download
|
||||||
|
</button>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
</template>
|
||||||
|
|
||||||
|
</div>
|
||||||
|
</main>
|
||||||
|
</div>
|
||||||
|
</body>
|
||||||
|
</html>
|
||||||
@@ -0,0 +1,156 @@
|
|||||||
|
/* ===== yt-scraper platform :: dark "data command center" ===== */
|
||||||
|
[x-cloak] { display: none !important; }
|
||||||
|
|
||||||
|
html, body {
|
||||||
|
background: #0a0a0f;
|
||||||
|
color: #e5e7eb;
|
||||||
|
font-family: ui-sans-serif, system-ui, -apple-system, "Segoe UI", Roboto, Helvetica, Arial, sans-serif;
|
||||||
|
-webkit-font-smoothing: antialiased;
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ambient background glow */
|
||||||
|
body::before {
|
||||||
|
content: "";
|
||||||
|
position: fixed;
|
||||||
|
inset: 0;
|
||||||
|
pointer-events: none;
|
||||||
|
z-index: 0;
|
||||||
|
background:
|
||||||
|
radial-gradient(60rem 60rem at 12% -10%, rgba(244, 63, 94, 0.10), transparent 60%),
|
||||||
|
radial-gradient(50rem 50rem at 110% 10%, rgba(220, 38, 38, 0.08), transparent 55%);
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ---- scrollbars ---- */
|
||||||
|
::-webkit-scrollbar { width: 10px; height: 10px; }
|
||||||
|
::-webkit-scrollbar-track { background: transparent; }
|
||||||
|
::-webkit-scrollbar-thumb { background: #27272a; border-radius: 8px; }
|
||||||
|
::-webkit-scrollbar-thumb:hover { background: #3f3f46; }
|
||||||
|
|
||||||
|
/* ---- glass surfaces ---- */
|
||||||
|
.glass {
|
||||||
|
background: rgba(24, 24, 27, 0.6);
|
||||||
|
backdrop-filter: blur(12px);
|
||||||
|
-webkit-backdrop-filter: blur(12px);
|
||||||
|
border: 1px solid #27272a;
|
||||||
|
border-radius: 1rem;
|
||||||
|
}
|
||||||
|
.glass-soft {
|
||||||
|
background: rgba(24, 24, 27, 0.45);
|
||||||
|
border: 1px solid rgba(39, 39, 42, 0.8);
|
||||||
|
border-radius: 1rem;
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ---- accent ---- */
|
||||||
|
.accent-grad { background-image: linear-gradient(to right, #f43f5e, #dc2626); }
|
||||||
|
.accent-grad:hover { background-image: linear-gradient(to right, #fb7185, #ef4444); }
|
||||||
|
.accent-text { color: #f43f5e; }
|
||||||
|
.accent-ring:focus { outline: none; box-shadow: 0 0 0 2px rgba(244, 63, 94, 0.5); }
|
||||||
|
|
||||||
|
/* ---- nav ---- */
|
||||||
|
.nav-item {
|
||||||
|
display: flex;
|
||||||
|
align-items: center;
|
||||||
|
gap: 0.75rem;
|
||||||
|
width: 100%;
|
||||||
|
padding: 0.6rem 0.85rem;
|
||||||
|
border-radius: 0.75rem;
|
||||||
|
color: #a1a1aa;
|
||||||
|
font-size: 0.9rem;
|
||||||
|
font-weight: 500;
|
||||||
|
transition: all 0.15s ease;
|
||||||
|
cursor: pointer;
|
||||||
|
}
|
||||||
|
.nav-item:hover { color: #fafafa; background: rgba(39, 39, 42, 0.6); }
|
||||||
|
.nav-item.active {
|
||||||
|
color: #fff;
|
||||||
|
background: linear-gradient(to right, rgba(244, 63, 94, 0.18), rgba(220, 38, 38, 0.06));
|
||||||
|
box-shadow: inset 2px 0 0 #f43f5e;
|
||||||
|
}
|
||||||
|
.nav-item .nav-ico { width: 1.1rem; height: 1.1rem; opacity: 0.9; }
|
||||||
|
|
||||||
|
/* ---- form controls ---- */
|
||||||
|
.field {
|
||||||
|
background: #09090b;
|
||||||
|
border: 1px solid #27272a;
|
||||||
|
border-radius: 0.6rem;
|
||||||
|
color: #e5e7eb;
|
||||||
|
padding: 0.5rem 0.7rem;
|
||||||
|
font-size: 0.875rem;
|
||||||
|
transition: border-color 0.15s ease, box-shadow 0.15s ease;
|
||||||
|
width: 100%;
|
||||||
|
}
|
||||||
|
.field:focus { border-color: #f43f5e; box-shadow: 0 0 0 2px rgba(244, 63, 94, 0.35); outline: none; }
|
||||||
|
.field::placeholder { color: #52525b; }
|
||||||
|
select.field { appearance: none; background-image: linear-gradient(45deg, transparent 50%, #71717a 50%), linear-gradient(135deg, #71717a 50%, transparent 50%); background-position: calc(100% - 18px) 55%, calc(100% - 13px) 55%; background-size: 5px 5px, 5px 5px; background-repeat: no-repeat; padding-right: 2rem; }
|
||||||
|
|
||||||
|
/* ---- buttons ---- */
|
||||||
|
.btn {
|
||||||
|
display: inline-flex; align-items: center; justify-content: center; gap: 0.5rem;
|
||||||
|
padding: 0.55rem 0.95rem; border-radius: 0.6rem; font-size: 0.875rem; font-weight: 600;
|
||||||
|
transition: all 0.15s ease; cursor: pointer; white-space: nowrap;
|
||||||
|
}
|
||||||
|
.btn-primary { color: #fff; }
|
||||||
|
.btn-ghost { background: rgba(39, 39, 42, 0.6); color: #d4d4d8; border: 1px solid #27272a; }
|
||||||
|
.btn-ghost:hover { background: rgba(63, 63, 70, 0.8); color: #fff; }
|
||||||
|
.btn-danger { background: rgba(220, 38, 38, 0.12); color: #fca5a5; border: 1px solid rgba(220, 38, 38, 0.4); }
|
||||||
|
.btn-danger:hover { background: rgba(220, 38, 38, 0.25); color: #fff; }
|
||||||
|
.btn:disabled { opacity: 0.5; cursor: not-allowed; }
|
||||||
|
|
||||||
|
/* ---- pills ---- */
|
||||||
|
.pill {
|
||||||
|
display: inline-flex; align-items: center; gap: 0.35rem;
|
||||||
|
padding: 0.2rem 0.6rem; border-radius: 999px; font-size: 0.72rem; font-weight: 600;
|
||||||
|
border: 1px solid transparent;
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ---- table ---- */
|
||||||
|
.tbl { width: 100%; border-collapse: separate; border-spacing: 0; }
|
||||||
|
.tbl th { text-align: left; font-size: 0.7rem; text-transform: uppercase; letter-spacing: 0.05em; color: #71717a; font-weight: 600; padding: 0.6rem 0.75rem; border-bottom: 1px solid #27272a; }
|
||||||
|
.tbl td { padding: 0.6rem 0.75rem; border-bottom: 1px solid rgba(39, 39, 42, 0.6); font-size: 0.85rem; vertical-align: middle; }
|
||||||
|
.row-hover:hover { background: rgba(244, 63, 94, 0.05); cursor: pointer; }
|
||||||
|
|
||||||
|
/* ---- status colors ---- */
|
||||||
|
.st-done { background: rgba(34, 197, 94, 0.12); color: #4ade80; border-color: rgba(34, 197, 94, 0.35); }
|
||||||
|
.st-error { background: rgba(239, 68, 68, 0.12); color: #f87171; border-color: rgba(239, 68, 68, 0.35); }
|
||||||
|
.st-pending{ background: rgba(234, 179, 8, 0.12); color: #facc15; border-color: rgba(234, 179, 8, 0.35); }
|
||||||
|
.st-running{ background: rgba(59, 130, 246, 0.12); color: #60a5fa; border-color: rgba(59, 130, 246, 0.35); }
|
||||||
|
.st-skipped{ background: rgba(161, 161, 170, 0.12); color: #a1a1aa; border-color: rgba(161, 161, 170, 0.35); }
|
||||||
|
.st-new { background: rgba(168, 85, 247, 0.12); color: #c084fc; border-color: rgba(168, 85, 247, 0.35); }
|
||||||
|
|
||||||
|
/* ---- progress bar ---- */
|
||||||
|
.progress-track { background: #18181b; border-radius: 999px; overflow: hidden; height: 0.6rem; border: 1px solid #27272a; }
|
||||||
|
.progress-fill { height: 100%; background-image: linear-gradient(to right, #f43f5e, #dc2626); transition: width 0.3s ease; }
|
||||||
|
|
||||||
|
/* ---- dropzone ---- */
|
||||||
|
.dropzone {
|
||||||
|
border: 2px dashed #3f3f46;
|
||||||
|
border-radius: 1rem;
|
||||||
|
background: rgba(24, 24, 27, 0.4);
|
||||||
|
transition: all 0.15s ease;
|
||||||
|
cursor: pointer;
|
||||||
|
}
|
||||||
|
.dropzone:hover { border-color: #f43f5e; background: rgba(244, 63, 94, 0.05); }
|
||||||
|
.dropzone.drag { border-color: #f43f5e; background: rgba(244, 63, 94, 0.10); transform: scale(1.005); }
|
||||||
|
|
||||||
|
/* ---- spinner ---- */
|
||||||
|
.spin { animation: spin 0.9s linear infinite; }
|
||||||
|
@keyframes spin { to { transform: rotate(360deg); } }
|
||||||
|
|
||||||
|
/* ---- wordcloud ---- */
|
||||||
|
.cloud-word { display: inline-block; margin: 0.35rem 0.4rem; line-height: 1; cursor: default; transition: opacity 0.15s ease; }
|
||||||
|
.cloud-word:hover { opacity: 0.75; }
|
||||||
|
|
||||||
|
/* ---- tabs ---- */
|
||||||
|
.tab { padding: 0.5rem 0.95rem; border-radius: 0.6rem; font-size: 0.85rem; font-weight: 600; color: #a1a1aa; cursor: pointer; transition: all 0.15s ease; }
|
||||||
|
.tab:hover { color: #fff; }
|
||||||
|
.tab.active { color: #fff; background: rgba(244, 63, 94, 0.15); box-shadow: inset 0 0 0 1px rgba(244, 63, 94, 0.4); }
|
||||||
|
|
||||||
|
/* log panel */
|
||||||
|
.log-panel { background: #050507; border: 1px solid #27272a; border-radius: 0.75rem; font-family: ui-monospace, "JetBrains Mono", "Cascadia Code", monospace; font-size: 0.75rem; line-height: 1.4; }
|
||||||
|
|
||||||
|
/* thumbnail */
|
||||||
|
.thumb { border-radius: 0.5rem; object-fit: cover; background: #18181b; }
|
||||||
|
|
||||||
|
/* fade-up transition for views */
|
||||||
|
.fade-enter-active { transition: all 0.22s ease; }
|
||||||
|
.fade-enter-from { opacity: 0; transform: translateY(6px); }
|
||||||
@@ -0,0 +1,4 @@
|
|||||||
|
@echo off
|
||||||
|
chcp 65001 >nul
|
||||||
|
powershell -NoProfile -ExecutionPolicy Bypass -File "%~dp0scripts\start-server.ps1" %*
|
||||||
|
exit /b %errorlevel%
|
||||||
@@ -0,0 +1,4 @@
|
|||||||
|
@echo off
|
||||||
|
chcp 65001 >nul
|
||||||
|
powershell -NoProfile -ExecutionPolicy Bypass -File "%~dp0scripts\stop-server.ps1" %*
|
||||||
|
exit /b %errorlevel%
|
||||||
@@ -0,0 +1,47 @@
|
|||||||
|
<!DOCTYPE html>
|
||||||
|
<html lang="es">
|
||||||
|
<head>
|
||||||
|
<meta charset="utf-8">
|
||||||
|
<meta name="viewport" content="width=device-width, initial-scale=1">
|
||||||
|
<title>yt-scraper · Gallery</title>
|
||||||
|
<style>
|
||||||
|
:root { color-scheme: dark; }
|
||||||
|
body { background: #0a0a0f; color: #e5e7eb; font-family: system-ui, -apple-system, "Segoe UI", Roboto, sans-serif; margin: 0; padding: 2rem; }
|
||||||
|
h1 { font-size: 1.6rem; margin: 0 0 .25rem; }
|
||||||
|
h1 span { color: #f43f5e; }
|
||||||
|
.meta { color: #71717a; margin-bottom: 2rem; font-size: .9rem; }
|
||||||
|
.grid { display: grid; grid-template-columns: repeat(auto-fill, minmax(280px, 1fr)); gap: 1.25rem; }
|
||||||
|
.card { background: #18181b; border: 1px solid #27272a; border-radius: 14px; overflow: hidden; transition: transform .15s, border-color .15s; }
|
||||||
|
.card:hover { transform: translateY(-3px); border-color: #f43f5e; }
|
||||||
|
.card img { width: 100%; aspect-ratio: 16/9; object-fit: cover; background: #0a0a0f; display: block; }
|
||||||
|
.body { padding: .85rem 1rem 1rem; }
|
||||||
|
.title { font-weight: 600; font-size: .95rem; line-height: 1.3; margin: 0 0 .5rem; display: -webkit-box; -webkit-line-clamp: 2; -webkit-box-orient: vertical; overflow: hidden; }
|
||||||
|
.title a { color: inherit; text-decoration: none; }
|
||||||
|
.title a:hover { color: #fb7185; }
|
||||||
|
.tags { display: flex; flex-wrap: wrap; gap: .3rem; margin-top: .5rem; }
|
||||||
|
.tag { background: #27272a; color: #a1a1aa; font-size: .72rem; padding: .15rem .5rem; border-radius: 999px; }
|
||||||
|
.row { font-size: .78rem; color: #71717a; display: flex; gap: 1rem; }
|
||||||
|
.dur { color: #f43f5e; font-variant-numeric: tabular-nums; }
|
||||||
|
</style>
|
||||||
|
</head>
|
||||||
|
<body>
|
||||||
|
<h1>yt-scraper <span>·</span> Gallery</h1>
|
||||||
|
<div class="meta">{{ videos | length }} videos{% if channel_id %} · canal {{ channel_id }}{% endif %}</div>
|
||||||
|
<div class="grid">
|
||||||
|
{% for v in videos %}
|
||||||
|
<div class="card">
|
||||||
|
{% if v.thumbnail %}<img src="{{ v.thumbnail }}" alt="" loading="lazy">{% endif %}
|
||||||
|
<div class="body">
|
||||||
|
<div class="title"><a href="{{ v.url }}" target="_blank" rel="noopener">{{ v.title or v.video_id }}</a></div>
|
||||||
|
<div class="row">
|
||||||
|
<span>{{ (v.upload_date[:4] ~ '-' ~ v.upload_date[4:6] ~ '-' ~ v.upload_date[6:8]) if v.upload_date and v.upload_date|length == 8 else (v.upload_date or '') }}</span>
|
||||||
|
{% if v.duration %}<span class="dur">{{ (v.duration // 3600) ~ ':' ~ '%02d' % ((v.duration // 60) % 60) ~ ':' ~ '%02d' % (v.duration % 60) if v.duration >= 3600 else ((v.duration // 60) ~ ':' ~ '%02d' % (v.duration % 60)) }}</span>{% endif %}
|
||||||
|
{% if v.view_count %}<span>{{ "{:,}".format(v.view_count) }} views</span>{% endif %}
|
||||||
|
</div>
|
||||||
|
{% if v.tags %}<div class="tags">{% for t in v.tags[:5] %}<span class="tag">{{ t }}</span>{% endfor %}</div>{% endif %}
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
{% endfor %}
|
||||||
|
</div>
|
||||||
|
</body>
|
||||||
|
</html>
|
||||||
@@ -0,0 +1,36 @@
|
|||||||
|
---
|
||||||
|
video_id: {{ video_id }}
|
||||||
|
title: {{ title | quote_yaml }}
|
||||||
|
channel: {{ channel_name | quote_yaml }}
|
||||||
|
channel_id: {{ channel_id | default('') }}
|
||||||
|
channel_url: {{ channel_url | default('') }}
|
||||||
|
upload_date: {{ upload_date | default('') }}
|
||||||
|
duration: {{ duration | default(0) }}
|
||||||
|
url: {{ url }}
|
||||||
|
transcript_lang: {{ transcript_lang | default('') }}
|
||||||
|
transcript_src: {{ transcript_src | default('') }}
|
||||||
|
views: {{ view_count | default('') }}
|
||||||
|
likes: {{ like_count | default('') }}
|
||||||
|
tags: {{ tags | to_json if tags else '[]' }}
|
||||||
|
thumbnail: {{ thumbnail | default('') }}
|
||||||
|
---
|
||||||
|
|
||||||
|
# {{ title }}
|
||||||
|
|
||||||
|
> [Ver en YouTube]({{ url }})
|
||||||
|
|
||||||
|
{% if description %}{{ description }}
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
{% endif %}
|
||||||
|
## Transcripcion
|
||||||
|
|
||||||
|
{% for section in sections if section.segments %}
|
||||||
|
### {{ section.title }} ({{ section.start | format_timestamp }})
|
||||||
|
|
||||||
|
{% for seg in section.segments %}
|
||||||
|
**{{ seg.start | format_timestamp }}** · {{ seg.text }}
|
||||||
|
{% endfor %}
|
||||||
|
|
||||||
|
{% endfor %}
|
||||||
Vendored
+34
@@ -0,0 +1,34 @@
|
|||||||
|
{
|
||||||
|
"events": [
|
||||||
|
{
|
||||||
|
"tStartMs": 0,
|
||||||
|
"dDurationMs": 2000,
|
||||||
|
"segs": [
|
||||||
|
{"utf8": "Hola "},
|
||||||
|
{"utf8": "mundo"}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"tStartMs": 2000,
|
||||||
|
"dDurationMs": 3000,
|
||||||
|
"segs": [
|
||||||
|
{"utf8": "Esta es "},
|
||||||
|
{"utf8": "la segunda frase"}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"tStartMs": 5000,
|
||||||
|
"dDurationMs": 0,
|
||||||
|
"segs": [
|
||||||
|
{"utf8": "\n"}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"tStartMs": 6000,
|
||||||
|
"dDurationMs": 2000,
|
||||||
|
"segs": [
|
||||||
|
{"utf8": "Tercera frase"}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
Vendored
+10
@@ -0,0 +1,10 @@
|
|||||||
|
WEBVTT
|
||||||
|
|
||||||
|
00:00:00.000 --> 00:00:02.500
|
||||||
|
Hola desde VTT
|
||||||
|
|
||||||
|
00:00:02.500 --> 00:00:05.000
|
||||||
|
Segundo bloque <c.colorE1E1E1>con tag</c>
|
||||||
|
|
||||||
|
00:00:06.000 --> 00:00:09.000
|
||||||
|
Tercer bloque final
|
||||||
@@ -0,0 +1,58 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from yt_scraper.chapters import align_chapters, chapters_from_info, Chapter, Section
|
||||||
|
from yt_scraper.parse import Segment
|
||||||
|
|
||||||
|
|
||||||
|
def test_align_no_chapters():
|
||||||
|
segs = [Segment(0.0, 1.0, "a"), Segment(1.0, 2.0, "b")]
|
||||||
|
sections = align_chapters(segs, [])
|
||||||
|
assert len(sections) == 1
|
||||||
|
assert sections[0].title == "Transcripcion"
|
||||||
|
assert len(sections[0].segments) == 2
|
||||||
|
|
||||||
|
|
||||||
|
def test_align_with_chapters():
|
||||||
|
segs = [
|
||||||
|
Segment(0.0, 1.0, "intro"),
|
||||||
|
Segment(5.0, 6.0, "mid"),
|
||||||
|
Segment(20.0, 21.0, "end"),
|
||||||
|
]
|
||||||
|
chapters = [
|
||||||
|
Chapter(title="Intro", start_time=0.0, end_time=10.0),
|
||||||
|
Chapter(title="Desarrollo", start_time=10.0, end_time=30.0),
|
||||||
|
]
|
||||||
|
sections = align_chapters(segs, chapters)
|
||||||
|
assert len(sections) == 2
|
||||||
|
assert sections[0].title == "Intro"
|
||||||
|
assert len(sections[0].segments) == 2
|
||||||
|
assert sections[1].title == "Desarrollo"
|
||||||
|
assert len(sections[1].segments) == 1
|
||||||
|
assert sections[1].segments[0].text == "end"
|
||||||
|
|
||||||
|
|
||||||
|
def test_align_pre_chapter_segments():
|
||||||
|
segs = [Segment(0.0, 1.0, "antes")]
|
||||||
|
chapters = [Chapter(title="Cap 1", start_time=5.0, end_time=10.0)]
|
||||||
|
sections = align_chapters(segs, chapters)
|
||||||
|
assert len(sections) == 2
|
||||||
|
assert sections[0].title == "Inicio"
|
||||||
|
assert sections[0].segments[0].text == "antes"
|
||||||
|
|
||||||
|
|
||||||
|
def test_chapters_from_info():
|
||||||
|
info = {
|
||||||
|
"chapters": [
|
||||||
|
{"title": "Cap 1", "start_time": 0, "end_time": 60},
|
||||||
|
{"title": "Cap 2", "start_time": 60, "end_time": 120},
|
||||||
|
]
|
||||||
|
}
|
||||||
|
chapters = chapters_from_info(info)
|
||||||
|
assert len(chapters) == 2
|
||||||
|
assert chapters[0].title == "Cap 1"
|
||||||
|
assert chapters[1].start_time == 60.0
|
||||||
|
|
||||||
|
|
||||||
|
def test_chapters_from_info_empty():
|
||||||
|
assert chapters_from_info({}) == []
|
||||||
|
assert chapters_from_info({"chapters": []}) == []
|
||||||
@@ -0,0 +1,157 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import csv
|
||||||
|
import io
|
||||||
|
import json
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from yt_scraper.store import Store, VideoRef
|
||||||
|
from yt_scraper.parse import Segment
|
||||||
|
from yt_scraper import segments, export, analysis
|
||||||
|
from yt_scraper.render import build_filename_stem
|
||||||
|
|
||||||
|
|
||||||
|
SAMPLE_MD = """---
|
||||||
|
video_id: abc123
|
||||||
|
title: "Vaporwave and nostalgia"
|
||||||
|
channel: Alpha
|
||||||
|
upload_date: 20240115
|
||||||
|
duration: 600
|
||||||
|
views: 1234
|
||||||
|
likes: 56
|
||||||
|
tags: ["vaporwave", "nostalgia"]
|
||||||
|
thumbnail: http://img.test/abc.jpg
|
||||||
|
---
|
||||||
|
|
||||||
|
# Vaporwave and nostalgia
|
||||||
|
|
||||||
|
> [Ver en YouTube](https://www.youtube.com/watch?v=abc123)
|
||||||
|
|
||||||
|
## Transcripcion
|
||||||
|
|
||||||
|
### Intro (00:00)
|
||||||
|
|
||||||
|
**00:15** · Bienvenidos al video sobre vaporwave.
|
||||||
|
**00:32** · El vaporwave es una estetica muy interesante.
|
||||||
|
|
||||||
|
### Historia (01:15)
|
||||||
|
|
||||||
|
**01:15** · El vaporwave nacio en los 2010s.
|
||||||
|
**02:40** · dragon ball es un anime clasico.
|
||||||
|
"""
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture()
|
||||||
|
def store(tmp_path):
|
||||||
|
s = Store(tmp_path / "t.db")
|
||||||
|
s.upsert_channel("UC1", "@alpha", "Alpha", 1)
|
||||||
|
s.upsert_videos([VideoRef("abc123", "UC1", "Vaporwave and nostalgia", "https://y/watch?v=abc123", "20240115", 600)])
|
||||||
|
s.mark_done("abc123", "data/markdown/Alpha/x.md", "es", "auto", True)
|
||||||
|
return s
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_markdown_metadata():
|
||||||
|
parsed = segments.parse_markdown(SAMPLE_MD)
|
||||||
|
assert parsed.metadata["video_id"] == "abc123"
|
||||||
|
assert parsed.metadata["title"] == "Vaporwave and nostalgia"
|
||||||
|
assert parsed.metadata["channel"] == "Alpha"
|
||||||
|
assert len(parsed.segments) == 4
|
||||||
|
assert abs(parsed.segments[0].start - 15.0) < 0.01
|
||||||
|
assert "vaporwave" in parsed.segments[0].text.lower()
|
||||||
|
chapters = parsed.chapters
|
||||||
|
assert chapters[0]["title"] == "Intro"
|
||||||
|
assert chapters[1]["title"] == "Historia"
|
||||||
|
|
||||||
|
|
||||||
|
def test_backfill_from_markdown(store, tmp_path):
|
||||||
|
md = tmp_path / "md" / "Alpha" / "vapor.md"
|
||||||
|
md.parent.mkdir(parents=True)
|
||||||
|
md.write_text(SAMPLE_MD, encoding="utf-8")
|
||||||
|
n = segments.backfill_from_markdown(store, tmp_path / "md")
|
||||||
|
assert n == 1
|
||||||
|
v = store.get_video("abc123")
|
||||||
|
assert v.view_count == 1234
|
||||||
|
assert v.like_count == 56
|
||||||
|
assert json.loads(v.tags) == ["vaporwave", "nostalgia"]
|
||||||
|
assert v.thumbnail == "http://img.test/abc.jpg"
|
||||||
|
assert v.segments_json is not None
|
||||||
|
segs = store.get_segments("abc123")
|
||||||
|
assert len(segs) == 4
|
||||||
|
# search now works
|
||||||
|
hits = store.search_segments("vaporwave")
|
||||||
|
assert len(hits) >= 1
|
||||||
|
# idempotent
|
||||||
|
n2 = segments.backfill_from_markdown(store, tmp_path / "md")
|
||||||
|
assert n2 == 0
|
||||||
|
|
||||||
|
|
||||||
|
def test_ts_helpers():
|
||||||
|
assert segments.ts_to_seconds("01:30") == 90.0
|
||||||
|
assert segments.ts_to_seconds("1:02:03") == 3723.0
|
||||||
|
|
||||||
|
|
||||||
|
def test_export_json_csv(store, tmp_path):
|
||||||
|
store.store_segments("abc123", [Segment(0.0, 2.0, "hello vaporwave")])
|
||||||
|
store.update_video_metadata("abc123", view_count=1234, tags=["vaporwave"])
|
||||||
|
j = export.export_json(store, "UC1", tmp_path / "out.json")
|
||||||
|
data = json.loads(j.read_text(encoding="utf-8"))
|
||||||
|
assert len(data["videos"]) == 1
|
||||||
|
assert data["videos"][0]["video_id"] == "abc123"
|
||||||
|
|
||||||
|
c = export.export_csv(store, "UC1", tmp_path / "out.csv")
|
||||||
|
with open(c, encoding="utf-8") as f:
|
||||||
|
reader = csv.DictReader(f)
|
||||||
|
rows = list(reader)
|
||||||
|
assert rows[0]["video_id"] == "abc123"
|
||||||
|
assert rows[0]["view_count"] == "1234"
|
||||||
|
|
||||||
|
|
||||||
|
def test_export_srt(store, tmp_path):
|
||||||
|
segs = [Segment(1.5, 3.5, "first line"), Segment(5.0, 7.0, "second line")]
|
||||||
|
store.store_segments("abc123", segs)
|
||||||
|
out = export.export_srt_video(store, "abc123", tmp_path / "x.srt")
|
||||||
|
assert out is not None
|
||||||
|
text = out.read_text(encoding="utf-8")
|
||||||
|
assert "1\n" in text
|
||||||
|
assert "00:00:01,500 --> 00:00:03,500" in text
|
||||||
|
assert "first line" in text
|
||||||
|
assert "second line" in text
|
||||||
|
|
||||||
|
|
||||||
|
def test_export_html(store, tmp_path):
|
||||||
|
store.update_video_metadata("abc123", thumbnail="http://t.test/x.jpg", tags=["vaporwave"])
|
||||||
|
out = export.export_html(store, "UC1", tmp_path / "g.html")
|
||||||
|
html = out.read_text(encoding="utf-8")
|
||||||
|
assert "abc123" in html
|
||||||
|
assert "Gallery" in html
|
||||||
|
|
||||||
|
|
||||||
|
def test_word_frequency(store):
|
||||||
|
segs = [
|
||||||
|
Segment(0, 1, "vaporwave vaporwave nostalgia nostalgia nostalgia"),
|
||||||
|
Segment(2, 3, "dragon dragon anime"),
|
||||||
|
]
|
||||||
|
store.store_segments("abc123", segs)
|
||||||
|
freq = analysis.word_frequency(store, "UC1", top=5)
|
||||||
|
assert freq["nostalgia"] == 3
|
||||||
|
assert freq["vaporwave"] == 2
|
||||||
|
assert freq["dragon"] == 2
|
||||||
|
# stopwords filtered
|
||||||
|
assert "y" not in freq
|
||||||
|
assert "the" not in freq
|
||||||
|
|
||||||
|
|
||||||
|
def test_term_timeline(store):
|
||||||
|
segs = [Segment(0, 1, "vaporwave mentioned here")]
|
||||||
|
store.store_segments("abc123", segs)
|
||||||
|
tl = analysis.term_timeline(store, "vaporwave", "UC1")
|
||||||
|
assert len(tl) == 1
|
||||||
|
assert tl[0][0] == "2024-01"
|
||||||
|
|
||||||
|
|
||||||
|
def test_build_filename_stem():
|
||||||
|
# callers pass already-normalized dates; stem uses them verbatim
|
||||||
|
stem = build_filename_stem("2024-01-01", "Hola Mundo!", "{upload_date}_{slug}")
|
||||||
|
assert stem == "2024-01-01_hola-mundo"
|
||||||
@@ -0,0 +1,59 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from yt_scraper.parse import Segment, parse_json3, parse_vtt, merge_adjacent
|
||||||
|
|
||||||
|
FIXTURES = Path(__file__).parent / "fixtures"
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_json3_basic():
|
||||||
|
raw = (FIXTURES / "sample.json3").read_text(encoding="utf-8")
|
||||||
|
segs = parse_json3(raw)
|
||||||
|
assert len(segs) >= 2
|
||||||
|
assert segs[0].start == 0.0
|
||||||
|
assert "Hola" in segs[0].text
|
||||||
|
assert "segunda" in segs[0].text
|
||||||
|
assert segs[-1].start == 6.0
|
||||||
|
assert "Tercera" in segs[-1].text
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_json3_skips_empty_events():
|
||||||
|
raw = (FIXTURES / "sample.json3").read_text(encoding="utf-8")
|
||||||
|
segs = parse_json3(raw)
|
||||||
|
texts = [s.text for s in segs]
|
||||||
|
assert all(t.strip() for t in texts)
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_vtt_basic():
|
||||||
|
raw = (FIXTURES / "sample.vtt").read_text(encoding="utf-8")
|
||||||
|
segs = parse_vtt(raw)
|
||||||
|
assert len(segs) == 2
|
||||||
|
assert segs[0].start == 0.0
|
||||||
|
assert "Hola" in segs[0].text
|
||||||
|
assert "tag" in segs[0].text
|
||||||
|
assert segs[1].start == 6.0
|
||||||
|
assert "Tercer" in segs[1].text
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_vtt_strips_tags():
|
||||||
|
raw = (FIXTURES / "sample.vtt").read_text(encoding="utf-8")
|
||||||
|
segs = parse_vtt(raw)
|
||||||
|
assert "<c." not in segs[1].text
|
||||||
|
|
||||||
|
|
||||||
|
def test_merge_adjacent():
|
||||||
|
segs = [
|
||||||
|
Segment(start=0.0, end=1.0, text="a"),
|
||||||
|
Segment(start=1.0, end=2.0, text="b"),
|
||||||
|
Segment(start=5.0, end=6.0, text="c"),
|
||||||
|
]
|
||||||
|
merged = merge_adjacent(segs, max_gap=0.0)
|
||||||
|
assert len(merged) == 2
|
||||||
|
assert "a b" in merged[0].text
|
||||||
|
assert merged[1].text == "c"
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_json3_invalid_json():
|
||||||
|
assert parse_json3("not json") == []
|
||||||
|
assert parse_json3('{"events": []}') == []
|
||||||
@@ -0,0 +1,77 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from yt_scraper.render import build_filename_stem, format_timestamp, quote_yaml
|
||||||
|
from yt_scraper.chapters import Section
|
||||||
|
from yt_scraper.parse import Segment
|
||||||
|
|
||||||
|
TEMPLATES = Path(__file__).parent.parent / "templates"
|
||||||
|
|
||||||
|
|
||||||
|
def test_format_timestamp_minutes():
|
||||||
|
assert format_timestamp(0.0) == "00:00"
|
||||||
|
assert format_timestamp(65.0) == "01:05"
|
||||||
|
assert format_timestamp(3599.0) == "59:59"
|
||||||
|
|
||||||
|
|
||||||
|
def test_format_timestamp_hours():
|
||||||
|
assert format_timestamp(3661.0) == "01:01:01"
|
||||||
|
|
||||||
|
|
||||||
|
def test_quote_yaml_plain():
|
||||||
|
assert quote_yaml("hola") == "hola"
|
||||||
|
|
||||||
|
|
||||||
|
def test_quote_yaml_special():
|
||||||
|
result = quote_yaml('hello "world"')
|
||||||
|
assert result.startswith('"')
|
||||||
|
assert result.endswith('"')
|
||||||
|
assert '\\"' in result
|
||||||
|
|
||||||
|
|
||||||
|
def test_quote_yaml_none():
|
||||||
|
assert quote_yaml(None) == '""'
|
||||||
|
|
||||||
|
|
||||||
|
def test_build_filename_stem():
|
||||||
|
stem = build_filename_stem("20240115", "Mi Video Genial!")
|
||||||
|
assert stem.startswith("20240115")
|
||||||
|
assert "mi-video-genial" in stem.lower()
|
||||||
|
|
||||||
|
|
||||||
|
def test_build_filename_stem_no_date():
|
||||||
|
stem = build_filename_stem(None, "Titulo")
|
||||||
|
assert "unknown" in stem
|
||||||
|
assert "titulo" in stem.lower()
|
||||||
|
|
||||||
|
|
||||||
|
def test_render_markdown(tmp_path):
|
||||||
|
template_file = tmp_path / "test.j2"
|
||||||
|
template_file.write_text(
|
||||||
|
"---\ntitle: {{ title | quote_yaml }}\n---\n# {{ title }}\n\n"
|
||||||
|
"{% for s in sections %}## {{ s.title }}\n{% for seg in s.segments %}**{{ seg.start | format_timestamp }}** {{ seg.text }}\n{% endfor %}{% endfor %}",
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
out_dir = tmp_path / "out"
|
||||||
|
sections = [
|
||||||
|
Section(
|
||||||
|
title="Intro",
|
||||||
|
start=0.0,
|
||||||
|
segments=[Segment(start=0.0, end=1.0, text="Hola"), Segment(start=5.0, end=6.0, text="Mundo")],
|
||||||
|
)
|
||||||
|
]
|
||||||
|
from yt_scraper.render import render_markdown
|
||||||
|
|
||||||
|
md_path = render_markdown(
|
||||||
|
template_file,
|
||||||
|
out_dir,
|
||||||
|
"20240115_test",
|
||||||
|
{"title": "Test Video", "sections": sections},
|
||||||
|
)
|
||||||
|
assert md_path.exists()
|
||||||
|
content = md_path.read_text(encoding="utf-8")
|
||||||
|
assert "# Test Video" in content
|
||||||
|
assert "## Intro" in content
|
||||||
|
assert "**00:00** Hola" in content
|
||||||
|
assert "**00:05** Mundo" in content
|
||||||
@@ -0,0 +1,169 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from yt_scraper.store import Store, VideoRef
|
||||||
|
from yt_scraper.parse import Segment
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture()
|
||||||
|
def store(tmp_path):
|
||||||
|
return Store(tmp_path / "test.db")
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture()
|
||||||
|
def seeded_store(store):
|
||||||
|
store.upsert_channel("UC1", "@alpha", "Alpha", 2)
|
||||||
|
store.upsert_channel("UC2", "@beta", "Beta", 1)
|
||||||
|
refs = [
|
||||||
|
VideoRef("v1", "UC1", "Hello world vaporwave", "https://y/watch?v=v1", "20240101", 120),
|
||||||
|
VideoRef("v2", "UC1", "Dragon ball nostalgia", "https://y/watch?v=v2", "20240201", 200),
|
||||||
|
VideoRef("v3", "UC2", "Vaporwave aesthetics", "https://y/watch?v=v3", "20240301", 300),
|
||||||
|
]
|
||||||
|
store.upsert_videos(refs)
|
||||||
|
return store
|
||||||
|
|
||||||
|
|
||||||
|
def test_migration_idempotent(tmp_path):
|
||||||
|
db = tmp_path / "a.db"
|
||||||
|
s1 = Store(db)
|
||||||
|
s2 = Store(db) # second instantiation should not error
|
||||||
|
# new columns exist
|
||||||
|
cols = {r["name"] for r in s1._connect().execute("PRAGMA table_info(videos)")}
|
||||||
|
for c in ("view_count", "tags", "segments_json", "chapters_json", "thumbnail"):
|
||||||
|
assert c in cols
|
||||||
|
# new tables exist
|
||||||
|
names = {r["name"] for r in s1._connect().execute("SELECT name FROM sqlite_master WHERE type='table'")}
|
||||||
|
for t in ("transcript_segments", "cookies_meta", "scrape_jobs"):
|
||||||
|
assert t in names
|
||||||
|
# fts table
|
||||||
|
fts = {r["name"] for r in s1._connect().execute("SELECT name FROM sqlite_master WHERE type='table' AND name LIKE '%fts%'")}
|
||||||
|
assert "transcript_fts" in fts
|
||||||
|
|
||||||
|
|
||||||
|
def test_store_and_search_segments(seeded_store):
|
||||||
|
store = seeded_store
|
||||||
|
segs = [
|
||||||
|
Segment(0.0, 2.0, "vaporwave is an aesthetic"),
|
||||||
|
Segment(3.0, 5.0, "dragon ball is a classic anime"),
|
||||||
|
]
|
||||||
|
store.store_segments("v1", segs)
|
||||||
|
store.store_segments("v3", [Segment(0.0, 2.0, "retro vaporwave vibes")])
|
||||||
|
|
||||||
|
hits = store.search_segments("vaporwave")
|
||||||
|
assert len(hits) == 2
|
||||||
|
vids = {h.video_id for h in hits}
|
||||||
|
assert vids == {"v1", "v3"}
|
||||||
|
|
||||||
|
# channel-scoped search
|
||||||
|
hits_c = store.search_segments("vaporwave", channel_id="UC2")
|
||||||
|
assert len(hits_c) == 1
|
||||||
|
assert hits_c[0].video_id == "v3"
|
||||||
|
|
||||||
|
# get_segments ordered
|
||||||
|
seg_rows = store.get_segments("v1")
|
||||||
|
assert len(seg_rows) == 2
|
||||||
|
assert seg_rows[0].idx == 0
|
||||||
|
|
||||||
|
|
||||||
|
def test_store_segments_overwrites(seeded_store):
|
||||||
|
store = seeded_store
|
||||||
|
store.store_segments("v1", [Segment(0.0, 1.0, "old")])
|
||||||
|
store.store_segments("v1", [Segment(0.0, 1.0, "new"), Segment(2.0, 3.0, "second")])
|
||||||
|
rows = store.get_segments("v1")
|
||||||
|
assert len(rows) == 2
|
||||||
|
assert rows[0].text == "new"
|
||||||
|
|
||||||
|
|
||||||
|
def test_update_video_metadata(seeded_store):
|
||||||
|
store = seeded_store
|
||||||
|
store.update_video_metadata("v1", view_count=1000, like_count=50, tags=["music", "retro"], thumbnail="http://t.png")
|
||||||
|
v = store.get_video("v1")
|
||||||
|
assert v.view_count == 1000
|
||||||
|
assert v.like_count == 50
|
||||||
|
assert json.loads(v.tags) == ["music", "retro"]
|
||||||
|
assert v.thumbnail == "http://t.png"
|
||||||
|
|
||||||
|
|
||||||
|
def test_query_videos_filters(seeded_store):
|
||||||
|
store = seeded_store
|
||||||
|
rows, total = store.query_videos(channel_id="UC1")
|
||||||
|
assert total == 2
|
||||||
|
rows2, total2 = store.query_videos(min_duration=250)
|
||||||
|
assert total2 == 1
|
||||||
|
assert rows2[0].video_id == "v3"
|
||||||
|
rows3, total3 = store.query_videos(q="dragon")
|
||||||
|
assert total3 == 1
|
||||||
|
assert rows3[0].video_id == "v2"
|
||||||
|
|
||||||
|
|
||||||
|
def test_cookie_vault(store):
|
||||||
|
from yt_scraper import cookies
|
||||||
|
sample = (
|
||||||
|
"# Netscape HTTP Cookie File\n"
|
||||||
|
".youtube.com\tTRUE\t/\tTRUE\t1800287621\tSID\tabc\n"
|
||||||
|
".youtube.com\tTRUE\t/\tTRUE\t1800287621\tSAPISID\tdef\n"
|
||||||
|
".youtube.com\tTRUE\t/\tFALSE\t1800287621\tPREF\tf1=2\n"
|
||||||
|
)
|
||||||
|
cid = cookies.import_text(store, sample, label="test", cookie_dir=str(Path(store.db_path).parent / "ck"))
|
||||||
|
vault = cookies.list_vault(store)
|
||||||
|
assert len(vault) == 1
|
||||||
|
assert vault[0].has_session is True
|
||||||
|
assert vault[0].cookie_count == 3
|
||||||
|
# set active + resolve
|
||||||
|
cookies.set_active(store, cid)
|
||||||
|
path = cookies.resolve_active_path(store, cookie_dir=str(Path(store.db_path).parent / "ck"))
|
||||||
|
assert path and Path(path).exists()
|
||||||
|
# delete
|
||||||
|
cookies.delete(store, cid, cookie_dir=str(Path(store.db_path).parent / "ck"))
|
||||||
|
assert cookies.list_vault(store) == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_cookie_invalid_rejected(store):
|
||||||
|
from yt_scraper import cookies
|
||||||
|
with pytest.raises(ValueError):
|
||||||
|
cookies.import_text(store, "not a cookie file", label="bad", cookie_dir=str(Path(store.db_path).parent / "ck"))
|
||||||
|
|
||||||
|
|
||||||
|
def test_cookie_only_one_active(store):
|
||||||
|
from yt_scraper import cookies
|
||||||
|
cdir = str(Path(store.db_path).parent / "ck2")
|
||||||
|
sample = "# Netscape\n.youtube.com\tTRUE\t/\tTRUE\t1800287621\tSID\tx\n"
|
||||||
|
c1 = cookies.import_text(store, sample, label="a", cookie_dir=cdir)
|
||||||
|
c2 = cookies.import_text(store, sample, label="b", cookie_dir=cdir)
|
||||||
|
cookies.set_active(store, c1)
|
||||||
|
cookies.set_active(store, c2)
|
||||||
|
actives = [c for c in cookies.list_vault(store) if c.is_active]
|
||||||
|
assert len(actives) == 1
|
||||||
|
assert actives[0].id == c2
|
||||||
|
|
||||||
|
|
||||||
|
def test_jobs_crud(store):
|
||||||
|
store.create_job("job1", "UC1", {"limit": 5})
|
||||||
|
store.update_job("job1", status="running", total=5)
|
||||||
|
store.update_job("job1", completed=3)
|
||||||
|
j = store.get_job("job1")
|
||||||
|
assert j.status == "running"
|
||||||
|
assert j.total == 5
|
||||||
|
assert j.completed == 3
|
||||||
|
store.update_job("job1", status="done", completed=5, finished=True)
|
||||||
|
j2 = store.get_job("job1")
|
||||||
|
assert j2.status == "done"
|
||||||
|
assert j2.finished_at is not None
|
||||||
|
|
||||||
|
|
||||||
|
def test_dashboard_aggregates(seeded_store):
|
||||||
|
store = seeded_store
|
||||||
|
store.store_segments("v1", [Segment(0.0, 2.0, "vaporwave hello")])
|
||||||
|
store.update_video_metadata("v1", view_count=500, tags=["music", "vaporwave"])
|
||||||
|
store.update_video_metadata("v2", tags=["music", "anime"])
|
||||||
|
d = store.dashboard()
|
||||||
|
assert len(d["channels"]) == 2
|
||||||
|
ch1 = [c for c in d["channels"] if c["channel_id"] == "UC1"][0]
|
||||||
|
assert ch1["video_count_db"] == 2
|
||||||
|
assert ch1["total_views"] == 500
|
||||||
|
tags = {t["tag"]: t["count"] for t in d["top_tags"]}
|
||||||
|
assert tags.get("music") == 2
|
||||||
Reference in New Issue
Block a user