OCR de imagenes con Mistral: ImageProcessor (png/jpg/webp/gif/bmp/tiff), modelo mistral-ocr-latest y recuperacion de fuentes unknown

This commit is contained in:
urieljareth
2026-09-13 22:39:39 -06:00
parent 239a7510c6
commit 15cf00cc42
12 changed files with 79 additions and 14 deletions
+1
View File
@@ -18,6 +18,7 @@
- `KnowledgeManager` owns filesystem layout and SQLite metadata; processors should not write final Markdown directly.
- `IngestionService` selects processors by `source_type` and wraps output with Markdown frontmatter for LLM ingestion.
- Text and Markdown are read directly; `.docx` uses `python-docx`; PDFs use PyMuPDF text extraction unless `use_ocr=true` is passed. PDF uploads can also pass `page_ranges` like `1-5,8,10-12`; ranges apply to both standard extraction and Mistral OCR.
- Images (png, jpg, jpeg, webp, gif, bmp, tiff) are always processed with Mistral OCR (`ImageProcessor`), no `use_ocr` flag needed. Sources stored with `source_type="unknown"` (uploaded before a format was supported) are re-detected by extension on reprocess.
- OCR uses Mistral `https://api.mistral.ai/v1/ocr` with `MISTRAL_OCR_MODEL`, defaulting to `mistral-ocr-latest`.
- Audio uses Deepgram `https://api.deepgram.com/v1/listen?model=nova-3&smart_format=true&language=es`.
- Video processing requires `ffmpeg`; audio is extracted to `data/tmp/`, transcribed with Deepgram, then deleted.