OCR de imagenes con Mistral: ImageProcessor (png/jpg/webp/gif/bmp/tiff), modelo mistral-ocr-latest y recuperacion de fuentes unknown
This commit is contained in:
@@ -18,6 +18,7 @@
|
||||
- `KnowledgeManager` owns filesystem layout and SQLite metadata; processors should not write final Markdown directly.
|
||||
- `IngestionService` selects processors by `source_type` and wraps output with Markdown frontmatter for LLM ingestion.
|
||||
- Text and Markdown are read directly; `.docx` uses `python-docx`; PDFs use PyMuPDF text extraction unless `use_ocr=true` is passed. PDF uploads can also pass `page_ranges` like `1-5,8,10-12`; ranges apply to both standard extraction and Mistral OCR.
|
||||
- Images (png, jpg, jpeg, webp, gif, bmp, tiff) are always processed with Mistral OCR (`ImageProcessor`), no `use_ocr` flag needed. Sources stored with `source_type="unknown"` (uploaded before a format was supported) are re-detected by extension on reprocess.
|
||||
- OCR uses Mistral `https://api.mistral.ai/v1/ocr` with `MISTRAL_OCR_MODEL`, defaulting to `mistral-ocr-latest`.
|
||||
- Audio uses Deepgram `https://api.deepgram.com/v1/listen?model=nova-3&smart_format=true&language=es`.
|
||||
- Video processing requires `ffmpeg`; audio is extracted to `data/tmp/`, transcribed with Deepgram, then deleted.
|
||||
|
||||
Reference in New Issue
Block a user