Offline Document Intelligence Studio
One FastAPI service that OCRs documents and runs local-LLM summarisation, extraction and question answering, with no cloud AI API.
Overview
A FastAPI app with a small HTML dashboard and six routers: OCR, chat, summarise, extract, predict and retrieval Q&A. LLM work is delegated to a local Ollama server, and the whole thing ships as a Docker image and a Windows executable.
Problem
Sensitive documents shouldn't have to leave the machine. The goal was OCR, summaries, field extraction and Q&A over uploaded files using only local models.
Architecture
- DashboardJinja2 + plain HTML/JS front end served by FastAPI.
- Routers/ocr /chat /summarize /extract /predict /rag, each a thin layer over a service.
- ServicesOCR, LLM, summary, extraction, prediction and retrieval logic.
- OllamaExternal local server called over HTTP, configurable through environment variables.
- PackagingDockerfile, compose file and a PyInstaller-aware path setup for the Windows executable.
AI pipeline
- 01
Input
Image or PDF upload.
- 02
Preprocessing
PDF pages at 300 dpi, then grayscale, resize, denoise, deskew and threshold with OpenCV.
- 03
OCR
Tesseract text extraction.
- 04
Retrieval
500-character chunks indexed with TF-IDF; top matches by cosine similarity.
- 05
LLM
Ollama answers only from the retrieved context, or summarises and extracts fields as JSON.
- 06
Output
Text, summary, extracted fields or a grounded answer in the dashboard.
- OCR with OpenCV preprocessing (denoise, deskew, threshold)
- Local LLM (Ollama, llama3.2) for chat, summaries and JSON extraction
- TF-IDF retrieval for document Q&A
- RandomForest classifier (Iris, as an assignment component)
Engineering
- Router and service separation
- Endpoints stay thin; OCR, LLM, retrieval and prediction each live in their own service module.
- Environment-driven config
- The Ollama URL, Poppler path and Tesseract command are environment variables, and a cloud-demo mode flag switches off what can't run in the hosted version.
- Honest hosting write-up
- The repo documents why Ollama can't run on a free Hugging Face Space, and what the hosted demo therefore omits.
- Frozen-app support
- Path handling accounts for PyInstaller's runtime layout so the same code runs as a script or a packaged exe.
Challenges
- A local LLM doesn't fit the free Hugging Face Space limits, so the hosted demo runs without the model-backed features.
- Packaging for Windows meant handling frozen-app resource paths.
Results & limits
Measured
- No evaluation numbers are recorded. Testing was manual.
Known limits
- Retrieval is lexical TF-IDF, not neural embeddings.
- The hosted demo is English OCR only; the Docker image installs English language data.
- Built to a course brief. The Iris classifier is a stand-in with no connection to the documents.
Stack
- FastAPI
- Python
- Tesseract
- OpenCV
- scikit-learn
- Ollama
- Docker
- Jinja2
Demo
Hugging Face Space. It may take a minute to wake, and the Ollama-backed features are off in the hosted version.