05 / Document AI
- 500+
- Documents processed a day
- 98%
- Extraction accuracy
- 95%
- Classification accuracy
The problem
Document intake was manual keying: slow, and error-prone at volume.
The approach
- 01
An OCR pipeline on Tesseract, with OpenCV preprocessing for noise removal and deskew before any text is read.
- 02
A TensorFlow CNN that classifies incoming documents into 10+ types.
- 03
Async processing on Redis and Celery, handling 100+ documents concurrently.
- 04
A React drag-and-drop upload interface with real-time progress.
Stack
TensorFlow was my production ML stack here, around 2022. My recent work is LLM and agent systems rather than classical model training.