DocuStruct AI
PDF-to-Validated-Data Platform

portfolio

Project Overview

Python

FastAPI

PostgreSQL

SQLAlchemy

Tesseract OCR

Docker

Built a document AI pipeline that turns invoice PDFs into validated structured data. Extracts text with PyMuPDF and Tesseract OCR fallback, validates fields against JSON Schema, scores confidence per field, and routes low-confidence results to a human-review queue before JSON/CSV export. Persists the full document lifecycle and corrections in PostgreSQL.

Contact Me

© 2026 Dileep Reddy Battu. All rights reserved.