DocuStruct AI
PDF-to-Validated-Data Platform

Built a document AI pipeline that turns invoice PDFs into validated structured data. Extracts text with PyMuPDF and Tesseract OCR fallback, validates fields against JSON Schema, scores confidence per field, and routes low-confidence results to a human-review queue before JSON/CSV export. Persists the full document lifecycle and corrections in PostgreSQL.