About this document
Legal PDF Metadata Extraction System by sanketsonar2002 is a document available to read on EtoBox.
The document outlines a Metadata Extraction System designed to automate the extraction of structured metadata from legal PDFs using a hybrid approach that combines SpaCy-based Named Entity Recognition and Regex. It details the technology stack, including Python, Flask, and Docker, and describes the model architecture with a custom NER model trained on legal text. The system outputs metadata in various formats, including JSON and HTML, for integration and user interface display.
- Author
- sanketsonar2002
- Language
- EN