Skip to content

Opening book details…

About this document

Legal PDF Metadata Extraction System by sanketsonar2002 is a document available to read on EtoBox.

The document outlines a Metadata Extraction System designed to automate the extraction of structured metadata from legal PDFs using a hybrid approach that combines SpaCy-based Named Entity Recognition and Regex. It details the technology stack, including Python, Flask, and Docker, and describes the model architecture with a custom NER model trained on legal text. The system outputs metadata in various formats, including JSON and HTML, for integration and user interface display.

Author
sanketsonar2002
Language
EN