Factual package intelligence from PyPI
Best open-source document to markdown extractor for LLM training data. Convert PDF, Word, PowerPoint, Excel, images, URLs to clean markdown, JSON, HTML locally. Alternative to Unstructured, Docling, Marker, MarkItDown, MinerU, PaddleOCR, Tesseract
pip install document-data-extractor
PyPI declares 29 unique dependency rules for this release. Environment markers are shown when supplied by the project.
>=9.0.0>=1.17.0>=0.11.6>=2.25.0>=4.9.0>=4.6.0>=0.8.11>=0.6.21>=1.11>=3.0.0>=1.3.0<2.0.0,>=1.21.0>=1.23.0>=0.1.0>=1.7.0>=0.23.0>=4.64.0<4.50.0,>=4.20.0 when python_version <= "3.8"<0.20.0,>=0.13.0 when python_version <= "3.8">=65.0.0 when python_version <= "3.8">=0.37.0 when python_version <= "3.8">=4.20.0 when python_version >= "3.9">=0.13.0 when python_version >= "3.9"The compact report shows 25 of 29 declarations. The interactive dependency graph loads the complete metadata.
document-data-extractor publishes 1 wheel and 1 source archive for version 1.0.4. Wheel platform tags: any.
Declared Python classifiers: 3.10, 3.11, 3.12, 3.8, 3.9.
PyPI lists 5 releases with files. The first dated release is ; 5 releases fall within the 365 days preceding the latest dated release. The current release files were uploaded on . The preceding dated release was .