Factual package intelligence from PyPI

document-data-extractor 1.0.4 package intelligence

Best open-source document to markdown extractor for LLM training data. Convert PDF, Word, PowerPoint, Excel, images, URLs to clean markdown, JSON, HTML locally. Alternative to Unstructured, Docling, Marker, MarkItDown, MinerU, PaddleOCR, Tesseract

pip install document-data-extractor

Current version
1.0.4
Python requirement
>=3.8
License
MIT
Distribution type
Pure Python wheel
Release files
2 (1 wheels, 1 source)
Download size
385.1 KiB
Release history
5 releases with files
Median release cadence
0 days

document-data-extractor dependencies

PyPI declares 29 unique dependency rules for this release. Environment markers are shown when supplied by the project.

The compact report shows 25 of 29 declarations. The interactive dependency graph loads the complete metadata.

Compatibility and release files

document-data-extractor publishes 1 wheel and 1 source archive for version 1.0.4. Wheel platform tags: any.

Declared Python classifiers: 3.10, 3.11, 3.12, 3.8, 3.9.

Release activity and sources

PyPI lists 5 releases with files. The first dated release is ; 5 releases fall within the 365 days preceding the latest dated release. The current release files were uploaded on . The preceding dated release was .