Converts PDFs into AI-ready structured outputs (Markdown, JSON with bounding boxes, HTML) for RAG and accessibility workflows; offers deterministic local parsing plus a hybrid AI mode for complex tables, OCR, formulas, and auto-tagging previews.
Extends RAG beyond text: parses PDFs and Office files containing images, tables, equations, and charts, then queries them through one multimodal knowledge graph. Built on LightRAG, it replaces separate parsing and retrieval tools.
A collection of ready-to-run Hugging Face Jobs OCR scripts that add a markdown column (or structured JSON) to image datasets, with model switching, layout detection, server-mode serving, and per-model options for table/form extraction.
Parses PDF resumes into structured JSON using LLMs, enriches profiles with GitHub signals, and outputs explainable category scores, evidence, bonuses and deductions. Runs fully local with Ollama or via Google Gemini; designed for reproducible, fairness-constrained resume scoring in hiring workflows.
Provides a reproducible, deduplicated corpus of text extracted from PDFs for LLM pretraining—about 3 trillion tokens from ~475 million documents in 1733 language-script pairs. Includes OCR and text extraction pipelines, per-page language IDs, MinHash deduplication, and is released under ODC‑By 1.0.
Converts images and PDFs into structured Markdown, HTML, or JSON while preserving layout, handling tables, math, handwriting, charts, and chemistry diagrams across 90+ languages. Runs locally via HuggingFace or against a vLLM server.
Generates summaries from URLs, YouTube videos, podcasts, PDFs, and local audio or video files. Backend-agnostic by design: the same pipeline drives local coding CLIs (Claude, Codex, Gemini) or hosted API providers (OpenAI, Google, xAI).
Fetches multi-source content (webpages, YouTube, PDFs, WeChat, paywalled articles, podcasts), uploads it to Google NotebookLM, and generates outputs such as podcasts, PPTs, mind maps, or quizzes. Differentiators: automatic paywall-bypass pipeline, Claude Code Skill integration, and CLI + MCP components for WeChat and document scraping.
Provides a physical reconstruction benchmark of OmniDocBench v1.5 by producing five real-world photographic variants (Scanning, Warping, Screen‑Photography, Illumination, Skew) for each of 1,355 pages, inheriting original ground-truth to enable controlled, scenario-wise evaluation of document parsing robustness.
Multimodal OCR and document-understanding toolkit for recognizing complex layouts, tables, formulas and code. Uses Multi-Token Prediction and stable RL for better training; ships as a 0.9B-parameter model with a Python SDK and deployment guides for vLLM, SGLang and Ollama.
Rust library for fast PDF classification and position-aware text extraction that converts native-text PDFs to structured Markdown without OCR. Offers per-page OCR routing, multi-column and table detection, and Python/Node.js/WebAssembly bindings for low-latency local pipelines.
A fast, local document parser that extracts spatial text with bounding boxes from PDFs and other formats. Bundles Tesseract OCR and supports HTTP OCR servers, multi-language bindings (Rust, Node, Python, WASM) and screenshot generation; best for lightweight local pipelines but less suited to very complex or heavily scanned documents.