Fetches multi-source content (webpages, YouTube, PDFs, WeChat, paywalled articles, podcasts), uploads it to Google NotebookLM, and generates outputs such as podcasts, PPTs, mind maps, or quizzes. Differentiators: automatic paywall-bypass pipeline, Claude Code Skill integration, and CLI + MCP components for WeChat and document scraping.
Provides a physical reconstruction benchmark of OmniDocBench v1.5 by producing five real-world photographic variants (Scanning, Warping, Screen‑Photography, Illumination, Skew) for each of 1,355 pages, inheriting original ground-truth to enable controlled, scenario-wise evaluation of document parsing robustness.
Paired brain MRI scans and radiology text annotations for multimodal vision–language research. Provides image-level labels and image–text pairs suited for VQA, classification, and image-to-text tasks; CC BY-NC-SA 4.0 and ~10K–100K samples — research/non-commercial use.
Runs local AI models on Apple Silicon as an OpenAI‑compatible server, emphasizing low latency, prompt caching, and reliable tool-calling. Optimized for M1–M4 Macs with multimodal support and drop‑in compatibility for IDEs and agent frameworks.
Converts scene intent into production-ready Seedance 2.0 prompts, reference-role mappings, and IP-safe rewrites for multimodal (text/image/audio/video) video generation. Ships as a modular agent-skill OS with multilingual examples, troubleshooting tools, and pro filmmaker handoff artifacts.
An instruction‑tuned Gemma 4 E4B multimodal model on Hugging Face that accepts text, images and audio and generates text; notable for 128K long context support, built-in thinking mode, and an on‑device‑friendly E4B architecture under an Apache‑2.0 license.
Community fine-tuned multimodal Qwen3.5-9B using Claude 4.6 distilled data to change the model's 'thinking' behavior; offers an uncensored 'heretic' flavor with image-text-to-text I/O, benchmark comparisons, and deployment notes for inference frameworks.
Provides an annotated multimodal human-motion dataset for language-to-action and robotics research, with BVH and MuJoCo files plus recordings targeted at Unitree-G1 and NVIDIA-SOMA platforms. Covers locomotion, gestures, dance and object interaction with English annotations and 100K–1M samples.
Generate text, images, video, audio and action/robot trajectories from combined text, image, video, audio and action inputs. A Mixture-of-Transformers omnimodal foundation model (Cosmos3‑Nano, 16B params) focused on Physical AI (robotics, AV, simulation) and optimized for NVIDIA GPU runtimes.
Provides 10 million synchronized egocentric experience episodes with structured 3D/4D multimodal annotations — 2.88B RGB frames, 720M depth frames, 576M pose/mocap frames and ~1PB total. Designed for embodied AI, robotics, and multimodal pretraining; research-only, gated access.
Instruction-tuned Mixture-of-Experts multimodal model that generates text from text+image inputs while activating a 4B subset of parameters for faster inference; supports a 256K context window, multilingual vision-language tasks, and is available under Apache-2.0.
Large-scale mid-training corpora for multimodal models: 10,809 ~60s video shards, caption splits (30s/60s/180s/>10min), 84 spatial-reasoning shards, and CSV mappings to source YouTube IDs. Small Parquet preview configs are provided for schema inspection.