A 20B-parameter MMDiT diffusion model that generates and edits images with accurate embedded text, including dense Chinese and English typography. Handles complex multi-line layouts and identity-preserving edits while keeping text legible.
Self-supervised vision foundation model producing dense, patch-level features that transfer to classification, segmentation, depth, and detection with a frozen backbone. Spans ViT-S (21M) to ViT-7B (6.7B params), plus ConvNeXt and satellite variants.
Generate parametric 3D CAD models from natural language and images in the browser, with real-time preview and exports to STL/SCAD. Runs client-side via OpenSCAD WebAssembly, extracts adjustable parameters, and integrates Anthropic for shape generation—suited for rapid prototyping and 3D-printing workflows.
Generates explorable, 3D-consistent virtual worlds from a single image or short video. Includes official implementations of Lyra‑1 (feed‑forward 3D/4D scene generation via video-diffusion self-distillation) and Lyra‑2 (long-horizon, explorable generative 3D worlds). Best for research and creative prototyping; requires substantial GPU compute.
Provides a curated collection of hands-on tutorials, workflows and auxiliary files for training and using generative-model tooling (Stable Diffusion, Flux, WAN). Key items include a WAN 2.1 LoRA training tutorial and an articles collection covering DreamBooth, LoRA, LyCORIS and SDXL.
Converts images and PDFs into structured Markdown, HTML, or JSON while preserving layout, handling tables, math, handwriting, charts, and chemistry diagrams across 90+ languages. Runs locally via HuggingFace or against a vLLM server.
Automatically generates complete short-form videos from a single topic: drafts script with an LLM, produces AI images/video, synthesizes multilingual TTS (including voice cloning), adds background music, and composes the final video. Supports local ComfyUI/RunningHub or direct model APIs and customizable templates.
Turn plain-English requests into editable draw.io diagrams: the model writes the underlying draw.io XML, which renders live in an embedded canvas. Upload images, PDFs, or text to replicate, refine through chat, and roll back via version history.
Generates real-time, infinite-length portrait video from one reference image on a 12GB GPU. Combines implicit facial signals and 3D keypoints with step-distilled diffusion and autoregressive micro-chunk streaming for low-latency live use.
Converts images (and other conditions) into high-fidelity, fully textured 3D assets using a 4B-parameter generative model and a field‑free sparse voxel format (O‑Voxel). Handles arbitrary topology, PBR materials, and near real-time mesh/voxel conversions; requires Linux and an NVIDIA GPU with >=24GB memory.
Generates natively editable PPTX from PDFs, DOCX, URLs, or Markdown — producing real PowerPoint shapes, text boxes, and charts (not images). An open-source, model-agnostic, local-first pipeline that integrates with multiple AI editors while keeping your data on-device.
Detects and surgically removes Google's SynthID watermark from images using multi-resolution spectral analysis and a resolution-aware codebook; provides a V3 bypass with high PSNR and strong phase-coherence reduction. Research-focused and intended for analysis/defense, not misrepresentation.