High-resolution image and video generation codebase and models that run with far lower compute and memory than typical diffusion systems. Uses linear-attention DiT variants, aggressive latent compression, and inference-scaling to support text-to-image (up to 4K), fast one/few-step generation, and efficient video pipelines.
Generates high-quality, editable 3D assets from text or images and decodes to radiance fields, 3D Gaussians, or textured meshes. Ships pretrained models up to 2B parameters, a 500K asset dataset and training code; best used with image conditioning and a ≥16GB NVIDIA GPU.
Optimizes and tests AI prompts in the browser, comparing original and rewritten versions side by side against any connected model. Runs fully client-side—keys go straight to the provider—and ships as web app, Chrome extension, and desktop builds.
Provides curated ComfyUI workflow templates and subgraph blueprints that package reusable node graphs, preview assets, and publishing pipelines for image/video generation. Includes a browsable Astro site with i18n, CI-driven sync/publish scripts, and PyPI packaging for easy distribution.
Real-time DETR detector on a DINOv2 backbone, covering detection, segmentation, and keypoints. Ships in six sizes (Nano to 2XL), beats YOLO on the COCO speed-accuracy curve, and transfers better to non-COCO real-world domains.
Collects 40+ importable n8n workflows from the AI Agents A-Z YouTube channel, each tied to one video episode — spanning content generation, social-media posting, and short-video and narrated-story pipelines, plus companion Docker MCP/REST servers.
Runs open-source LLMs and multimodal models entirely on mobile devices for offline, private inference. Offers Agent Skills, Thinking Mode, Ask Image, audio scribe, model management and benchmarks, with Gemma 4 and Hugging Face integration.
Real-time 3D Gaussian Splatting renderer for web apps using THREE.js. Integrates splat and mesh rendering with a Rust + Wasm component, supports major splat formats (.PLY, .SPZ, .SOG) and targets broad WebGL2 support for mobile-friendly dynamic scenes.
GPU-accelerated, non-destructive RAW photo editor designed for fast, low-footprint desktop workflows (<20MB). Built with Rust/WGPU/Tauri, it offers real-time 32-bit GPU processing, AI-assisted masking and optional ComfyUI integration for generative edits, plus presets and batch export.
Detects, segments, and tracks every instance of an open-vocabulary concept in images and video from a text phrase or visual exemplar, not just one object per prompt. An 848M-param model reaching ~75-80% of human accuracy across 270K concepts.