Discover the Best AI Resources
Curated essentials, no noise — just what matters
Provides GGUF-quantized builds of the Qwythos-9B-v2 LLM for local runtimes, with multiple quant levels, optional MTP-enabled variants, a 1,048,576-token context window, and an optional BF16 vision projector for multimodal use.
A 9B-parameter Qwen3.5-based multimodal model tuned to preserve chain-of-thought reasoning while eliminating repetition loops; restores native multi-token prediction, supports 1,048,576-token context, and targets research/red-team use.
Uses large-scale text-to-video generative pretraining to create GenCeption, a feed-forward perception model that performs diverse vision tasks from text instructions—depth, surface normals, camera pose, referring segmentation, and 3D keypoints—often matching or surpassing specialized models while requiring far less task-specific data.
Explores unsupervised visual pretraining on visually rich documents to improve language-model intelligence; shows visual-pretrained models outperform text-only counterparts on the same corpora. Key aspects: direct use of images/layouts (no OCR-only pipeline), scalable across backbones and benchmarks.
Reconstructs 4D dynamic human scenes from sparse, low-overlap multi-camera captures by decoupling background synthesis and human modeling. Synthesizes hundreds of camera-controlled background views with a video diffusion model, initializes deformable Gaussian humans via cross-view identity and triangulated keypoints, then applies motion-adaptive recursive enhancement to reduce artifacts.
Generates minute-scale, temporally coherent dance videos from full music tracks using a hierarchical two-stage approach: global keyframe planning plus local temporal refinement; suitable when long-range musical structure and rhythmic continuity matter.
Proposes Riemannian Isometric Policy Optimization (RIPO) to fix exploration collapse in PPO-style RL for LLMs by aligning policy updates with the policy manifold's Riemannian geometry, improving exploration–exploitation balance and optimization stability across competition benchmarks.
Generates a new camera viewpoint from a reference video: an IC‑LoRA adapter for LTX‑Video 2.3 that re‑renders the same scene from a requested discrete camera angle while preserving subject and content. Trained on synthetic multi‑view data, proof‑of‑concept with limited viewpoint range and best for small, chained angle shifts.
Provides a deliberative Agent OS layer for robots that handles scene-conditioned planning, context-isolated skill execution, multi-stage verification, persistent multi-modal graph memory, and edge–cloud collaboration. Introduces EmbodiedWorldBench (16 scenes, 200+ tasks) and a failure-driven self-evolution loop; shows improved task success and strong memory benchmark scores.
Unifies high-level visual-language reasoning and low-level control for visual navigation by decoupling cognition and control: a slow vision-language reasoner produces pixel goals with explicit chain-of-thought, and a fast action expert converts those anchors into continuous waypoints for robust urban and indoor navigation.
A GGUF-format Qwen3.6 35B base model image-text-to-text release repaired via tensor-level SVD/scale correction and packaged with Hermes agent tweaks; multimodal (vision + text), MoE architecture, ready for GGUF runtimes like llama.cpp.
A GGUF-distributed Qwen3.6 35B MoE model variant repaired with a