AIAny
AI Model2025
Icon for item

TRELLIS.2 — Native and Compact Structured Latents for 3D Generation

Converts images (and other conditions) into high-fidelity, fully textured 3D assets using a 4B-parameter generative model and a field‑free sparse voxel format (O‑Voxel). Handles arbitrary topology, PBR materials, and near real-time mesh/voxel conversions; requires Linux and an NVIDIA GPU with >=24GB memory.

Introduction

TRELLIS.2 tackles a practical bottleneck in 3D content creation: converting 2D inputs into photoreal, production-ready 3D assets without costly remeshing or manual retouch. Its core insight is replacing continuous iso-surface fields with a compact, field-free sparse voxel (O‑Voxel) and structured latents, which lets a single large diffusion-like pipeline represent complex topology, internal cavities, and full PBR textures in a compact, efficiently processed form.

What Sets It Apart
  • O‑Voxel + compact structured latents: maps textured meshes to a sparse, field-free voxel latent that preserves open surfaces, non-manifold geometry, and internal structures — so you avoid lossy conversions and retain artist-level topology.

  • Large-scale image→3D pipeline (4B params) with staged sparse VAEs and DiT-based flow models: this enables high-resolution outputs (up to 1536³ tested) with practical runtimes on H100/A100 — so you can get production-quality meshes and PBR textures in seconds to minutes rather than hours.

  • End-to-end tooling for inference and training with conversion utilities (O‑Voxel, CuMesh, FlexGEMM): provides both pretrained inference (Hugging Face model) and full training recipes, so teams can run inference out-of-the-box or scale training on Objaverse-XL style datasets.

Who It's For and Tradeoffs

Great fit if you need image-to-3D or shape-conditioned texture generation for games, VFX, or asset libraries and can provide high-end NVIDIA GPUs and Linux infrastructure. It reduces manual remeshing and texture baking work while producing PBR-ready GLB exports.

Look elsewhere if you require lightweight, real-time generation on consumer hardware (the 4B model and CUDA-accelerated toolchain expect >=24GB GPUs) or if you prefer surface-field representations (e.g., pure SDF/NeRF workflows) — TRELLIS.2 prioritizes topological fidelity and photoreal PBR over minimal runtime memory footprint.

Where It Fits

Positioned between research diffusion/DiT generative models and production asset pipelines: it targets teams that need scalable, high-quality 3D generation with training capability, not just small demo models. The packaged converters and CUDA-optimized utilities make it a practical bridge from large-model research to studio asset production.

Information

  • Websitegithub.com
  • OrganizationsMicrosoft
  • AuthorsJianfeng Xiang, Xiaoxue Chen, Sicheng Xu, Ruicheng Wang, Zelong Lv, Yu Deng, Hongyuan Zhu, Yue Dong, Hao Zhao, Nicholas Jing Yuan …
  • Published date2025/11/26

Categories

More Items

Hugging Face
AI Model2026

Post-trained multimodal Qwen3.8-27B variant that uses alternating SFT and RLOO to reduce pathological long reasoning tails; ships multiple quantization tiers (BF16, FP8, NVFP4, INT8, INT4, GGUF), supports MTP and DFlash2 speculative decoding, and includes detailed benchmark and runtime recommendations.

Hugging Face
AI Model2026

A 2-bit quantized GGUF of Qwen3.8-27B that fits under 8 GB and runs on stock llama.cpp while preserving function/tool-calling behavior; includes an optional small vision add-on and is tuned for agent/tool workflows (Apache-2.0).

Hugging Face
AI Image2026

Runs text-to-image generation and instruction-guided image editing in 8 denoising steps. An accelerated checkpoint of Qwen-Image-2.1 that preserves the same 7B visual generator, native RGBA support, Diffusers QwenImage21Pipeline compatibility, a saved 8-step sampling schedule (CFG=1), and prefix KV cache reuse for multi-reference editing.