Evaluates schema-guided structured extraction from documents: given a document and a JSON schema, systems must return a schema-valid JSON with page-and-box grounding. Covers 370 documents (4,869 pages) across 8 business domains and 67 document types; scores value accuracy, word/page grounding, and long-list completeness.
A small public sample of egocentric human demonstration video with synchronized 3D hand and body pose annotations for imitation learning and embodied-AI research. Delivered in Parquet and common multimodal packages (LeRobot, MCAP) for schema inspection before requesting gated access to larger EgoSuite releases.
Provides raw, unscripted first-person household video footage for training vision and embodied AI models. Released incrementally on Hugging Face in WebDataset shards with metadata parquets under Apache‑2.0; current raw tier contains ~7,834 hours (≈397k videos).
Turns adapter placement for PEFT on YOLO-family real-time detectors into an auditable constraint-planning problem that emits budgeted target-module plans or calibrated refusals; shows planner-selected RS-LoRA improves mAP and cuts peak training memory in evaluated detectors.
A 29.6B-parameter multimodal causal language model with a dedicated ViT-G/14 perception encoder for running agentic, tool-using, multimodal reasoning locally on consumer hardware. Offers 4-bit quantized weights and a DFlash drafter for speculative decoding to reduce memory and speed up generation.
Runs a quantized, locally executable 29.6B multimodal causal language model optimized for agentic workflows. Includes a perception encoder for image+text input, 4-bit quantized weights for 24–32GB devices, a DFlash drafter for speculative decoding, and robust tool-call support.
Multimodal vision-language model optimized for on-device image+text tasks: image captioning, full-page OCR with layout annotation, grounding/bounding-box prediction, and function calling. Built on the LFM2.5-2.6B backbone with a SigLIP2 NaFlex 400M vision encoder and tuned for low-latency, low-memory edge inference.
Post-training distribution-level objective that augments static Fréchet-distance losses with an adversarially learned representation and a real-feature whitening step to stabilize min–max optimization and avoid trivial feature amplification; targets one-step image generator post-training.
Builds an editable, persistent 3D world state to drive iterative previsualization for film, games, and design — enabling local edits and recombinations instead of one-shot video regeneration. Uses separate stages for state construction, state evolution, and state access, with render-feedback camera refinement.
Externalizes persistent scene state into a camera-indexed world bank and designs a long-horizon teacher whose sparse-attention supervision is distilled into a three-step student, enabling responsive, low-latency interactive long-horizon video generation with bounded denoiser context.
Provides an FP8-post-trained 27B multimodal causal language model with a native vision encoder, large-context support (262,144 native, extensible to 1,000,000), controllable thinking-mode reasoning, and compatibility with common inference engines for deployment.
Predicts future video frames conditioned on an observed frame, a language instruction, and a sequence of end-effector poses and gripper states for robot manipulation. Uses per-arm SE(3) geometric encoding (PRoPE-style), a lightweight depth branch, SAM3 masks with a frozen V-JEPA teacher, and distribution-matching distillation for efficient, consistent action-conditioned rollouts.