Installation-oriented dataset that packages ComfyUI-ready files and instructions for running MiniMax H3 locally — includes pruned/INT8/BF16 checkpoints, matching Qwen3-VL text encoders, video/audio VAEs, and official ComfyUI workflow templates for joint audio+video generation.
Human-annotated text dataset that labels perceived “AI slop” with a continuous human slop_score (-1 / 0 / +1) plus provenance metadata (source_dataset, source_row_id, content_hash). Collected via Bench Labs SlopFinder from public datasets for training classifiers and studying subjective perception.
Provides a human-verified benchmark of 1,927 heterogeneous articulated 3D objects with part-level articulation semantics and intrinsic physical-property annotations for evaluating physical grounding and simulation readiness. Includes URDF assemblies, aligned point clouds, per-part JSON annotations, and a curated evaluation protocol; licensed CC BY-NC 4.0 (non-commercial).
Snapshot delivery of arXiv metadata, submission files and rendered documents in multiple Parquet configs (metadata, paper_text, latex, source, pdf, ps). Includes ~3.15M papers, full-text TeX assemblies and indexes to fetch large assets for training, retrieval and analysis.
Transforms scientific code repositories into executable, agent-learnable environments that support task generation, execution, and scientific verification. Agent-guided repository transformation produces verified interaction trajectories used to train the PhAI-IDE model family. Intended for research on agent learning, scientific-code repair, and training RL/SFT models.
Trains vision, text, tabular, recommendation, and medical imaging models through a layered PyTorch API. High-level learners cover common workflows, while lower layers stay available when researchers need custom behavior.
Memory layer that lets AI agents remember users and context across sessions.