Provides live Codex-CLI agent run traces from GPT-5.6 Sol capturing coding, debugging, security reviews, and harness/seed workflows in cumulative next-action prefixes — suitable for supervised fine-tuning and analysis of tool-using coding agents.
Provides layered code pretraining corpora (L2 ~400B tokens, L3 ~150B tokens) across 11 languages by filtering ~192M public GitHub repositories into standardized files, algorithmically relevant selections, and implementation-grounded programming exercises. Includes per-file metadata (role, algo relevance, quality) and serialized task records; released under Apache-2.0.
Runs locally on constrained devices to turn text into guaranteed-parsable JSON tool calls, typed structured extractions, or sentence embeddings. Delivered as a single compact weights file (8–29 MB) with a laddered 2–20-layer design, low-bit quantisation and calibrated confidence scores for on-device apps.
Provides a programming model and distributed runtime that preserves parent–child lineage and ordered variable-cardinality expansions for 1→M→1 dataflow pipelines, enabling completion-driven cross-input GPU batching and ordered gathers for foundation-model data preparation. Demonstrates multi-GPU scaling speedups and lower end-to-end time versus Ray Data and Daft.