Browser-based, client-side video editor for multi-track editing, GPU-accelerated preview and local exports without uploading files; leverages WebCodecs/WebGPU and includes an AI upscaling option.
Filtered subset of the OPUS 4.6 parallel corpus that isolates reasoning-related translation examples and removes 979 refusals, providing a cleaner 3,000×-filtered dataset for training or evaluating NLP models focused on reasoning in translation.
Browser-based visual editor and learning hub for RDF/OWL ontologies (targeted at Microsoft Fabric IQ): interactive graph exploration, a searchable catalogue, an embeddable viewer, RDF/XML import/export, and a natural-language→ontology preview — all as a zero-backend static site.
Provides a 1,000-row sample user–item interaction Parquet for the TAAC2026 recommendation task, using a flat column layout with 120 top-level columns (IDs, labels, user/item int & dense features, and four-domain behavioral sequences). Updated 2026-04-10.
Curated 100K subset of geometrically diverse CAD construction sequences sampled from a 1M agentically synthesized corpus — each item includes executable CadQuery scripts, 8 rendered views, STL/STEP exports, and precomputed DINOv3 embeddings for retrieval and benchmarking.
Provides paired before/after satellite images with question–answer annotations for semantic change understanding. Includes Yes/No and multiple-choice formats, delivered in Hugging Face datasets (streaming-friendly), suited for remote-sensing multimodal VQA and semantic change captioning research.
Clinical question-answering model for psychological support in obesity weight-management. Integrates UK Biobank population evidence to produce clinically interpretable, stigma-aware responses that help clinicians identify distress, prompt screening, and suggest appropriate referrals.
An ~18B frankenmerge text-generation model that stacks two 32-layer Qwen3.5-based finetunes and ships as a 9.2GB Q4_K_M GGUF for efficient local inference. A 1000-step QLoRA heal reduces layer-boundary code corruption and targets coding, reasoning, multilingual chat, and 12–16GB GPU compatibility.
Generates English text matching pre-1931 style — a 13B language model trained on ~260B tokens of pre-1931 English, useful for historical-language generation and stylistic research. An instruction-tuned variant exists for interactive tasks.
Instruction-tuned 13B LLM post-trained on 260B tokens of pre-1931 English and finetuned with online DPO (LLM-as-judge) to improve instruction-following; suited for period-style English generation and etiquette/letter-writing formats, but not optimized for contemporary factual updates.
Benchmark dataset for evaluating clinician-facing chat assistants: physician-authored conversations plus rubric items, use-case and difficulty labels, specialty metadata, and a built-in canary to reduce benchmark contamination. Hosted on Hugging Face under an MIT license.