Detects and redacts personally identifiable information (PII) in user-typed text on-device, replacing sensitive values with stable placeholders before any data leaves the browser. Uses a small quantized ONNX token-classification model plus deterministic recognizers for structured identifiers, and applies a policy-driven keep-set for coarse geography.
Introduces AISPA, a user-centric framework to audit system prompts in LLM applications, and applies it to 3,249 instructions from 88 commercial products to classify protective versus problematic instructions. Highlights design variability, growing prompt length/protection, persistent problematic directives, and calls for transparency and oversight.
Evaluates how large language models fabricate user attributes in personalization and whether model self-monitoring is a reliable signal. Introduces MirageBench (150 personas, 6 personalization tasks, judge-validated faithfulness taxonomy) and a 12-model leaderboard revealing pervasive over-inference and a 'Self-Monitoring Inversion'.
Hugging Face dataset for the MVA Hackathon 2026 containing pediatric rare-disease genomic data (~85 GB across 11 files). Access requires accepting dataset conditions; intended for genomic ML, variant analysis, and hackathon submissions, with notable storage and privacy constraints.
Provides aggregated, privacy-preserving cluster outputs from three external research teams' analyses of ~250k Claude/Claude Code conversations; includes per-team CSVs (Stanford, Oxford, METR) for studying human–AI collaboration and model behavior without raw conversations.
Provides 4.5 billion TikTok video records with captions, timestamps, music IDs and engagement counts for research; split across 27 zstd-compressed Parquet files (~289 GB) and sampled via TikTok's mobile API; released for research-use only with privacy and ToS caveats.
Compresses a 27B-class multimodal model into end-to-end ternary weights to run 27B reasoning on-device: 5.9–8.6 GB deployed footprint, 262K-token context, ~98.2% of FP16 benchmark performance; ships MLX and GGUF packs and runs on Apple MLX and CUDA.