AIAny
Icon for item

Does Learning Protein Folding Generalize to Broader Reasoning?

Turns solved protein structures into FoldingCorpus and Fold2Reason — a post-training recipe that supervises an LLM with discrete structural Q&A plus continuous 3D geometry to improve spatial, graph and scientific reasoning; reports consistent gains across 10 benchmarks.

Introduction

Why this matters now Learning from text alone often teaches surface answers rather than spatial or topological logic. Protein structures are unusual: one solved coordinate set yields thousands of precisely checkable spatial facts. This work asks whether that non-linguistic, structure‑dense supervision can teach general reasoning behaviors in large language models and demonstrates a concrete post‑training recipe that produces measurable transfer.

Key Findings
  • Convert structures into training signals: FoldingCorpus provides 14,400 QA records derived from ~1,200 proteins (1,000 train / 100 dev / 100 test) via 12 deterministic structural operators. A frozen coordinate/distogram decoder supplies continuous geometry supervision during post‑training.
  • Fold2Reason recipe: post-trains a base LLM (Qwen3.5-9B in experiments) using the discrete QA targets through the model language head and a shared geometry constraint via a frozen decoder and LoRA‑adapted representations.
  • Quantitative effects: structure prediction scores on FoldBench rise by 2.7–3.5× versus baseline Qwen3.5-9B; on a 10-dataset general reasoning suite (General-10) macro accuracy increases from 45.09% to 48.33% (+3.23 pp), with positive mean changes on all datasets.
  • Ablations and controls: most transfer comes from FoldingCorpus discrete supervision; geometry adds modest, concentrated gains in spatial and local structural readouts. Random/synthetic/shuffled-structure controls produce much smaller or negative effects.
Who it's for & tradeoffs

Great fit if you care about whether non‑text scientific data can serve as practical post‑training supervision to improve spatial/graph/scientific reasoning in LLMs. It is especially relevant for researchers exploring multimodal fine‑tuning, dataset engineering for reasoning transfer, and LLM adapters (LoRA) workflows. Look elsewhere if you need an out‑of‑the‑box production model upgrade for unrelated tasks (the method requires additional post‑training infrastructure, protein-derived data engineering, and careful control experiments). Gains are model‑ and task‑dependent rather than universally large.

Where it fits

This paper sits between dataset engineering and model post‑training studies: it reframes a solved scientific prediction task (protein folding) as a reusable supervision source for general reasoning, complementing work on protein LMs, geometric learning, and post‑training recipes for LLMs.

Brief methodology note

The pipeline creates deterministic structural operators from solved coordinates to generate QA supervision, routes discrete answers to the LLM language head, and concurrently constrains shared representations with a frozen geometry decoder during training; downstream evaluation uses only the adapted language model, removing protein inputs and the geometry decoder to test behavioral transfer.

Information

  • Websitearxiv.org
  • OrganizationsShanghai Jiao Tong University, Fudan University, Shanghai Innovation Institute, Northeastern University
  • AuthorsYong Liu, Zhanpeng Shi, Yizhou Dang, Zhongyue Zhang, Xiaoliang Shi, Zhijian Wei, Shuangjia Zheng
  • Published date2026/09/30

More Items

Hugging Face

A 180,000-row labeled dataset of agent decision steps for training fast decision models that choose whether to call a tool, which tool to pick, and whether arguments are complete. Provides leakage-safe splits, typed questions, canonical state/request fields, and ready-to-use Parquet format for low-latency routers.

Hugging Face

Provides a cleaned, section-chunked bilingual (Arabic + English) Wikipedia-derived corpus plus a curated Egyptian-history subset for LLM pretraining and SFT. Includes large pretrain and finetune splits, article-level eval holdouts, Parquet format, CC-BY-SA-4.0, and Arabic orthography caveats.

Hugging Face

Provides a sanitized, labeled SOC capture and merged provenance graph for intrusion-detection research, including 2,011,674 live signals, 51,371 incident graphs, MITRE ATT&CK mappings, and deterministic attack reports.