Why this matters now Learning from text alone often teaches surface answers rather than spatial or topological logic. Protein structures are unusual: one solved coordinate set yields thousands of precisely checkable spatial facts. This work asks whether that non-linguistic, structure‑dense supervision can teach general reasoning behaviors in large language models and demonstrates a concrete post‑training recipe that produces measurable transfer.
Key Findings
- Convert structures into training signals: FoldingCorpus provides 14,400 QA records derived from ~1,200 proteins (1,000 train / 100 dev / 100 test) via 12 deterministic structural operators. A frozen coordinate/distogram decoder supplies continuous geometry supervision during post‑training.
- Fold2Reason recipe: post-trains a base LLM (Qwen3.5-9B in experiments) using the discrete QA targets through the model language head and a shared geometry constraint via a frozen decoder and LoRA‑adapted representations.
- Quantitative effects: structure prediction scores on FoldBench rise by 2.7–3.5× versus baseline Qwen3.5-9B; on a 10-dataset general reasoning suite (General-10) macro accuracy increases from 45.09% to 48.33% (+3.23 pp), with positive mean changes on all datasets.
- Ablations and controls: most transfer comes from FoldingCorpus discrete supervision; geometry adds modest, concentrated gains in spatial and local structural readouts. Random/synthetic/shuffled-structure controls produce much smaller or negative effects.
Who it's for & tradeoffs
Great fit if you care about whether non‑text scientific data can serve as practical post‑training supervision to improve spatial/graph/scientific reasoning in LLMs. It is especially relevant for researchers exploring multimodal fine‑tuning, dataset engineering for reasoning transfer, and LLM adapters (LoRA) workflows. Look elsewhere if you need an out‑of‑the‑box production model upgrade for unrelated tasks (the method requires additional post‑training infrastructure, protein-derived data engineering, and careful control experiments). Gains are model‑ and task‑dependent rather than universally large.
Where it fits
This paper sits between dataset engineering and model post‑training studies: it reframes a solved scientific prediction task (protein folding) as a reusable supervision source for general reasoning, complementing work on protein LMs, geometric learning, and post‑training recipes for LLMs.
Brief methodology note
The pipeline creates deterministic structural operators from solved coordinates to generate QA supervision, routes discrete answers to the LLM language head, and concurrently constrains shared representations with a frozen geometry decoder during training; downstream evaluation uses only the adapted language model, removing protein inputs and the geometry decoder to test behavioral transfer.