AIAny
Icon for item

Foundations of Large Language Models

A concise textbook-style book that explains foundational concepts and techniques for large language models, covering pre-training, generative models, prompting, alignment, inference, and reasoning. Structured as self-contained chapters for readers with some ML/NLP background or those seeking a principled introduction to LLM foundations.

Introduction

The rapid rise of large language models shifted NLP from task-specific, supervised pipelines to a pretrain-then-adapt paradigm. This book distills that methodological shift into a compact, chaptered treatment that emphasizes the core ideas you need to understand why current LLMs work and how key components interact.

Key Findings
  • Pre-training foundations: explains common self-supervised objectives and why scale and data diversity drive emergent capabilities, so readers can judge trade-offs between objectives and corpus design.
  • Generative model mechanics: walks through decoder/encoder-decoder choices, scaling laws, and strategies to handle long contexts, so you understand architectural and training implications for generation quality.
  • Prompting and instruction methods: surveys prompting families (static prompts, chain-of-thought, automatic prompt design) and their practical limits, so readers can pick prompting approaches aligned with task constraints.
  • Alignment and fine-tuning: covers instruction tuning and human-feedback alignment techniques (including SFT/RLHF-style ideas), clarifying what alignment achieves and where it fails.
  • Inference and reasoning: discusses runtime inference strategies and reasoning techniques, highlighting how model capabilities interact with decoding and chain-of-thought approaches.
Who it's for and trade-offs

Great fit if you want a principled, compact reference that explains why common LLM practices (pretraining objectives, architecture choices, prompting, alignment) work and how they connect. It is accessible without deep prior specialization, but assumes basic ML/NLP and Transformer familiarity. Look elsewhere if you need exhaustive state-of-the-art survey papers, implementation recipes, or hands-on tutorials with code—this work emphasizes foundations and conceptual clarity over step-by-step engineering.

Where it fits

Positions itself between a classroom textbook and a curated set of research notes: useful for students, researchers new to LLMs, and engineers who need conceptual grounding before diving into implementation or benchmarking.

Information

  • Websitearxiv.org
  • OrganizationsNLP Lab, Northeastern University, NiuTrans Research
  • AuthorsTong Xiao, Jingbo Zhu
  • Published date2025/01/16

More Items

Serves token-level routed LLM inference by dispatching requests to per-model asynchronous subservers and using delayed-batching scheduling to reduce admission latency and step desynchronization. Exposes a request-centric route-send-receive API and reports 2.01–64.15× decoding throughput gains versus single-LLM servers.

Converts historical interaction traces into a reusable, queryable “worldbook” and runs a language-based world model agent (Trace2Env) as the environment for LLM agents — enabling stateful, grounded simulation with improved next-observation fidelity and long-horizon consistency.

Analyzes why self-evolving reasoning models collapse under repeated self-training and proposes R-Quest: a feedback-driven pipeline that trains solvers to reject invalid questions and uses a frozen base model to detect task-level repetition, filtering training data to sustain multi-round gains.