AIAny
Icon for item

cwc-workshops

Collection of hands-on workshop materials and sample code from Anthropic's "Code with Claude" series, covering Claude Managed Agents, memory (Dreaming Service), eval-driven agent development, and multi-agent patterns. Not maintained and not accepting contributions.

Introduction

Most agent tutorials show a handful of examples; this repo collects full, runnable workshop code that walks from an amnesiac baseline to production-style managed agents and eval-driven iteration. The value is concrete recipes and graded evals you can run to see how prompt, memory, and orchestration changes affect real outputs.

What Sets It Apart
  • Practical workshop artifacts: complete sample projects (Streamlit incident dashboard, PPTX-generating agent, Deal Desk multi-agent demo, agent-battle game harness) rather than abstract diagrams, so you can inspect prompt variants, skills, and evaluation harnesses.
  • Focus on agent engineering primitives: memory stores and Dreaming Service for cross-session consolidation, Skills + callable_agents decomposition, and MCP-based integrations for streaming and tool gating — each taught with incremental exercises and programmatic evals.
  • Eval-driven workflow examples: an explicit grader setup (programmatic metrics + LLM-as-judge) to iterate agent designs against a 10-task suite, showing measurable deltas instead of ad-hoc changes.
Who it's for and trade-offs

Great fit if you want hands-on examples for building and evaluating LLM-based agents tied to the Anthropic/Claude ecosystem, or if you need concrete patterns for memory, multi-agent coordination, and grading outputs. Look elsewhere if you need a maintained, production-grade SDK or vendor-agnostic libraries: examples assume Anthropic platform concepts and keys, some demos call platform services, and the repo is marked not maintained and not accepting contributions.

Where it fits

Serves as a pragmatic companion to API docs and SDKs: use it to prototype agent architectures, test eval-led prompt changes, or study memory/dreaming patterns. It complements official SDK samples but is not a drop-in production framework — treat it as workshop material and learning code rather than a maintained product.

Information

  • Websitegithub.com
  • OrganizationsAnthropic PBC
  • Authorsjeffcurry-ant, michael-cohen-io, mroknich, felixfbecker, rodrigo-olivares, TanveerMittal
  • Published date2026/05/06

More Items

Hugging Face

Provides 5.5K+ self-contained data-analysis RL tasks: each row bundles a real tabular dataset, a question, and a deterministically-gradable gold answer. Verified from jupyter-agent notebooks; splits for training, held-out testing, and quick eval; intended for prompting, fine-tuning, and agent RL.

Hugging Face
AI Model2026

A 9B agentic multimodal SFT checkpoint distilled from Qwen3.5-9B for coding, general agent tasks, visual coding and cybersecurity. Provided by Xiaomi MiMo as a research seed (77.4B-token SFT mix) to bootstrap agentic RL and tool-use experiments.

Hugging Face
AI Model2026

Preview agentic language model for research and engineering workflows that turns research questions into executable, verifiable workflows via tool use and long-context reasoning; built on a 744B-parameter MoE (GLM-5.2) with MIT-licensed BF16 and FP8 checkpoints.