Agents & Tool Use

Function calling, ReAct loops, MCP, agent memory architectures and evaluation harnesses.

15concepts
169flashcards
117minutes of reading
  1. 01 Agent Evaluation Harnesses Single-output accuracy says nothing about an agent that takes thirty steps; evaluating agents means scoring trajectories, environment state, and reliability across runs. advanced 10m 5 cards
  2. 02 Agent Memory Architectures An agent whose only memory is its context window is amnesiac between sessions; persistent memory is the architecture that decides what to keep, where, and how to retrieve it. advanced 10m 5 cards
  3. 03 Agentic AI and ReAct From single tool calls to multi-step agents that plan, act, observe, and recover from errors. advanced 9m 6 cards
  4. 04 Code Execution as a Tool Interface Instead of calling tools one at a time through the context window, the agent writes code against tool APIs in a sandbox, which cuts both tool-definition overhead and intermediate results out of the token budget. advanced 7m 18 cards
  5. 05 Computer-Use Agents: Operating a GUI Through Pixels How agents that click and type on a real desktop differ from tool-calling agents, why GUI grounding is the bottleneck, and what OSWorld measured that API benchmarks cannot. advanced 8m 15 cards
  6. 06 Durable Agent Execution and Recovery Long-running agents fail mid-task for mundane reasons, so the loop needs checkpointed state, idempotent side effects, and the ability to resume from a step rather than restart from the prompt. advanced 6m 15 cards
  7. 07 Long-Horizon Agent Reliability Why per-step accuracy compounds into task failure, how METR's time-horizon metric reframes agent capability, and which architectural moves actually raise the exponent. advanced 8m 15 cards
  8. 08 Orchestrator-Worker Subagent Architectures A lead agent decomposes a task and spawns subagents with clean context windows that explore in parallel and return compressed summaries, buying breadth and context isolation at a large token cost and a coordination risk. advanced 7m 15 cards
  9. 09 Sandboxing and Least Privilege for Agents Why agent security has to be enforced outside the model, how capability scoping and human-in-the-loop gates work, and what the CaMeL design proves about the limits of prompting. advanced 8m 15 cards