RAG, Memory & Context

Continual Learning and Scaling

Last updated 2026-08-13

Key points

What it is

  • Continual learning is a way for AI to keep improving from new experiences without forgetting old knowledge, like a human learning from mistakes.
  • It's about adapting and compressing experiences into reusable structures for future behavior, avoiding "catastrophic forgetting" (losing old abilities when learning new ones).
  • It's not the same as model fine-tuning (adjusting weights after training) or just scaling up model size, which only creates a "world's smartest novice" (a system that can attempt any problem but never accumulates skill).
  • The goal is to turn failures into replayable tests, apply updates, and prove the fixes help without breaking past performance.

How to use it

  • Capture feedback when your AI agent fails, turn that failure into a replayable learning environment (a test you can rerun), and practice it.
  • Apply an optimization step, like calling a function called "Rely Optimize," to adjust the agent based on that feedback, without creating regression (breaking past performance).
  • Test each update to prove it helps the failure and breaks nothing that already worked, then repeat lifelong for compounding gains.
  • Use a "sandwich" approach: try harness engineering (improving the system around the AI) first, fine-tune to break through a ceiling, then do more harness engineering.

Watch out for

  • Avoid the "sunk cost fallacy" (assuming continual learning must layer on top of frozen checkpoints because we trained models a certain way).
  • Don't just look at total reward (a single measure of performance) alone, as it confounds continual learning ability with base model strength.
  • Measure reward gain (how much better you did on a task versus what your base model could do initially), cost, and ability to retain prior information without forgetting.
  • Treat continual learning as a first-order requirement (a key need from the start) and design for it, not just an add-on.

Tools named

  • Rely Optimize (a function for adjusting an AI agent based on feedback), Pareto frontiers (trade-off curves for measuring performance)

Lesson 1: What is Continual Learning and Scaling and why it matters

Continual learning (a model improving from new experience without forgetting old knowledge) is the bridge from raw intelligence to true expertise. Without it, scaling a model only creates what Yu Su calls "the world's smartest novice"—a system that can attempt any problem but never accumulates skill, brute-forcing its way through every task. Parth Asawa defines continual learning as "sample efficient online learning that is stable over long horizons," meaning the model must retain prior information while adapting to new data.

For AI agents, this works like human learning: agents operate in environments, produce trace data (records of actions taken), and then integrate that feedback back into their state to improve future behavior. Soheil Feizi emphasizes that this improvement must be "verifiable"—each update is tested, gains are measured, and nothing that already works breaks, avoiding regression (a drop in performance on old tasks).

Why does this matter for scaling? As Ronak Malde explains, we've exhausted the gains from pre-training (initial training on massive internet data) and benchmark scaling—the real unlock is continual learning. Without it, you face catastrophic forgetting (losing old abilities when learning new ones), a problem researchers have studied for decades. The practical insight is that continual learning isn't just model fine-tuning; useful updates can happen in the harness and memory layer. As Vivek Trivedy notes, agents produce more data than ever before, so managing that data at scale and integrating it back into agent state is the core challenge. This is why continual learning matters: it's how AI moves from smart to experienced, without breaking what already works.

Sources

Lesson 2: How to use Continual Learning and Scaling: step-by-step

Continual learning (a system that keeps improving from new experience) starts with a simple definition: it is adaptive compression of experience into reusable structures for future behavior. Without it, scaling your model only gives you the "world's smartest novice" — super smart, but it brute-forces every new problem without accumulating expertise.

To use continual learning step by step, begin by capturing feedback. When your agent fails, lift that failure into a replayable learning environment (a test you can rerun). This turns the failure into a task you can practice. Next, apply an optimization step, like calling a function called "Rely Optimize," to adjust the agent based on that feedback. The key is to do this without creating regression (breaking past performance). Test each update to prove it helps the failure and breaks nothing that already worked, then repeat lifelong for compounding gains.

For evaluation, don't just look at total reward. Total reward alone confounds continual learning ability with base model strength. Instead, measure reward gain — how much better you did on the fifth task versus what your base model could do initially — plus cost and ability to retain prior information without forgetting. These metrics sit on Pareto frontiers (trade-off curves), so no single number defines success.

In practice, language models today are trained as frozen checkpoints and were never designed to be continual learners. One hypothesis is we're stuck in a sunk cost fallacy: because we trained models the current way, we force continual learning methods to work on top. If designed from first principles, continual learning might be one single phase of training, not a patch.

A promising direction is on-policy distillation, an algorithm (a step-by-step procedure) that keeps reinforcement learning's power while adding task distribution, sometimes surpassing RL itself.

Sources

Lesson 3: Best practices and pitfalls

Continual learning (updating an AI from new experience) is not the same as model fine-tuning (adjusting weights after training). Many useful updates happen in the harness and memory layer—the surrounding system that captures and replays tasks. If you only scale raw model size, you get the "world's smartest novice": brilliant but unable to accumulate expertise, brute-forcing every problem.

A core pitfall is the "sunk cost fallacy." Because we trained models a certain way, we assume continual learning must layer on top of frozen checkpoints. But these models were never designed as continual learners. Instead, treat continual learning as a first-order requirement—design for it from the start.

Evaluation is equally flawed. Total reward alone confounds learning ability with base model strength. Measure reward gain, cost, and base capability on separate Pareto frontiers (trade-off curves), not one metric. A model that cannot update its priors (initial beliefs) from new data fails meaningfully.

Best practice: verifiable continual learning. Turn any failure into a replayable test, apply a "rely optimize" update, then prove the fix helps and breaks nothing that already worked. Each update is tested, every gain measured, and improvements compound over time.

In practice, use a "sandwich" approach: try harness engineering first, fine-tune to break through a ceiling, then do more harness engineering. The goal is that a mistake happens once and never again—so capture failures as tasks and continually replay them without regression. For reinforcement, treat it as a data-mining problem: agents act in an environment, and learning from those actions is continual.

Sources