Open Continual Reinforcement Learning

Intelligence
that keeps
learning.

Building the foundations for autonomous agents that learn, adapt, and grow through a lifetime of experience.

A LIFETIME, NOT A TRAINING PHASE
An agent learning through continual interaction with a changing world Paths connect observations in a larger world. Experience flows to an agent, and actions return to the world. An expanding pattern represents knowledge built over time. This is a conceptual illustration. Agent EXPERIENCE ACTION THE WORLD observe · learn · act
Small agent. Open-ended world.CONCEPT / 001
01 Lifelong interaction02 Learning under constraints03 Reproducible evidence

The question that brings us together

What if learning
never stopped?

An intelligent agent should keep learning from the world it inhabits.

We are building toward AI whose capabilities develop through ongoing interaction: discovering useful state, making predictions, choosing actions, and revising what it knows as the world changes.

OpenCRL brings together the environments, algorithms, evaluation tools, and teaching materials needed to study that process. We draw inspiration from the Alberta Plan and a wider tradition of reinforcement learning, while keeping our tools open to different agent architectures.

From a research question to a shared toolkit

The OpenCRL ecosystem

Five projects.
One shared direction.

Small, composable projects.
Built to be understood, replaced, and extended.

01 / ENVIRONMENTS & EVALUATIONAlpha

WildGym

A common ground for continual learning. Connect changing worlds to agents, and evaluate adaptation across a lifetime.

  • Environment adapters
  • Streaming interaction
  • Frozen evaluation
Inside the alpha
The alpha includes 26 runnable cases, upstream environment adapters, and independent frozen evaluations. Integration checks are distinct from full benchmark reproductions.

Next contribution: add an adapter with explicit observation, action, timing, and reset semantics.

02 / AGENT RESEARCHPrototype

OpenOaK

An inspectable agent composition connecting state, prediction, control, and planning—with room for new ideas in every component.

  • Replaceable components
  • Bounded models
  • Dyna-style planning
Inside the prototype
A minimal composition of linear control, reward prediction, a bounded experience model, and fixed action repetition. Learned representations and option discovery are future research; this is not a complete OaK implementation.

Next contribution: introduce a component and test its learning, freezing, and checkpoint behavior.

03 / SHARED FOUNDATIONSAlpha

OpenCRL Core

Minimal protocols for agents and environments. Share an interaction contract while keeping learning algorithms independent.

  • Agent / environment APIs
  • Lifecycle semantics
  • Checkpoints
Inside the alpha
Shared input types, adapters, state handling, and plugin entry points support the other projects. Agents act; environments step. Environment resets and the end of an agent's learning lifetime are separate concerns.

Next contribution: exercise the contract with an independent agent or a new interaction setting.

04 / LEARNING TOGETHERFirst chapters

Experience Book

Learn continual RL by building it. Follow executable lessons from an agent's lifetime to TD control, evaluation, and composition.

  • Four runnable lessons
  • Small experiments
  • Shared interfaces
Inside the first chapters
The initial lessons cover lifetime and reset semantics, linear one-step Q-learning, frozen comparisons, and a minimal OpenOaK composition. They are teaching examples, not paper reproductions.

Next contribution: explain one mechanism with a small runnable example and an observable failure case.

05 / RESEARCH WORKFLOWSComing soon

RL Research Workbench

A research handbook and experiment workflows for human and AI researchers. Turn questions into testable hypotheses, reproducible experiments, and evidence for the next research decision.

  • Research handbook
  • Experiment workflows
  • Evidence audits
Planned for release
RL Research Workbench is planned as a future OpenCRL project. It will bring together research methodology, experiment protocols, and tools for tracing results to their sources. Connections to WildGym and the other projects are planned.

Experience Book teaches continual RL through runnable lessons; Workbench will support the process of investigating new ideas, including failed experiments and unresolved questions.

Public source releases are in preparation. The first four projects describe the current alpha; RL Research Workbench is planned for a future release.

Our research direction

Experience in.
Lasting capability out.

Progress means better future prediction and control—not just a larger history of the past.

01 / INTERACT

Learn in a changing world.

Study continuous experience, partial observations, and changing dynamics. Treat boundaries and resets as properties to explain.

02 / ADAPT

Build useful knowledge.

Investigate state construction, plasticity, prediction, planning, and temporal abstraction under explicit resource budgets.

03 / EVALUATE

Make improvement visible.

Measure adaptation, retention, and transfer. Account for interaction and computation, and preserve the evidence behind every claim.

A place to test ideas

Beyond a
single episode.

Different settings expose different challenges. WildGym connects them without treating them as interchangeable benchmarks.

When yesterday's strategy stops working.

Track adaptation as dynamics, rewards, or context change. Separate recovery from retention, and report which changes are visible to the agent.

ADAPTERS & SETTINGS CARL · NS-Gym · COOM sequences

Start with the foundations

A shared reading desk.

Ideas that inform our work.
Read, question, implement, and extend.

Build with us

Big questions.
Small, useful contributions.

Bring an environment, a baseline, a careful reproduction, or a lesson that makes one idea clearer. Help make continual RL easier to study—and easier to build on.

Find OpenCRL on GitHub
01

Connect a worldDocument its signals, timing, and lifecycle.

02

Test an ideaShare a baseline, a failure case, and reproducible evidence.

03

Make it understandableTurn a mechanism into a small, executable lesson.