First, let us be clear: what follows is a vision, not a capability we have today. This is a forward-looking roadmap β a direction we are building toward, but have not yet reached. We want to discuss a concept that matters to us: an agent that learns from its own work.
Today: starting from scratch
Today, an intelligent agent begins each conversation almost anew. It may have access to memory, but it does not learn from the quality of its own work yesterday. If one of its proposals was rejected, that rejection does not become a lasting lesson. If a pattern exists in its mistakes, the agent cannot see it. The agent works, but it does not reflect on that work.
The vision: meta-cognition
We are exploring a layer of meta-cognition: the ability for an agent to evaluate its own output. In this vision, the agent maintains a lasting picture of its behaviour β facts about how it has acted, notes from human-reviewer corrections, and patterns in its history of accepted and rejected proposals. The agent reads this record and gradually adjusts its behaviour. Each human correction, instead of being forgotten, becomes a lesson.
Why we havenβt built it yet
If this concept is valuable, why isnβt it here today? We have deliberately deferred it. At the current stage, the complexity and cost of this layer are not justified by the benefit it brings. Our priority is establishing solid foundations β a system that is reliable and understandable today. Meta-cognition is a layer that builds upon those foundations, not before them. This is a conscious decision, not an oversight.
Why we talk about it anyway
If we have not built it, why discuss it now? Because architectural direction matters. We want to be clear about what we are building today and what we aspire to tomorrow β without blurring the line between the two. The self-learning colleague is an aspiration we believe in and leave room for in our current design, but we do not sell it as todayβs reality.
Putting it together
The difference between a tool and a colleague is often just this: a tool is identical each time it runs, while a colleague improves with time. We are building toward this vision β an agent that learns from its own work. But until the day it arrives, we will call it what it is: a forward-looking roadmap, not a promise for today.