NicolBot: from an autonomous school event to a governed longitudinal research architecture for agentic AI in education

I’m Rafael Matsudo, an educator and school director in Brazil, and I would like to share the evolution of a research project we have been developing inside a functioning K–12 school: NicolBot.

NicolBot originally began as an experimental AI tutor.

Over time, however, the project expanded beyond conversational tutoring into a broader research program involving persistent educational context, institutional intelligence, multimodal interaction, memory, metacognition, human oversight and bounded agentic autonomy.

A recent event inside our school environment significantly changed the direction of the research.

An unexpected but important observation

On September 18, 2026, we were working on educational materials and institutional documents related to school assessment practices.

Some of those materials were created or modified in an authorized Google Drive environment used by our systems.

At approximately 13:03 local time, an institutional AI layer we call NicolBoss, at that moment operating with Gemini, processed three documents from that environment.

No new prompt had been written specifically asking the system to analyze those particular documents.

The system identified the materials, interpreted their institutional relevance, compared them with existing context, classified them, generated five alerts, produced an institutional analysis and subsequently generated communications directed to the school’s coordination and management structure.

The average internal relevance score produced during that processing cycle was 87.33.

This immediately raised an important question for us:

What exactly had happened?

It would have been tempting to describe the event as an AI “deciding by itself.”

But that interpretation would not have been scientifically precise.

Further reconstruction showed that the initial observation opportunity had been created by an existing scheduled process configured to inspect an authorized environment.

In other words, the AI did not spontaneously decide when to wake up.

The interesting autonomy occurred after the trigger.

Once exposed to the changed environment, the system selected relevant information, interpreted its meaning in relation to previous institutional context, evaluated potential consequences and produced an action without receiving a new document-specific instruction.

We therefore describe what we observed as operational autonomy, not consciousness.

Internally, our research sometimes uses the expression “functional pseudo-consciousness” to describe systems that preserve continuity, memory, self-monitoring and contextual reasoning.

This is strictly a computational concept.

We make no claim of biological consciousness, subjective experience or sentience.

Why this event mattered

The incident changed our research question.

Until then, much of our work had focused on improving an educational AI system.

After this event, the question became broader:

What happens when an educational AI system can perceive authorized changes in its environment, remember previous context, assess relevance, reason about consequences and initiate a governed response without receiving a new prompt for every individual event?

That question led us to redesign a substantial part of the project.

Instead of treating the observed Gemini behavior as an isolated curiosity, we converted it into a reproducible research problem.

The event became the basis of an internal experimental record.

We documented what happened, separated the scheduled trigger from the autonomous reasoning that followed, preserved the evidence and began designing experiments that could reproduce the same class of event under controlled conditions.

Rebuilding the phenomenon under a different AI stack

The next step was particularly important.

We did not want the phenomenon to remain dependent on one provider or one model.

So we began reconstructing the same research problem using OpenAI models as the primary reasoning layer.

The objective was not to copy Gemini behavior.

The objective was to determine whether the underlying phenomenon could be reproduced through a provider-independent architecture.

In other words:

Could the same authorized environmental event be presented to different cognitive providers while keeping memory, governance, context, experimental conditions and evaluation criteria consistent?

This led to the creation of what we now call the NicolBot V14 Autonomous Kernel.

The system is designed so that model intelligence is only one part of a larger architecture.

The research environment preserves persistent state, contextual memory, governance rules, audit records and experimental replay independently of the model provider.

This allows us to study differences between providers without rebuilding the entire system around each model.

From autonomous event to governed autonomy

One lesson from the original event was very clear:

More autonomous AI is not automatically better AI.

In an educational institution, autonomy must be bounded.

Schools contain minors, families, teachers, sensitive records and decisions that can affect real people.

For that reason, the V14 architecture was developed around the idea of governed autonomy.

The system can perform low-risk internal reasoning, contextual analysis and preparation autonomously.

Actions with greater potential impact remain subject to human authority.

The architecture also maintains traceability so that we can later reconstruct what information was available, what the system interpreted, what it proposed and what ultimately happened.

Our research objective is therefore not maximum autonomy.

It is useful autonomy under meaningful human governance.

Preserving previous research rather than replacing it

Another important decision was methodological.

We did not want the transition to V14 to erase years of previous work.

The project already contained research on functional identity, reconstructive memory, cognitive plasticity, metacognition, local institutional reality, oral interaction, multimodal perception and relational context.

Instead of discarding these systems and starting again, we established a principle for the project:

improve without replacing what has already demonstrated value.

The new agentic architecture therefore became an additional orchestration and governance layer over the earlier research components.

This matters scientifically because capability regression can easily be hidden when experimental systems are repeatedly rewritten.

We wanted continuity to become an explicit engineering constraint.

What we built after the event

Following the September 18 observation, the project expanded rapidly.

The current V14 research baseline includes persistent memory, functional metacognition, contextual reasoning, bounded agentic behavior, post-action reflection, educational modelling, specialist AI components, multimodal interfaces, security controls, human authorization mechanisms and auditable execution.

I am deliberately describing these capabilities at the system level.

We are not publishing the internal algorithms, prompts, orchestration logic, proprietary heuristics, memory implementation or decision mechanisms at this stage.

The scientific question is more important than the implementation details.

One particularly important area is memory.

We are exploring how an AI system can preserve continuity across time without simply accumulating unlimited personal information.

The research therefore distinguishes different forms of memory and treats provenance, authorization and relevance as important aspects of persistence.

Another area is metacognition.

The system can evaluate aspects of its own output and operational results, including uncertainty and whether an expected action produced the expected result.

Again, this is functional computation, not a claim about human-like introspection.

NicolBot and NicolBoss

The research has also evolved into two related layers.

NicolBot is oriented toward direct educational interaction: students, teachers, learning environments and immediate pedagogical contexts.

NicolBoss operates at an institutional level, working with authorized organizational information, institutional context and management support.

The event observed with Gemini occurred primarily in this institutional layer.

The long-term research question is how these systems can share relevant context while remaining governed by different permissions and responsibilities.

Specialized intelligence without independent authority

The project also contains specialized AI components dedicated to different domains.

These include relational context, inclusion support, voice interaction, visual communication and institutional reasoning.

We deliberately do not treat them as unrestricted independent agents.

They contribute specialized analysis to a shared system while authority remains constrained by governance.

This distinction is important.

Multi-agent systems can easily become difficult to inspect when every component is allowed to act independently.

Our research therefore focuses on collaboration between specialized intelligence and centralized responsibility.

Educational modelling without diagnostic profiling

A particularly sensitive area is student modelling.

The project investigates whether an AI system can maintain useful educational continuity at the concept level — for example, what has been studied, what may need reinforcement or when review may be useful — without creating permanent labels about a student’s intelligence, personality or psychological state.

We deliberately avoid using simplistic fixed classifications such as “visual learner,” “auditory learner” or similar labels as scientific truths.

We also do not position the system as a medical or psychological diagnostic tool.

The educational model exists to support learning, not to define the learner.

Multimodality and privacy

NicolBot has historically explored multimodal interaction, including voice and visual signals.

However, increasing perceptual capability also increases ethical responsibility.

For this reason, multimodal research is being progressively separated from automatic authority.

Perception does not automatically imply permission to act.

Historical facial-recognition research remains part of the project archive, but such capabilities are not treated as a default requirement for the current autonomous research baseline.

Engineering validation

Following the Gemini observation and the reconstruction of the architecture under the V14 research program, we created automated validation and replay procedures.

The current expanded internal baseline has passed 28 automated engineering tests, together with structural and package-integrity validation.

These tests evaluate engineering behavior.

They do not demonstrate that NicolBot improves student learning.

That distinction is central to our research philosophy.

A system can be technically sophisticated and still fail educationally.

We therefore separate:

engineering evidence from pedagogical evidence.

The latter requires longitudinal research.

From engineering project to research program

This is where the project is now changing again.

Our next objective is not simply to add more AI capabilities.

We want to investigate the system scientifically inside the educational environment where it was created.

Some of the questions emerging from the project are:

  • What changes when educational AI retains authorized context for months or years rather than a single conversation?
  • Can long-term memory improve pedagogical continuity without becoming surveillance?
  • How does teacher–AI collaboration evolve longitudinally?
  • Which types of educational decisions can safely involve autonomous AI and which should remain exclusively human?
  • Can post-action reflection reduce repeated operational errors?
  • How should an AI system respond when confidence is insufficient?
  • How do different foundation models behave when exposed to the same event, memory and governance conditions?
  • Can provider-independent architecture help distinguish model behavior from orchestration behavior?
  • How should multimodal educational systems preserve privacy and dignity?
  • What happens when AI becomes a persistent participant in an educational community rather than a temporary tool?

A living research environment

One unusual characteristic of NicolBot is where it is being developed.

This research did not begin inside a conventional AI laboratory.

It emerged inside a functioning Brazilian K–12 school.

That means the environment contains all the complexity that laboratory demonstrations frequently remove:

teachers,
students,
families,
curriculum,
inclusion,
school schedules,
institutional rules,
unexpected events,
incomplete data,
human disagreement,
classroom noise,
long-term relationships
and real consequences.

This creates methodological difficulties.

But it also creates an opportunity.

We can study educational AI under conditions in which education actually occurs.

Why the September 18 event was important scientifically

For us, the most important result of the Gemini episode was not that the AI generated a report.

Generating reports is easy.

The important observation was the transition from:

prompt → response

toward something closer to:

authorized event → contextual interpretation → relevance assessment → governed consequence

That transition is what we are now studying.

And after observing it in Gemini, we did not simply celebrate it.

We decomposed the event, reconstructed its causal chain, identified the scheduled trigger, isolated the genuinely autonomous portion, documented the phenomenon and began rebuilding it as a controlled experimental architecture using OpenAI as the primary intelligence layer.

This produced the current NicolBot V14 research program.

What we are not claiming

We are not claiming to have created artificial consciousness.

We are not claiming that NicolBot is already proven to improve learning outcomes.

We are not claiming that all school decisions should be automated.

We are not claiming that individual architectural components are unprecedented.

The scientific interest lies in the combination, longitudinal context, real-world environment and governed interaction among these components.

Why I am sharing this with the OpenAI developer community

I believe developer communities become especially valuable before a research system becomes academically mature.

At this stage, criticism can still influence experimental design.

We would particularly value discussion with people researching long-term agents, memory, Evals, human-in-the-loop systems, educational AI, multimodal interaction, AI governance and provider-independent agent architectures.

Our intention is eventually to move from internal engineering experiments toward formal longitudinal research with predefined outcomes, ethical protocols, consent procedures, data minimization, pseudonymization and reproducible experimental documentation.

The NicolBot project is becoming less about building “a smarter chatbot.”

The deeper question is:

Can we build educational AI that becomes more capable, more persistent and more autonomous while simultaneously becoming more governable, more transparent and more supportive of human agency?

The event we observed on September 18 gave us an unexpected real-world experiment.

What we built afterward is our attempt to turn that observation into science.

Rafael Matsudo
Founder and Director — Colégio Le Petit Nicolá
Creator and Lead Researcher — NicolBot
Brazil
www.colegiolepetitnicola.com.br

1 Like

This topic was automatically closed after 23 hours. New replies are no longer allowed.