Can Understanding Be Delegated?
What AI Can Do for Our Thinking—and What We Still Have to Construct for Ourselves
1. The Delegation Paradox
Over the years, I have watched technology take over more and more of the work that once demanded our direct attention.
We stopped doing arithmetic by hand. We stopped memorising every phone number. We stopped carrying printed maps. In enterprise work, we learned to trust systems to remember transactions, apply rules, generate reports and move information between teams.
Most of this felt like progress because it was.
The purpose of a good tool is often to remove unnecessary effort.
Research on cognitive offloading describes this more formally: we routinely use external tools and actions to reduce the amount of information processing we need to perform internally. Philosophers have gone even further, asking whether parts of our thinking can sometimes extend into the tools and environments around us. [1]
Generative AI takes that familiar pattern somewhere more interesting.
It does not merely remember something for us.
It can read a long document and tell us what matters.
It can compare competing arguments.
It can propose an architecture, structure a strategy, explain a technical problem, synthesise evidence or draft a recommendation before we have fully formed one ourselves.
That is enormously useful.
I have experienced this myself. There are moments when I give an AI system a problem that might once have taken me an afternoon to structure, and within minutes it returns something coherent, organised and surprisingly close to what I was trying to reach.
The first reaction is often relief.
The second, if we are paying attention, is a more interesting question:
How much of the thinking that produced this answer did I actually do?
That question matters because there is a difference between delegating an operation and delegating part of the process through which we make sense of the operation.
If I use a calculator to multiply two numbers, I can still understand what multiplication means.
The calculator performs the calculation inside a conceptual structure I already possess.
But imagine something different.
Suppose I ask an AI system to read a complex business case, identify the important evidence, decide which constraints matter, reconcile conflicting information and recommend a course of action.
The system has not merely saved me time.
It has helped construct the interpretation.
That is a different kind of delegation.
The question is no longer simply:
Can AI perform cognitive work for us?
Clearly, it can.
The more difficult question is:
Can AI perform that work while leaving us with enough internal understanding to judge what it has produced?
This distinction became clearer to me through years of working with enterprise systems and technical teams.
I have seen people operate sophisticated systems extremely well when the process behaves as expected.
Then one assumption changes.
A data dependency breaks.
A business rule that nobody had noticed suddenly matters.
The procedure is no longer enough, and the person has to reconstruct why the system works the way it does.
That is often where the difference between following a representation and understanding a system becomes visible.
AI may amplify the same difference.
We can receive the summary without tracing the argument.
We can receive the recommendation without weighing the evidence.
We can receive the architecture without understanding why one constraint mattered more than another.
We may even be able to explain the result convincingly because the language and structure have already been prepared for us.
And when the result is good, there may be no immediate reason to notice what is missing.
This is the central paradox of cognitive delegation:
We can delegate cognitive work without necessarily being able to delegate the understanding that makes its results judgeable.
The problem is not that intelligence has moved outside our heads.
Human thinking has always relied on tools, books, colleagues, institutions and accumulated knowledge.
Nor is the problem that AI makes work easier.
In many situations, that is exactly why we want it.
The problem appears when the work we delegate includes some of the same sense-making processes through which we might otherwise have built enough understanding to evaluate the answer.
That creates a dependency we do not always see.
Delegating work is powerful when we retain a reliable way of judging what comes back.
But what happens when judging the delegated answer requires knowledge we might have developed only by doing some of the delegated thinking ourselves?
That is the question at the centre of this essay.
And it leads to a condition that I suspect many of us have already experienced without having a name for it:
we can begin using an explanation before we have truly made it our own.
That is where Proxy Understanding begins.
2. Proxy Understanding
One of the strange things about an AI-generated explanation is that it arrives already finished.
It usually does not show us the uncertainty, discarded paths or half-formed reasoning that might have existed along the way.
It arrives polished.
Structured.
Confident.
Ready to use.
That convenience is part of what makes these systems so valuable.
But it also changes the experience of understanding.
In ordinary learning, we often build understanding gradually.
We encounter a problem.
We misunderstand part of it.
We ask a better question.
Something does not fit.
We revise our view.
Over time, the relationships between ideas begin to settle into something we can use independently.
AI can compress that journey.
Sometimes that is exactly what we want.
But compression creates an interesting possibility: we may acquire the representation of understanding faster than we acquire the structure underneath it.
Research on metacognition has shown that people can misjudge what they have learned, and that fluency can sometimes influence judgments of understanding. [2]
I have noticed versions of this in professional conversations for years.
Someone presents a process, an architecture or a recommendation with complete confidence.
The explanation sounds right.
The terminology is correct.
The slides are polished.
Then someone asks a slightly different question.
Not a trick question.
Just:
Why does this assumption hold?
Or:
What happens if this dependency changes?
And suddenly the fluency disappears.
That does not mean the person knows nothing.
Usually they know quite a lot.
They may understand the recommendation, remember the steps and even be able to reproduce the explanation accurately.
What they may not yet possess is enough of the underlying relational model to reconstruct the reasoning when circumstances move outside the original frame.
This is what I mean by Proxy Understanding.
Proxy Understanding is the condition in which we can use, repeat or navigate an externally supplied representation without having constructed enough of the underlying relational model to independently explain, challenge or adapt it.
The idea is deliberately narrower than simply ‘not understanding’.
Proxy Understanding is not ignorance.
It is not even necessarily poor performance.
In fact, that is part of what makes it difficult to notice.
A person can perform well while relying on a representation they did not fully construct.
They can use a dashboard.
Follow an architecture.
Repeat an argument.
Apply a recommendation.
Everything may work perfectly while the surrounding conditions remain stable.
The limitation becomes visible when something changes.
A good mental model is not merely a collection of facts.
It contains relationships.
Why one variable affects another.
Which assumptions are carrying the argument.
What can change without breaking the system.
What cannot.
Research on mental models has long examined how people reason by constructing internal representations of situations rather than relying only on isolated propositions or facts. [3]
That idea matters here because understanding becomes most valuable when the world does not behave exactly as expected.
In enterprise work, I have seen this difference repeatedly.
A documented process tells us what normally happens.
Understanding tells us what to look for when it does not.
A technical specification tells us what the system is supposed to do.
Understanding helps us recognise which dependency may have failed when the behaviour changes.
A business rule tells us what decision should normally follow.
Understanding allows us to question whether the assumptions behind that rule still make sense.
AI can provide extraordinarily useful representations of all of these things.
It can often make the relationships clearer.
It can help us see patterns we missed.
It can even challenge our initial interpretation.
But if we simply inherit the finished explanation, we may hold the answer without holding enough of the structure that gives the answer meaning.
That is why Proxy Understanding can feel so convincing.
We have the vocabulary.
We have the conclusion.
We may even have the confidence.
What we do not necessarily have is the internal map.
This connects directly to Understanding Drift.
In the first essay, I described the growing distance that can emerge between what we are able to express and what we have actually constructed internally.
Proxy Understanding is one way that distance can form.
The explanation may genuinely help us.
It may even be correct.
But if the explanation becomes a substitute for constructing enough of the relationships ourselves, expression can begin to outrun comprehension.
The risk is not that every AI-generated explanation is shallow.
Many are useful.
Some are excellent.
The risk is that usefulness can hide dependency.
And that dependency often stays invisible until the environment changes.
Then the deeper question appears:
Can I still explain this when the wording is gone?
Can I adapt it when one assumption changes?
Can I recognise when the recommendation no longer fits the situation?
If the answer is no, we may not yet possess the understanding we thought we had.
We may possess something adjacent to it.
Something useful.
Something operational.
Something that can get us started.
But not yet something fully our own.
That distinction leads naturally to the next question.
Because AI does not always weaken construction.
Sometimes it does the opposite.
Sometimes it helps us build a stronger mental model than we would have built alone.
The real issue, then, is not whether AI participates in our thinking.
It is how it participates.
Is it helping us construct understanding?
Or quietly standing in for the construction?
That is the difference between a scaffold and a substitute.
3. Scaffold or Substitute
The more I have worked with AI, the less useful I find the question:
Is using AI good or bad for thinking?
It is too broad.
A calculator can save effort without weakening mathematical understanding.
A good teacher can explain something faster than we would have discovered it alone.
A colleague can point out an assumption we missed.
A search engine can save hours of retrieval without doing the judgement for us.
AI belongs somewhere in that same family of cognitive support.
But it is unusually powerful because it can move much further into the process.
It can do more than retrieve.
It can frame.
Compare.
Interpret.
Draft.
Synthesise.
Recommend.
That is why the distinction between scaffolding and substitution matters.
Sometimes AI helps us think.
Sometimes it does enough of the thinking that we no longer need to construct as much ourselves.
Those two experiences can look very similar from the outside.
The output arrives quickly in both cases.
The difference is what happened inside the person using it.
I have seen this in my own work.
There are times when I ask AI to challenge an argument I have already formed.
It gives me a counterexample.
I push back.
It gives another.
I reconsider one assumption.
Then I rewrite the argument in my own words.
In that situation, the tool is not replacing my thinking.
It is creating friction in the right places.
It is helping me see relationships I had not noticed.
That is scaffolding.
The final understanding is still being constructed by me, even though AI participated heavily in the process.
But there is another kind of interaction.
I can ask the system to read a topic I barely know, structure the problem, identify the important concepts, compare the evidence and produce a polished conclusion.
Then I can take that conclusion and move forward.
That may be efficient.
It may even be accurate.
But the cognitive journey is different.
The system has not merely supported the construction.
It has performed part of the construction on my behalf.
That is substitution.
The distinction is not moral.
Substitution is not automatically bad.
There are many things we should happily allow technology to do for us.
Few of us want to manually perform every calculation, sort every record, search every database or format every document.
The value of technology has always included the ability to remove work that does not deserve our full attention.
The important question is more specific:
What cognitive work remains for the human to do?
Educational research gives us useful language here.
Scaffolding has long been used to describe support that helps a learner perform beyond what they could manage alone while still participating in the learning process.
More recent work on human–AI learning asks a related question: when does intelligent assistance augment human learning, and when does it begin to replace parts of the activity through which learning would normally occur? [4]
The answer is not simple.
Prior knowledge matters.
Task design matters.
The way the tool is used matters.
An expert and a novice can receive exactly the same AI-generated answer and experience it very differently.
For the expert, the answer may connect immediately to an existing mental model.
They can recognise what is missing.
They can test the assumptions.
They can disagree intelligently.
The AI output becomes another input into an already-developed structure.
For the novice, the same output may become the structure.
That difference matters.
Research on the expertise-reversal effect is a useful reminder that support does not work identically for people with different levels of prior knowledge. [5]
I have seen versions of this in training environments.
An experienced practitioner can look at a proposed solution and quickly say:
Something is wrong here.
They may not immediately be able to explain why.
But years of accumulated relationships, exceptions and consequences are already present in their mental model.
A newcomer may see the same solution and think:
It looks complete.
Both are looking at the same artefact.
Only one has enough internal structure to interrogate it.
This is why AI can be both extraordinarily helpful and quietly deceptive.
It can reduce unnecessary friction.
But some friction is not unnecessary.
Sometimes the confusion, comparison, failed attempt or difficult question is part of how the model gets built.
That does not mean struggle is inherently valuable.
There is no virtue in making learning harder than it needs to be.
But there is a difference between removing waste and removing construction.
AI can help us distinguish those two if we use it deliberately.
Instead of asking:
Give me the answer.
What assumptions am I missing?
What would make this argument fail?
Show me two competing interpretations.
Ask me questions until you can see where my understanding is weak.
Do not solve this yet. Help me structure the problem.
Those are very different interactions.
The same technology is involved.
But the cognitive role changes.
In one case, AI becomes a producer of finished representations.
In the other, it becomes a partner in construction.
That distinction becomes especially important in professional environments.
Organisations naturally optimise for speed.
If a tool can draft the report, summarise the meeting, structure the proposal, inspect the log and prepare the recommendation in a fraction of the time, the efficiency gain is obvious.
What is less visible is what happens to the human learning that used to occur inside those tasks.
A junior analyst once had to struggle through messy data before understanding why a metric mattered.
A new consultant had to listen carefully to conflicting stakeholders before learning how organisations really behave.
An engineer had to diagnose failures before recognising which dependencies were actually fragile.
If AI increasingly performs the first synthesis, the work may become faster.
But some opportunities to build judgement may also become thinner.
That does not mean we should preserve inefficient work simply because it once taught us something.
It means we need to become more deliberate about where learning now happens.
If the old path to expertise is being compressed, we may need to design a new one.
This is where the scaffold–substitute distinction becomes more than an educational idea.
It becomes an organisational design problem.
AI can take over more of the task.
But organisations still need people who can recognise when the task has been done badly.
That creates a tension.
The more cognition we externalise, the more important it becomes to know which internal capabilities we still need to preserve.
And that takes us to the most important question in the essay so far:
What do we need to understand in order to judge what has been delegated?
That is where The Evaluation Dependency begins.
4. The Evaluation Dependency
The more we delegate, the more important another question becomes:
How do we know when the delegated result is good enough?
That sounds obvious.
But in practice, it is where much of the difficulty begins.
If an AI system writes a sentence badly, we can usually see it.
If it produces invalid code, a compiler may reject it.
If it retrieves the wrong figure, we may be able to check the source.
Those are relatively visible failures.
But many of the decisions we make at work are not like that.
A strategy can sound convincing while resting on a weak assumption.
An architecture can look elegant while ignoring a dependency that only becomes important under pressure.
A recommendation can be well structured and still frame the problem incorrectly.
A summary can accurately reflect the documents it was given and still miss the issue that matters most to the people involved.
The difficult failures are often not obvious errors.
They are omissions.
Misplaced emphasis.
Context that never entered the prompt.
Trade-offs that were never surfaced.
Questions that were never asked.
I have seen this kind of problem in enterprise work many times, long before generative AI arrived.
A report can be technically correct and still be operationally misleading.
A dashboard can display every requested metric and still encourage the wrong decision.
A system design can satisfy the documented requirement while failing to account for how people actually work around the process.
The challenge is not simply producing information.
It is judging whether the information still makes sense in context.
AI makes this challenge more important because the output can be unusually persuasive.
The wording is clean.
The structure is complete.
The reasoning appears orderly.
That surface quality can make us feel that evaluation has already happened.
But presentation is not validation.
To evaluate an answer properly, we often need to ask questions that sit underneath the answer:
What assumptions is this conclusion carrying?
What evidence was treated as important?
What was left out?
What would make this recommendation fail?
What happens if the environment changes?
Who experiences the consequences if this interpretation is wrong?
Those are not formatting questions.
They require understanding.
Research on epistemic vigilance has long examined how people assess information received from others rather than accepting it automatically. Research on automation has also shown that human oversight can become vulnerable when people place too much confidence in apparently reliable systems. [6] [7]
AI does not remove that problem.
It can make it harder to notice.
This is where I think a deeper dependency emerges.
Imagine that I delegate not only the final drafting of an answer, but also the earlier stages:
the problem framing,
the evidence selection,
the comparison,
the interpretation,
and the synthesis.
The system then returns a polished recommendation.
Now I have to evaluate it.
But what do I use to evaluate it?
If the very reasoning I would normally have used to construct my own mental model was also delegated, I may discover that the answer requires knowledge I never had the opportunity to build.
That is what I mean by Evaluation Dependency.
Delegation becomes epistemically risky when the operations we outsource are also the operations we would need in order to understand why the resulting answer should—or should not—be trusted.
This is not an argument against delegation.
It is an argument for noticing what evaluation requires.
Sometimes evaluation is easy because the result can be checked independently.
Sometimes it is difficult because the evaluator must already understand the domain well enough to recognise what does not fit.
That difference is crucial.
I have experienced this even in relatively ordinary AI use.
If I ask for help rewriting something I already understand, I can usually judge the result quickly.
I know the subject.
I know what I mean.
I can see when the wording has changed the meaning.
But if I ask for a recommendation in an area where I have very little background, the situation changes.
The answer may look equally polished.
My confidence in evaluating it should not.
That sounds simple when stated directly.
In practice, though, fluency makes it easy to forget.
The better the system becomes at producing plausible outputs, the easier it becomes to confuse quality of presentation with quality of judgement.
And the danger is not always that the AI gives us a completely false answer.
More often, the answer may be mostly right.
That can be harder.
A completely wrong answer invites scrutiny.
A mostly right answer can carry one assumption quietly through the entire decision.
If we do not know where to look, we may never see it.
This leads to a question I now find more useful than asking whether an AI system is ‘accurate’ in general:
If the AI is wrong, can I recognise the error without already possessing much of the understanding that the AI helped me avoid constructing?
That question changes how we think about delegation.
It shifts attention away from what the AI can do and toward what the human still needs to know.
And it reveals why some tasks are much safer to delegate than others.
Not because they are necessarily easier.
Not because they are less intellectually impressive.
But because their failures are easier to see.
That distinction gives us the next part of the framework.
If evaluation depends on how visible failure is, then we need a better way to think about where delegation should stop, where it can safely continue, and where stronger human understanding is still required.
That is the Epistemic Delegation Boundary.
5. The Epistemic Delegation Boundary
Once we see Evaluation Dependency clearly, a second question follows:
Which kinds of cognitive work can we delegate safely, and which kinds still require substantial human understanding?
At first, it is tempting to answer this by task type.
Writing might feel delegable.
Strategy might feel less so.
Coding may seem technical enough to automate.
Judgement may seem inherently human.
But the more I look at real work, the less useful those labels become.
Some coding tasks are easy to verify.
Some writing tasks carry enormous interpretive risk.
Some complex calculations can be checked almost mechanically.
Some apparently simple recommendations depend on years of contextual knowledge.
So I do not think the right dividing line is simply:
easy versus difficult
or even:
technical versus human
A more useful distinction is this:
How visible is a meaningful failure to someone who did not construct the underlying model?
That question changes the picture.
Consider a few examples.
If AI generates code with a syntax error, the compiler may reject it.
If it converts one data format into another incorrectly, validation rules may expose the problem.
If it retrieves a statistic, we may be able to check the primary source.
If it expands a repetitive pattern, we can often inspect the result against explicit rules.
In these cases, much of the evaluation happens outside the evaluator's head.
The environment gives us a test.
The system fails.
The schema breaks.
The source disagrees.
We do not necessarily need to reconstruct every cognitive step the AI took in order to know that something went wrong.
That is lower evaluation dependency.
Now consider a different class of work.
Suppose AI helps frame a business problem.
It may produce a clear and sensible problem statement.
But perhaps it has framed the issue around cost when the real constraint is trust.
Or around efficiency when the real issue is regulation.
Or around technology when the failure is actually organisational.
Nothing necessarily ‘breaks’.
The framing can look perfectly reasonable.
The error lies in what the frame excludes.
That is a very different kind of failure.
The same applies to interpreting conflicting evidence.
A system can summarise both sides accurately and still weight them badly.
It can identify every major argument and still misunderstand which uncertainty matters most.
It can produce a recommendation that is internally coherent but poorly fitted to the environment in which the decision must actually be made.
There may be no red warning.
No failed test.
No compiler message.
No single sentence we can point to and say, ‘That is obviously wrong.’
The failure is visible only to someone who understands enough of the surrounding system to notice that something important does not fit.
That is higher evaluation dependency.
I think of these two conditions as ends of an Epistemic Delegation Boundary.
Not a hard line.
A continuum.
At one end, the output is largely externally verifiable.
At the other, evaluation increasingly depends on internally held knowledge, context, judgement and experience.
A simplified version looks like this:
The important point is that delegability is not an inherent property of the task label.
Summarisation is a good example.
Summarising a straightforward factual document may carry relatively low evaluation dependency if the reader can easily compare the summary with the source.
But summarising a complex regulation, a body of conflicting research, or a politically sensitive consultation may require much deeper understanding.
The task label is identical.
The evaluation burden is not.
I have seen similar differences in enterprise technology.
A system can tell us that a job failed.
That may be easy.
Understanding why it failed can be much harder.
Was it bad data?
A timing dependency?
A security change?
A process assumption?
An integration issue?
A business rule that nobody documented properly?
The visible error is often only the surface.
The real diagnosis requires a model.
This is why I think the boundary matters.
It gives us a better question than:
Can AI do this task?
In many cases, the answer will increasingly be yes.
The more useful question is:
What would I need to understand in order to recognise a consequential failure?
That is a governance question.
It is also a learning question.
And increasingly, it is an organisational design question.
Because if we delegate a high-evaluation-dependency task to AI, we need to be confident that someone still holds enough of the relevant model to judge the result.
Sometimes that will be an experienced practitioner.
Sometimes it will be peer review.
Sometimes independent testing.
Sometimes formal controls.
Sometimes a second system checking the first.
The point is not that humans must personally reconstruct every step.
The point is that evaluation capability must exist somewhere.
That matters because ‘human in the loop’ can sound reassuring while meaning very little.
A human presence does not automatically create effective oversight.
If the person in the loop lacks the context, authority, or internal model required to challenge the system, governance exists only on paper.
That observation has become increasingly important to me.
As organisations adopt more capable AI systems, we may be tempted to measure governance by whether a person approved the output.
But approval is not the same as evaluation.
And evaluation is not the same as understanding.
The more consequential the decision, the more important those distinctions become.
So before delegating a cognitive process, I would ask two questions:
How much time will this save?
And then the more important one:
What would be required to recognise a meaningful failure—and does the person overseeing the work still hold enough of the mental model needed to see it?
That is the boundary I care about.
Not because it tells us where AI must stop forever.
But because it tells us where human understanding still carries responsibility.
And once we see the problem that way, the final question becomes unavoidable:
How should we manage our own understanding in a world where more and more cognition can be externalised?
That is where Understanding Stewardship begins.
6. Understanding Stewardship
If the Epistemic Delegation Boundary tells us where delegation becomes more demanding, then the final question is what we do about it.
I do not think the answer is to retreat from AI.
That would misunderstand both the technology and the history of human progress.
We have always built tools to extend ourselves.
We write things down because memory is limited.
We use calculators because attention is valuable.
We build databases because no individual can carry the operational memory of an organisation.
We use search because knowledge is too distributed for any one person to hold.
AI belongs within that same long story of extension.
But it adds something new.
It can now participate not only in remembering and retrieving, but in interpreting, structuring and recommending.
That means we need to become more deliberate about what we are externalising.
For me, that is what Understanding Stewardship means.
Understanding Stewardship is the deliberate practice of deciding which parts of cognition we are willing to externalise, while preserving enough internally constructed understanding to remain capable of judgement, adaptation and responsibility.
The word stewardship matters.
It does not imply ownership in the sense of controlling everything ourselves.
A steward looks after something valuable so that it remains usable over time.
Understanding deserves the same care.
In organisations, this may mean resisting a very natural temptation.
If AI can do the first draft, the first analysis, the first synthesis and perhaps even the first recommendation, efficiency encourages us to let it.
Sometimes that will be exactly the right decision.
But we should also ask what those earlier stages used to teach us.
A junior analyst does not become experienced only by reading excellent final reports.
They learn by deciding what matters.
They make weak assumptions.
Someone challenges them.
They discover that the data does not support the story they expected.
They revise.
A young engineer does not develop judgement only by seeing correct architectures.
They see failures.
They misunderstand dependencies.
They learn why one constraint matters more than another.
A consultant does not understand organisations simply by receiving perfect summaries of stakeholder interviews.
They learn through contradiction.
Through ambiguity.
Through noticing that what people say and what the process actually rewards are sometimes different things.
Much of expertise is built inside these encounters.
Research on expertise development emphasises the importance of structured practice, feedback and repeated adaptation rather than experience alone. [8]
AI may make some of them less necessary.
That is good.
But if we remove too many of them without creating new opportunities for construction, we may become more productive while quietly weakening the path through which judgement develops.
That possibility deserves attention.
Not panic.
Design.
Perhaps some organisations will begin treating certain forms of reasoning almost like supervised practice.
Let AI produce a recommendation, but ask the practitioner to predict the answer first.
Let AI summarise the evidence, but require the analyst to identify the assumptions.
Use AI to generate alternatives, then ask the team to explain which one fails and why.
Let a junior professional work with an AI assistant, but periodically remove the assistance and see whether the underlying model remains.
This is not about making work artificially difficult.
It is about ensuring that convenience does not accidentally consume the conditions through which expertise grows.
The same principle applies individually.
I use AI every day.
And one habit I increasingly value is noticing what kind of help I am asking for.
Sometimes I want speed.
Sometimes I want a second pair of eyes.
Sometimes I want an explanation.
Sometimes I want disagreement.
Sometimes I want the system to do something completely so that I do not have to think about it again.
All of those can be legitimate.
But they are not cognitively equivalent.
There are moments when I now stop and ask:
Do I want the answer, or do I want to understand this?
That question sounds simple.
I think it may become increasingly important.
When internal understanding matters, we need to participate in the construction, not only inspect the finished result.
I may ask the AI not to give me the conclusion yet.
I may ask it to challenge my reasoning.
I may try to explain the idea back without looking at the response.
I may ask what would make the recommendation wrong.
Or I may deliberately spend another ten minutes with the problem before asking for help.
Not because independent struggle is morally superior.
It is not.
But because sometimes the internal model is the thing I am actually trying to build.
That distinction is easy to forget when answers are cheap.
And answers are becoming very cheap.
Understanding is not.
Understanding still requires connection.
Context.
Revision.
Judgement.
And sometimes responsibility.
This is why I do not think the central challenge of AI is whether machines will eventually ‘think better than us.’
That question is interesting, but it can distract from a more immediate one.
What kinds of thinking do we still need to be capable of ourselves?
Not because humans must remain intellectually independent from machines.
We never were independent from our tools, our teachers or one another.
The goal is not independence.
It is epistemic sovereignty.
By that I mean something modest but important:
enough internally constructed understanding to question, adapt and take responsibility for what external intelligence produces.
That may become one of the most valuable human capabilities in an AI-rich world.
Not remembering more than the machine.
Not calculating faster.
Not producing the first draft.
But knowing when the answer does not fit.
Knowing which assumption needs another look.
Knowing when the situation has changed.
Knowing when another person will carry the consequence of a decision.
Knowing when confidence is not yet justified.
And being able to say:
I understand enough of this to take responsibility for what happens next.
That is a much higher standard than simply being able to use AI well.
It is also, I think, a more hopeful one.
Because it does not ask us to compete with artificial intelligence.
It asks us to become more deliberate about what we want intelligence to do for us.
AI can research for us.
It can draft for us.
It can search for us.
It can compare, summarise, simulate and recommend.
It can extend our reach far beyond what an individual could manage alone.
We should use that capability.
But when we ask it to understand for us without reconstructing enough of that understanding ourselves, we risk stepping out of the driver's seat of our own minds.
The future, then, may not belong to those who resist AI.
Nor simply to those who use it most aggressively.
It may belong to those who learn where to delegate, where to participate, and where understanding itself must remain part of the work.
That is the responsibility behind Understanding Stewardship.
And it leaves us with one final question.
If judgement depends on mental models we have constructed internally, then:
What are those models actually made of?
How do we build them?
How do experience, explanation, attention, error, reflection and context become something we can later call understanding?
And what happens to that architecture when fluent answers arrive before the model has had time to form?
That is the next question.
Not whether AI can give us better answers.
But how we continue building the structures inside ourselves that allow us to know what those answers mean.
That is the construction we must examine next.
Notes & Sources
This essay is an original synthesis informed by established work in cognitive offloading, metacognition, mental models, human–AI learning, expertise and automation. The sources below are included as light intellectual anchors rather than as a formal academic bibliography.
[1] Risko, E. F., & Gilbert, S. J. (2016). “Cognitive Offloading.” Trends in Cognitive Sciences, 20(9), 676–688. View source
[2] Bjork, R. A., Dunlosky, J., & Kornell, N. (2013). “Self-Regulated Learning: Beliefs, Techniques, and Illusions.” Annual Review of Psychology, 64, 417–444. View source
[3] Johnson-Laird, P. N. (1983). Mental Models: Towards a Cognitive Science of Language, Inference, and Consciousness. View source
[4] Molenaar, I. (2022). “Towards Hybrid Human–AI Learning Technologies.” European Journal of Education, 57, 632–645. View source
[5] Kalyuga, S., Ayres, P., Chandler, P., & Sweller, J. (2003). “The Expertise Reversal Effect.” Educational Psychologist, 38(1), 23–31. View source
[6] Sperber, D., et al. (2010). “Epistemic Vigilance.” Mind & Language, 25(4), 359–393. View source
[7] Parasuraman, R., & Manzey, D. H. (2010). “Complacency and Bias in Human Use of Automation: An Attentional Integration.” Human Factors, 52(3), 381–410. View source
[8] Ericsson, K. A. (2006). “The Influence of Experience and Deliberate Practice on the Development of Superior Expert Performance.” In The Cambridge Handbook of Expertise and Expert Performance. View source
About the Framework
The terms Proxy Understanding, Evaluation Dependency, Epistemic Delegation Boundary and Understanding Stewardship are used here as part of the conceptual synthesis developed in this essay. They are positioned alongside, rather than as substitutes for, the established research traditions referenced above.