Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

MindOS Learning Manual: Modality State | When the Eyes Are Doing Two Jobs at Once

MindOS · Modality State · Observe Visual Competition → Test Representation Demand → Shift Transient Words When Useful → Preserve Accessibility and Key Text → Integrate → Reconstruct → Fade Presentation Support → Transfer → Return

Wait, What? Sometimes Reading the Explanation Can Make the Diagram Harder to Understand

A learner is watching an animation of a scientific process.

The diagram is moving. Labels are changing. Arrows are appearing. At the same time, a paragraph at the bottom of the screen explains what is happening.

Every part is correct.

But the learner keeps looking down to read the words and then back up to find what changed in the diagram. By the time they locate the right visual element, the animation has already moved on.

The problem is not necessarily poor attention, weak reading, or lack of effort.

The eyes may simply be doing two demanding jobs at once.

Quick Answer

Owned Learner Job: when essential visual information and essential explanatory words compete for visual processing at the same time, determine whether moving some transient verbal explanation into spoken narration reduces that competition without removing information the learner still needs to see, reread, decode, or access.

The RFE is not “audio is better than text.” It is:

Match the presentation mode to the actual processing bottleneck, then remove presentation support until the learner can reconstruct and use the idea independently.

The Core Problem: Visual Competition

Imagine a learner studying a moving diagram while also reading a full explanation on screen.

Both sources require visual attention. The learner must repeatedly switch between them, keep earlier information active, locate the corresponding part of the picture, and integrate the two sources before the next information arrives.

Under some conditions, replacing the explanatory paragraph with spoken narration can reduce that competition. The eyes can stay with the visual representation while the words arrive through the auditory channel.

This pattern is commonly called the modality principle or instructional modality effect.

This Is Not “Use More Senses”

The useful idea is not that learning improves whenever more senses are activated.

MindOS asks a narrower question:

Are two essential information streams competing for the same processing route at the same time?

If yes, changing modality may help. If no, narration may add little, or may even make the material harder to control.

The Owned Boundary: This Is Not Spatial-Contiguity State

Spatial-Contiguity State owns a different problem: related visual information is too far apart, so the learner spends effort searching between locations.

Modality State asks whether essential words and graphics are both demanding the visual system at the same moment.

A lesson can satisfy spatial contiguity and still create modality competition. For example, perfectly integrated labels can still become visually dense when a rapidly changing animation and a large explanatory paragraph must be processed together.

The Owned Boundary: This Is Not Segmenting State

Segmenting State slows or divides the information stream so the learner can process one meaningful unit before the next arrives.

Modality State changes where the verbal information is processed. Segmenting changes when the next information arrives.

Sometimes the learner needs both. Sometimes one is enough.

The Owned Boundary: This Is Not Pretraining State

Pretraining State stabilises the main components before a complex explanation begins.

If the learner cannot recognise the parts, changing narration will not repair the missing component model. Teach the parts first.

The Evidence Is Stronger Under Some Conditions Than Others

A 2005 meta-analysis by Paul Ginns synthesised 43 independent modality effects and found a clear average advantage for presenting graphics visually while related verbal explanation was spoken rather than printed. The effect was especially important under system-paced conditions and when the material had high element interactivity.

A larger 2016 meta-analysis included 91 empirical studies. It found small average advantages for narration over visual text for both retention and transfer, but the effect was much larger under particular conditions: system-paced presentations, dynamic pictures, and shorter learning materials.

A 2025 meta-analysis of Richard Mayer’s multimedia-learning research also found the modality principle among the stronger design effects in that corpus, while emphasising that effects across multimedia principles vary substantially by medium, learner group and learning outcome.

The practical conclusion is therefore conditional:

Narration is most plausible when the learner must inspect changing or complex visual information while the verbal explanation is also transient and time-sensitive.

Observable Learner Signatures

  • The learner repeatedly looks away from the diagram to read explanatory text and then struggles to relocate the relevant visual element.
  • They understand the text alone and the picture alone, but lose the relation when both are presented simultaneously.
  • Performance improves when the same visual is narrated instead of accompanied by a long on-screen paragraph.
  • The learner pauses or rewinds because they cannot read and inspect the visual quickly enough.
  • They miss changes in an animation while reading captions or explanatory text.
  • They can follow narrated dynamic material but cannot reconstruct it later without the presentation.
  • They need certain technical terms visible even though narration helps with the main explanation.
  • They perform better with text when they need to reread unfamiliar language or symbols.

These are observations, not diagnoses. Weak language knowledge, poor hearing, inaccessible media, missing prerequisite knowledge, attention drift, weak visual literacy and fast pacing can produce similar behaviour.

Discrimination Test 1: Same Content, Different Modality

Keep the explanation and visuals the same. Present one version with on-screen explanatory text and another with spoken explanation.

Then test:

  • retention;
  • explanation;
  • transfer;
  • where the learner looked or paused;
  • whether reconstruction improves.

If narration helps substantially, visual competition becomes a plausible weak link.

Discrimination Test 2: Narration or Simply Slower Pacing?

Give the learner the text version again, but allow full self-pacing and pauses.

If the text condition recovers when the learner controls the pace, the earlier bottleneck may have been transient information rather than modality itself.

This matters because the 2016 meta-analysis found larger modality effects under system-paced conditions than learner-paced ones.

Discrimination Test 3: Does the Learner Need the Words to Stay Visible?

Some information is poorly suited to narration-only presentation.

  • technical terms;
  • equations;
  • symbolic expressions;
  • unfamiliar names;
  • new vocabulary;
  • precise definitions;
  • information the learner must compare repeatedly.

If the learner repeatedly asks to hear a word again or cannot hold a complex symbolic expression in speech, keep the critical text visible.

Discrimination Test 4: Is Audio Actually Accessible?

Narration is not a universal accessibility solution.

Learners who are Deaf or hard of hearing need captions or other accessible equivalents. Learners studying in a second language may benefit from visible text. A noisy environment may make audio unreliable. Some learners may need transcripts because spoken information disappears too quickly.

Accessibility is not an optional exception to a learning principle. If the learner cannot access the channel, there is no learning advantage to distribute across it.

Discrimination Test 5: Is the Visual Itself Doing Real Work?

If the screen contains decorative images while the real lesson is verbal, narration may not solve a meaningful load problem.

Ask what the learner must actually inspect:

  • a causal animation?
  • a diagram?
  • a graph?
  • a geometric transformation?
  • a worked solution?
  • a map?
  • a visual comparison?

If the visual representation is not carrying essential information, the modality principle may not be the relevant owner.

The Smallest Useful Intervention

Do not convert every word on the screen into speech.

Use narration for the transient explanatory stream while preserving short visual anchors:

  • key labels;
  • equations;
  • technical vocabulary;
  • step numbers;
  • state names;
  • critical conditions;
  • optional captions or transcript access.

This creates a cleaner division of labour: the visual channel can track structure while the auditory channel carries explanation, but the learner still has persistent anchors where persistence matters.

The MindOS Modality Protocol

Step 1 — Name the Target Cognitive Operation

What must the learner ultimately do: explain a mechanism, interpret a graph, follow a transformation, reconstruct a proof, compare structures or select a method?

Step 2 — Identify the Essential Visual Work

What must the learner look at while the explanation is happening?

Step 3 — Identify the Essential Verbal Work

Which words explain what the visual means, why it changes, or how the learner should interpret it?

Step 4 — Decide What Must Remain Visible

Keep stable visual anchors for terms, equations, labels and information that benefits from rereading.

Step 5 — Move Only the Transient Explanation

Use narration for explanatory sentences that would otherwise compete with the visual representation.

Step 6 — Add Learner Control Where Possible

Pause, replay, captions, transcript access and speed control allow the learner to recover when the narration is too transient.

Step 7 — Stop and Reconstruct

Close or pause the presentation. Ask the learner to explain the mechanism, redraw the structure or predict what happens next.

Step 8 — Reduce Presentation Support

Remove narration prompts, reduce highlighted labels, or present a new diagram without guided explanation.

Step 9 — Change the Representation

Use a new graph, diagram, wording or surface example. The learner should recognise the underlying relation without needing the original multimedia package.

Step 10 — Return After Delay

Several days later, ask for explanation or transfer before replaying the instructional media.

Worked Example: Science Animation

A learner studies an animation showing how pressure changes move blood through the heart. The screen also contains a full paragraph describing each stage.

The learner keeps missing valve changes while reading.

Repair:

  • keep valve names and chamber labels visible;
  • move the causal explanation into narration;
  • allow pause/replay;
  • after each stage, stop and ask the learner to predict the next pressure/valve state;
  • later show a fresh heart diagram without narration and ask for the complete explanation.

The learner has not succeeded merely because the video felt easier. Success appears when the mechanism can be reconstructed without the multimedia support.

Worked Example: Mathematics

A teacher demonstrates a geometric transformation while a long written explanation appears beside the moving diagram.

Instead, keep the transformation notation, coordinates and key labels visible while the explanation is spoken. Pause after each transformation and ask the learner to predict the next position.

Then remove the narration and give a new transformation problem. The learner must interpret the notation and perform the operation independently.

Worked Example: English and Language Learning

The boundary changes for language learning.

If the learning target includes decoding unfamiliar words, spelling, vocabulary or second-language comprehension, visible text may be essential rather than redundant. Caption research in second-language learning shows that captions can support comprehension and vocabulary under many conditions.

MindOS therefore does not remove captions simply because narration exists. The learner’s actual language job decides the modality design.

Accessibility Boundary: Captions Are Not a Mistake

Multimedia-learning principles are sometimes misread as a reason to remove captions. That is unsafe and educationally incomplete.

The World Wide Web Consortium’s Web Accessibility Initiative recommends captions and transcripts so audio and video remain accessible across different sensory and situational needs. U.S. accessibility guidance likewise requires captions for many forms of synchronized media.

The correct educational question is not:

“Should captions exist?”

It is:

“What presentation options let this learner access the information while minimising unnecessary competition?”

Accessibility can require multiple equivalent routes. The learner should be able to choose the route that works.

How Do We Know?

Several evidence layers matter.

Evidence Boundary

  • The modality effect is not a universal rule that narration always beats text.
  • Average effects are strongly moderated by pacing, dynamic versus static visuals, duration, complexity and learner characteristics.
  • Self-paced text can reduce or eliminate some disadvantages seen in system-paced visual-text conditions.
  • Technical vocabulary, equations, symbols and unfamiliar words may need persistent visual presentation.
  • Second-language learners may benefit from captions or visible text.
  • Deaf and hard-of-hearing learners require accessible alternatives to narration-only design.
  • A visually easier lesson does not prove learning unless retention, explanation or transfer improves.
  • Claims about separate processing channels are theoretical explanations of observed effects, not direct classroom measurements of a learner’s brain.
  • The learner’s final capability must be tested after multimedia support is reduced.

Common Misconceptions

  • “Audio is always better than text.” False. The advantage is conditional.
  • “Never show captions with narration.” False. Accessibility and language-learning needs can make captions essential.
  • “The more channels we use, the better.” False. Extra information can create redundancy and distraction.
  • “If a learner likes audio, use narration.” Preference alone does not establish the correct learning operation.
  • “If narration helps, the learner has a visual-processing problem.” No clinical inference follows from this educational comparison.
  • “A narrated animation proves understanding.” Only an independent return test can show that.

Technology Rule: Who Performed the Target Cognitive Operation?

Video platforms, AI tutors and learning applications can automatically narrate, caption, highlight and animate instructional content.

That can improve access and reduce unnecessary load. It can also hide whether the learner is still building the model.

Ask:

  • Did the tool merely move words to a better channel?
  • Did the learner still identify the causal relation?
  • Did the learner predict the next state?
  • Can they redraw or explain the model when narration stops?
  • Can they handle a new diagram?

Technology succeeds when it reduces avoidable presentation cost while leaving the learner responsible for the target thinking.

Staged Practice

  1. Supported presentation: visual representation + concise narration + persistent key labels.
  2. Active pause: learner retrieves what changed.
  3. Prediction: learner predicts the next state before narration continues.
  4. Reduced narration: remove some explanatory sentences and require learner completion.
  5. Visual-only reconstruction: learner explains the diagram without narration.
  6. New representation: learner applies the same model to a changed diagram or problem.
  7. Delayed return: learner reconstructs before replaying the media.

Scaffold Fade

At first, the lesson may distribute information across narration and visuals carefully. Later, the learner should need less guided explanation.

Fade in this order when appropriate:

  • full explanation;
  • short narration cues;
  • labels only;
  • unlabelled visual;
  • new representation;
  • independent explanation or solution.

The learner should not become dependent on the voice that originally made the lesson easier to follow.

Transfer Test

Present a new complex visual with no narration. Ask the learner to identify where verbal support would help, what information must remain visible, and then explain the system independently.

Transfer is present when the learner understands the design problem and the underlying concept, rather than merely preferring one media format.

Delayed Independent Return Test

Several days later, test the learner before replaying the lesson.

  • Can they reconstruct the mechanism?
  • Can they explain why each visual state matters?
  • Can they use the idea in a new problem?
  • Can they identify the correct technical terms?
  • Can they perform without narration?

If the learner needs the original voice track to rebuild the idea, the presentation succeeded as support but the learning has not yet become independent.

Examination Implications

Most written examinations remove narration completely.

Therefore a learner who studied successfully through narrated graphics must eventually practise with static diagrams, written questions, equations and unfamiliar representations.

The modality scaffold belongs in learning. Examination readiness requires the learner to reconstruct the underlying model when the scaffold is absent.

Parent and Tutor Teaching Guide

When a child says, “I cannot read this and watch the diagram at the same time,” do not assume they need to concentrate harder.

  • “Which part of the screen do your eyes need to follow?”
  • “Which words need to stay visible?”
  • “Could I explain the moving part aloud while you watch?”
  • “Would pausing solve the problem instead?”
  • “Do you need captions or the transcript as well?”
  • “Now close the media. What happened?”
  • “Can you explain the same idea from this new diagram?”

The goal is not to discover a permanent “visual learner” or “auditory learner.” The goal is to match support to the processing demand of this task, then reduce the support.

MindOS Direction Graph

Complex visual + explanatory words → learner knows components? → no: Pretraining → yes → visual search cost? → Spatial Contiguity → information arriving too fast? → Segmenting → visual and verbal streams competing? → test Modality → preserve accessibility/key text → integrate → reconstruct → reduce narration → new representation → delayed independent return.

If the learner still cannot coordinate the elements after the presentation is simplified, route to Working Memory Load or Chunking. If the learner attends to irrelevant details, use Relevance-Filtering or Signaling. If the support becomes unhelpful as expertise grows, use Expertise-Reversal State.


MindOS rule: do not ask one processing channel to carry two demanding streams when a cleaner division would help. But never confuse a better presentation with a better learner. The final receipt is understanding that survives when the presentation support is reduced.