Wait, What? A Visual Text Is Not a Picture With Words Added
A Secondary 3 student looks at a poster and describes what can be seen: a bottle, a child, a large headline, a logo. The answer is accurate but weak because the student has catalogued objects rather than explained how the text makes meaning.
A multimodal text is a coordinated meaning system. Image, caption, colour, scale, typography, placement, white space, sequencing, symbols and written language can support, qualify or even contradict one another. Reading therefore requires more than noticing features. The learner must explain what relationship those features create for a particular audience and purpose.
Do not ask only, “What is in the image?” Ask, “What has been made important, what relationship has been built, and how is the viewer being positioned?”
Quick Answer
The Secondary 3 multimodal-reading engine is:
PURPOSE → AUDIENCE → SALIENCE → WORDS → IMAGE → LAYOUT → RELATIONSHIP → PERSPECTIVE → EVIDENCE → EFFECT.
- Purpose: what is the text trying to achieve?
- Audience: whose attention, belief or action matters?
- Salience: what does the viewer notice first and why?
- Words: what claim, tone or instruction is expressed verbally?
- Image: what does the visual represent, imply or symbolise?
- Layout: how are elements ordered, grouped and separated?
- Relationship: do image and words reinforce, extend, complicate or contradict each other?
- Perspective: whose viewpoint is normalised or foregrounded?
- Effect: how does the design shape the viewer’s understanding or response?
Owned Secondary 3 Learning Job
This guide owns one job: helping Secondary 3 students read visual and multimodal texts as designed arguments rather than collections of separate features.
It belongs to the Secondary English Learning Hub and extends the Reading, Inference, Language Effect and Summary Guide. Its Batch 3 companions cover continuous writing, listening and note-taking, and oral communication.
The Current G3 Endpoint
For the 2027 Singapore-Cambridge Secondary Education Certificate G3 English Language syllabus K300, Paper 2 comprehension includes visual text, while Paper 1 Situational Writing involves a given situation with a visual text. This means visual information is not peripheral: students must be able to combine verbal and non-verbal information and explain how meaning is built across modes.
Official reference: SEAB 2027 G3 English Language Syllabus K300.
Start With Purpose and Audience
Before analysing colour or font size, determine the likely communication job.
- Is the text warning, selling, informing, inviting, recruiting, reassuring or persuading?
- Who is expected to respond?
- What action or belief would count as success?
- What concern or desire does the design appear to target?
A bold red headline means little by itself. Its effect depends on what the text is trying to do and to whom.
Salience: What Does the Viewer Notice First?
Salience is created through size, contrast, placement, isolation, colour, focus and unusual imagery. The most salient element often acts as an entry point into the message.
Ask:
- What is largest?
- What is central?
- What contrasts strongly with the background?
- What is isolated by white space?
- What human face, gesture or object attracts attention?
- What appears first in the likely viewing path?
Then connect the salience to purpose. “The headline is large” is observation. “The oversized headline makes the warning the viewer’s first task, before the smaller explanatory detail is read” is analysis.
Worked Example 1: Scale and Consequence
Imagine an environmental poster. A single takeaway cup appears tiny at the bottom of the page. Above it is a huge mound of disposable cups stretching almost to the headline: “One cup does not stay one cup.”
Weak answer: The poster uses a small cup and a big pile to show waste.
Stronger answer: The extreme difference in scale turns one ordinary individual purchase into a visibly large collective consequence. By placing the single cup below the accumulated pile, the design guides the viewer from personal action to system-level waste, reinforcing the warning that repeated small choices add up.
Words and Image Can Reinforce Each Other
When words and image carry the same direction of meaning, they reinforce.
Example: a road-safety poster says “Look twice” beside an image in which a cyclist is almost hidden behind a car pillar. The wording instructs while the image demonstrates why the instruction matters.
Words and Image Can Extend Each Other
Sometimes one mode supplies information the other does not.
Headline: “Your evening routine travels into tomorrow.” Image: a student falling asleep in class while a glowing phone sits beside the bed in a smaller inset image.
The headline gives the general cause-and-effect idea; the paired images make the connection between late-night phone use and next-day fatigue concrete.
Words and Image Can Create Tension
Multimodal texts can be powerful when the words and image do not agree literally.
A campaign poster might show a traffic jam beneath the cheerful phrase “Freedom of the road”. The tension can invite the viewer to question whether widespread car dependence actually creates freedom in a congested city.
Do not assume contradiction means poor design. It may be deliberate irony.
The Image–Words Relationship Test
| Relationship | Question to ask |
|---|---|
| Reinforce | How do both modes push the same message? |
| Extend | What new information does one mode add? |
| Specify | How does the image make an abstract claim concrete? |
| Contrast | What difference between words and image creates meaning? |
| Qualify | How does one mode limit or complicate the other? |
| Sequence | How does the viewer move from one mode to the next? |
Layout Is Argument Architecture
Layout determines what appears to belong together and what feels separate. Proximity can create categories. Columns can compare. Arrows can imply sequence or cause. A top-to-bottom design can move from problem to solution. A before-and-after layout can build contrast.
Ask not merely “Where is it placed?” but “What relationship does that placement create?”
Worked Example 2: Before and After
A public-health infographic places two kitchens side by side. The left side is labelled “after cooking” and shows food scraps in a general waste bin. The right side is labelled “before throwing” and shows edible leftovers boxed for later use, while peelings go into a compost container.
The parallel layout creates a direct comparison between two routines. Because corresponding objects occupy similar positions, viewers can see that the desired behaviour is not a completely different lifestyle but a change in the final disposal decision.
Typography Carries Hierarchy and Tone
Font size, weight, style and spacing can signal which text is headline, instruction, evidence or fine print. Typography can also contribute to tone.
A playful rounded typeface may suit a youth event but undermine a serious emergency warning. A compressed all-capitals phrase may feel urgent or forceful. Small light text can intentionally become secondary.
Do not analyse typography in isolation. Explain the fit between typographic choice and communication job.
Colour Needs Context
Students sometimes rely on memorised colour meanings: red means danger, blue means calm, green means nature. These associations can be useful, but they are not universal rules.
Analyse colour comparatively:
- What contrasts?
- What belongs to the same colour family?
- Which element receives the strongest saturation?
- Does colour separate categories?
- Does it match an organisation or campaign identity?
- What does it do in this specific design?
Framing and Cropping Control What Exists for the Viewer
A close-up creates intimacy and removes surrounding context. A wide shot can make a person seem small relative to the environment. Cropping can exclude causes, alternatives or other people.
Ask: What can the viewer see, and what cannot the viewer see? Perspective is partly built through selection.
Angle and Distance Can Shape Power
A low angle may make a figure appear dominant. A high angle can make a subject appear small or vulnerable. Direct eye contact can create engagement. A distant figure may appear anonymous or representative rather than individual.
Again, avoid rigid rules. Explain how angle or distance works with the rest of the design.
Perspective: Who Gets to Define the Problem?
A multimodal text always selects a frame. A campaign about public transport could show crowded trains, efficient movement, elderly accessibility, environmental benefits or worker schedules. Each choice foregrounds a different problem and stakeholder.
Perspective questions include:
- Whose experience is shown?
- Whose experience is absent?
- Which stakeholder is positioned as expert, victim, customer or decision-maker?
- What assumption is treated as normal?
- What alternative framing could produce a different judgement?
Worked Example 3: Perspective in a Campaign
Poster A about cycling shows a healthy young adult riding on an empty scenic path. Poster B shows a parent cycling with a child on a protected urban lane near shops and public transport.
Poster A frames cycling mainly as leisure and personal wellbeing. Poster B frames it as ordinary urban mobility and family access. Neither image is neutral: each selects a different idea of who cycles and why.
Numbers and Charts Are Also Designed
Graphs can look objective while still making choices about scale, category, time period and emphasis.
- What is being measured?
- What time period is selected?
- Does the vertical axis begin at zero?
- Are categories comparable?
- Is one number highlighted more strongly?
- Does the caption interpret the data for the viewer?
Students do not need advanced statistics to notice that design choices can magnify or minimise apparent differences.
Worked Example 4: Data and Headline
A chart shows a programme’s participation rising from 42% to 48%. The headline says “Participation Surges”.
A careful reader separates observation from framing. Participation increased by six percentage points. Whether “surges” is justified depends on context, baseline and comparison. The headline contributes an evaluative interpretation rather than simply reporting the number.
Call to Action: What Does the Text Want Next?
Many persuasive visual texts end with a behaviour: donate, register, scan, attend, stop, report, share, vote in a school poll or change a habit.
Analyse how the call to action is made easy or urgent. Is there a QR code, deadline, short imperative, benefit, social proof or reduced number of steps?
Audience Positioning
Texts can position the audience as responsible citizens, smart consumers, caring parents, members of a school community, potential victims or people capable of making change.
Second-person pronouns such as you, inclusive pronouns such as we, direct gaze, testimonials and familiar settings can all contribute to positioning.
Comparison Between Two Visual Texts
When comparing two texts, match the dimension. Do not analyse Text A completely and then Text B completely without connecting them.
| Dimension | Text A | Text B |
|---|---|---|
| Audience | Current users | Potential new users |
| Tone | Urgent warning | Optimistic invitation |
| Image strategy | Consequences foregrounded | Benefits foregrounded |
| Call to action | Stop or avoid | Join or adopt |
A good comparison sentence makes the relationship visible: “While Text A motivates through the cost of inaction, Text B motivates through the benefits of participation.”
The Evidence Rule: Point to the Design
Visual-text answers should be evidence-bounded just like written-text answers.
Do not say “The poster is aimed at teenagers” only because it looks colourful. Identify features: student models, school setting, informal second-person language, app-style interface, event timing after school. The conclusion becomes defensible because multiple clues converge.
The Multimodal Evidence Triangle
- Question ↔ Interpretation: does the answer address the requested purpose, audience or effect?
- Interpretation ↔ Feature: does the named design feature support the interpretation?
- Feature ↔ Context: does the feature have that effect in this particular text?
Earliest Weak-Link Diagnosis
| Visible problem | Likely earliest weak link | Repair |
|---|---|---|
| Lists objects | Relationship analysis | Ask what the feature does for purpose |
| Names colour meanings mechanically | Context | Compare colour with surrounding elements |
| Says “large font attracts attention” only | Salience-to-purpose link | Explain why that message must be noticed first |
| Misidentifies audience | Evidence ownership | Collect several audience clues |
| Describes graph without framing | Data interpretation | Separate measured change from headline claim |
| Comparison becomes two analyses | Matched dimension | Use one criterion across both texts |
| Invents intention | Evidence boundary | State only what design features reasonably support |
A Multimodal Reading Practice Cycle
- Look for five seconds and record what you noticed first.
- Identify probable purpose and audience.
- Read all words carefully.
- Describe the main image without interpretation.
- Analyse how image and words relate.
- Map layout and viewing path.
- Identify one perspective or assumption.
- Choose three evidence-rich features.
- Write one answer explaining effect.
- Change the audience and redesign one feature mentally.
The Redesign Test
One of the best ways to prove that a student understands a design choice is to change it.
Ask: What would happen if the child in the poster were replaced by an elderly commuter? If the headline moved to the bottom? If the bright call-to-action button became the smallest element? If the photograph became a chart?
The redesign test exposes the function of the original choice.
The Audience Rotation Test
Take one campaign—reducing food waste, promoting exercise, encouraging reading—and redesign it for:
- Primary school students;
- Secondary students;
- parents;
- working adults.
What changes in image choice, statistics, tone, layout and call to action? Audience becomes visible as a design constraint.
Student Multimodal-Reading Receipt
- What is the text trying to achieve?
- Who is it trying to move?
- What do I notice first?
- Why was that element made salient?
- What do the words claim?
- What does the image add?
- How are elements grouped or sequenced?
- What perspective is foregrounded?
- What evidence supports my audience judgement?
- Does the design motivate through fear, benefit, belonging, responsibility or another route?
- If comparing, am I using the same dimension for both texts?
Parent and Tutor Teaching Guide
When a student gives a feature-only answer, ask: “So what?” If the student says “The headline is large”, ask “Why does that matter here?” If the student says “Red shows danger”, ask “What other element makes danger the relevant interpretation?” The goal is to move from feature naming to relationship and purpose.
Use real everyday texts: transport posters, school announcements, product packaging, public-health graphics, event advertisements and charity campaigns. Remove brand names if necessary and ask the student to infer audience and purpose from design evidence.
Then run the redesign test. If the student can predict how meaning changes when one design choice changes, the analysis is becoming causal rather than descriptive.
Unfamiliar Transfer Challenge
Give the student three unfamiliar multimodal texts on the same topic: a poster, an infographic and a social-media-style campaign tile. Ask for:
- purpose and audience for each;
- first point of salience;
- one image–words relationship;
- one layout effect;
- one perspective or assumption;
- one comparison across all three;
- one redesign recommendation for a different audience.
Then change the topic entirely. Transfer is proven when the same analytical questions still work.
Common Secondary 3 Multimodal Traps
- Describing objects instead of explaining function.
- Using memorised colour meanings without context.
- Calling every image symbolic.
- Ignoring written language while analysing pictures.
- Ignoring image while analysing the headline.
- Assuming the largest element is automatically the most important without explaining why.
- Inventing audience from one weak clue.
- Treating charts as neutral and captions as irrelevant.
- Comparing two texts without a matched dimension.
- Giving a personal reaction instead of explaining intended effect.
Useful eduKate Sengkang Routes
- How Students Read Images, Captions, Layout and Words as One Multimodal Text
- How Author Purpose Shapes Language, Evidence and Structure
- How Comparison Turns Two Texts Into Analytical Judgement
- How Tone, Attitude and Viewpoint Shape Interpretation
Evidence Boundary
The multimodal-analysis labels in this guide are eduKate learning scaffolds, not official examination answer formulas. Actual question wording governs what must be explained. Use the framework to connect visible design evidence to defensible interpretations of purpose, audience and effect.
Continue Secondary 3 English Learning Guide Batch 3
- Continuous Writing: Narrative, Expository, Reflective and Argument Control
- Listening, Note-Taking, Inference and Structure
- Oral Communication: Planned Response and Spoken Interaction
- Return to the Secondary English Learning Hub
The Quiet Return
A visual text becomes readable when the student stops naming features and starts reconstructing the design decisions behind them.
See what is present. Notice what is made important. Connect image and words. Ask whose perspective is being built. Explain what the design is trying to make the viewer do next.