PSLE-SCI-REALITY-0148
Wait, What? Twenty Coloured Tracks Can Describe One Storm
A weather graphic appears on a news page. One tropical weather system is marked near the coast. From its current position, twenty coloured lines stretch across the map in slightly different directions. Some curve north. Some bend east. A few travel much farther west.
A learner looks at the tangle and says, “There must be twenty storms coming.”
No. The twenty lines may all be forecast members: different model calculations exploring possible futures for the same weather system.
Weather forecasting begins from observations of a huge atmosphere that can never be measured perfectly at every point. Small differences in the estimated starting state can grow as the forecast moves forward. Forecast centres therefore run models more than once, making small changes to starting conditions or model configurations. The resulting family of forecasts is called an ensemble.
When those possible tracks are drawn together, the map can look like tangled noodles. That is why people sometimes call it a spaghetti plot.
The Reality Lab habit is: when you see many forecast lines, first ask what each line represents before counting physical events.
Quick Answer
- Identify the observed weather system at the starting time.
- Read the legend: are the coloured lines different ensemble members, different forecast models, different starting times or actual observed tracks?
- If they are ensemble members, treat them as alternative model solutions for the same evolving system—not as separate storms.
- Look at the spread. Tight clustering can indicate greater agreement among model members; wider spread indicates less agreement about the future path or state.
- Do not treat the ensemble mean or the most crowded line as a guaranteed future.
- Do not interpret spread as a simple probability map unless the official product explicitly defines it that way.
- As new observations arrive, expect later forecasts to change. Updating is part of forecasting, not evidence that the earlier forecast was dishonest.
The Exact Learner Job This Page Owns
This page owns one real-world evidence-transfer job: evaluating an ensemble or spaghetti forecast graphic whose many model lines are mistaken for many separate real weather systems or many guaranteed future tracks.
It does not become a general meteorology manual, a storm-safety service or a mathematical probability textbook. Existing PSLE Science owners retain models, prediction, uncertainty, graph reading and evidence evaluation. Reality Lab applies those skills to a specific public science object that learners may see in weather reports.
- Reality Lab Vol No.041: “The Forecast Says 30°C” — Is That One Certain Future?
- Reality Lab Vol No.133: “The Forecast Cone Gets Wider” — Is the Storm Itself Getting Bigger?
- Reality Lab Vol No.145: “The Weather Radar Is Red Here” — Is It Definitely Raining That Hard at Ground Level?
- Reality Lab Vol No.044: “The Model Fits the Data” — Does That Prove the Explanation?
Original Reality Lab Case: Storm K and the Twenty Futures
This is an original fictional teaching case. It does not reproduce a live storm forecast and should not be used for real-world hazard decisions.
At 8:00 a.m., fictional Storm K is observed at one location. A forecasting centre runs twenty model members. All twenty begin near the analysed position of Storm K, but they do not remain together.
| Forecast lead time | What the members look like | Scientific interpretation |
|---|---|---|
| 12 hours | Most tracks are tightly grouped | The model members broadly agree on the near-term path |
| 48 hours | Tracks form two clusters | Two different future pathways appear plausible within this ensemble |
| 96 hours | Tracks spread widely | Uncertainty in the longer-range path has grown |
How many real storms are represented at the starting time? One: Storm K.
How many model futures are shown? Twenty.
This distinction is the heart of the article. One observed system can have many modelled futures.
Observed, Modelled and Claimed
| Layer | Example |
|---|---|
| Observed | A weather system has a measured and analysed state at the forecast starting time. |
| Model input | The forecast system constructs several slightly different starting states or model configurations. |
| Model output | Each ensemble member produces one possible future evolution. |
| Representation | The member tracks are drawn together on one map. |
| Unsupported claim | Every line is a separate storm that physically exists. |
| Better claim | The lines display a set of model possibilities for the future of the weather system under the ensemble design. |
Why Not Run the Model Just Once?
If scientists knew the exact state of every part of the atmosphere and had a perfect model, one run might seem sufficient. Real forecasting does not have those conditions.
Observations come from weather stations, satellites, aircraft, ships, balloons, radar and other systems. They provide enormous amounts of information, but there are still gaps and measurement limits. Forecast models also simplify a world containing processes across many scales.
NOAA ensemble guidance explains that ensembles use multiple forecasts to sample uncertainty in initial conditions and model behaviour. If slightly different reasonable starting states lead to very different futures, that is scientifically useful information. It tells the forecaster that the future is sensitive to uncertainty in the starting state or model.
The Starting-Condition Check: Tiny Differences Can Grow
Imagine two model members begin with almost the same atmospheric state. One has a pressure value slightly different in one region; another has a slightly altered wind field within the range of plausible analysis uncertainty.
For the first few hours, their forecasts may remain close. Later, the paths may separate. This does not mean one computer “changed its mind” randomly. It can be a consequence of a complex system in which small starting differences grow through time.
A learner should therefore resist a common mistake: different forecast members are not necessarily disagreements caused by incompetence; their disagreement is often deliberately generated to reveal forecast sensitivity.
The Model Check: Members Can Differ for More Than One Reason
Ensembles are not all constructed in exactly the same way. Some vary initial conditions. Some vary parts of the model. Some combine forecasts from different modelling systems. Public “spaghetti” graphics can also mix several independent forecast models rather than members of one formal ensemble.
That makes the legend essential. Before interpreting the tangle, ask:
- Are these members from one ensemble system?
- Are they different forecast models from different centres?
- Are some tracks from older forecast runs?
- Are any lines observations rather than forecasts?
- Does the product show a control run or ensemble mean separately?
Two spaghetti plots can look similar while representing different evidence structures.
The Spread Check: What Does It Mean When the Lines Fan Out?
In a well-designed ensemble, the spread among members gives information about forecast uncertainty. If the members remain close together, they agree more strongly on that aspect of the forecast. If they diverge widely, the ensemble is showing more uncertainty.
But “wide spread” does not mean “the storm has become physically wider”. This is the same category error that Reality Lab Vol No.133 repairs for forecast cones. The graphic is showing uncertainty about a future state, not necessarily the physical size of the weather system.
Nor should you assume that the real future must fall inside the outermost member. An ensemble samples possible outcomes using a particular system. Reality can still lie outside the sampled spread.
The Cluster Check: What If the Lines Split Into Groups?
Sometimes ensemble members form two or more clusters. Perhaps one group turns the storm north while another carries it west. That pattern can be more informative than one smooth average line drawn between them.
Suppose ten members pass north of an island and ten pass south. The average of all twenty might run directly over the island even if almost no individual member takes that path. The mean can therefore describe the centre of the ensemble mathematically without corresponding to one actual member forecast.
The learner habit is: check the members before treating an ensemble average as a literal track.
The Mean Check: Is the Average Track the “Most Likely” Track?
Not automatically.
An ensemble mean combines member forecasts, but the average can smooth out important structure. If the ensemble is tightly clustered, the mean may be a useful summary. If the ensemble has several distinct clusters, the mean may fall into a region where few members actually go.
Official forecast systems use far more information than “draw the average line and call it the answer”. Forecasters examine observations, model performance, ensemble behaviour, known biases, changes between forecast cycles and other evidence.
The Probability Check: Does 12 of 20 Mean Exactly a 60% Chance?
It is tempting to count ensemble members as though they were twenty independent fair lottery tickets. If 12 tracks pass through one region, a learner may say, “The chance is exactly 60%.”
That can be too simple. Ensemble members can share the same model, assumptions and biases. They are generated by a designed prediction system, not by twenty independent universes sampled perfectly at random. Some official ensemble products are calibrated statistically so member frequencies can support probabilistic interpretation, but the meaning depends on the forecast system and product.
For Primary 5/6, the safe rule is: member counts show how the ensemble behaves; use an official probability only when the product explicitly defines one.
The Independence Check: Twenty Lines Are Not Twenty Independent Experiments
Many ensemble members share the same forecasting model, observation network and data-assimilation system. Their differences are intentionally structured variations around a common system.
This means the ensemble is excellent for exploring uncertainty, but a cluster of members is not the same as twenty fully independent research teams reaching the same conclusion from unrelated evidence.
That distinction connects to the broader Reality Lab rule: counting repeated outputs is not always the same as counting independent evidence.
The Time Check: Why Does Spread Usually Grow Farther Into the Future?
Short-range forecasts begin close to the latest observations and analysis. As the forecast moves farther ahead, uncertainty in the starting state and model processes can grow. Ensemble tracks often spread farther apart at longer lead times.
That does not mean long-range forecasting is useless. It means the forecast should communicate decreasing certainty appropriately. The scientific value is not only in predicting one outcome; it is also in revealing when the prediction becomes less constrained.
The Update Check: Why Can Tomorrow’s Spaghetti Plot Look Different?
Forecast systems are run again as new observations arrive. If the atmosphere develops differently from what earlier model members expected, later runs begin from a newer estimate of reality. Tracks can shift, clusters can disappear and spread can shrink or grow.
A changed forecast is not automatically a failed forecast. Forecasting is a repeated process:
- observe;
- estimate the current state;
- model possible futures;
- compare new observations with earlier predictions;
- update.
This is scientific prediction behaving as a living evidence system.
The Representation Check: Colours Can Make One Member Look Special
A graphic designer may draw one line thick and bright while the others are pale. What does the highlighted line mean? It could be the control member, the ensemble mean, an official forecast, a particular model, or simply the website creator’s preferred track.
The colour alone does not establish scientific authority. Read the legend.
Similarly, a map cropped tightly around one cluster can make the spread look larger or smaller. Scale, map projection and geographic window influence visual perception even when the forecast data are unchanged.
Alternative Explanation 1: The Lines Are Different Models, Not One Formal Ensemble
Some public spaghetti plots collect tracks from several forecast models. The interpretation still involves model uncertainty, but the lines may not be exchangeable members of one ensemble. Counting them like identical votes can be misleading.
Alternative Explanation 2: Some Lines Are Older Forecast Runs
A graphic can overlay forecasts issued at different times. An old run and a new run do not begin from exactly the same information. If the page fails to label issuance times, a viewer may mistake model evolution for simultaneous uncertainty.
Alternative Explanation 3: One Track Is the Official Forecast
An official centre may publish a human-produced forecast track alongside ensemble guidance. The official forecast is a synthesis, not simply “one more spaghetti member”. If the graphic mixes them, the legend should make that distinction clear.
Alternative Explanation 4: The Real Storm Has Changed
Sometimes later forecasts shift because the observed system itself changes: its structure, steering environment or surrounding weather pattern evolves. A forecast update can therefore reflect both better observations and a genuinely changing physical system.
What Evidence Would Strengthen an Ensemble-Forecast Claim?
- The legend states exactly what each line represents.
- The forecast issuance time is shown.
- The observed starting position is distinguished from forecast positions.
- The source names the ensemble or model system.
- Official probability products are used when a probability claim is made.
- Ensemble spread is shown without pretending it is the physical storm size.
- Forecast updates remain available so readers can see how uncertainty changed.
- An official forecasting agency explains the product rather than a social-media repost supplying an invented interpretation.
What Would Weaken It?
- A post counts every line as a separate real storm.
- The legend is missing.
- Old and new forecast cycles are mixed without timestamps.
- The ensemble mean is presented as a guaranteed path.
- The outermost lines are treated as a hard boundary that reality cannot cross.
- Member frequency is converted into an exact probability without checking the product definition.
- Wide spread is described as the storm physically expanding.
- One dramatic outlier line is promoted while the rest of the ensemble is hidden.
Worked Case 1: Nineteen Lines East, One Line West
A screenshot highlights the single western outlier and says, “Model predicts direct hit.” The full ensemble shows nineteen members travelling east. The outlier should not be deleted merely because it is unusual, but neither should it be presented as the whole forecast. A responsible interpretation shows the entire ensemble and explains the outlier’s place within it.
Worked Case 2: Two Equal Clusters
Ten members turn north and ten turn south. The ensemble mean runs straight between them. A learner says, “The mean track is the path most members chose.” That is false. No member may follow the average path. The two clusters are the important structure.
Worked Case 3: Tight Early, Wide Late
At 24 hours, all twenty members are close. At 120 hours, they cover a large region. This does not show that the storm grows from a narrow object into a giant one. It shows that the predicted future location becomes less tightly constrained in this ensemble.
Worked Case 4: Twelve of Twenty Cross an Island
A social post says, “Exactly 60% chance.” The count is 12/20, but the scientific meaning depends on the ensemble’s design and calibration. The safer student statement is: “In this ensemble, 12 of 20 members cross the island; consult the official probabilistic forecast for an actual probability statement.”
Worked Case 5: Forecast Changes Overnight
Yesterday’s tracks spread across a wide region. Today’s newer run clusters farther north. Someone says the computer “got caught lying”. A better interpretation is that new observations and a later starting state changed the modelled futures. Forecast skill should be judged systematically across many cases, not by treating normal updates as deception.
Tempting Reasoning That Fails
- “Twenty lines mean twenty storms.” They can be twenty model futures for one storm.
- “The thick line must be the true path.” Its meaning depends on the legend.
- “The average track is guaranteed.” An average is a summary and may not match any member.
- “Wide spread means the storm itself is huge.” Spread represents forecast disagreement, not necessarily physical size.
- “If models disagree, forecasting is useless.” The disagreement is itself information about uncertainty.
- “More members on one side always equal the exact probability.” Probability interpretation depends on ensemble design and calibration.
- “A changed forecast means the earlier scientists were dishonest.” Updating with new evidence is part of forecasting.
Model and Measurement Limits
An ensemble does not contain every physically possible future. It samples possibilities according to its model, resolution, initial-condition method and perturbation design. Shared model biases can affect many members together.
Ensemble spread can also be too narrow or too wide relative to real forecast errors. Forecast centres therefore evaluate and calibrate ensemble systems over many past cases.
A spaghetti plot is therefore best treated as a window into model uncertainty, not as a perfect map of all futures.
How Far Can the Conclusion Travel?
If ensemble members cluster strongly around one future, we can say the model system shows greater agreement about that outcome than when the members are widely spread. If the members split into clusters, we can say the model system supports multiple distinct pathways.
We cannot automatically conclude that:
- each line is a separate physical storm;
- the real future must follow one of the lines exactly;
- the ensemble mean is the most likely path;
- the outer lines form a guaranteed boundary;
- member counts are exact probabilities;
- model agreement proves the physical mechanism is fully understood.
PSLE-Style Transfer Case
A computer model is used to predict the path of one floating object in a changing current. The scientists run the model eight times using slightly different starting estimates because the exact current is uncertain. Eight predicted paths appear on the screen.
Question: Why is it incorrect to say that eight floating objects were observed?
Reasoned answer: Only one real object was being predicted. The eight lines are alternative model outputs produced from slightly different plausible starting conditions. They show uncertainty about the future path, not eight observed objects.
Explained Practice
Practice A: All members are close at 12 hours but widely spread at 96 hours. What changed most clearly? Forecast agreement decreased with lead time.
Practice B: Fifteen members form one cluster and five form another. Can you erase the five? No. They may represent a lower-supported but still plausible pathway within the ensemble.
Practice C: The ensemble mean lies between two clusters and no member follows it. Is the mean “wrong”? Not mathematically. It is a summary, but treating it as a literal member track would misrepresent the ensemble structure.
Practice D: Yesterday’s run differs from today’s. What should you check first? The issue times, new observations and whether the physical system or model analysis changed.
Practice E: A screenshot has twenty lines but no legend. Can you safely call it an ensemble? No. Find the source and product definition first.
Delayed Independent Return: The L-I-N-E-S Check
- L — Legend: What does each line actually represent?
- I — Issue time: Were the lines generated at the same forecast cycle?
- N — Number of real systems: How many physical storms or weather systems were actually observed?
- E — Ensemble spread: Are the members tightly clustered, split into groups or widely spread?
- S — Scope of claim: Are you describing model uncertainty, an official probability or a guaranteed future?
Parent and Tutor Teaching Guide
Put one coin on a sheet of paper to represent the real storm. From the coin, draw ten pencil paths using tiny changes in starting direction. Ask the learner, “How many coins are there?” One. “How many possible paths did we draw?” Ten. This separates the physical object from the model possibilities immediately.
Next draw two clusters of paths and ask the learner to draw their arithmetic average. If the average falls between the clusters, ask whether any actual pencil path followed it. This teaches why a summary can be useful without being a literal forecast member.
Finally, hide the legend and ask the learner what they can safely infer. The correct answer should shrink. That is the lesson: a scientific graphic is not fully interpretable without its provenance and definitions.
Authoritative Sources
- Singapore Examinations and Assessment Board — 2026 PSLE Science Syllabus
- Ministry of Education, Singapore — 2023 Primary Science Teaching and Learning Syllabus
- NOAA Physical Sciences Laboratory — Ensemble Forecast Description
- NOAA/NWS Weather Prediction Center — Ensemble Forecast Training
- NOAA/NWS — Hydrologic Ensemble Forecasting and Uncertainty
- NOAA/NWS Weather Prediction Center — Experimental Ensemble Guidance Tools
NOAA ensemble guidance describes ensembles as collections of model forecasts used to estimate uncertainty and possible future states rather than as observations of multiple real systems. The current Singapore Primary Science framework asks learners to make predictions, interpret information, evaluate methods, consider uncertainty and communicate reasoning. Ensemble graphics are a direct real-world example of those habits.
The Quiet Return
A forecast line is not a storm.
It is a model’s statement about where a storm might go.
When twenty lines fill the map, do not count twenty storms. Count one scientific question being explored twenty ways: given what we know now, how could the future unfold?